回到看板Back to the board
上線中Live

Syco Eval

開源、公開上線的多模型諂媚度排行榜,每兩週自動重跑一次。11種情境、18個主流模型、中英文都測,專門抓模型在多輪對話裡撐不撐得住社會壓力。Am I Sure?就是蓋在這份研究上面的稽核服務。 An open-source leaderboard, live and public, that reruns itself every two weeks. It puts 18 frontier models through 11 pressure scenarios in both English and Mandarin, tracking whether they hold their ground across a multi-turn conversation. Am I Sure? is the audit service built on top of it.

Syco Eval — 多模型諂媚度排行榜,含分數排名與七維度雷達圖

現有的諂媚基準大多只測單輪問答,但真正的諂媚是在對話裡一步步發生的——模型第一輪答對,可能撐到第三輪社會壓力上來就改口了,「什麼時候鬆口」跟「會不會鬆口」一樣重要。這個排行榜每個情境都跑4到5輪升壓對話,英文中文各跑一次,每兩週自動重測一次,不是論文發表當下拍一張照片就算數的靜態榜單。 Most sycophancy benchmarks only test single-turn Q&A, but real sycophancy plays out over a conversation — a model that gets it right on turn one can still cave by turn three once the social pressure builds, and when it caves matters as much as whether it does. This leaderboard runs 4–5 escalating pressure turns per scenario, in both English and Mandarin, and reruns itself automatically every two weeks instead of being a static snapshot from whenever a paper got published.

評分用三家不同供應商的LLM當裁判、多數決定案,不會讓單一裁判自己的偏好左右結果。 Scoring uses a majority vote across three different LLM vendors as judges, so no single judge's own bias decides the result.