Benchmark health
Catalog discovery, scraper implementation, admission, and rank eligibility are separate stages. This page reports each without turning missing publication into a claim that a benchmark is dead.
What we can answer today
Each card is a kind of work people use AI for. Open one to see how the options compare.
Evidence current as of 40 days ago
Coding (general)
ExploratoryRank coding models and agents on general software tasks: code generation, repository-level SWE, and terminal work
Reached: Admitted
7 implemented · 7 admitted · 0 rank-eligible
MCP tool orchestration
ExploratoryCoordinate MCP tools to complete a task
Reached: Admitted
3 implemented · 3 admitted · 0 rank-eligible
General knowledge QA
ExploratoryAnswer broad knowledge questions accurately
Reached: Admitted
3 implemented · 3 admitted · 0 rank-eligible
RAG and retrieval
ExploratoryRetrieve relevant context and ground an answer
Reached: Admitted
13 implemented · 13 admitted · 0 rank-eligible
Coding (frontend)
ExploratoryRank models and agents on building and iterating web interfaces from a specification
Reached: Admitted
1 implemented · 1 admitted · 0 rank-eligible
Reasoning
ExploratorySolve novel multi-step reasoning tasks
Reached: Admitted
3 implemented · 3 admitted · 0 rank-eligible
Factuality
ExploratoryProduce correct and grounded factual claims
Reached: Admitted
3 implemented · 3 admitted · 0 rank-eligible
Writing
ExploratoryProduce high-quality creative and long-form prose from an open prompt, judged on human-preference writing quality.
Reached: Admitted
2 implemented · 2 admitted · 0 rank-eligible
Security code review
ExploratoryFind security vulnerabilities in real application code and fix them without breaking legitimate behavior
Reached: Admitted
4 implemented · 2 admitted · 0 rank-eligible
Code review
ExploratoryReview a code change and catch the real bugs it introduces without burying them in invalid or low-value comments
Reached: Admitted
2 implemented · 1 admitted · 0 rank-eligible
16 more kinds of work, not ranked yetShow
EvalRank understands these questions but does not yet have enough independent evidence to rank them. They are listed so the gaps are visible.
Function and tool calling
Not rankedEmit correct schema-valid tool calls
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Web browsing and navigation
Not rankedRetrieve and act on live web content
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Computer use
Not rankedOperate a graphical interface to complete a task
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Deep research
Not rankedSynthesize multiple sources with traceable citations
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Customer support agent
Not rankedResolve user support issues end to end
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Enterprise and CRM workflow
Not rankedExecute business workflows across enterprise systems
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Mathematical reasoning
Not rankedSolve quantitative or symbolic problems
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Long-term memory
Not rankedPersist and recall useful information across sessions
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Finance
Not rankedPerform domain-grounded financial reasoning and workflows
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Legal
Not rankedPerform domain-grounded legal reasoning and drafting
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Medical
Not rankedPerform domain-grounded clinical reasoning and question answering
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Multilingual
Not rankedMaintain quality across languages and translation tasks
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Vision and multimodal
Not rankedReason over images, audio, or video
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
SRE incident response
Not rankedDiagnose and repair live service or infrastructure incidents
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Professional deliverables
Not rankedCreate review-ready professional work products from a complete brief, domain context, and reference files.
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Computational research reproduction
Not rankedReproduce published computational results by implementing or executing experiments from papers, code, data, and environments.
Reached: In catalog
0 implemented · 0 admitted · 0 rank-eligible
Health generated Aug 27, 2026, 12:00 AM UTC