Pick the right AI for the job
Compare models on real benchmark evidence, see how strong that evidence is, and save the exact reasoning behind your choice.
What we can answer today
Each card is a kind of work people use AI for. Open one to see how the options compare.
Evidence current as of 40 days ago
Coding (general)
ExploratoryRank coding models and agents on general software tasks: code generation, repository-level SWE, and terminal work
Reached: Admitted
MCP tool orchestration
ExploratoryCoordinate MCP tools to complete a task
Reached: Admitted
General knowledge QA
ExploratoryAnswer broad knowledge questions accurately
Reached: Admitted
RAG and retrieval
ExploratoryRetrieve relevant context and ground an answer
Reached: Admitted
Coding (frontend)
ExploratoryRank models and agents on building and iterating web interfaces from a specification
Reached: Admitted
Reasoning
ExploratorySolve novel multi-step reasoning tasks
Reached: Admitted
Factuality
ExploratoryProduce correct and grounded factual claims
Reached: Admitted
Writing
ExploratoryProduce high-quality creative and long-form prose from an open prompt, judged on human-preference writing quality.
Reached: Admitted
Security code review
ExploratoryFind security vulnerabilities in real application code and fix them without breaking legitimate behavior
Reached: Admitted
Code review
ExploratoryReview a code change and catch the real bugs it introduces without burying them in invalid or low-value comments
Reached: Admitted
16 more kinds of work, not ranked yetShow
EvalRank understands these questions but does not yet have enough independent evidence to rank them. They are listed so the gaps are visible.
Function and tool calling
Not rankedEmit correct schema-valid tool calls
Reached: In catalog
Web browsing and navigation
Not rankedRetrieve and act on live web content
Reached: In catalog
Computer use
Not rankedOperate a graphical interface to complete a task
Reached: In catalog
Deep research
Not rankedSynthesize multiple sources with traceable citations
Reached: In catalog
Customer support agent
Not rankedResolve user support issues end to end
Reached: In catalog
Enterprise and CRM workflow
Not rankedExecute business workflows across enterprise systems
Reached: In catalog
Mathematical reasoning
Not rankedSolve quantitative or symbolic problems
Reached: In catalog
Long-term memory
Not rankedPersist and recall useful information across sessions
Reached: In catalog
Finance
Not rankedPerform domain-grounded financial reasoning and workflows
Reached: In catalog
Legal
Not rankedPerform domain-grounded legal reasoning and drafting
Reached: In catalog
Medical
Not rankedPerform domain-grounded clinical reasoning and question answering
Reached: In catalog
Multilingual
Not rankedMaintain quality across languages and translation tasks
Reached: In catalog
Vision and multimodal
Not rankedReason over images, audio, or video
Reached: In catalog
SRE incident response
Not rankedDiagnose and repair live service or infrastructure incidents
Reached: In catalog
Professional deliverables
Not rankedCreate review-ready professional work products from a complete brief, domain context, and reference files.
Reached: In catalog
Computational research reproduction
Not rankedReproduce published computational results by implementing or executing experiments from papers, code, data, and environments.
Reached: In catalog
Health generated Aug 27, 2026, 12:00 AM UTC