Skip to main content

Evals

Atomic’s model-selection docs are keyed to two live external eval sources rather than a hand-maintained table of scores. This page lists each eval, what it measures, the measured numbers for the models in Atomic’s catalog, and when to reference it for a given workflow role — so an agent authoring a workflow can pick a model for a task type from evidence rather than from an aggregate rank.
No single benchmark is the source of truth. Validate these inputs against Atomic’s own workflow evals, whose task distribution is closer to the work you intend to run. Artificial Analysis was re-fetched on 2026-09-05, including its September 4 Intelligence Index revision; the per-evaluation scores, leaderboard rows and Coding Agent Index rows below were read from the rendered charts on that date. The DeepSWE leaderboard rows below retain their 2026-09-03 compilation date and were read from the live page on 2026-09-05. Every number is a rounded value as displayed by the source; unrounded values and confidence intervals live on the linked pages.

The two sources at a glance

Pick by task type

Start here when a stage needs a model. Each row names the benchmark that measures the task type, the top measured picks and the cheapest pick that stays close, using the tables further down. “Measured” means the exact configuration named; a different effort level or agent is a different row on the source. Three cross-cutting reads from the numbers:
  • Fable 5.1 (max, default fallback) is the broadest model — first or tied-first on Terminal-Bench, AA-Briefcase, GDPval, SciCode, HLE and Omniscience accuracy — but it is the most expensive per AA task ($6.12) and its non-hallucination rate is 27%, below Opus 5 (39%) and Astra (49–55%). It is absent from the DeepSWE snapshot.
  • GPT-6 Astra is the document and terminal specialist — it leads GDP.pdf at every effort level, sits within a point of the Terminal-Bench leader, and its DeepSWE Best row solves 74% in 29 steps — but it trails the Anthropic rows by 5 points on AA-Briefcase and 8–9 on GDPval, where Muse Spark 1.3 max also leads it by 7, and it trails Meta, xAI and Z.AI by 9–11 points on 𝜏³-Banking.
  • The OpenAI budget tier is accurate but overconfident. Luna and Sol score 39–49% on HLE and 81–90% on Terminal-Bench, yet answer wrongly rather than abstain 91–93% of the time when they do not know. Pair them with verification tool nodes; do not use them for uncited research summaries.

DeepSWE — coding-agent performance

DeepSWE is the closest public proxy for what Atomic actually does. Tasks are written from scratch (not scraped from PRs), so no model has seen the solutions; solutions require substantially more code than SWE-bench-style suites; and verifiers test behavior rather than implementation.
  • Current snapshot: DeepSWE v1.1, 113 tasks across 91 repositories and 5 languages, updated September 3, 2026. The site reports 28 measured models and displays 21 leaderboard rows by default, out of 70 published model/effort configurations.
  • Metric: pass@1, plus average cost per task, output tokens, and agent steps.
  • When to reference: default weighting for debugger, worker, and any code-writing role. This is the table that drives Model Selection and Pareto Efficiency.
  • Watch: cost and step count, not just score — a model that passes but takes 268 steps (e.g. sonnet-5) is a poor worker even at a good pass rate, and the two accuracy leaders sit at opposite ends of that axis: Gemini 3.8 Flash leads the highest-published-effort reading the linked pages use at 166 average steps, while the live default Best view’s leader, GPT-6 Astra [xhigh], averages 29.

Live leaderboard, Best view

DeepSWE’s default table is the Best view: the best-scoring effort configuration per model. These are the 21 rows it displayed on 2026-09-05 for the September 3, 2026 snapshot. Model Selection instead tabulates the highest published effort per model, so four rows differ there (gpt-6-astra [max], claude-fable-5 [max], grok-4.6 [xhigh], gemini-3.7-flash [high]). Confidence intervals are DeepSWE’s displayed ±. What the three charts say together:
  • Accuracy is flat at the top. Three models display 74% and a fourth 73%, all inside each other’s confidence intervals. Choose among them on cost and steps, not score.
  • Cost spans two orders of magnitude at the same score. Luna [max] and Gemini 3.8 Flash [high] reach 67% and 74% for 0.61and0.61 and 2.36; Opus 5 [max] and Fable 5 [xhigh] reach 74% and 70% for 11.84and11.84 and 13.41. Sonnet 5 [max] is the outlier to avoid: 54% for $26.40 and 268 steps.
  • Steps predict wall time and tool-call load. Astra [xhigh] (29) and Sol [max] (61) finish in a third of the steps that Gemini 3.8 Flash [high] (166) or DeepSeek V4 Pro [max] (155) need. For a worker loop that pays per tool call or that a reviewer must audit, prefer the low-step row at the same accuracy.
  • The cheap tier is honest about its ceiling. GLM-5.3-Flash [max] 63% at 0.24andDeepSeekV4Flash[max]530.24 and DeepSeek V4 Flash [max] 53% at 0.46 are the only rows under $1 besides Luna; they are budget workers, not judgment gates.

Artificial Analysis: current measures

Intelligence Index v4.2

The September 4, 2026 announcement and current methodology, retrieved 2026-09-05, identify Artificial Analysis Intelligence Index v4.2. It adds AA-Briefcase and GDP.pdf, removes GPQA Diamond from this index, upgrades AA-LCR to v1.1, improves SciCode grading, and rebalances the weights. It also revises GDPval-AA v2 and AA-Briefcase Elo sampling and anchoring. Do not compare scores across index revisions as though only the models changed. The announcement describes v5 as upcoming, not current. The ten evaluations and their contributions are: This is primarily a text-based, English-language suite, not a universal measure of multimodal or multilingual quality. The additional evaluations, such as AutomationBench-AA, AA-AnalystAgent and ITBench-AA, can be better matches for SaaS workflows, spreadsheet analysis or incident diagnosis. Their presence on the site does not make them Intelligence Index components. GPQA Diamond also remains visible separately and in the Engineering capability index.

Headline leaderboard rows for catalog models

From the LLM leaderboard, retrieved 2026-09-05. Cost is AA’s weighted cost per Intelligence Index task (confirmed against the model-page label), not a token price. Speed is output tokens per second on the default 10k-input workload. Latency is AA’s time to first token, which for a streaming reasoning model can be the first reasoning token; end-to-end is seconds to a 500-token answer including thinking. means AA does not report the value. Two things the latency columns make obvious that the index hides: max effort on OpenAI and Anthropic models costs three to eight minutes before the first answer token (Astra max 464 s, Fable 5.1 max 267 s, Terra max 245 s, Luna max 173 s), and the fast interactive tier is Opus 5 high or xhigh (23–34 s), Astra medium or low (7–25 s), Muse Spark 1.3 (19–43 s) and the Flash models (1–16 s). Pick effort for an interactive session from this column, not from the index.

Per-evaluation scores for catalog models

Read from the “Intelligence Evaluations” charts on the AA model pages (GPT-6 Astra, GLM-5.3-Flash, Gemini 3.7 Flash, Claude Sonnet 5, GPT-5.6 Terra) on 2026-09-05. AA-Briefcase and GDPval-AA v2 are Elo scales; the model pages display them as (Elo − 500) / 2000, so 58% is Elo 1666 and 63% is Elo 1769. Everything else is a pass rate. Bold marks the column leader among these rows. Agentic and coding evaluations Reasoning, knowledge and document evaluations How to read the per-evaluation tables:
  • Effort buys different things on different evaluations. Raising Astra from medium to max moves AA-Briefcase from 48% to 53% and GDP.pdf from 30% to 33%, but Terminal-Bench is flat at 88–90% across every level, including low. Sol high beats Sol max on SciCode. Do not assume the top effort is the best row for a coding stage; check the column.
  • Knowledge-work agents and coding agents are different skills. Gemini 3.8 Flash scores 88% on Terminal-Bench and 74% on DeepSWE, yet 35% on AA-Briefcase — the lowest of these rows. GLM-5.3-Flash is the opposite shape: 59% on GDPval, at the level of Opus 5 high, for $0.18 per Index task. Route by the column that matches the stage.
  • Non-hallucination is a family trait, not an intelligence signal. The Z.AI, Meta and xAI models abstain at 66–72%; Anthropic models sit at 27–61%; OpenAI’s Sol, Luna and Terra and both DeepSeek rows sit at 5–12%. For an uncited research summary or a “does this API exist” question, a 7% model needs a verification tool node behind it regardless of its index score.
  • Long context is not a differentiator at the top. Every row except Kimi K3 (89%) sits at 79–85% on AA-LCR v1.1. Choose long-context stages on cost and on the accuracy or document columns instead.

Additional evaluations for catalog models

Not Intelligence Index components. AA measures only some models on each; a blank means AA had no result for that configuration on 2026-09-05, not a zero. GPQA Diamond and MMMU-Pro are included because they remain on the model pages; GPQA is saturated (89–96% across every row) and no longer discriminates. Two rows worth knowing: Gemini 3.7 Flash leads AutomationBench-AA (63%) and AA-AnalystAgent (60%) among measured rows, so it is the SaaS-automation and spreadsheet candidate despite its weak AA-Briefcase; and Sol max leads IFBench (73%) and ITBench-AA (56%), which makes it the strict-format and incident-diagnosis candidate in the OpenAI family.

Coding Agent Index v1.4 is a different comparison

The Artificial Analysis Coding Agent Index evaluates named agent + model + settings combinations, not interchangeable base-model rows. Its methodology, retrieved 2026-09-05, identifies v1.4, current since August 2026. It equally weights DeepSWE, Terminal-Bench v2.1 and SWE-Atlas-QnA. Those components contain 113, 89 and 124 tasks respectively, each with three attempts per task. Per-evaluation pass@1 averages attempts within a task, then tasks within an evaluation. Reward-hacked Terminal-Bench attempts receive zero. Cost and execution time instead pool task attempts across the suite. Cost uses pay-per-token API pricing, including supported cache charges, not subscription-plan prices. Execution time is measured wall time; missing telemetry is excluded from the relevant average, not treated as zero. Agent defaults apply unless the row specifies other settings. The fourteen rows on the rendered leaderboard, read 2026-09-05: Three reads from the agent table:
  • SWE-Atlas-QnA is where the Anthropic rows separate. Fable 5.1 and Opus 5 score 55–56% on repository-understanding questions against 33–52% for every other row; on DeepSWE and Terminal-Bench they are inside the pack. If a stage is mostly reading and explaining code rather than patching it, that column is the one to weight.
  • Muse Spark 1.3 is the value row. Muse Code + Muse Spark 1.3 (max) ties Opus 5 on the index for $1.58 per task, and leads DeepSWE inside AA’s harness at 68. Its xhigh row halves wall time to 12.8 minutes for four index points.
  • Cheap and fast is a real trade. Codex + Luna (max) at 57 costs 0.29andfinishesin8.0minutes;itsitsthreepointsbehindFable5.1onDeepSWEandlosesitsgaponSWEAtlasQnAandTerminalBenchinstead.Codex+DeepSeekV4Flashat50costs0.29 and finishes in 8.0 minutes; it sits three points behind Fable 5.1 on DeepSWE and loses its gap on SWE-Atlas-QnA and Terminal-Bench instead. Codex + DeepSeek V4 Flash at 50 costs 0.06 but trails on all three components.
AA’s DeepSWE component uses the DeepSWE dataset with the named agent. It is not the same experiment as Datacurve’s mini-swe-agent leaderboard, and the two disagree: inside AA’s harness Muse Code + Muse Spark 1.3 (68) edges Codex + Astra (67), while Datacurve’s Best view has Astra [xhigh] at 74% and has not published Muse Spark 1.3 at all (its Muse Spark 1.2 [xhigh] row sits at 55%). Neither its component score nor its composite belongs in the DeepSWE frontier. Earlier versions of these docs referred to a base-model Coding Index and Agentic Index. Neither is listed in the current capability directory or capability methodology inspected on 2026-09-05. We therefore do not assign them a current version or silently rename either to Coding Agent Index. Use the named coding and agentic evaluations above instead.

Professional capability indices

The current directory lists Finance & Accounting, Strategy & Ops, Legal, Healthcare & Medical, Engineering, and Economics. The capability methodology specifies domain-dependent components and weights, rather than a single shared formula. It displays no version identifier. Use the matching domain index when its task mix fits your work.

Price, task cost and latency

Read the definitions and API performance methodology, retrieved 2026-09-05, before comparing efficiency charts:
  • Token prices are USD per million native tokens. AA’s blended price assumes cache-hit, input and output tokens in a 7:2:1 ratio. That synthetic mix is not your workflow’s bill.
  • Intelligence Index cost per task uses actual token consumption, provider prices and typical measured cache hit rates, weighted by the index’s evaluation weights. It is neither the total cost of running the suite nor DeepSWE dollars per task. The leaderboard’s $ column above is this value.
  • Output speed uses standardized o200k_base tokens after the first chunk. The default workload is 10k input tokens; the usual displayed result is the median over 72 hours. The 100k workload instead uses a 14-day median. These are API measurements, not coding-agent completion times.
  • Time to first token can mean the first reasoning token. Time to first answer token includes thinking time. Compare these separately from output speed when interactive latency matters.
  • The homepage’s Intelligence Index Time per Task estimates weighted decode time from output tokens and speed; it excludes TTFT and overhead. Do not call it measured end-to-end wall time. The Coding Agent Index execution-time metric does measure wall time.

Role to benchmark map

See Model Selection for a small dated shortlist and production effort guidance. Benchmark settings are measurement configurations, not instructions to raise every role’s effort.

Keeping the docs fresh

  1. Record each source’s retrieval date separately from its publication or snapshot date. Follow the rendered charts and methodology, not just an old article’s score.
  2. Preserve exact model, reasoning configuration, agent, benchmark version and units. A changed index or agent can change the ranking without a new model release.
  3. Say unmeasured on the named benchmark and date. Missing text extraction is not evidence of absence; inspect the rendered page. Never transfer a predecessor’s score.
  4. Check the configured catalog and live provider access separately. These docs do not change runtime routing or model defaults.
  5. AA’s per-evaluation numbers are only in client-rendered Recharts bar charts, so a plain HTTP fetch returns headings without values. To refresh them, open the model page in a headless browser, scroll the whole page so every chart animates in, then read each chart’s foreignObject labels (model names, in bar order) alongside its svg text nodes (values, in the same order). The DeepSWE leaderboard and the AA evaluation leaderboards (AA-Briefcase, GDPval-AA v2) render as text and fetch cleanly.