AnthropicNot designated as open-weight in LiveBench metadata

Claude Sonnet 5 xHigh Effort

Model profile and sourced results included in the Kronos AI Index. Snapshot values are frozen at the calculation time shown below.

Performance data source: LiveBench. Kronos compares the seven category scores with the highest model LiveBench designates open-weight in each category.

Last LiveBench check: Sep 27, 2026, 12:01 PM UTC · Data status: Stale

Open Frontier Delta

-8.0 pp

LiveBench percentage points relative to the LiveBench-designated open-weight frontier.

Coverage

100%

Eligible

LiveBench Overall

76.04%

Source reference; not the Kronos headline metric.

Evidence Confidence

100

Separate from coverage; based on the evidence rubric.

Methodology version

1.3.0

Calculated Sep 25, 2026, 2:12 PM UTC

Weights

Not designated as open-weight in LiveBench metadata

Model listing status: active.

LiveBench category deltas

LiveBench category scores and differences from the LiveBench-designated open-weight frontier
CategoryModel scoreOpen-weight baselineCategory deltaSubtasks
Reasoning88.69%Kimi K390.67%-2.0 pp4 / 4
Coding80.68%Smaug Agentic82.47%-1.8 pp2 / 2
Agentic Coding59.39%DeepSeek V4.1 Flash Max Effort · 2026-09-1077.27%-17.9 pp3 / 3
Mathematics92.94%DeepSeek V4 Pro 081395.09%-2.1 pp4 / 4
Data Analysis71.74%Smaug Agentic79.9%-8.2 pp3 / 3
Language74.97%Kimi K385.53%-10.6 pp3 / 3
Instruction Following63.86%Qwen 3.8 Flash Next77.11%-13.2 pp4 / 4

LiveBench release 2026-06-25 · commit 050d326a7378fd5c70018e85700f32a910c304df · checked Sep 25, 2026, 2:12 PM UTC · model metadata SHA-256 f96c19bf94e15e6ad2cd03b59e4d7efe176292834820058b44c58c0c1098dfef.

Model metadata

Provider
Anthropic
Open-weight verification
Not designated by LiveBench
License
Not provided
Release date
Not provided
Context window
Not provided
Input price
Not provided / 1M tokens
Output price
Not provided / 1M tokens
Output speed
Not provided
Median time to first token
Not provided
Reasoning mode
Not provided

Automatically imported from LiveBench release 2026-06-25. Baseline eligibility follows the release-pinned LiveBench openweight flag; independent model-license verification remains unverified.

Benchmark contributions and open-weight baselines

Category score details are shown in the LiveBench category table above.

Archived v1.1 Artificial Analysis evidence

No supporting Artificial Analysis evidence is recorded in this snapshot.

Excluded benchmarks

No benchmark exclusions are recorded in this snapshot.

Score history

Immutable Open Frontier Delta and historical Frontier Uplift snapshots over time
Snapshot timeMethodMetricCoverageConfidence
Sep 25, 2026, 2:12 PM UTC1.3.0-8.0 pp100%100

Sources and provenance

Model and weight provenance, pinned LiveBench release files, and any historical evidence links are listed here.

AGI Capability Beta

Insufficient evidence

Breadth, depth, reliability, autonomy, adaptation, and robustness are not combined into a score in this release.

  • BreadthInsufficient evidence
  • DepthInsufficient evidence
  • ReliabilityInsufficient evidence
  • AutonomyInsufficient evidence
  • AdaptationInsufficient evidence
  • RobustnessInsufficient evidence

Methodology and limitations