Methodology · v1.3.0

Open Frontier Delta

This index compares model scores from a pinned LiveBench release with the highest-scoring model that LiveBench designates open-weight in each category. It is a sourced benchmark comparison, not a universal measure of intelligence.

Primary sources: the LiveBench release repository, the LiveBench leaderboard, and the LiveBench dataset datasheet.

Seven fixed categories

Version 1.3 uses the seven categories published by LiveBench. The category labels and task mapping come from the release's machine-readable category file; Kronos does not map them into the former four-domain panel.

  • Reasoningreasoning
  • Codingcoding
  • Agentic Codingagentic-coding
  • Mathematicsmathematics
  • Data Analysisdata-analysis
  • Languagelanguage
  • Instruction Followinginstruction-following

Category scores and headline formula

For each model and category, Kronos takes the arithmetic mean of that release's available subtask scores. This matches LiveBench's category averaging. Each of the seven category deltas then receives equal weight in the headline.

Dm,c = Sm,c − Oc

Dm = (1 / 7) × Σc Dm,c

  • Sm,c is model m's LiveBench score in category c.
  • Oc is the highest score in category c among rows whose exact model key has openweight: true in LiveBench's release-pinned model metadata.
  • Scores are expressed in LiveBench percentage points. Zero means parity; +4.2 means 4.2 points above the open frontier; −3.1 means 3.1 points below it.
  • The headline requires a model score and a source-designated open-weight baseline in all seven categories. Partial category values and coverage remain visible, but the headline stays blank.

Automatic open-weight baseline

The ingestion reads the openweight flag from LiveBench's src/Table/modelLinks.js at the same pinned source commit as the score release. It automatically selects the highest-scoring flagged row separately in each category. Exact LiveBench model keys and profiles are created from the same source metadata, so no manual baseline approval is needed.

This designation is LiveBench's classification. Kronos does not infer or verify each model's license from the flag; the public model verification field remains separate. A different model may lead each category.

Release provenance and change detection

The adapter follows the latest date in LiveBench's official release list and fetches the score CSV, category JSON, and model metadata at the same full Git commit. It stores the release ID, commit SHA, SHA-256 file hashes, and check time with its observations and immutable snapshots. It does not fetch optional cost files or scrape leaderboard HTML.

A daily check with unchanged score, category, and model metadata hashes creates no duplicate observations or snapshots. A new release or corrected source file creates new observations and an immutable snapshot. Unknown categories, malformed scores, duplicate model identities, or a category without a scored LiveBench-designated open-weight model hold the batch for review. Published records link to the source release files and model metadata.

LiveBench Overall, coverage, and confidence

LiveBench Overall is shown as a source reference using the release's available category averages. It is not the Kronos headline. Coverage is the share of the seven categories with both a model score and a LiveBench-designated open-weight baseline. For v1.3, High evidence confidence means the source rows passed validation and are pinned to official files; it is not a statistical probability or an evaluation confidence interval.

LiveBench's public score table does not include repeated-run uncertainty for every displayed model and category. Kronos does not estimate missing uncertainty or fill missing category values with zero.

Methodology history and AGI Capability Beta

v1.0 and v1.1 snapshots remain immutable and keep their stored Frontier Uplift values and historical formulas. v1.2 snapshots remain preserved as recorded. v1.3 writes Open Frontier Delta snapshots using automatic, release-specific LiveBench baseline selection; it does not recalculate earlier records or create Artificial Analysis performance observations.

AGI Capability Beta remains unscored. Breadth, depth, reliability, autonomy, adaptation, and robustness are not collapsed into an unsupported index value.

Evidence Confidence rubric

Older manually reviewed records may use the High, Medium, or Low rubric below. The v1.3 LiveBench ingestion assigns High only after validating and pinning the source files. Confidence is reported separately from coverage and does not change Open Frontier Delta.

high

Most relevant factors meet their High criteria; the source is independently checkable and material limitations are minor.

medium

Evidence is usable with at least one disclosed limitation; no critical provenance or comparability gap remains.

low

A critical source, configuration, contamination, reproducibility, or ceiling gap remains. Low-confidence evidence is excluded from scoring.

High, Medium, and Low criteria for each evidence confidence factor
FactorHighMediumLow
Evaluator independenceEvaluator is organizationally independent of the model provider and discloses relevant conflicts.Evaluation is provider-run or co-authored, with enough detail to identify that limitation.Evaluator identity or material conflicts are unclear, or the provider claim cannot be independently checked.
Source accessibilityThe result and the method are publicly accessible at a stable source URL.The result is public, but some supporting method or artifact is restricted or incomplete.The result cannot be inspected in a public source, or the cited source does not support it.
ReproducibilityA named harness, public benchmark version, and sufficient run instructions or artifacts permit an independent rerun.The benchmark and major procedure are clear, but a material reproduction detail is missing.The procedure is too incomplete to reproduce or compare responsibly.
Benchmark version clarityDataset, split, and benchmark version are explicit and match the registered benchmark.The benchmark is identifiable, but one version or split detail remains uncertain.The benchmark version or evaluated split is unknown or materially ambiguous.
Evaluation configuration completenessModel identifier, prompt or agent setup, tools, sampling, and scoring configuration are documented where applicable.The model and major settings are known, with a non-critical configuration gap disclosed.Important model, prompt, tool, sampling, or scoring settings are absent.
Repeated-run evidenceRun count and dispersion or confidence bounds are reported, or deterministic evaluation is justified.A single score is reported with a clear evaluation record and its uncertainty is acknowledged.Run count or stochasticity is unknown and the reported precision is not supported.
Contamination assessmentThe source describes a credible contamination check or a suitably private/held-out test set.Known contamination concerns are discussed and judged manageable with a stated rationale.Material contamination risk is unknown, high, or not addressed.
Human or ceiling provenanceThe human cohort, ceiling type, protocol, score, and source are explicit and comparable.The empirical ceiling is sourced, but cohort or protocol comparability is limited and disclosed.The ceiling is absent, inferred from prose, theoretical, or otherwise not supported as an empirical human result.

View the current index