DeepSeek V4 Flash 0731
Model profile and sourced results included in the Kronos AI Index. Snapshot values are frozen at the calculation time shown below.
Performance data source: LiveBench. Kronos compares the seven category scores with the highest model LiveBench designates open-weight in each category.
Last LiveBench check: Sep 27, 2026, 12:01 PM UTC · Data status: Stale
Open Frontier Delta
-9.8 pp
LiveBench percentage points relative to the LiveBench-designated open-weight frontier.
Coverage
100%
Eligible
LiveBench Overall
74.17%
Source reference; not the Kronos headline metric.
Evidence Confidence
100
Separate from coverage; based on the evidence rubric.
Methodology version
1.3.0
Calculated Sep 25, 2026, 2:12 PM UTC
Weights
LiveBench-designated open weight
Model listing status: active.
LiveBench category deltas
| Category | Model score | Open-weight baseline | Category delta | Subtasks |
|---|---|---|---|---|
| Reasoning | 86.63% | Kimi K390.67% | -4.0 pp | 4 / 4 |
| Coding | 74.98% | Smaug Agentic82.47% | -7.5 pp | 2 / 2 |
| Agentic Coding | 46.77% | DeepSeek V4.1 Flash Max Effort · 2026-09-1077.27% | -30.5 pp | 3 / 3 |
| Mathematics | 86.79% | DeepSeek V4 Pro 081395.09% | -8.3 pp | 4 / 4 |
| Data Analysis | 79.33% | Smaug Agentic79.9% | -0.6 pp | 3 / 3 |
| Language | 79.18% | Kimi K385.53% | -6.4 pp | 3 / 3 |
| Instruction Following | 65.52% | Qwen 3.8 Flash Next77.11% | -11.6 pp | 4 / 4 |
LiveBench release 2026-06-25 · commit 050d326a7378fd5c70018e85700f32a910c304df · checked Sep 25, 2026, 2:12 PM UTC · model metadata SHA-256 f96c19bf94e15e6ad2cd03b59e4d7efe176292834820058b44c58c0c1098dfef.
Model metadata
- Provider
- DeepSeek
- Open-weight verification
- unverified
- License
- Not provided
- Release date
- Not provided
- Context window
- Not provided
- Input price
- Not provided / 1M tokens
- Output price
- Not provided / 1M tokens
- Output speed
- Not provided
- Median time to first token
- Not provided
- Reasoning mode
- Not provided
Automatically imported from LiveBench release 2026-06-25. Baseline eligibility follows the release-pinned LiveBench openweight flag; independent model-license verification remains unverified.
Benchmark contributions and open-weight baselines
Category score details are shown in the LiveBench category table above.
Archived v1.1 Artificial Analysis evidence
No supporting Artificial Analysis evidence is recorded in this snapshot.
Excluded benchmarks
No benchmark exclusions are recorded in this snapshot.
Score history
| Snapshot time | Method | Metric | Coverage | Confidence |
|---|---|---|---|---|
| Sep 25, 2026, 2:12 PM UTC | 1.3.0 | -9.8 pp | 100% | 100 |
Sources and provenance
Model and weight provenance, pinned LiveBench release files, and any historical evidence links are listed here.
- LiveBench model metadata at 2026-06-25 source commit (opens in a new tab)Checked Sep 25, 2026
- DeepSeek source page (opens in a new tab)Checked Sep 25, 2026
- Hugging Face model page (opens in a new tab)Checked Sep 25, 2026
- LiveBench 2026-06-25 score table (opens in a new tab)Checked Sep 25, 2026
- LiveBench 2026-06-25 category map (opens in a new tab)Checked Sep 25, 2026
AGI Capability Beta
Insufficient evidenceBreadth, depth, reliability, autonomy, adaptation, and robustness are not combined into a score in this release.
- BreadthInsufficient evidence
- DepthInsufficient evidence
- ReliabilityInsufficient evidence
- AutonomyInsufficient evidence
- AdaptationInsufficient evidence
- RobustnessInsufficient evidence
