Current two-Spark ranking
All rows use v6.7.1-challenge at grader 11d21bf. Scores compare complete deployments; checkpoint, context, KV precision, and speculative mode are shown per row.
8 shown
| # | Deployment | TrueScore | Quality | Calibration | Reliability | Median latency |
|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash 0731NVFP4 · 1M context · DSpark K5 · | 87.6 | 91.0 | 82.2 | 84.0 | 4.44s |
| 2 | Solar Open2 250BNVFP4 · 64K context · BF16 KV · Spec off · | 86.7 | 85.0 | 93.7 | 81.8 | 5.13s |
| 3 | Laguna M.1NVFP4 · 32K context · NVFP4 KV · Spec off · | 86.1 | 79.7 | 100.0 | 87.5 | 6.95s |
| 4 | Qwen 3.5 397B-A17BINT4 AutoRound · 262K context · FP8 KV · Spec off · | 84.5 | 83.6 | 83.8 | 87.8 | 3.98s |
| 5 | Nemotron 3 Super 120B-A12BNVFP4 · 262K context · FP8 KV · Spec off · | 82.2 | 76.9 | 85.6 | 93.9 | 3.94s |
| 6 | MiMo V2.5NVFP4 · 1M context · NVFP4 KV · MTP1 · | 81.9 | 81.1 | 81.0 | 83.3 | 3.19s |
| 7 | Step 3.7 FlashNVFP4 · 262K context · FP8 KV · Native MTP3 · | 81.2 | 80.7 | 85.7 | 82.6 | 6.32s |
| 8 | Inkling SmallNVFP4 · 262K context · Reasoning 0 · Spec off · | 80.6 | 75.5 | 87.3 | 84.1 | 2.25s |
Model-generated 3D tests
The exact scored artifacts from each qualified deployment.
VIS-05 · Walk / run cycle
Fixed benchmark contract
v6.7.1-challenge, grader11d21bf- 20 challenge scenarios, three repeats
- Thinking off, temperature 0.3
- Golden gate and endpoint preflights passed
- Zero transport errors across all eight reports
Reading the board
This cohort measures deployed systems rather than base-model intelligence alone. Quantization, context length, KV precision, and speculative decoding differ by recipe. Throughput was skipped, so median turn latency is reported without a tokens-per-second claim.
