Current one-Spark ranking
All rows use the repaired v6.7.1-challenge contract at grader 11d21bf. Older runs are intentionally omitted because different methodologies are not comparable.
| # | Deployment | TrueScore | Quality | Calibration | Reliability | Median latency |
|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash 0731 DS4 · 1M context · Spec off · | 92.0 | 93.8 | 92.7 | 86.8 | 5.98s |
| 2 | DeepSeek V4 Flash 0731 DS4 · 1M context · DSpark · | 86.7 | 88.7 | 87.1 | 79.6 | 5.62s |
| 3 | Qwopus AWQ vLLM · 65K context · MTP-K1 · | 85.8 | 78.1 | 100.0 | 91.9 | 6.86s |
| 4 | Aeon Ultimate MM NVFP4 vLLM · 65K context · Spec off · | 85.3 | 82.7 | 93.7 | 83.8 | 10.23s |
| 5 | Gemma 4 26B-A4B Q4_K_M llama.cpp · 32K context · Spec off · | 84.6 | 79.5 | 88.9 | 92.3 | 0.95s |
| 6 | KAT-Coder V2.5 Dev NVFP4 vLLM · 65K context · Spec off · | 84.0 | 76.6 | 98.4 | 83.6 | 1.65s |
| 7 | Qwen 3.5 9B BF16 vLLM · 32K context · MTP-K1 · | 83.1 | 74.8 | 100.0 | 83.4 | 3.36s |
| 8 | Qwopus AWQ vLLM · 65K context · Spec off · | 83.1 | 78.6 | 93.7 | 84.2 | 10.06s |
| 9 | Qwen 3.5 9B BF16 vLLM · 32K context · Spec off · | 81.9 | 71.1 | 100.0 | 90.5 | 5.78s |
| 10 | Nemotron 3 Nano Omni Q4_K_M llama.cpp · 32K context · Spec off · | 81.0 | 78.6 | 84.8 | 78.9 | 1.57s |
| 11 | Qwen 3.6 27B NVFP4 vLLM · 65K context · Spec off · | 80.5 | 77.4 | 81.0 | 91.2 | 8.48s |
| 12 | Agents A1 NVFP4 vLLM · 65K context · Spec off · | 79.9 | 67.7 | 100.0 | 87.0 | 2.37s |
| 13 | Qwythos 9B vLLM · 32K context · Spec off · | 79.8 | 69.3 | 98.0 | 86.8 | 6.74s |
| 14 | Qwen 3.6 27B Obliteratus Q5_K_M llama.cpp · 32K context · Spec off · | 79.6 | 79.8 | 71.8 | 91.8 | 9.00s |
| 15 | Qwen 3.6 27B NVFP4 vLLM · 65K context · DFlash-k8 · | 79.5 | 80.8 | 73.1 | 81.6 | 3.30s |
| 16 | Aeon Ultimate MM NVFP4 vLLM · 65K context · MTP-K1 · | 79.0 | 75.8 | 81.9 | 85.0 | 6.78s |
| 17 | Qwable 5 27B Coder Q4_K_M llama.cpp · 32K context · Spec off · | 78.1 | 77.2 | 77.9 | 81.2 | 8.01s |
| 18 | Holo 3.1 35B-A3B Q4_K_M llama.cpp · 32K context · Spec off · | 77.5 | 73.7 | 73.7 | 91.3 | 0.62s |
| 19 | Huihui Qwen 3.6 35B-A3B Q4_K_M llama.cpp · 32K context · Spec off · | 77.2 | 71.9 | 81.0 | 84.3 | 1.41s |
| 20 | Ornith 1.0 35B NVFP4 vLLM · 65K context · Spec off · | 77.2 | 78.6 | 68.2 | 82.0 | 2.49s |
| 21 | Qwen 3.6 27B PiTune Q4_K_M llama.cpp · 32K context · Spec off · | 76.8 | 76.8 | 70.7 | 85.8 | 7.76s |
| 22 | Qwen 3.6 35B Heretic NVFP4 vLLM · 65K context · Spec off · | 76.8 | 74.4 | 81.0 | 73.2 | 2.40s |
| 23 | HauHauCS 35B NVFP4 vLLM · 65K context · Spec off · | 76.5 | 80.2 | 61.7 | 82.6 | 2.34s |
| 24 | Nemotron 3 Nano Aeon NVFP4 vLLM · 65K context · Spec off · | 75.5 | 76.8 | 67.0 | 77.9 | 1.06s |
| 25 | Qwen 3.6 35B-A3B NVFP4 vLLM · 65K context · Spec off · | 73.1 | 74.6 | 61.8 | 78.9 | 1.33s |
| 26 | DeepSeek V4 Flash IQ2_XXS llama.cpp · 32K context · Spec off · | 68.5 | 59.4 | 87.3 | 70.1 | 15.10s |
| 27 | Ornith 1.0 35B NVFP4 vLLM · 65K context · DFlash-K8 · | 68.3 | 71.4 | 50.8 | 77.1 | 1.37s |
| 28 | Bonsai 27B Q1 llama.cpp · 32K context · Spec off · | 66.1 | 57.5 | 75.3 | 72.9 | 2.20s |
| 29 | MiniCPM-V 4.6 BF16 vLLM · 32K context · Spec off · | 58.9 | 46.7 | 61.0 | 87.6 | 0.84s |
Model-generated 3D tests
Compare the exact qualified artifacts behind the visual scores. Each deployment generated both animations from the same prompts; only the selected pair is loaded.
VIS-05 · Walk / run cycle
Current two-Spark ranking
Eight TP2 deployments completed the same v6.7.1-challenge contract with zero transport errors. Checkpoint, context, KV precision, and speculative mode are shown per row.
| # | Deployment | TrueScore | Quality | Calibration | Reliability | Median latency |
|---|---|---|---|---|---|---|
| 1 | DeepSeek V4 Flash 0731NVFP4 · 1M context · DSpark K5 · | 87.6 | 91.0 | 82.2 | 84.0 | 4.44s |
| 2 | Solar Open2 250BNVFP4 · 64K context · BF16 KV · Spec off · | 86.7 | 85.0 | 93.7 | 81.8 | 5.13s |
| 3 | Laguna M.1NVFP4 · 32K context · NVFP4 KV · Spec off · | 86.1 | 79.7 | 100.0 | 87.5 | 6.95s |
| 4 | Qwen 3.5 397B-A17BINT4 AutoRound · 262K context · FP8 KV · Spec off · | 84.5 | 83.6 | 83.8 | 87.8 | 3.98s |
| 5 | Nemotron 3 Super 120B-A12BNVFP4 · 262K context · FP8 KV · Spec off · | 82.2 | 76.9 | 85.6 | 93.9 | 3.94s |
| 6 | MiMo V2.5NVFP4 · 1M context · NVFP4 KV · MTP1 · | 81.9 | 81.1 | 81.0 | 83.3 | 3.19s |
| 7 | Step 3.7 FlashNVFP4 · 262K context · FP8 KV · Native MTP3 · | 81.2 | 80.7 | 85.7 | 82.6 | 6.32s |
| 8 | Inkling SmallNVFP4 · 262K context · Reasoning 0 · Spec off · | 80.6 | 75.5 | 87.3 | 84.1 | 2.25s |
Two-Spark model-generated 3D tests
The exact scored artifacts from each qualified TP2 deployment.
Four-Spark full-suite result
GLM 5.2 completed the larger v6.7.1-full contract with 76 scenarios, two repeats, a clean golden gate, and zero transport errors. This full-suite result is kept separate from the 20-scenario challenge rankings above.
| # | Deployment | TrueScore | Quality | Calibration | Reliability | Median latency |
|---|---|---|---|---|---|---|
| 1 | GLM 5.2 QuantTrioInt4/Int8 mix · 316K context · vLLM TP4 · MTP-K5 · thinking off | 95.0 | 93.1 | 99.4 | 97.3 | 4.34s |
GLM model-generated visual tests
All five exact artifacts from the qualified run. Open any result in its own tab for the full animation.
Creative generation archive
Solar system, galaxy, DNA, and Frogger outputs from earlier SparkBench visual suites.
Archived generation
Qualification contract
- 12/12 golden gate and tool-call preflight
- Controlled repeats at temperature 0.3
- Thinking off where supported
- Zero transport errors and complete artifacts
- Clean quarantine status
Reading the score
TrueScore is quality-dominant: quality 55%, calibration 25%, reliability 15%, and speed 5%. These are deployment results, so engine, quantization, context, and speculative mode are part of each measured configuration.
One-, two-, and four-Spark results are kept in separate cohorts because their hardware and scenario contracts differ.
