SparkBench on NVIDIA DGX Spark
NVIDIA DGX Spark · Current Cohorts

SparkBench v6.7.1

One, two, and four Sparks. Qualified results only.

29qualified deployments
40attempted configurations
20 × 3scenarios and repeats
DGX Spark · TP1

Current one-Spark ranking

All rows use the repaired v6.7.1-challenge contract at grader 11d21bf. Older runs are intentionally omitted because different methodologies are not comparable.

29 shown
#DeploymentTrueScoreQualityCalibrationReliabilityMedian latency
1 DeepSeek V4 Flash 0731 DS4 · 1M context · Spec off · 92.0 93.8 92.7 86.8 5.98s
2 DeepSeek V4 Flash 0731 DS4 · 1M context · DSpark · 86.7 88.7 87.1 79.6 5.62s
3 Qwopus AWQ vLLM · 65K context · MTP-K1 · 85.8 78.1 100.0 91.9 6.86s
4 Aeon Ultimate MM NVFP4 vLLM · 65K context · Spec off · 85.3 82.7 93.7 83.8 10.23s
5 Gemma 4 26B-A4B Q4_K_M llama.cpp · 32K context · Spec off · 84.6 79.5 88.9 92.3 0.95s
6 KAT-Coder V2.5 Dev NVFP4 vLLM · 65K context · Spec off · 84.0 76.6 98.4 83.6 1.65s
7 Qwen 3.5 9B BF16 vLLM · 32K context · MTP-K1 · 83.1 74.8 100.0 83.4 3.36s
8 Qwopus AWQ vLLM · 65K context · Spec off · 83.1 78.6 93.7 84.2 10.06s
9 Qwen 3.5 9B BF16 vLLM · 32K context · Spec off · 81.9 71.1 100.0 90.5 5.78s
10 Nemotron 3 Nano Omni Q4_K_M llama.cpp · 32K context · Spec off · 81.0 78.6 84.8 78.9 1.57s
11 Qwen 3.6 27B NVFP4 vLLM · 65K context · Spec off · 80.5 77.4 81.0 91.2 8.48s
12 Agents A1 NVFP4 vLLM · 65K context · Spec off · 79.9 67.7 100.0 87.0 2.37s
13 Qwythos 9B vLLM · 32K context · Spec off · 79.8 69.3 98.0 86.8 6.74s
14 Qwen 3.6 27B Obliteratus Q5_K_M llama.cpp · 32K context · Spec off · 79.6 79.8 71.8 91.8 9.00s
15 Qwen 3.6 27B NVFP4 vLLM · 65K context · DFlash-k8 · 79.5 80.8 73.1 81.6 3.30s
16 Aeon Ultimate MM NVFP4 vLLM · 65K context · MTP-K1 · 79.0 75.8 81.9 85.0 6.78s
17 Qwable 5 27B Coder Q4_K_M llama.cpp · 32K context · Spec off · 78.1 77.2 77.9 81.2 8.01s
18 Holo 3.1 35B-A3B Q4_K_M llama.cpp · 32K context · Spec off · 77.5 73.7 73.7 91.3 0.62s
19 Huihui Qwen 3.6 35B-A3B Q4_K_M llama.cpp · 32K context · Spec off · 77.2 71.9 81.0 84.3 1.41s
20 Ornith 1.0 35B NVFP4 vLLM · 65K context · Spec off · 77.2 78.6 68.2 82.0 2.49s
21 Qwen 3.6 27B PiTune Q4_K_M llama.cpp · 32K context · Spec off · 76.8 76.8 70.7 85.8 7.76s
22 Qwen 3.6 35B Heretic NVFP4 vLLM · 65K context · Spec off · 76.8 74.4 81.0 73.2 2.40s
23 HauHauCS 35B NVFP4 vLLM · 65K context · Spec off · 76.5 80.2 61.7 82.6 2.34s
24 Nemotron 3 Nano Aeon NVFP4 vLLM · 65K context · Spec off · 75.5 76.8 67.0 77.9 1.06s
25 Qwen 3.6 35B-A3B NVFP4 vLLM · 65K context · Spec off · 73.1 74.6 61.8 78.9 1.33s
26 DeepSeek V4 Flash IQ2_XXS llama.cpp · 32K context · Spec off · 68.5 59.4 87.3 70.1 15.10s
27 Ornith 1.0 35B NVFP4 vLLM · 65K context · DFlash-K8 · 68.3 71.4 50.8 77.1 1.37s
28 Bonsai 27B Q1 llama.cpp · 32K context · Spec off · 66.1 57.5 75.3 72.9 2.20s
29 MiniCPM-V 4.6 BF16 vLLM · 32K context · Spec off · 58.9 46.7 61.0 87.6 0.84s

Model-generated 3D tests

Compare the exact qualified artifacts behind the visual scores. Each deployment generated both animations from the same prompts; only the selected pair is loaded.

VIS-04 · Overtake race

VIS-05 · Walk / run cycle

Current two-Spark ranking

Eight TP2 deployments completed the same v6.7.1-challenge contract with zero transport errors. Checkpoint, context, KV precision, and speculative mode are shown per row.

8qualified deployments
0%transport errors
20 × 3scenarios and repeats
DGX Spark · TP2
8 shown
#DeploymentTrueScoreQualityCalibrationReliabilityMedian latency
1DeepSeek V4 Flash 0731NVFP4 · 1M context · DSpark K5 · 87.691.082.284.04.44s
2Solar Open2 250BNVFP4 · 64K context · BF16 KV · Spec off · 86.785.093.781.85.13s
3Laguna M.1NVFP4 · 32K context · NVFP4 KV · Spec off · 86.179.7100.087.56.95s
4Qwen 3.5 397B-A17BINT4 AutoRound · 262K context · FP8 KV · Spec off · 84.583.683.887.83.98s
5Nemotron 3 Super 120B-A12BNVFP4 · 262K context · FP8 KV · Spec off · 82.276.985.693.93.94s
6MiMo V2.5NVFP4 · 1M context · NVFP4 KV · MTP1 · 81.981.181.083.33.19s
7Step 3.7 FlashNVFP4 · 262K context · FP8 KV · Native MTP3 · 81.280.785.782.66.32s
8Inkling SmallNVFP4 · 262K context · Reasoning 0 · Spec off · 80.675.587.384.12.25s

Two-Spark model-generated 3D tests

The exact scored artifacts from each qualified TP2 deployment.

VIS-04 · Overtake race

VIS-05 · Walk / run cycle

Four-Spark full-suite result

GLM 5.2 completed the larger v6.7.1-full contract with 76 scenarios, two repeats, a clean golden gate, and zero transport errors. This full-suite result is kept separate from the 20-scenario challenge rankings above.

95.0TrueScore
97.4%Pass@1 and Pass@K
76 × 2scenarios and repeats
DGX Spark · TP4
#DeploymentTrueScoreQualityCalibrationReliabilityMedian latency
1GLM 5.2 QuantTrioInt4/Int8 mix · 316K context · vLLM TP4 · MTP-K5 · thinking off95.093.199.497.34.34s
25.61tok/s · 985-token prompt
24.83tok/s · 7,733-token prompt
23.59tok/s · 30,861-token prompt
55.9aggregate tok/s · C8

GLM model-generated visual tests

All five exact artifacts from the qualified run. Open any result in its own tab for the full animation.

VIS-01 · Solar system

100.0Open

VIS-02 · Spiral galaxy

100.0Open

VIS-03 · DNA helix

95.8Open

VIS-04 · Overtake race

49.0Open

VIS-05 · Walk / run cycle

91.5Open

Creative generation archive

Solar system, galaxy, DNA, and Frogger outputs from earlier SparkBench visual suites.

73 generations

Archived generation

Open

Qualification contract

  • 12/12 golden gate and tool-call preflight
  • Controlled repeats at temperature 0.3
  • Thinking off where supported
  • Zero transport errors and complete artifacts
  • Clean quarantine status

Reading the score

TrueScore is quality-dominant: quality 55%, calibration 25%, reliability 15%, and speed 5%. These are deployment results, so engine, quantization, context, and speculative mode are part of each measured configuration.

One-, two-, and four-Spark results are kept in separate cohorts because their hardware and scenario contracts differ.