SparkBench on NVIDIA DGX Spark
NVIDIA DGX Spark · Current Cohort

SparkBench v6.7.1

Two Sparks. One contract. Qualified results only.

8qualified deployments
0%transport errors
20 × 3scenarios and repeats
DGX Spark · TP2

Current two-Spark ranking

All rows use v6.7.1-challenge at grader 11d21bf. Scores compare complete deployments; checkpoint, context, KV precision, and speculative mode are shown per row.

8 shown
#DeploymentTrueScoreQualityCalibrationReliabilityMedian latency
1DeepSeek V4 Flash 0731NVFP4 · 1M context · DSpark K5 · 87.691.082.284.04.44s
2Solar Open2 250BNVFP4 · 64K context · BF16 KV · Spec off · 86.785.093.781.85.13s
3Laguna M.1NVFP4 · 32K context · NVFP4 KV · Spec off · 86.179.7100.087.56.95s
4Qwen 3.5 397B-A17BINT4 AutoRound · 262K context · FP8 KV · Spec off · 84.583.683.887.83.98s
5Nemotron 3 Super 120B-A12BNVFP4 · 262K context · FP8 KV · Spec off · 82.276.985.693.93.94s
6MiMo V2.5NVFP4 · 1M context · NVFP4 KV · MTP1 · 81.981.181.083.33.19s
7Step 3.7 FlashNVFP4 · 262K context · FP8 KV · Native MTP3 · 81.280.785.782.66.32s
8Inkling SmallNVFP4 · 262K context · Reasoning 0 · Spec off · 80.675.587.384.12.25s

Model-generated 3D tests

The exact scored artifacts from each qualified deployment.

VIS-04 · Overtake race

VIS-05 · Walk / run cycle

Fixed benchmark contract

  • v6.7.1-challenge, grader 11d21bf
  • 20 challenge scenarios, three repeats
  • Thinking off, temperature 0.3
  • Golden gate and endpoint preflights passed
  • Zero transport errors across all eight reports

Reading the board

This cohort measures deployed systems rather than base-model intelligence alone. Quantization, context length, KV precision, and speculative decoding differ by recipe. Throughput was skipped, so median turn latency is reported without a tokens-per-second claim.