Local agent benchmark · thinking off
DeepSeek V4 Flash edges Qwen3.8-27B
The same 76 SparkBench v6.7.1 scenarios, repeated three times at temperature 0.2 and top-p 0.95.
93.01DeepSeek V4 FlashOfficial 0731 · vLLM TP2 · DSpark K5 · 2 DGX Sparks
90.94Qwen3.8-27BNVFP4 · SGLang · MTP3 · 1 DGX Spark
93.38DeepSeek quality
97.31Qwen calibration
2.20sDeepSeek median latency
~93.9Reliability tie
Who won where
DeepSeek led execution-heavy work. Qwen was substantially stronger at knowing when not to act.
| Domain | Qwen3.8-27B | DeepSeek V4 Flash | Winner |
|---|---|---|---|
| Robustness | 100.00 | 70.69 | Qwen |
| Safety | 97.30 | 96.85 | Qwen |
| Agentic | 94.06 | 96.30 | DeepSeek |
| Classification | 100.00 | 100.00 | Tie |
| Code | 80.90 | 94.59 | DeepSeek |
| Composition | 100.00 | 100.00 | Tie |
| Instruction | 87.42 | 95.79 | DeepSeek |
| Long context | 55.88 | 55.88 | Tie |
| Planning | 95.49 | 96.85 | DeepSeek |
| Structured output | 96.14 | 97.56 | DeepSeek |
| Tool use | 77.45 | 84.63 | DeepSeek |
| Visual | 81.23 | 84.62 | DeepSeek |
Hardware caveat: this compares each model's recommended local deployment, not equal-hardware efficiency. Qwen used one DGX Spark; DeepSeek used two.
Audit note
Two SQLite graders crashed under Apple's Xcode Python and created false zeroes for both models. All saved SQL responses were regraded under Homebrew Python without regenerating model output; all 12 affected repeats passed. Every other score, latency, token count, and artifact is unchanged.