Local agent benchmark · thinking off

DeepSeek V4 Flash edges Qwen3.8-27B

The same 76 SparkBench v6.7.1 scenarios, repeated three times at temperature 0.2 and top-p 0.95.

93.01DeepSeek V4 FlashOfficial 0731 · vLLM TP2 · DSpark K5 · 2 DGX Sparks
90.94Qwen3.8-27BNVFP4 · SGLang · MTP3 · 1 DGX Spark
93.38DeepSeek quality
97.31Qwen calibration
2.20sDeepSeek median latency
~93.9Reliability tie

Who won where

DeepSeek led execution-heavy work. Qwen was substantially stronger at knowing when not to act.

DomainQwen3.8-27BDeepSeek V4 FlashWinner
Robustness100.0070.69Qwen
Safety97.3096.85Qwen
Agentic94.0696.30DeepSeek
Classification100.00100.00Tie
Code80.9094.59DeepSeek
Composition100.00100.00Tie
Instruction87.4295.79DeepSeek
Long context55.8855.88Tie
Planning95.4996.85DeepSeek
Structured output96.1497.56DeepSeek
Tool use77.4584.63DeepSeek
Visual81.2384.62DeepSeek

Hardware caveat: this compares each model's recommended local deployment, not equal-hardware efficiency. Qwen used one DGX Spark; DeepSeek used two.

Audit note

Two SQLite graders crashed under Apple's Xcode Python and created false zeroes for both models. All saved SQL responses were regraded under Homebrew Python without regenerating model output; all 12 affected repeats passed. Every other score, latency, token count, and artifact is unchanged.