Qwen Flash-Next FP8 · vLLM TP2
82,830 completion tokens · 2,073.2s wall time · finish: stop
The canonical voxel-pagoda prompt, interpreted by Qwen Flash-Next and Ling Flash VL. Orbit both gardens and compare the detail—not just generation speed.
Left-to-right panel order matches the labels above. These are browser recordings of generated code, not recordings of model inference. On mobile, rotate your device or use fullscreen to inspect details.
82,830 completion tokens · 2,073.2s wall time · finish: stop
22,139 completion tokens · 862.2s wall time · finish: stop
Both requests enabled thinking, requested xhigh, used temperature 0.7 and an explicit 120,000-token output budget. An identical effort label does not establish equivalent reasoning policy across model families.
Both used two DGX Sparks, but different model architectures and serving engines. Qwen’s recipe used MTP; Ling’s did not. No HTML repairs are recorded for the displayed pagoda artifacts.
Reasoning was recorded as characters in the run metadata, not a verified per-arm reasoning-token split. Completion-token counts include reasoning. The video is a cropped excerpt of the existing comparison.
Completion counts include reasoning. Reasoning character counts are not token counts. Saved wall times use each run’s original measurement boundary; no cross-model speedup or causal quantization claim is made here.
The Ling output is the same saved generation used in both the Qwen vs Ling and Nex vs Ling comparisons, not an independent second sample.