Read the recipe.
Check the evidence.
Public Weschera repositories, including recipes, tools and forks. Descriptions below are repository-supplied historical summaries—not newly validated performance claims. Read the pinned setup, limitations, license and upstream credits before reproducing a result.
34 public repositories
GLM-5.3-Flash-NVIDIA-NVFP4-2x-DGX-Spark
Official NVIDIA GLM-5.3-Flash NVFP4 on 2x DGX Spark: measured DFlash2 gains, 120K retrieval, raw evidence and honest quality caveats.
GLM-5.3-Flash-EXL3-recipe-regression-2026-09
Evidence pack: GLM-5.3-Flash EXL3 TrueScore 90.9 (Mia recipe pinned b5ab8091) vs 76.8 (repo main 9c0794b) on 2x DGX Spark — chat_template.jinja regression, full transcripts
Ling-3.0-flash-VL-2x-DGX-Sparks
Day-0 SGLang TP2 recipe: Ling-3.0-flash-VL FP8 (vision MoE) on 2x NVIDIA DGX Spark, 128K context, measured speeds + honest negatives
qwen-flash-next-uncensored-fp8-2x-dgx-spark
Qwen3.8-Flash-Next Uncensored FP8 on 2x DGX Spark: vLLM TP2, MTP k=3, genuine 128K qualification. Measured: 51 tok/s code.
spark-x25-4b-dgx-recipe
Spark-X2.5-4B BF16 on one DGX Spark: 128K recipe, pagoda video, raw evidence and honest limitations
qwen38-27b-mlx-mtp-recipe
M4 Max 128GB: Qwen3.8-27B 8-bit MLX native MTP recipe, measured DFlash2 comparison and concurrency caveats
Qwen3.8-Flash-Next-NVFP4-2x-DGX-Spark
NVIDIA official Qwen3.8-Flash-Next NVFP4 on 2x DGX Spark: vLLM TP2 + MTP-3 spec decode, 89.7/100 Spark Bench, 58 tok/s structured. Includes 3 vLLM patches.
GLM-5.3-Flash-Unsloth-1x-DGX-Spark
GLM-5.3-Flash on ONE DGX Spark: Unsloth UD-IQ3_XXS + llama.cpp MTP. Spark Bench TrueScore 93.4 (thinking OFF) — beats our 2-Spark EXL3 TP2. Validated serve config, thinking-off template patch, full evidence.
Qwen3.8-Flash-Next-1x-DGX-Spark
Qwen3.8-Flash-Next (125B MoE, 6B active) on a single NVIDIA DGX Spark: llama.cpp qwen4exp/mtp + Unsloth UD-Q4_K_XL, measured MTP draft sweep, 47 tok/s code
DeepSeek-V4-Flash-Vision-Exp-ds4-Mac-Studio
DeepSeek V4 Flash Vision-Exp on a 128GB Mac Studio with antirez/ds4 (Metal + DSpark): verified recipe, 29.8 tok/s thinking-on pagoda run, honest single/2x DGX Spark comparison
qwen38-flash-next-omlx-mac
Verified Qwen3.8 Flash Next oMLX Lightning MTP deployment guide for Apple Silicon
spark-bench
Published v6.8.0 uncapped benchmark: 76 scenarios across 12 domains. Pinned release, historical results and serving recipes.
dual-model-visual-harness
Reproducible side-by-side visual coding harness for OpenAI-compatible model endpoints
glm53-flash-2spark-tp2
GLM-5.3-Flash NVFP4 (vcruz305 quant) on 2x DGX Spark — vLLM TP2 + native MTP, 128K ctx: measured throughput, MTP acceptance, long-gen instability findings, battle-tested launch scripts
glm53-flash-one-spark
Public project; inspect the repository for scope and status.
qwen38-flashnext-dgx-spark
Qwen3.8-Flash-Next NVFP4 on 2x DGX Spark (SGLang TP2 + NEXTN): 47 tok/s single-stream, 119 tok/s x8, SparkBench 86.3. Credits MiaAI-Lab for the SM121 kernel approach.
qwen38-flashnext-single-spark
Qwen3.8-Flash-Next on ONE DGX Spark - 262K context, llama.cpp + unsloth GGUF
ornith-9b-mac-mini
Public project; inspect the repository for scope and status.
Qwen3.8-27B-NVFP4-DFlash2-DGX-Spark
Reproducible Qwen3.8-27B NVFP4 + DFlash2 recipe and qualification evidence for one NVIDIA DGX Spark
Qwen3.8-27B-oMLX-MTP-Mac
Qwen 3.8 27B at 48-65 tok/s on a Mac Studio M4 Max — measured oMLX native-MTP recipe, same-day A/B vs mlx-vlm+MTP sidecar, raw benchmark JSON included
GLM-5.2-QuantTrio-4x-DGX-Spark
GLM-5.2 (753B-class MoE, QuantTrio Int4-Int8Mix) on 4x NVIDIA DGX Spark — vLLM TP4 over RoCE, 316K context, ~30 tok/s, TrueScore 91.5. Battle-tested launch scripts, verify-before-launch, and troubleshooting from real outages.
wesche-com
wesche.com — Raul Wesche tattoo artist site
reachy-companion
Public project; inspect the repository for scope and status.
Qwen3.8-27B-DGX-Spark-Quant-Ladder
Qwen3.8-27B on DGX Spark: measured quant ladder (bf16/FP8/NVFP4/GGUF), native MTP tuning, thinking-effort ops, and a verify-what-you-launched script
Atlas-DGX-Spark-Quickstart
Build Atlas from source on a DGX Spark and serve Qwen3.8-27B at 30+ tok/s — battle-tested steps, verified numbers, real troubleshooting
Qwen3.8-2.4T-A95B-UD-Q1_0-4x-DGX-Spark
Exact 64K llama.cpp + RoCEv2 RDMA + native MTP3 recipe for the 397GB Qwen3.8 2.4T-A95B UD-Q1_0 GGUF across four NVIDIA DGX Sparks.
atlas
Pure Rust Inference Engine
Qwen3.6-35B-DSpark8-DGX-Spark
Reproducible Qwen3.6-35B NVFP4 + DSpark8 recipe and concurrency evidence for one NVIDIA DGX Spark.
DeepSeek-V4-Flash-0731-DSpark-2x-DGX-Spark
Qualified vLLM TP2 + DSpark K7 recipe for official DeepSeek V4 Flash 0731 on two DGX Sparks
Laguna-S-2.1-NVFP4-1x-DGX-Spark
poolside Laguna S-2.1 (118B/8B MoE) on ONE NVIDIA DGX Spark - validated NVFP4 recipe + day-0 benchmark: TrueScore 86.5, agentic 99.3, 19 tok/s single / 84 tok/s @ c8. Needs vLLM >= 0.25 (older stacks emit gibberish with tools).
qwen-sglang-dgx-spark
Deploy Qwen3.6-35B on a DGX Spark (GB10) with SGLang v0.5.15 for long-context multi-agent serving. Includes a reproducible SGLang-vs-vLLM comparison.
DeepSeek-v4-Flash-DSpark-1M-NVFP4-KV-2x-DGX-Spark
DeepSeek V4 Flash DSpark on 2x DGX Spark — NVFP4 KV cache, 1M context, speculative decoding. Based on MiaAI-Lab dual-Spark packaging.
Qwopus3.6-27B-Q4_K_M-DGX-Spark
Qwopus 27B (Q4_K_M) on NVIDIA DGX Spark (GB10) via llama.cpp — MTP speculative decoding, 49-scenario TrueScore benchmark
specserve
Lightweight web GUI for managing vLLM model serving with speculative decoding
No matching research. Try another model or filter.