Public research directory

Read the recipe.
Check the evidence.

Public Weschera repositories, including recipes, tools and forks. Descriptions below are repository-supplied historical summaries—not newly validated performance claims. Read the pinned setup, limitations, license and upstream credits before reproducing a result.

GLM-5.3-Flash-NVIDIA-NVFP4-2x-DGX-Spark

Official NVIDIA GLM-5.3-Flash NVFP4 on 2x DGX Spark: measured DFlash2 gains, 120K retrieval, raw evidence and honest quality caveats.

recipe ↗

GLM-5.3-Flash-EXL3-recipe-regression-2026-09

Evidence pack: GLM-5.3-Flash EXL3 TrueScore 90.9 (Mia recipe pinned b5ab8091) vs 76.8 (repo main 9c0794b) on 2x DGX Spark — chat_template.jinja regression, full transcripts

recipe ↗

Ling-3.0-flash-VL-2x-DGX-Sparks

Day-0 SGLang TP2 recipe: Ling-3.0-flash-VL FP8 (vision MoE) on 2x NVIDIA DGX Spark, 128K context, measured speeds + honest negatives

recipe ↗

qwen-flash-next-uncensored-fp8-2x-dgx-spark

Qwen3.8-Flash-Next Uncensored FP8 on 2x DGX Spark: vLLM TP2, MTP k=3, genuine 128K qualification. Measured: 51 tok/s code.

recipe ↗

spark-x25-4b-dgx-recipe

Spark-X2.5-4B BF16 on one DGX Spark: 128K recipe, pagoda video, raw evidence and honest limitations

recipe ↗

qwen38-27b-mlx-mtp-recipe

M4 Max 128GB: Qwen3.8-27B 8-bit MLX native MTP recipe, measured DFlash2 comparison and concurrency caveats

recipe ↗

Qwen3.8-Flash-Next-NVFP4-2x-DGX-Spark

NVIDIA official Qwen3.8-Flash-Next NVFP4 on 2x DGX Spark: vLLM TP2 + MTP-3 spec decode, 89.7/100 Spark Bench, 58 tok/s structured. Includes 3 vLLM patches.

recipe ↗

GLM-5.3-Flash-Unsloth-1x-DGX-Spark

GLM-5.3-Flash on ONE DGX Spark: Unsloth UD-IQ3_XXS + llama.cpp MTP. Spark Bench TrueScore 93.4 (thinking OFF) — beats our 2-Spark EXL3 TP2. Validated serve config, thinking-off template patch, full evidence.

recipe ↗

Qwen3.8-Flash-Next-1x-DGX-Spark

Qwen3.8-Flash-Next (125B MoE, 6B active) on a single NVIDIA DGX Spark: llama.cpp qwen4exp/mtp + Unsloth UD-Q4_K_XL, measured MTP draft sweep, 47 tok/s code

recipe ↗

DeepSeek-V4-Flash-Vision-Exp-ds4-Mac-Studio

DeepSeek V4 Flash Vision-Exp on a 128GB Mac Studio with antirez/ds4 (Metal + DSpark): verified recipe, 29.8 tok/s thinking-on pagoda run, honest single/2x DGX Spark comparison

recipe ↗

qwen38-flash-next-omlx-mac

Verified Qwen3.8 Flash Next oMLX Lightning MTP deployment guide for Apple Silicon

recipe ↗

spark-bench

Published v6.8.0 uncapped benchmark: 76 scenarios across 12 domains. Pinned release, historical results and serving recipes.

tool ↗

dual-model-visual-harness

Reproducible side-by-side visual coding harness for OpenAI-compatible model endpoints

tool ↗

glm53-flash-2spark-tp2

GLM-5.3-Flash NVFP4 (vcruz305 quant) on 2x DGX Spark — vLLM TP2 + native MTP, 128K ctx: measured throughput, MTP acceptance, long-gen instability findings, battle-tested launch scripts

recipe ↗

glm53-flash-one-spark

Public project; inspect the repository for scope and status.

recipe ↗

qwen38-flashnext-dgx-spark

Qwen3.8-Flash-Next NVFP4 on 2x DGX Spark (SGLang TP2 + NEXTN): 47 tok/s single-stream, 119 tok/s x8, SparkBench 86.3. Credits MiaAI-Lab for the SM121 kernel approach.

recipe ↗

qwen38-flashnext-single-spark

Qwen3.8-Flash-Next on ONE DGX Spark - 262K context, llama.cpp + unsloth GGUF

recipe ↗

ornith-9b-mac-mini

Public project; inspect the repository for scope and status.

recipe ↗

Qwen3.8-27B-NVFP4-DFlash2-DGX-Spark

Reproducible Qwen3.8-27B NVFP4 + DFlash2 recipe and qualification evidence for one NVIDIA DGX Spark

recipe ↗

Qwen3.8-27B-oMLX-MTP-Mac

Qwen 3.8 27B at 48-65 tok/s on a Mac Studio M4 Max — measured oMLX native-MTP recipe, same-day A/B vs mlx-vlm+MTP sidecar, raw benchmark JSON included

recipe ↗

GLM-5.2-QuantTrio-4x-DGX-Spark

GLM-5.2 (753B-class MoE, QuantTrio Int4-Int8Mix) on 4x NVIDIA DGX Spark — vLLM TP4 over RoCE, 316K context, ~30 tok/s, TrueScore 91.5. Battle-tested launch scripts, verify-before-launch, and troubleshooting from real outages.

recipe ↗

wesche-com

wesche.com — Raul Wesche tattoo artist site

fork ↗

reachy-companion

Public project; inspect the repository for scope and status.

tool ↗

Qwen3.8-27B-DGX-Spark-Quant-Ladder

Qwen3.8-27B on DGX Spark: measured quant ladder (bf16/FP8/NVFP4/GGUF), native MTP tuning, thinking-effort ops, and a verify-what-you-launched script

recipe ↗

Atlas-DGX-Spark-Quickstart

Build Atlas from source on a DGX Spark and serve Qwen3.8-27B at 30+ tok/s — battle-tested steps, verified numbers, real troubleshooting

recipe ↗

Qwen3.8-2.4T-A95B-UD-Q1_0-4x-DGX-Spark

Exact 64K llama.cpp + RoCEv2 RDMA + native MTP3 recipe for the 397GB Qwen3.8 2.4T-A95B UD-Q1_0 GGUF across four NVIDIA DGX Sparks.

recipe ↗

atlas

Pure Rust Inference Engine

fork ↗

Qwen3.6-35B-DSpark8-DGX-Spark

Reproducible Qwen3.6-35B NVFP4 + DSpark8 recipe and concurrency evidence for one NVIDIA DGX Spark.

recipe ↗

DeepSeek-V4-Flash-0731-DSpark-2x-DGX-Spark

Qualified vLLM TP2 + DSpark K7 recipe for official DeepSeek V4 Flash 0731 on two DGX Sparks

recipe ↗

Laguna-S-2.1-NVFP4-1x-DGX-Spark

poolside Laguna S-2.1 (118B/8B MoE) on ONE NVIDIA DGX Spark - validated NVFP4 recipe + day-0 benchmark: TrueScore 86.5, agentic 99.3, 19 tok/s single / 84 tok/s @ c8. Needs vLLM >= 0.25 (older stacks emit gibberish with tools).

recipe ↗

qwen-sglang-dgx-spark

Deploy Qwen3.6-35B on a DGX Spark (GB10) with SGLang v0.5.15 for long-context multi-agent serving. Includes a reproducible SGLang-vs-vLLM comparison.

recipe ↗

DeepSeek-v4-Flash-DSpark-1M-NVFP4-KV-2x-DGX-Spark

DeepSeek V4 Flash DSpark on 2x DGX Spark — NVFP4 KV cache, 1M context, speculative decoding. Based on MiaAI-Lab dual-Spark packaging.

recipe ↗

Qwopus3.6-27B-Q4_K_M-DGX-Spark

Qwopus 27B (Q4_K_M) on NVIDIA DGX Spark (GB10) via llama.cpp — MTP speculative decoding, 49-scenario TrueScore benchmark

recipe ↗

specserve

Lightweight web GUI for managing vLLM model serving with speculative decoding

tool ↗