An independent AI lab by Raúl WescheLOCAL COMPUTE / OPEN EVIDENCE

Run it.
Break it.
Show it.

Big models on real hardware. Playable worlds, cloud vs local comparisons, and the evidence behind what actually works.

Explore the experiments
DGX SPARK · APPLE SILICON · CLOUD CONTROLS
Qwen-generated red voxel pagoda surrounded by green gardens, a koi pond and cherry blossoms
A world, written by a model.Voxel Pagodas · explore the model comparison archive
15 experimentsSelected worlds & comparisons
4× DGX SparkLocal multi-node research
Mac StudioApple Silicon + orchestration
Open artifactsFailures included
01 / Featured experiment

Same prompt.
Different water.

Astra vs DeepSeek V4.1 Cloud vs Local. Watch the comparison, then try the originals.

03 / Recipes & serving

Our recipes.
Your hardware.

Reproduce the setups: DGX Spark and Apple Silicon, quantization, native MTP and DFlash. Configurations, upstream credits and measured limitations included.

01

GLM Flash, official NVIDIA NVFP4

2× DGX Spark / DFlash2 / diagnostic recipe

Measured acceleration and long-input retrieval, with output-quality caveats. The drafter carries a noncommercial license; this is not production certification.

02

Qwen Flash-Next, FP8 on two Sparks

DGX Spark / vLLM / native MTP

The local serving recipe behind recent long-thinking visuals. Configuration, runtime evidence and practical deployment limits—not a context-free speed claim.

03

Ling-3.0 Flash VL, day-zero serving

2× DGX Spark / FP8 / SGLang

A vision-native model on a development serving stack. Includes a sustained pagoda generation and the setup needed to reproduce it.

04

Native MTP vs DFlash on Mac

Mac Studio / Qwen3.8-27B / MLX

A short-prompt acceleration study with a 4-bit/8-bit follow-up. Serialization and correctness caveats matter as much as the measured throughput.

05

When a serving recipe changes the score

GLM-5.3 Flash EXL3 / Mia recipe / historical evidence

Two pinned recipe revisions, full transcripts and a template-regression investigation. Environment and grader differences are disclosed in the evidence pack.

06

DeepSeek Vision-Exp on Apple Silicon

Mac Studio / DwarfStar / DSpark

A native C + Metal deployment, pagoda artifact and repair record. Configured context capacity is separate from filled-context performance.

07

The Qwen quantization ladder

DGX Spark / NVFP4 · FP8 · GGUF · BF16

Explore precision, engine and speculative-decoding trade-offs. These are deployment comparisons, not a claim that only quantization changed.

08

Reproduce a visual comparison

Matched prompts / raw streams / browser artifacts

The portable harness for preserving model responses, comparing generated HTML and keeping repairs separate from the original output.

04 / Spark-Bench

Test the model.
Keep the evidence.

Our open-source benchmark for local, agentic workloads: executable code, tool use, multi-turn workflows and rendered visual outputs.

Published release / v6.8.0

The benchmark

76 scenarios across 12 domains. The uncapped release preserves responses, grading reasons and run provenance. This link is pinned to the published release, not a moving development branch.

Browse v6.8.0 ↗

Same-version evidence

Qwen vs GLM

The published v6.8.0 comparison: Qwen3.8-Flash-Next and GLM-5.3-Flash, with two repeats across the full suite. Read the per-repeat results, transcripts and exclusions together.

Open the evidence ↗

Historical board / v6.7.1

The Spark leaderboard

Published one-, two- and four-Spark cohorts, with their original methodology labels and generated artifacts. Keep these historical scores separate from v6.8.0 results.

Explore the leaderboard ↗

Published scope: v6.8.0 at 125ba16. An operator benchmark for specific model-and-recipe combinations—not a universal model ranking. Browse the release’s serving recipes ↗

05 / The hardware

A small lab.
A lot of possibility.

Local hardware for ownership and experimentation. Cloud models for comparison. No single machine tells the whole story.

NVIDIA Grace Blackwell

4× DGX Spark

Model serving, multi-node inference, quantization and speculative decoding. One-, two-, and four-node experiments.

Apple Silicon

Mac Studio

Persistent agents, orchestration, browser validation and media production, plus a separate history of MLX and llama.cpp experiments.

Research archive

UGREEN NAS

Verified backups, checkpoint transfers and preserved evidence. Active Spark serving uses local NVMe—not network storage.

Hardware overview, not a live availability dashboard. Runtime, precision, context and topology belong to each individual experiment.

06 / The evidence standard

Keep the failures.
Show the receipts.

A beautiful output is only half the story. The useful part is knowing what generated it, what broke, and what changed before you saw it.

01

Name the whole setup.

Model, quant, engine, hardware, context, thinking policy and tool harness. Tokens per second without a denominator is not enough.

02

Separate original from repaired.

Preserve model output. Label display fixes and follow-up attempts. A gallery thumbnail is not proof of an untouched first attempt.

03

Test what actually renders.

Open the page, exercise the controls, inspect the motion. A passing syntax check is not a working simulation.

04

Don’t turn unlike runs into a ranking.

Different harnesses, timing boundaries or thinking settings get a caveat—not a manufactured speedup. Failed runs are results too.

What is still planned?

Automated twice-daily Visual Arena runs, a unified controller with node leasing, and a repeatable Quant Factory remain roadmap work—not claimed production capabilities. This page features published experiments and sourced research, not completion promises.

Try it. Inspect it. Make it better.

The worlds are playable. The research is public. Bring your own skepticism.

Open SparkBench