Quant Factory
Turn source checkpoints into practical local deployments and measure the quality cost instead of guessing.
- Q8 / Q5 / Q4 controls
- Hashes + toolchain provenance
- Quality, speed, memory, size
Quants, visual stress tests, and reproducible inference across Apple Silicon and NVIDIA DGX Spark. The hardware changes. The evidence standard does not.
The DGX cluster is one part of the system. Mac Studio testing, quantization research, model archiving, browser validation, and public evidence all belong under the same lab identity.
Turn source checkpoints into practical local deployments and measure the quality cost instead of guessing.
Give local models the same uncapped creative task, preserve every token, and test what actually renders.
Treat engine, quant, context, and speculative decoding as part of the measured system.
No artificial hardware hierarchy. Apple Silicon, NVIDIA Grace Blackwell, and network storage each answer a different part of the local-model question.
Local inference, MLX and llama.cpp experiments, grading, browser validation, orchestration, and media production.
NVFP4 and FP8 serving, long-context tests, tensor-parallel deployments, speculative decoding, and multi-node research.
Source checkpoints, verified quants, immutable outputs, hash manifests, videos, and evidence that survives the active job.
These are not screenshots generated for a post. Each scene began as one model response, was preserved as HTML, opened in a real browser, and published with repairs disclosed.

The same uncapped solar-system brief on two-Spark TP2 and four-Spark TP4 deployments.
Open experiment ↗
Two long-reasoning local models build complete interactive universes.
Open experiment ↗
Metallic paper folds, transforms, flies, and breathes fire.
Open experiment ↗
Two models turn a calm museum object into a living supercell.
Open experiment ↗A result is only useful when the checkpoint, quant, runtime, request, raw stream, browser behavior, and any intervention can all be traced.
Model revision, license, tokenizer, and hashes.
Quant, engine, context, topology, and thinking policy.
Matched request, uncapped output, synchronized release.
Raw stream, reasoning, answer, usage, and original HTML.
Browser errors, shaders, motion, framing, and grading.
Publish evidence and keep display repairs separate.
Model output is never silently repaired. If a small display correction is needed, the untouched artifact remains available and the repair gets its own copy, diff, badge, and disclosure.
A quant must survive tokenizer checks, tool use, instruction following, long-context retrieval, coding, visual generation, and comparison with a higher-quality control.
The public work already exists. The next step is connecting it with a safe controller, node leases, twice-daily jobs, and a consistent review package.
SparkBench evidence, visual battles, reproducibility guides, and deployment notes.
One job registry across Mac Studio, DGX Spark, NAS, and multiple agents.
Permanent control prompt, rotating challenge, browser QA, video, and Telegram review.
Verified GGUF and native NVIDIA lanes with quant-tax reports and NAS archiving.
The portable visual harness and benchmark history are public. Machine-specific infrastructure stays private; methods and results do not.