Build before claiming.
Guides start from hardware that actually runs the model, not from a spec sheet. If a system is new to the lab, it is labelled as planned — not benchmarked.
The technical workbench
I build systems, investigate how they work, measure them, and publish the evidence. The dual-RTX-3090 workstation is the baseline Local AI test system; this page tracks the tools and experiments around it.
What the lab works on
Local AI is the deepest area right now; the research and software around it stay public alongside it.
Hardware, quantization, runtimes, benchmarks, and interactive explanations for local inference.
Explore local ai →ResearchReproducible experiments on market regimes, filters, and execution — including the ideas that fail against data.
Explore research →SoftwareLocal-first projects for data, measurement, and research workflows.
Explore software →How the work is done
The same loop runs whether the system is a GPU, a trading strategy, or a data pipeline.
Guides start from hardware that actually runs the model, not from a spec sheet. If a system is new to the lab, it is labelled as planned — not benchmarked.
Every number on the site is one of three things: measured on the lab hardware, estimated from published data, or a planned experiment. The label stays on the page.
Ideas that fail against the data get written up too. A negative result is still a result, and the notebook includes them.
Current environment
The current Local AI test system. It runs the standardized llama.cpp context and concurrency corpus, CPU-offload experiments, model evaluation, and the software used to publish the results.
These systems are incoming and have not been tested here. Published-data analysis exists, but no DGX Spark result on this site is presented as a lab measurement.
Read the platform guide →On the bench
A working inventory rather than a product roadmap.
PP and TG across 0–32K, selected long-context runs through 256K, and one-, two-, and four-stream throughput on dual RTX 3090s.
Explore the corpus →Inference runtimes, benchmark harnesses, remote-access tooling, data pipelines, Astro, TypeScript, and Python.
Long-context behavior, multi-GPU placement, CPU offload, speculative decoding, model serving, and repeatable ways to separate measured results from estimates.
Browse Local AI →Planned
No results are published before the hardware is in hand and the work is done.
Two NVIDIA DGX Spark systems are incoming. They are not part of the tested environment yet, so no benchmarks or hands-on results from them appear on the site. Existing coverage of the platform, based on published data, lives in the DGX Spark guide.
Parameters and model size, context windows, attention and mixture-of-experts, Linux installation, model storage, and security and privacy are listed as planned guides on the Local AI hub.