The technical workbench

Build systems. Investigate them. Publish the evidence.

I build systems, investigate how they work, measure them, and publish the evidence. The dual-RTX-3090 workstation is the baseline Local AI test system; this page tracks the tools and experiments around it.

What the lab works on

Three areas, one method.

Local AI is the deepest area right now; the research and software around it stay public alongside it.

How the work is done

The subject changes. The standard does not.

The same loop runs whether the system is a GPU, a trading strategy, or a data pipeline.

Approach

Build before claiming.

Guides start from hardware that actually runs the model, not from a spec sheet. If a system is new to the lab, it is labelled as planned — not benchmarked.

Evidence

Measured, estimated, or planned.

Every number on the site is one of three things: measured on the lab hardware, estimated from published data, or a planned experiment. The label stays on the page.

Integrity

Failures stay public.

Ideas that fail against the data get written up too. A negative result is still a result, and the notebook includes them.

Current environment

The baseline and what comes next.

Baseline · in use

Threadripper Pro, 2 × RTX 3090 24GB.

The current Local AI test system. It runs the standardized llama.cpp context and concurrency corpus, CPU-offload experiments, model evaluation, and the software used to publish the results.

Incoming · untested

2 × NVIDIA DGX Spark.

These systems are incoming and have not been tested here. Published-data analysis exists, but no DGX Spark result on this site is presented as a lab measurement.

Read the platform guide →

On the bench

Recent work, tools, and current experiments.

A working inventory rather than a product roadmap.

Recent benchmark work

Context scaling and concurrency.

PP and TG across 0–32K, selected long-context runs through 256K, and one-, two-, and four-stream throughput on dual RTX 3090s.

Explore the corpus →
Software and tools

llama.cpp, Linux, and analysis code.

Inference runtimes, benchmark harnesses, remote-access tooling, data pipelines, Astro, TypeScript, and Python.

Current experiments

Local inference under real constraints.

Long-context behavior, multi-GPU placement, CPU offload, speculative decoding, model serving, and repeatable ways to separate measured results from estimates.

Browse Local AI →

Planned

What is next is labelled as planned.

No results are published before the hardware is in hand and the work is done.

Hardware

DGX Spark hands-on.

Two NVIDIA DGX Spark systems are incoming. They are not part of the tested environment yet, so no benchmarks or hands-on results from them appear on the site. Existing coverage of the platform, based on published data, lives in the DGX Spark guide.

Guides

Planned fundamentals and software guides.

Parameters and model size, context windows, attention and mixture-of-experts, Linux installation, model storage, and security and privacy are listed as planned guides on the Local AI hub.