Harris Oldroyd · Public Notebook

Harris Oldroyd

Local AI, systematic research, and the systems behind both.

A public research notebook built from the hardware up: running models locally, testing trading ideas against data, and building the software the work needs.

Start here

3 entry points

What I work on

Three areas, one notebook.

Local AI is the deepest area right now; the research and software around it stay public alongside it.

Local AI

The deepest area of the notebook.

Hardware, models, runtimes, and benchmarks for running AI locally — with a learning hub for going deeper.

Learning hub

Learn how local AI actually works.

Learning paths for fundamentals, hardware, models, and software, with interactive explainers for tokenization, attention, and the KV cache.

The selected work below follows local systems from hardware and quantization through model behavior, experiments, and measured results.

From prompt to next token

Text so far

The best pet is a

Back revisits the previous generation's Pick step. Changing that choice starts a new continuation from there.

Featured

Work worth starting with.

Six selected pieces span hardware, fundamentals, model architecture, interpretability, and measured local inference.

Hardware

I Built a 48 GB VRAM Local AI Workstation with Two RTX 3090s

A dual RTX 3090 local AI workstation with 48 GB of distributed VRAM, Threadripper Pro, WRX80, 128 GB ECC RAM and enough expansion room for larger multi-GPU experiments.

Read the piece →
research

I Abliterated Qwen3.5-9B and Traced the Refusal Through the Weights

I weakened refusal in Qwen3.5-9B and found that most of the change came from one simple update in layer 15.

Read the piece →
Model analysis

Qwen 3.8 Flash Explained: The 51B N-Gram Memory Architecture

An in-depth guide to Qwen3.8-Flash-Next and its 51B n-gram embedding table, updated with dual-RTX-3090 ExLlamaV3 results: hot expert placement, working MTP, and the memory tradeoff behind long-context inference.

Read the piece →
Fundamentals

LLM Quantization Without the Alphabet Soup

A beginner's guide to local LLM quantization: what bits per weight means, how popular GGUF formats compare, when to choose Q2 through Q8, and why Q4 is a strong starting point.

Read the piece →
Model analysis

Qwen 3.8 27B: The Practical Sweet Spot for Local AI

An in-depth guide to Qwen3.8-27B for local inference. We examine its Artificial Analysis Intelligence Index standing, native multi-token prediction (MTP), weights versus KV cache quantization, parallel context slots, and firsthand daily driver observations on a dual-GPU workstation.

Read the piece →
Benchmark

RTX 3090 for Local AI Inference

A practical look at the RTX 3090 for local AI workloads, covering specs, VRAM, bandwidth, TFLOPS, used pricing, power limiting, LLM performance, image generation, and why it remains one of the best-value GPUs for running open models at home.

Read the piece →

All areas

Recent work across the notebook.

The latest from every area — Local AI, research, and software.

Elsewhere in the notebook

The wider research practice.

Quant research, software projects, and technical writing each have their own section.