Text so far
The best pet is a
Harris Oldroyd · Public Notebook
A public research notebook built from the hardware up: running models locally, testing trading ideas against data, and building the software the work needs.
Start with tokenization, memory, attention, and the KV cache before choosing a local stack.
Compare Ollama, llama.cpp, and vLLM, then connect a model to the hardware and workflow you want.
Compare real inference measurements first; use estimates separately when you are still planning.
What I work on
Local AI is the deepest area right now; the research and software around it stay public alongside it.
Hardware, quantization, runtimes, benchmarks, and interactive explanations for local inference.
Explore local ai →ResearchReproducible experiments on market regimes, filters, and execution — including the ideas that fail against data.
Explore research →SoftwareLocal-first projects for data, measurement, and research workflows.
Explore software →Local AI
Hardware, models, runtimes, and benchmarks for running AI locally — with a learning hub for going deeper.
Learning paths for fundamentals, hardware, models, and software, with interactive explainers for tokenization, attention, and the KV cache.
The selected work below follows local systems from hardware and quantization through model behavior, experiments, and measured results.
The best pet is a
GPT-2 BPE tokens. ␠ marks a leading space included in the token.
Prompt IDs use the GPT-2 vocabulary (r50k_base). Continuations use a small example distribution.
[0.2, −0.7, …]Embeddings represent meaning. Attention mixes in useful context from the other tokens before the next-token scores are formed.
Sampled result: dog
Temperature reshapes the distribution; top-p keeps the smallest cumulative set at or above p, then renormalizes it. Zero temperature chooses the highest-probability candidate. A fixed draw keeps this toy example reproducible.
The best pet is a dog
Back revisits the previous generation's Pick step. Changing that choice starts a new continuation from there.
Featured
Six selected pieces span hardware, fundamentals, model architecture, interpretability, and measured local inference.
A dual RTX 3090 local AI workstation with 48 GB of distributed VRAM, Threadripper Pro, WRX80, 128 GB ECC RAM and enough expansion room for larger multi-GPU experiments.
Read the piece →researchI weakened refusal in Qwen3.5-9B and found that most of the change came from one simple update in layer 15.
Read the piece →Model analysisAn in-depth guide to Qwen3.8-Flash-Next and its 51B n-gram embedding table, updated with dual-RTX-3090 ExLlamaV3 results: hot expert placement, working MTP, and the memory tradeoff behind long-context inference.
Read the piece →FundamentalsA beginner's guide to local LLM quantization: what bits per weight means, how popular GGUF formats compare, when to choose Q2 through Q8, and why Q4 is a strong starting point.
Read the piece →Model analysisAn in-depth guide to Qwen3.8-27B for local inference. We examine its Artificial Analysis Intelligence Index standing, native multi-token prediction (MTP), weights versus KV cache quantization, parallel context slots, and firsthand daily driver observations on a dual-GPU workstation.
Read the piece →BenchmarkA practical look at the RTX 3090 for local AI workloads, covering specs, VRAM, bandwidth, TFLOPS, used pricing, power limiting, LLM performance, image generation, and why it remains one of the best-value GPUs for running open models at home.
Read the piece →All areas
The latest from every area — Local AI, research, and software.
Why does an LLM run fast on turn one and crawl as context expands? We break down why memory bandwidth bottlenecks local inference, calculate theoretical speed limits, analyze measured dual-RTX 3090 context sweeps from 0K to 256K, and examine practical settings to keep generation fast.
DeepSeek V4.1 Flash introduces 748B total parameters, a 196B engram table, 8B prefill / 16B decode asymmetric routing, native vision, and an 890 B/token KV cache. Here is how the architecture works, how it delivers 100–200 tok/s at 150 million tokens per dollar, and why it's my new daily driver for code.
A technical and economic guide to HBM stack height, capacity, bandwidth, packaging yield, qualified high-speed supply, and the competing interpretations of lower-layer AI memory.
A practical review of GLM 5.3 Flash covering its 320B-A18B architecture, API economics against DeepSeek V4 Flash and Luna, coding behavior, quantized memory footprint, and my plans to run it locally on an M5 Ultra Mac Studio.
Elsewhere in the notebook
Quant research, software projects, and technical writing each have their own section.
Methodology, testing, failures, and iteration on systematic strategies.
Browse quant research →Software projectsProjects spanning quantitative research, automation, data engineering, and personal interests.
Browse software projects →Essays and notesLong-form writing wherever curiosity leads: technology, markets, AI, and decision-making.
Browse essays and notes →