fi-inference · September 2026

Own your intelligence, on two old servers and €500 of upgrades

I had a week off between jobs, two old Dell servers and a pair of AI subscriptions. The result is a 180B model at 33 tok/s, with no GPU.

Two Dell R530 servers stacked inside a wooden box
The two R530s in their box. Ten years old, with 288 GB of RAM between them.

TL;DR

  • Two 2016-era Dell R530s now run Qwen3.8-Flash-Next (180B total, 6B active, 4-bit) at 33 tok/s on one stream and 36 tok/s across four.
  • I already owned the servers. The upgrades cost under €500 in used parts.
  • Coding agents ran nearly 300 experiments in nine days to get there, most of them on a smaller sibling model first. On the target, the first version of my own engine managed 1.45 tok/s; stock llama.cpp, the number to beat, managed 5.3.
  • The whole stack, from the weights to my MCP tools, now runs in my basement.

I had about a week off between jobs. The DGX Spark I'd been using belonged to my previous employer, so it went back, and that left me with a problem: I had gotten hooked on having capable agents that run locally, private, always on and silent, and I no longer had the hardware for it.

What I did have was two Dell PowerEdge R530s from around 2016, switched off in my basement office. They weren't broken. They were loud.

Two Dell R530 servers stacked on a desk, front view
Before: two R530s, switched off because they were too loud to live with.

Then I did some arithmetic. The Spark reads memory at 273 GB/s, and for generating text that number more or less is the speed. Each R530 has two CPU sockets, and each socket has four memory channels of its own. With the right CPUs those channels run DDR4-2400: 76.8 GB/s per socket on paper, about 307 GB/s across all four sockets. That's the Spark's bandwidth and then some, just in four places instead of one. And the two boxes together have 288 GB of RAM, against the Spark's 128. Getting one token through four separate memory systems became the whole project.

At that point I was sure: an inference engine written for exactly this hardware and exactly one model could be very fast. The week turned into nine days, and it went better than I expected. I would have been happy with 20 tok/s.

What I builtHardware, model, engine, serving, fans. Skip it if you like; the story doesn't need it.
Hardware
Two Dell R530s with four Xeon E5-2695 v4s (72 cores) and 288 GB of DDR4-2400. A Mellanox ConnectX-3 InfiniBand card sits in each box, joined by one direct cable at 56 Gb/s. Ubuntu Server boots from an SSD in each box, and NVMe drives on PCIe adapter cards hold the weights and the prompt cache. Out went the old CPUs, the spinning disks and the SAS controller.
Model
Qwen3.8-Flash-Next at 4-bit (Q4_K_M): about 180B parameters in total, about 6B read per token, plus its multi-token-prediction head.
Engine
A custom C++ engine, written by coding agents. Experts are pinned to four memory domains (2 sockets × 2 boxes), activations are exchanged over RDMA, and the MTP head drives speculative decoding.
Serving
An OpenAI-compatible server for my agents: streaming chat completions with tool calls and separate reasoning output, the model's own chat template, reasoning effort from off to xhigh, and the usual sampling knobs (temperature, top-p, top-k, min-p). Up to four sessions at once, with extra slots waking only when they're needed, so a lone request runs as fast as on a single-stream server. 262k context, and a prompt cache on NVMe so a returning conversation skips its prefill.
Quiet
My own fan controller replaces Dell's. It follows the CPU temperatures, ramps gently instead of jumping, and was tuned in a simulator fitted to thermal and microphone measurements: from peaks over 60 dB(A) at my desk to 42 in normal use. In the finished acoustic box: 39 in normal use, 31–32 at idle.
Display
An iPad on the wall shows both servers live: temperatures, fans, decode and prefill speed, memory bandwidth per socket.
Agents
Opus 5.5 and Fable 5.1 (Claude Max 5x); Luna 6, Sol 6 and a little Astra (ChatGPT Plus).

I came out of it believing three things:

  1. Owning your intelligence is the thing.
  2. CPU inference on used hardware is a thing.
  3. Autoresearch works far beyond training scripts.

1Own your intelligence

This is why I do any of it. Sequoia's Own Your Intelligence makes the case for companies. For one person it's the same case with smaller numbers.

For one person, the sensitive data is a knowledge base that grows by more than 1,000 entries a month, my playbooks, my task manager, my own chat and agent front end, and an inventory of the homelab, all exposed as MCP servers. My coding agent, pi, works with them all day. Every prompt to a hosted model carries a slice of that out of the house. Now it doesn't have to: the model runs on weights I hold, on hardware I own. Nobody can swap it, meter it, rate-limit it, or read along.

Hosted frontier agents built the replacement, and they're still better at that kind of work. But the model they built it for doesn't need to be the best, only good enough to act on my behalf, and it is: Qwen3.8-Flash-Next scores 40 on the Artificial Analysis Intelligence Index, above GPT-5.6 Luna at maximum effort (37), OpenAI's cheap tier I would otherwise reach for.

30 35 40 45 50 55 60 $0.03 $0.10 $0.30 $1 $3 $10 cost per index task at API prices (log scale) index Claude Opus 5.5 (max): 58 at $5.98 per task (closed) Claude Opus 5.5 (xhigh): 56 at $3.46 per task (closed) Claude Sonnet 5.5 (max): 56 at $7.60 per task (closed) Claude Opus 5.5 (high): 54 at $1.82 per task (closed) Claude Fable 5.1 (max): 53 at $7.63 per task (closed) Claude Fable 5.1 (xhigh): 53 at $5.98 per task (closed) GPT-6 Astra (max): 53 at $3.26 per task (closed) GPT-6 Astra (xhigh): 52 at $2.31 per task (closed) Claude Sonnet 5.5 (xhigh): 52 at $2.74 per task (closed) GPT-6.1 Sol (max): 52 at $0.72 per task (closed) Claude Opus 5.5 (medium): 51 at $1.34 per task (closed) Claude Fable 5.1 (high): 51 at $3.91 per task (closed) GPT-6.1 Sol (xhigh): 51 at $0.39 per task (closed) GPT-6 Astra (high): 51 at $1.73 per task (closed) GPT-6.1 Sol (high): 50 at $0.32 per task (closed) GPT-6 Astra (medium): 50 at $1.54 per task (closed) Claude Fable 5.1 (medium): 49 at $2.98 per task (closed) Muse Spark 1.3 (max): 48 at $1.60 per task (closed) GPT-6.1 Sol (medium): 48 at $0.21 per task (closed) GPT-6 Sol (max): 48 at $1.05 per task (closed) Claude Fable 5.1 (low): 47 at $2.37 per task (closed) Claude Sonnet 5.5 (high): 47 at $1.08 per task (closed) Grok 4.7 (xhigh): 46 at $3.74 per task (closed) Grok 4.7 (high): 46 at $2.73 per task (closed) MiMo-V2.6-Pro: 46 at $0.13 per task (open weights) GPT-6 Astra (low): 46 at $0.82 per task (closed) Qwen3.8 Max (0902): 45 at $5.41 per task (closed) Muse Spark 1.3 (xhigh): 45 at $1.37 per task (closed) GLM-5.3 (max): 45 at $2.01 per task (open weights) GPT-6 Sol (xhigh): 44 at $0.52 per task (closed) Kimi K3 (max): 44 at $2.00 per task (open weights) Step 5 Preview: 44 at $0.72 per task (open weights) GPT-6 Sol (high): 43 at $0.38 per task (closed) Claude Opus 5.5 (low): 42 at $0.55 per task (closed) GPT-6.1 Sol (low): 42 at $0.13 per task (closed) GPT-5.6 Terra (max): 42 at $1.40 per task (closed) GLM-5.3-Flash: 42 at $0.25 per task (open weights) Gemini 3.8 Flash (high): 41 at $1.24 per task (closed) Claude Sonnet 5.5 (medium): 41 at $0.59 per task (closed) Qwen3.8 2.4T A95B: 40 at $2.16 per task (open weights) Gemini 3.8 Flash (medium): 40 at $0.93 per task (closed) DeepSeek V4.1 Flash (max): 39 at $0.27 per task (open weights) GPT-5.6 Terra (xhigh): 38 at $0.63 per task (closed) MiMo-V2.6-Flash: 38 at $0.06 per task (open weights) GPT-6 Luna (max): 37 at $0.07 per task (closed) DeepSeek V4 Pro 0813 (max): 36 at $0.67 per task (open weights) GLM-5.3 (low): 34 at $0.85 per task (open weights) GPT-5.6 Terra (high): 34 at $0.34 per task (closed) GPT-6 Sol (low): 34 at $0.13 per task (closed) GPT-6 Luna (xhigh): 34 at $0.04 per task (closed) Qwen3.8 27B (xhigh): 34 at $1.01 per task (open weights) DeepSeek V4 Flash 0731 (max): 34 at $0.22 per task (open weights) GPT-6 Luna (high): 32 at $0.03 per task (closed) Claude Opus 5.5 (max) GPT-6.1 Sol (xhigh) MiMo-V2.6-Pro GPT-6 Luna (high) GPT-5.6 Luna (max): 37 at $0.18 per taskGPT-5.6 Luna (max) 37 Qwen3.8-Flash-Next: 40 at $0.37 per taskQwen3.8-Flash-Next 40
Artificial Analysis Intelligence Index v4.3.2 against cost per index task at hosted API prices, as of 30 September 2026. Hover or tap a point for its value.

2CPU inference on used hardware is a thing

One formula explains most of it:

decode tok/s ≈ memory bandwidth ÷ bytes read per token

Everyone who runs models locally learns this eventually. It has a fun consequence, though. If you're generating one token at a time for one person, compute barely matters and memory channels do.

Off-the-shelf software doesn't use four memory systems at once, though. llama.cpp ran slower on two boxes than on one. The reason is NUMA:1 in my measurements, a socket reads its own memory more than twice as fast as its neighbour's. Anything that makes a socket fetch weights from the other socket pays that penalty on every token.

So my engine uses owned experts. A mixture-of-experts model has hundreds of small expert networks and uses only a few per token. In my engine each expert lives in exactly one socket's memory and never moves. When a token needs an expert, the token goes to the expert. Weights stay put; only activations travel, a few kilobytes per token.

For the link between the boxes, my upgrade research pointed at used Mellanox ConnectX-3 cards: cheap now, and fast. I'll admit I got a little giddy. This was my first InfiniBand, the datacenter interconnect Nvidia bought Mellanox for, because it's what ties large GPU pods together. Two cards and a half-metre cable: 56 Gb/s, direct, no switch. The bandwidth barely mattered, because a few kilobytes per token don't need a big pipe. Latency did. With RDMA, one box writes straight into the other box's memory, with no kernel and no TCP involved, and the other box notices by watching a flag. That takes about 0.9 microseconds. A ping over the same cable through the normal network stack takes about 190.2 I bought a fat pipe and got a short one, which was what I needed.

A short InfiniBand cable held in a hand
Rear of the two servers in blue light, with InfiniBand and Ethernet cables
Left: the half-metre cable. Right: the back of the box, with the InfiniBand link and the direct Ethernet cables.

The shopping list follows from the formula:

  • Memory channels come first. Swapping the Xeon E5-2640 v4s for used E5-2695 v4s lets the same DIMMs run at 2400 instead of 2133.
  • Interconnect latency comes next. See above; I'm still pleased about it.
  • Then the fewest active parameters for the quality you need. This one was less a purchase than a lucky fit. If the only models near the top of the index were dense, or read far more parameters per token, I wouldn't have tried any of this, on the Spark or on the R530s. Flash-Next reads about 6B parameters per token for its 40. DeepSeek V4 Flash, which I ran every day on the Spark at 2-bit, reads 13B for a 34. Qwen3.8 27B, dense on my dual-4090 workstation, reads all 27B for a 34. GLM-5.3-Flash scores higher, 42, but reads 18B, three times as much.
An Intel Xeon E5-2640 v4 on bubble wrap
One of the old E5-2640 v4s, on its way into storage.

All in, the upgrades came to under €500. Ten-year-old datacenter gear gets dumped by the pallet, and patience pays: two of the three used SSDs I bought had about 30 days of power-on time.

Measured
Decode, one stream, short prompt33 tok/s
Decode, one stream, 2.6k-token context29 tok/s
Four streams at once36 tok/s combinedeach stream bit-identical to running alone
Prefill, 2.6k-token prompt58 tok/s
Stock llama.cpp, same model5.3 tok/s
Context262k tokensprompt cache on NVMe
Power per box: idle / single-stream decode / heavy load84 / 170–180 / 266 Wdecode tuned for quiet, not speed
Electricity per million generated tokens≈ €1.20–1.60at about 40 ct/kWh

The llama.cpp line compares whole systems. It ran on one box without speculative decoding, while my engine uses both boxes and the model's own multi-token-prediction head. It had 36 threads, one per physical core. On the smaller development model I had swept NUMA modes and thread counts properly: nothing beat about 12 tok/s on one box, and two boxes over RPC dropped to 8.

Four streams barely beat one, and the formula says why: each stream calls on different experts, so four streams read almost four times the weights. It doesn't scale well, but 288 GB leaves room for the concurrency, and that's what subagents and parallel work need.

Prefill is the weak spot, as on any CPU. At 58 tok/s, a coding agent's 19k-token system prompt takes five to six minutes cold. The prompt cache on NVMe is what makes that livable: after a restart, a returning 15.5k-token conversation was back in 24 seconds instead of four to five minutes cold.

The electricity line is rough on purpose: both boxes at about 175 W each, German household prices averaging around 40 ct/kWh over the last three years, and output tokens only. That's more than a datacenter charges for the same model (about $0.47 per million output tokens for hosted Flash-Next), about what Luna costs, and far below frontier prices. The point was never saving money. It was running privately at a reasonable cost, and it does that.

3Autoresearch did the work

Karpathy's autoresearch puts a coding agent in a loop steered by one program.md: change something, measure, keep or discard, repeat. He built it for training runs. I pointed the same pattern at a systems problem.

My program.md sets three rules:

  • Pre-register every experiment with a prediction.
  • Prove correctness before speed, bit-exact against a reference.
  • Keep a change only if it beats the measured noise floor in paired A/B runs.
Autoresearch, generalisedA primer for running the loop on your own problem.

Karpathy's data engine at Tesla was a flywheel: the fleet surfaces failures, they get labelled, the net retrains, the fleet redeploys. It still needed people to label. In autoresearch the labeller is a benchmark, so the flywheel can run unattended all night. What keeps it turning:

evaluator · agent can't edit predict change verify measure keep/drop expected Δ code bit-exact paired A/B vs noise fail journal + git: every outcome, including fails the next session starts here
  1. The evaluator is the product. The agent optimises exactly what you measure, bugs and noise included. Make it fast, know its noise floor, and keep it out of the agent's reach. Spend your first day here, before touching the prompt.
  2. Keep memory in files. An append-only journal, fails included, plus one commit per kept change. Every session is disposable, and dead ends stay dead.
  3. Never let it finish. A target is a stop condition. Make the goal relative, such as beating the best so far, and the loop keeps spinning while you sleep.

For the mechanics, go to Karpathy: the README and his program.md.

I didn't start on the 180B. The first practical tests, and most of the engine, ran against Qwen3-Next-80B, the smallest model with the same structure; the move to the target came once the engine worked. The agents (Codex first, then Claude Code) wrote the C++ kernels, ran the A/Bs on both servers over ssh, and kept the journal. I set the goals, made the calls, and said no a lot. All of it ran on two consumer subscriptions.

Nine days: nearly 300 experiments, 600+ journal entries, about 1,300 commits, and, on the target model, 1.45 → 33 tok/s, from the engine's first run on it to today. Three changes did most of it:

  1. Worker teams per socket. 7.7x in a single step, still on the 80B.
  2. Speculative decoding plus an RDMA mailbox. Two boxes finally beat one: 13 → 20 tok/s.
  3. The grind. Dozens of 1–5% wins took it from 25 to 33.

A word about targets. I started at 20 tok/s, then it was 25, then 30, then 35. None of these numbers mean anything. I kept moving the target because an agent that reaches its goal does the reasonable thing: it stops, writes a nice summary, and waits for me. When it's on a roll, you don't want that.

0 10 20 30 40 180 200 220 240 260 experiment # (1–173 ran on the 80B) tok/s target 35 (earlier 20, 25, 30) stock llama.cpp, one box: 5.3 E175: 2.6 tok/s (not kept) E185: 5.3 tok/s (not kept) E187: 6.5 tok/s (not kept) E189: 8.5 tok/s (not kept) E190: 6.9 tok/s (not kept) E191: 10.0 tok/s (not kept) E192: 11.1 tok/s (not kept) E193: 11.2 tok/s (not kept) E195: 12.2 tok/s (not kept) E197: 7.8 tok/s (not kept) E201: 13.3 tok/s (not kept) E202: 14.7 tok/s (not kept) E203: 12.0 tok/s (not kept) E205: 15.9 tok/s (not kept) E210: 25.3 tok/s (not kept) E213: 25.5 tok/s (not kept) E215: 24.5 tok/s (not kept) E216: 25.6 tok/s (not kept) E218: 23.1 tok/s (not kept) E219: 25.7 tok/s (not kept) E220: 10.4 tok/s (not kept) E221: 30.7 tok/s (not kept) E224: 25.4 tok/s (not kept) E225: 25.6 tok/s (not kept) E226: 25.5 tok/s (not kept) E227: 22.4 tok/s (not kept) E229: 24.9 tok/s (not kept) E232: 25.5 tok/s (not kept) E235: 27.4 tok/s (not kept) E236: 24.0 tok/s (not kept) E238: 27.7 tok/s (not kept) E239: 27.3 tok/s (not kept) E241: 24.6 tok/s (not kept) E243: 27.7 tok/s (not kept) E247: 27.5 tok/s (not kept) E249: 27.8 tok/s (not kept) E252: 27.5 tok/s (not kept) E254: 27.4 tok/s (not kept) E257: 25.6 tok/s (not kept) E260: 28.3 tok/s (not kept) E261: 29.4 tok/s (not kept) E261: 29.7 tok/s (not kept) E262: 23.3 tok/s (not kept) E263: 29.1 tok/s (not kept) E265: 30.1 tok/s (not kept) E268: 31.8 tok/s (not kept) E273: 30.4 tok/s (not kept) E274: 31.4 tok/s (not kept) E275: 31.7 tok/s (not kept) E276: 30.9 tok/s (not kept) E220: 7.3 tok/s at 2.6k context E226: 13.7 tok/s at 2.6k context E227: 18.9 tok/s at 2.6k context E231: 22.2 tok/s at 2.6k context E234: 23.1 tok/s at 2.6k context E252: 24.8 tok/s at 2.6k context E261: 25.3 tok/s at 2.6k context E261: 25.7 tok/s at 2.6k context E264: 26.9 tok/s at 2.6k context E270: 29.9 tok/s at 2.6k context 29.9 at 2.6k E174: 1.45 tok/s, new best E178: 2.70 tok/s, new best E179: 3.78 tok/s, new best E180: 4.57 tok/s, new best E183: 4.95 tok/s, new best E184: 7.60 tok/s, new best E186: 9.00 tok/s, new best E188: 11.13 tok/s, new best E194: 12.27 tok/s, new best E196: 13.32 tok/s, new best E200: 15.85 tok/s, new best E204: 17.10 tok/s, new best E206: 20.45 tok/s, new best E207: 23.05 tok/s, new best E208: 24.03 tok/s, new best E209: 24.79 tok/s, new best E211: 25.90 tok/s, new best E212: 25.95 tok/s, new best E214: 27.56 tok/s, new best E233: 27.57 tok/s, new best E234: 27.73 tok/s, new best E248: 27.77 tok/s, new best E250: 28.02 tok/s, new best E252: 29.28 tok/s, new best E261: 30.37 tok/s, new best E264: 33.28 tok/s, new best E271: 33.48 tok/s, new best 33.5 best
Single-stream decode on Qwen3.8-Flash-Next, per experiment, as of 30 September. The numbering continues from the 80B development model. Every point was measured; grey ones didn't beat the best and were thrown away. Hover or tap a point for its value.

The same loop made the servers livable. They sit in the room I work in, and at my desk they peaked at over 60 dB(A). The fan controller brought normal use down to 42. The finished acoustic box took it to 39, and 31–32 at idle. The loop helped with all of it:

A phone sound meter reading 56.5 dB(A)
A sound meter display reading 39.1
Left: an early reading at my desk, 56.5 dB(A) and peaking at 61. Right: 39.1 today, in normal use, in the box.
  • Fans. iDRAC, Dell's management controller, did not recognise the new InfiniBand and NVMe adapter cards and assumed the worst: 8,000–11,000 rpm at idle. The agents measured it with a calibrated microphone, fitted a simulator to those measurements, and optimized a fan policy in it.
  • The acoustic box. We went back and forth on designs until a much simpler one won. I bought the parts, drew the final version in Shapr3D on my iPad, and handed the export back to the agents to check it and plan the build steps. Then I built it.
  • Everything around it. A hardware inventory, a porting guide for DeepSeek V4.1 Flash and GLM-5.3-Flash, and a review of this post.
Three large fans taped to the front of a server
Two fans mounted in a cardboard box on a server
Two of the designs we went through before the box: tape and Noctuas, then cardboard.
R530 R530 air in air out intake fans exhaust fans front rear 800 mm deep, 532 wide, 300 high
Build manual: floor and railsfloor and rails
Build manual: first wallfirst wall
Build manual: fan platesfan plates
Build manual: servers inservers in
The design that won, in section. Air comes in under the front, rises, crosses both servers front to back, rises again and leaves on top at the rear. Below: four steps from my build manual, rendered from the model.

What almost stopped me

  • The first numbers: llama.cpp across two nodes, RDMA included, was not encouraging.
  • A server that wouldn't boot after the CPU swap. It reported a DIMM error, and reseating didn't help. The cause was a speck of thermal paste on one CPU socket pin. One swipe fixed it for good.
  • The noise: at 60+ dB(A), I seriously doubted this could ever run always-on.
  • Wear: one DIMM now logs a corrected error at the same address every day. Running ten-year-old parts warmer to keep them quiet has a cost, and I don't know it yet.
  • The GPU temptation: halfway through, I priced out a spare RTX 3060 12 GB with 64 GB of DDR5, and a build around my dual-4090 workstation.
  • Woodworking: I'm not much of a woodworker. The box was an adventure, and it's finished.
An open CPU socket with old thermal paste
Inside of an unfinished wooden box
Left: mid CPU swap. Right: the box, mid build.

None of these ended it. Most of them became the next experiment.

A rack with the wooden server box and an iPad showing live server telemetry
Close-up of the iPad showing power, fan speed and temperatures for both servers
Its home now: the rack in my office. The display on the front shows both servers live: power, fans, temperatures.

Try this

Measure your box's memory bandwidth and divide it by the bytes your model reads per token. You won't reach that number. But it tells you in five minutes whether the hardware you already own is worth a week of your time.

Next for mine: I'm pointing the box at itself. The same loop will build and test a new custom inference engine for a new model, and this time the agents doing the research run on the box. No APIs.

Own your intelligence. Then let it build the next one.

  1. NUMA stands for non-uniform memory access. It means the memory attached to the other CPU is further away, and you pay for the trip.↩
  2. Strictly, the ping is a round trip and the RDMA write goes one way, so the fair ratio is closer to 100x than 200x. I'm still impressed.↩