Investors · Pre-seed

Open models that do more with less.

BTL
Lagos
Sep 2026

Bad Theory Labs trains open models and does the research that lets them do more with less compute, fewer bits and fewer tokens. We're raising a $1.5M pre-seed.

I run the lab from Lagos. We've had no institutional money and no permanent compute, so every result below came from rented GPUs and one laptop. The newest one, Interference Search, is the reason this page exists.

Interference Search

Language models think in a line. They write one token after another, and when an idea fails they back up in text and try again.

I watched Qwen3-1.7B fail at Countdown, the puzzle where you combine a few numbers to hit a target. It wasn't getting the arithmetic wrong. 45% of its failed attempts repeated an expression it had already ruled out itself.

On one problem it wrote the correct answer at token 1,313. Then it checked it nine more times, drifted into other ideas and ran out of budget without answering.

So I built Interference Search, which reasons over explicit states. Every live branch expands at once and the environment executes the moves. Branches that land on the same state merge into one, and a small trained judge drops the ones that can no longer reach the goal. The survivors move forward together, one level at a time.

On 30 hard Countdown problems:

SystemSolvedSequential steps
Qwen3-1.7B thinking in text, 1,500 tokens3 / 30one token at a time
Trained judge in one line of thought, 200 judged positions21 / 307.9 on the ones it solves
Interference Search, same judge, same 200 positions30 / 303

Given 1,500 judged positions, the single line also solves all 30, but it needs 23.7 sequential steps on average against 3.

When I replaced the trained judge with a linear probe on the 1.7B model's own hidden states, the search solved 15 of 30. The model generated zero tokens to get there. All of this ran on one M2 laptop with 16 GB of memory.

Why it goes past puzzles

Most of what a line of thought does is repeated work, and the share grows with the problem. At six numbers, 831,176 paths collapse into 13,229 distinct states. The ratio grew 3.2 times from four numbers to five, then 5.5 times from five to six.

The judge was trained on four- and five-number problems. On ten seven-number problems it cut the search 12.6 times without losing a solution.

I'll be straight about the limits. Everything is one seed. Countdown has an exact solver, which makes it a clean lab and nothing more.

On code, 30 MBPP problems the model had already failed, Interference Search solved 9 and best-of-N solved 8. That's noise. The real work now is making states and merging hold up where equivalence is harder to check, starting with code and terminal tasks.

I open-sourced the code, the trained judge, every raw result and the experiments that failed. Read the paper or get the code.

The models

Our releases have passed 210,000 downloads on Hugging Face, 192,760 of them for BTL-4 Compact.

Tinfield 1 is our newest release, an agentic model for terminal and software work built on Qwen3.8-Flash-Next. It has 177B parameters with 6.6B active per token, and our quantized builds run on a single 64 GB machine.

BenchmarkTinfield 1Qwen3.8-Flash-Next (base)Claude Opus 4.8
Terminal-Bench 4.033.029.023.6
DeepSWE v1.16258.759

Those scores are for the full weights. I haven't run the quantized builds on those benchmarks yet.

BTL-4 Compact put a 35B model into 9.96 GB at 2.30 bits per weight. It kept 94.1% of the full model's behaviour on our retention gate, 111 of 118 cases.

The finding behind it surprised me. I expected the damage at two bits to come from the output head, and it didn't. Almost all of it came from how the value range of each small group of weights was chosen.

Searching for a better range took false-premise rejection from 64.1% to 97.4%. It cost 12 seconds of GPU time.

Everything we shipped before that

We shipped our first model on June 22 and Tinfield 1 on September 21. These are the releases in between, plus the first one.

ReleasedModelWhat it isResultDownloads
Aug 5, 2026BTL-435B mixture-of-experts agent model, about 2.1B active per token, fine-tuned from Ornith-1.0-35B78.4% SWE-bench Verified, run by an outside lab. 73.5% BFCL v4 against 69.2% for the base, same harness, all 1,240 cases. 66.1% LiveCodeBench v6, 442 problems9,995
Aug 5, 2026Macaw2.6B assistant that runs entirely on a Mac, built on LFM2.5-2.6B97 hand-written, tested tools across Mail, Calendar, Files, the screen and the system. Chains steps and checks each one2,411
Jul 20, 2026BTL-3 CompactThe whole 27B text model in one 8.39 GB file, with our own CUDA and Metal runtimeKept 83 of the 90 tool behaviours the full model got right on a fresh 100-case gate. Smaller than an 8B model in FP162,827
Jul 20, 2026BTL-327B agent model for coding and tool use, post-trained on Qwen3.6-27B88.5% BFCL v4 on all 1,240 cases. 95.1% HumanEval. 91.2% on knowing when not to call a tool313
Jun 22, 2026BTL-2 Coder7B code-review adapter on Qwen2.5-Coder-7BStructured security findings: SQL injection, path traversal, auth bypass. Our first release71

Downloads are all-time counts from Hugging Face as of September 25, 2026. BFCL and LiveCodeBench were run in-house with the official scorers on full splits.

Our own SWE-bench runs on BTL-4 scored far lower, and the cause turned out to be our harness, which was broken. So I stopped trusting it for that benchmark and had an outside lab run BTL-4 instead. Their run came back at 78.4%.

Other research in progress

These are earlier-stage. I've written down where each one actually stands, because some of them aren't ready to claim.

ProjectWhat it doesWhere it stands
ESPLets a text-only model use a screen with no vision model, by probing the interface and streaming what changedThe rewrite takes 255 to 347 ms per action where the first version took 1.4 to 4.4 s, and sends the model 6 to 15 times fewer bytes. Its four registered hypotheses haven't been tested end to end yet
Research agent loopA 4B agent drives a retriever on BrowseComp-Plus, a hard web-research benchmarkBeat matched single-shot retrieval by +0.240 recall on a 5-question pilot. It found the same evidence with a third of the documents. Our retrieval harness reproduces the published baselines
One-bit modelsRecovery training for models squeezed to one or two bits per weightOn a small mixture-of-experts model, ternary weights recovered 55.0% top-1 agreement with the original against 46.3% for binary. Neither version knows when to stop generating yet
PrismOpen-source layer that adds planning, tools and verification around a small local modelMIT-licensed. Early results on a 4B model look positive, but they're one trial on 16 tasks, so I'm not quoting them until they're rerun

What I believe, and the business

Every result here came from spending less and checking that the model kept what mattered. I think the next big gains in reasoning will come from how models search. Interference Search is my first hard evidence for that.

Revenue today is close to nothing. Our products have made less than $200 in their lifetime, because almost all of my time has gone into the research and the models.

The plan has three parts:

  • Open weights, so people find and run our models.
  • Paid hosted access to the latest BTL models.
  • Paid private deployments for teams that need to run them on their own hardware.

Team and the raise

There are four of us: me as founder and research lead, a senior ML researcher, a research intern and one person running community and releases.

I'm raising a $1.5M pre-seed. It pays for compute for the next model and the first full-time hires. It also pays for taking Interference Search from Countdown into code and agent tasks, with the goal of a model that thinks in states from the start.

If this is interesting, I'd like 30 minutes with you. I'll walk you through the paper and run the search live.

Al-Ameen, Bad Theory Labs · hello@badtheorylabs.com

$1.5M pre-seed

Talk to the founderRead the paper