Research paper · September 2026

Where a Hybrid MoE
Spends Its Bytes

A measured compression allocation for a 177B hybrid model with sparse experts, linear attention, gated residual streams, and a 51B-parameter n-gram table.

176.94BModel scale
61.5 GiBSmallest file
34.67 GiBGPU-resident core
97.4%Private retention

The largest tensors were not the bytes a decoded token spent most often.

A token reads about 5.09 GiB from the compressed model. Linear-attention projections account for 40.6% of that traffic. Routed experts account for 15.5%, despite dominating stored weight size. That reversal changed the offload and precision plan.

The other surprise was the engram table. Its accesses are extremely concentrated, but its measured effect is not. The hottest 5% of rows serve 73.9% of held-out reads and recover only 28% of the table's output contribution when the remaining rows are removed. A simulated 2.5-bit code across all rows moved mean KL from 0.287 to 0.296.

Claim boundary

The engram measurements come from one agent-session distribution. The 2.5-bit table saving is derived from a restricted-code experiment; its CUDA path and full-model build remain untested. The private gate is a damage check, not a general quality score.

Measure the architecture before deciding where to cut.

One precision rule would have spent bytes on the wrong components. The final recipe follows measured traffic, output effect, tensor width, and edit exactness.

  1. 01

    Count bytes a token reads

    Stored size and decode traffic rank this architecture differently. Routed experts occupy most of the file but account for 15.5% of modeled per-token reads.

  2. 02

    Keep the dense spine resident

    Linear-attention projections account for 40.6% of modeled reads and hyper-connections another 12.4%, so the dense path stays on the GPU.

  3. 03

    Preserve row coverage

    Pruning the engram table to its hottest 5% of rows recovered only 28% of its measured output contribution. Lower precision across all rows caused much less damage.

  4. 04

    Search the two-bit range

    Importance-weighted range search reduced reconstruction error on the expert down projections from 51.9% to 11.7% without changing Q2_0's file layout.

  5. 05

    Verify the physical artifact

    The final merge was checked across 131 source shards and 530 LoRA targets before conversion, then measured as the exact GGUF users run.

Compact

72.04 GiB overall and 45.22 GiB GPU-resident, with expert gate and up projections at IQ2_XXS and expert down projections at IQ4_NL.

Mini

61.5 GiB overall and 34.67 GiB GPU-resident. Range-searched Q2_0 down projections cut another 10.55 GiB from the GPU side.

Engram

A 320-million-row lookup table where the tested pruning route lost most measured effect, while all-row low precision stayed close to the shipped build.

Exact edits

Per-channel scale migration remains exact through SwiGLU and the gated residual read. General rotation does not commute with the input-dependent gates.

Read the paper

The complete report is available as a PDF.

The paper includes the byte ledger, held-out engram ablations, exactness derivations, private retention checks, speed measurements, kernel cross-checks, limitations, and both final artifact sizes.