blog

Introducing StrataFlow

ProductOctober 5, 20266 min

Today we are sharing StrataFlow: a local inference engine that runs very large models on the hardware you already own, with no GPU required. The idea is simple to state and a little surprising the first time you see it work - a model far bigger than your RAM can run on a normal laptop or desktop CPU, inside a bounded memory budget, by streaming its weights from disk.

The big model, bounded RAM, no GPU idea

Mixture-of-Experts (MoE) models are huge on disk but only touch a small slice of their weights for any single token. StrataFlow takes advantage of that: it keeps the always-on trunk plus a bounded number of active experts resident, and streams the rest in from disk per token. The model lives on your SSD; only the parts a token actually uses are held in memory.

The question everyone asks is the obvious one: the model is 24 GB, how can it run on 12 GB (or 8 GB) of RAM? The answer is that RAM holds only the working set, not the whole model. Weights live in strata - layers of memory - and flow between them on demand:

   VRAM  (fastest, smallest)   <- hot experts, trunk, KV cache (optional GPU)
    ^v
   RAM   (fast, medium)        <- warm experts, resident trunk
    ^v
   SSD   (large, slowest)      <- the full model; cold experts stream in per token

For each token, only the trunk and the top-k active experts per layer are needed. StrataFlow keeps just that working set in RAM and streams the cold experts from disk. Everything else only has to be reachable on disk, not held in memory. You set the resident budget with a single knob (--expert-slots), and resident RAM stays bounded even as the on-disk model grows.

A consequence worth sitting with: more memory buys speed, not capability. The output is identical at every memory size. A bigger RAM budget just lets more experts stay cached, so fewer disk reads happen per token - the answer the model produces does not change. On an 8 GB machine the real constraints are therefore enough disk space for the model file, and your SSD read throughput, which sets tokens per second. RAM size does not cap what you can run.

How per-expert streaming actually works

StrataFlow runs its own forward pass over ggml. It is not a wrapper that hands your model to another engine - it builds and executes the compute graph itself, so it decides exactly which weights are in memory and when. That control is what makes true per-expert streaming possible.

Within a layer, the order matters. The router is read first. Only then are the top-k active experts made resident, right before the expert matmul runs. So resident expert memory is bounded to what a token actually uses, not the whole layer. A single tiered weight store spans VRAM, RAM and SSD with one shared caching policy - not three separate mechanisms bolted together - and the whole thing is fed by a streaming-friendly model format, .strata, built from standard GGUF. One .strata file holds the always-resident trunk plus aligned per-expert blobs so each expert can be read quickly.

To keep the cache warm, the router read-back feeds a predictor that prefetches the next layer's likely experts ahead of use. That raises the cache hit rate without changing results. And because placement is fiddly to tune by hand, StrataFlow profiles your machine on first run and decides what goes where - no hand-tuning placement flags.

What is measured, not asserted

We gate correctness on hard evidence. Every engine change has to reproduce the reference logits bit-for-bit (within floating-point noise), so StrataFlow's own forward pass matches a standard decode byte-for-byte. The bounded-RAM behavior is measured the same way: a 785 MiB model runs at about 132 MiB peak RSS, and a 1177 MiB model also runs in about 132 MiB. Peak memory stays flat as the on-disk model grows.

Quantized weights are validated too. Q8_0 and the K-quant families Q4_K/Q6_K are checked against the oracle for an identical greedy token sequence and stream through the .strata path. A benchmark harness reports real numbers - tokens per second, time to first token, peak RSS, resident weights and bytes per token - across a matrix of model sizes and memory budgets.

Where it stands, honestly

StrataFlow has a working engine core on CPU, green in CI on Linux, macOS and Windows. The bounded-RAM streaming mechanism is proven, and a real downloaded quantized model runs end to end through it - a dense TinyLlama on free Colab. We would rather be precise than impressive, so here is what is not done yet.

Running a real large MoE (as opposed to dense) end to end with captured numbers is demonstrated on Colab/real hardware, but not yet proven in our offline test sandbox - those large-model numbers are still being gathered. The IQ quantization families are not yet validated. Other MoE architectures (Qwen2-MoE, DeepSeek-MoE) are not yet implemented; only llama-arch MoE and dense llama are supported today. GPU backends are structurally reachable but need real GPU hardware to validate. And there is no tagged release and no OpenAI-compatible HTTP server front-end yet.

Try it, and the licensing

StrataFlow runs on CPU. The easiest way to see the big-model, bounded-RAM, no-GPU path on a real machine is the Colab guide (colab/README.md in the repository), which builds the engine, packs a model to .strata, and streams it with a bounded memory budget.

StrataFlow is source-available under the Coaade Source-Available License, Version 1.0. It is free for personal, non-commercial use - read it, learn from it, run it, and modify it for yourself. It is not for commercial use, reselling, rebranding, white-labeling, hosting as a service, or building a competing product; it is not an OSI open-source license and does not convert to Apache, MIT, or any other license over time. Companies, teams, and organizations, including for internal use, need a separate commercial license. The owner and licensor is Coaade Inc., a Delaware C corporation. StrataFlow is built on llama.cpp / ggml (MIT), and the repository contains no model weights. For commercial terms, contact contact@coaade.com.