press release

Coaade Launches StrataFlow: Run Very Large Models on CPU With No GPU Required

A source-available local inference engine that streams Mixture-of-Experts weights from disk, running models far bigger than available RAM inside a bounded memory budget.

Dover, Delaware - October 5, 2026 - Coaade Inc. today announced StrataFlow, a local inference engine for large Mixture-of-Experts (MoE) models that runs on the hardware people already own, with no GPU required. StrataFlow streams model weights from disk per token, so a model far bigger than a machine's RAM can run on a normal laptop or desktop CPU inside a bounded memory budget.

Large MoE models are huge on disk but only use a small slice of their weights for any single token. StrataFlow takes advantage of that structure: it keeps the always-on trunk plus a bounded number of active experts resident in memory, and streams the remaining cold experts in from the SSD as each token needs them. The model lives on disk; only the parts a token actually uses are held in memory.

The practical effect is that RAM holds only the working set, not the whole model. Users set the resident budget with a single knob, and resident memory stays bounded even as the on-disk model grows. More memory buys speed rather than capability - the output is identical at every memory size, because a larger RAM budget simply caches more experts and reduces disk reads per token. On a smaller machine the real constraints become disk space for the model file and SSD read throughput, which sets the token rate, rather than the amount of RAM.

Unlike tools that hand a model to a separate backend, StrataFlow runs its own forward pass over ggml. It builds and executes the compute graph itself, so it controls exactly which weights are in memory and when. Each layer's router is read first, and only the top-k active experts are made resident before the expert matmul runs, keeping resident expert memory bounded to what a token actually uses. A single tiered weight store spans VRAM, RAM and SSD under one caching policy, fed by a streaming-friendly model format called .strata that is built from standard GGUF. A predictor warms the next layer's likely experts into cache ahead of use to raise the hit rate without changing results.

StrataFlow's correctness and memory behavior are measured rather than asserted. Every engine change is gated on reproducing the reference engine's logits bit-for-bit within floating-point noise, so its own forward pass matches a standard decode. Peak memory stays flat as the model grows: a 785 MiB model and a 1177 MiB model both run at roughly 132 MiB peak RSS. Quantized weights in the Q8_0 and K-quant Q4_K/Q6_K families are validated against an oracle for an identical greedy token sequence, and a benchmark harness reports tokens per second, time to first token, peak RSS, resident weights and bytes per token across a range of model sizes and memory budgets.

"The question we hear most is how a 24 GB model can possibly run on 8 GB of RAM, and the honest answer is that you never needed to hold the whole model in memory at once," said a Coaade spokesperson. "StrataFlow keeps only the working set resident and streams the rest from disk, with the output identical at every memory size. We have proven the bounded-RAM mechanism and we are gathering real large-model numbers on real hardware."

Availability and what is next

StrataFlow has a working engine core on CPU that is green in continuous integration on Linux, macOS and Windows, covering the llama-architecture MoE family and dense llama. The easiest way to see the big-model, bounded-RAM, no-GPU path on a real machine is the Colab guide included in the repository, which builds the engine, packs a model to .strata, and streams it within a bounded memory budget. A real downloaded quantized model - a dense TinyLlama - runs end to end on free Colab.

Coaade is sharing its roadmap plainly. Running a real large MoE, as opposed to a dense model, end to end with captured numbers is demonstrated on Colab and real hardware but not yet proven in the company's offline test sandbox, and those large-model numbers are still being gathered. The IQ quantization families are not yet validated. Other MoE architectures such as Qwen2-MoE and DeepSeek-MoE are not yet implemented; only llama-arch MoE and dense llama are supported today. GPU backends are structurally reachable but need real GPU hardware to validate. There is no tagged release and no OpenAI-compatible HTTP server front-end yet.

About StrataFlow

StrataFlow is a local inference engine for large Mixture-of-Experts models that runs on CPU without a GPU, streaming experts from disk to run models larger than available RAM inside a bounded memory budget. It is source-available under the Coaade Source-Available License, Version 1.0: free for personal, non-commercial use, and not permitted for commercial use, reselling, rebranding, white-labeling, hosting as a service, or building a competing product. It is not an OSI open-source license and does not convert to Apache, MIT, or any other license over time. Companies, teams, and organizations, including for internal use, need a separate commercial license.

About Coaade

Coaade Inc. is a Delaware C corporation building specialized, local-first AI that keeps users' data on their own hardware. StrataFlow is built on llama.cpp / ggml (MIT), and the repository contains no model weights.

Media contact

Coaade Inc.
Email: contact@coaade.com