
Strata lets consumer PCs run a 125-billion-parameter AI model locally, using system RAM and GPU memory instead of cloud servers.
Strata is an open-source inference engine that allows the 125-billion-parameter Qwen3.8-Flash-Next AI model to run on consumer gaming PCs. The project supports Windows and Linux systems with compatible NVIDIA or AMD graphics cards and is designed to keep AI processing on the user’s machine rather than sending data to a remote server. Strata is released under the MIT licence, while the model files and some third-party components retain their respective licences.
The project addresses one of the main challenges of running large AI models locally: memory requirements. Strata distributes the workload across the GPU, system RAM and processor, allowing hardware with 12GB or more of VRAM and at least 32GB of RAM to run the model. The project recommends 64GB of RAM for running all supported model sizes, while around 80GB of free storage is required.
Strata uses a mixture-of-experts approach in which the model contains 24,576 experts, but only 10 are required for each word. The system keeps frequently used experts on the graphics card while storing the complete set in system RAM, with the processor handling additional computation. It also uses speculative decoding, where a smaller helper predicts upcoming tokens for the larger model to verify, reducing the time required to generate responses.
On an RTX 5070 with 12GB of VRAM, Ryzen 5 7600 and 64GB RAM, the project reports up to 94 tokens per second with its Q2_0 model configuration. An AMD RX 9070 XT with 16GB VRAM and 47GB RAM achieved up to 60 tokens per second in the same benchmark. Strata can also process images and expose local OpenAI- and Anthropic-compatible APIs, allowing coding assistants and other applications to connect to the locally running model.
The open-source project provides the inference engine, installation scripts, documentation, tests and server components, with support for multi-GPU configurations and an MCP server for AI-assisted setup and control. It also incorporates components from projects including llama.cpp and ggml. By combining model compression, GPU-RAM sharing and local inference, Strata provides developers with an open implementation for experimenting with large AI models on consumer hardware rather than depending entirely on cloud-based inference.
For more information, click here.














































































