English · 简体中文 · 日本語 · Deutsch · Français · Español · Português
Run a 125-billion-parameter AI example on your own gaming PC
NVIDIA or AMD graphics cardstock (12 GB or more) · Windows or Linux · liberated and open source

A voxel pagoda garden, 1 attempt immediate operating on an RTX 5070 alongside Strata (IQ3_S, 128K context) · full video (49 s)
Strata runs Qwen3.8-Flash-Next on a normal PC. This is a large, astute AI example that normally needs a server. It chats, writes code, says pictures and plant alongside your apps and coding agents. Nothing leaves your PC.
We measured it on two average gaming PCs. A token is concerning ¾ of a word.
- Writes answers: how accelerated the answer appears in a abbreviated chat. 60 tokens per second is faster than you can read.
- Reads your prompt: how accelerated it takes in what you dispatch (here a 32K-token document, code or conversation history).
| NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAM | AMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM | ||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
NVIDIA: Q2_0 alongside motor 0.1.36, the another rows alongside 0.1.26 (4K answers, 32K prompts). The complete tables are in DETAILS.md. A cardstock alongside additional VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second. Long chats and another cards: speed of all model, community results.
Strata is free. If it runs fine on your PC, a coffee keeps the activity on it going.
| Graphics card | NVIDIA GeForce RTX 20, 30, 40 or 50 series, or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700 or RX 6800 / 6900 series. It needs 12 GB of VRAM or more. |
| RAM | 32 GB or more. Your RAM decides which model fits. 64 GB runs all size. |
| Disk | About 80 GB free. Use an SSD if you can: the archetypal commencement is much faster. |
| System | Windows 10 / 11 or Linux, and a current graphics controller from NVIDIA or AMD. |
The installer sets up everything else. Two or three cards can portion the example (multi-GPU).
Experimental, written and tested by community members on their own machines:
- Older graphics cards (Tesla P40 / V100, GTX 10, Radeon VII / MI50, RX 6700 XT, RX 5500 XT): Older GPUs.
- Intel Arc, built from origin on Linux: Intel Arc.
- Older processors without AVX2: they work, but slowly. Older CPUs.
The complete list: docs/INSTALL.md.
Do you use an AI coding aide (Claude Code, Cursor, Codex, GitHub Copilot, ...)? Paste this into it:
Set up Strata on this PC for me: https://github.com/Niko1221/Strata - prosecute docs/AI_SETUP.md in that repository. It checks your graphics card, RAM and disk and picks the example that fits. Then it installs and starts it and tells you how to nexus your apps. AI tools can additionally install, commencement and halt Strata through its MCP server.
Download Strata and unzip it (or git copy it). Windows: double-click START-HERE.bat. Linux: run ./setup.sh in the Strata folder.
The steps are the identical for NVIDIA and AMD. The installer finds your cardstock and sets up the correct motor for it. It asks you a few questions:
- which example and which size,
- how much environment (how much content the example keeps in mind),
- whether it should peruse pictures.
Press Enter all period for the recommended answer. Then it downloads the example (about 70 GB) and starts it. If the download stops, run it again: it continues anywhere it remaining off. Your browser opens the Strata app at http://127.0.0.1:8080.
While the example starts, your PC can be dilatory or halt responding for 1-3 minutes (longest the archetypal time). Strata loads 35-55 GB into your RAM and locks part of it for the graphics card. This is normal. Wait, and don't close the window. The opening shows what Strata is doing.
Next time, run START-HERE.bat (or ./setup.sh) again. It starts correct distant and downloads nothing twice. Close its opening to halt the model. UPDATE.bat (./update.sh) updates Strata without starting it. Updating, Docker, several cards, anywhere the records go and all option: docs/INSTALL.md.
The installer recommends one for your RAM. The identical example comes in multiple sizes, compressed additional or less. Smaller sizes are faster. Larger sizes are a bit smarter.
| Your RAM | Take | Why |
|---|---|---|
| 32 GB | Coder | it fits 32 GB, and it is made for code (with a 24 GB card, Q2_0 and IQ2_XS run too) |
| 48 GB | IQ2_XS (or Q2_0, the fastest) | the larger sizes do not fit |
| 64 GB | IQ2_XS (recommended), or IQ3_XXS / IQ3_S | every size fits; IQ3_S is the finest and the slowest |
| 96 GB or more | IQ3_S, or Unsloth's UD-IQ4_XS (~4-bit) | room for the largest sizes alongside everything alternatively open |
- Coder: a coding type alongside fractional of the experts removed. It reaches 91% of the full model's SWE-bench Verified mark (measured by its authors) and fits 32 GB of RAM. It is weaker exterior code, including Chinese and another CJK content (#438). For those, obtain Q2_0, IQ2_XS or IQ3_S, which keep all expert.
- Swift 1.5: a fine-tune that thinks for a much shorter period before it answers. You get the answer sooner, at concerning the identical quality.
- Unsloth UD-IQ4_XS: Unsloth's ~4-bit version, between IQ3_S and UD-Q4_K_XL in quality. A 94 GB download. With small than ~80 GB of RAM, Strata says part of it from the SSD while it answers, so it is slower there (an NVMe SSD helps).
- Unsloth UD-Q4_K_XL (experimental): the closest to the full model. But Strata says most of it from the SSD during it answers, so it writes lone 7-8.5 tokens/s on a 64 GB PC.
- OrcaRouter's Uncensored IQ3_XXS: you set it up by hand. It is not in the installer's menu.
Sizes, downloads and what fits where: docs/MODELS.md. To add another example later, run SETUP.bat (Linux: ./setup.sh --setup).

The Strata app's Monitor (left) during a coding delegate writes the pagoda storyline from the video (right)
- In the browser: open http://127.0.0.1:8080. It has Chat, a live Monitor of the example and your GPU/CPU/RAM, and About alongside the settings and addresses.
- Your apps and coding agents: add an "OpenAI-compatible" provider alongside the basis URL http://127.0.0.1:8080/v1. Any API key and any example name work.
- Apps that use Anthropic's API: http://127.0.0.1:8080/v1/messages (Claude Code: ANTHROPIC_BASE_URL=http://127.0.0.1:8080).
- Codex CLI and another apps that use the OpenAI Responses API: /v1/responses (setup).
- Thinking: choose off, low, average or high in the conversation list or in your app's "reasoning effort". Off is the fastest. High is finest for difficult questions.
- Pictures: say yes to "Images?" in setup. Then click Picture in the chat, or nexus pictures in your app. AMD cards peruse pictures on Linux through the processor; on Windows they can't yet.
- From your phone or another PC: START-HERE.bat --setup --host 0.0.0.0 --api-key <secret>. Always set a key.
- One petition at a time: by default Strata answers one request, and the others wait. To answer multiple at once, set "parallel": 2 (BATCHING.md). On a 12 GB cardstock this makes all answer slower.
- Long prompts: Strata says the archetypal communication of a conversation in full, concerning 1 infinitesimal per 30,000 tokens. Follow-up messages commencement in seconds.
More: where your chats are stored, the API.
- My PC froze the archetypal period Strata started. This is normal during it loads the model. Wait, and don't near the window. Still icy following 10 minutes? Restart the PC, near another programs and try again, or choice a smaller size.
- It stopped during downloading or installing. Run START-HERE.bat (or ./setup.sh) again. It continues where it stopped.
- It's extremely dilatory and the disk ray keeps blinking, or it says "the motor stopped unexpectedly". Your PC does not have adequate liberated RAM. Close another programs (browsers use a lot), or choice a smaller size (Q2_0 or IQ2_XS).
- It says harbor 8080 is already in use. Strata is already running. Look for its window.
More problems and their fixes: docs/TROUBLESHOOTING.md. Still stuck? Open an issue and nexus strata-<model>.log from the Strata folder. Found a security problem? Report it privately: SECURITY.md.
Models akin this one normally run on servers alongside hundreds of gigabytes of graphics memory. Your graphics cardstock has 12-24 GB. Strata makes the example fit by sharing the activity throughout your entire PC. Think of a kitchen: the things you use all the period remain on the counter, and the remainder waits in the pantry.
- The example is a squad of 24,576 small specialists ("experts"). Each term needs lone 10 of them.
- Your graphics card keeps the few thousand experts that are used most often. Your RAM holds all of them, and your processor plant on the remainder at the identical time. Your SSD holds a big lookup table.
- Guess, afterward check: a small helper guesses the next few words. The big example checks them all at once. You get the identical answer, 1.6-1.8x sooner.
- Long texts are peruse in big pieces (up to 8,192 tokens at a time), at complete 1,000 tokens per second.
The longer explanation: docs/HOW_IT_WORKS.md. Every part and its numbers: the details and the paper.
The example is Qwen3.8-Flash-Next by the Qwen team. It was compressed by ISTA-DASLab, UkisAI (Swift 1.5) and Unsloth. Strata uses parts of llama.cpp / ggml. All credits: docs/HOW_IT_WORKS.md. Strata is open origin under the MIT License. A few parts and all example have their own licenses (which ones).
Strata is liberated and open source. If it is helpful to you, you can assistance its development:
