Run open models as accelerated as your hardware allows
Magnitude is an open origin conclusion motor for agents that optimizes itself for your exact hardware. It compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. One click connects the delegate you already use (Pi, OpenCode, Hermes, Codex, and more). Works on Apple Silicon, NVIDIA, AMD, or nothing but a CPU.
Download Magnitude for macOS, Windows, or Linux
⭐ Help us attain additional developers and develop the Magnitude community. Star this repo!
demo-9-29.mp4- Download Magnitude, instal it, and open the app.
- Choose a recommended example in Discover and download it.
- Connect your delegate in Connections and commencement using it.
The desktop app includes the dimension CLI. No distinct facility is needed.
- Up to 2x faster than llama.cpp: 92% faster decode on Metal, 19% on CUDA
- Tuned on your device: kernels are tuned on your hardware before a example runs
- Built for the finest models: hand-optimized kernels for famous open-weight families
- Memory that flexes: 27% small recollection per agent, freed whenever agents stop
- Fast concurrent sessions: sessions portion prefix caches to forestall slowdown
- Works alongside your agent: one click to nexus Pi, OpenCode, Hermes, Codex, and more
- Free, private, open source: no token costs, nothing leaves your machine, Apache 2.0
An open origin conclusion motor that optimizes itself for your hardware. It ships as a desktop app that runs open models and connects them to the delegate you already use.
They container kernels precompiled for broad classes of hardware. Magnitude compiles and tunes its kernels on your genuine equipment before a example runs, so they fit your exact chip. See the benchmarks against llama.cpp.
Any Apple Silicon, NVIDIA, or AMD GPU, or nothing but a CPU. There is no fixed minimum. Smaller machines run smaller models, and additional recollection lets you run larger ones.
macOS, Linux, and Windows.
See the complete catalog at magnitude.dev/models. We compose optimized kernels for the most famous open-weight families, which is how we attack generalist engines.
One click connects Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. Anything alternatively plant through the OpenAI-compatible API.
Yes. Prompts, files, and models remain on your machine. No net needed formerly a example is downloaded.