Launch HN: Magnitude (YC S25) – Self-optimizing conclusion motor for agents

Hacker News by 3 min read 153x views
Launch HN: Magnitude (YC S25) – Self-optimizing conclusion motor for agents

Share Post

Magnitude icon

Run open models as accelerated as your hardware allows

Download Magnitude Documentation Discord Follow Magnitude on Twitter GitHub Repo stars

Magnitude is an open origin conclusion motor for agents that optimizes itself for your exact hardware. It compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. One click connects the delegate you already use (Pi, OpenCode, Hermes, Codex, and more). Works on Apple Silicon, NVIDIA, AMD, or nothing but a CPU.

Download Magnitude for macOS, Windows, or Linux

⭐ Help us attain additional developers and develop the Magnitude community. Star this repo!

demo-9-29.mp4
  1. Download Magnitude, instal it, and open the app.
  2. Choose a recommended example in Discover and download it.
  3. Connect your delegate in Connections and commencement using it.

The desktop app includes the dimension CLI. No distinct facility is needed.

  • Up to 2x faster than llama.cpp: 92% faster decode on Metal, 19% on CUDA
  • Tuned on your device: kernels are tuned on your hardware before a example runs
  • Built for the finest models: hand-optimized kernels for famous open-weight families
  • Memory that flexes: 27% small recollection per agent, freed whenever agents stop
  • Fast concurrent sessions: sessions portion prefix caches to forestall slowdown
  • Works alongside your agent: one click to nexus Pi, OpenCode, Hermes, Codex, and more
  • Free, private, open source: no token costs, nothing leaves your machine, Apache 2.0

Up to 2x faster than llama.cpp

 9% faster prefill and 92% faster decode on Metal, 23% faster prefill and 19% faster decode on CUDA

An open origin conclusion motor that optimizes itself for your hardware. It ships as a desktop app that runs open models and connects them to the delegate you already use.

How is it faster than llama.cpp, Ollama, or LM Studio?

They container kernels precompiled for broad classes of hardware. Magnitude compiles and tunes its kernels on your genuine equipment before a example runs, so they fit your exact chip. See the benchmarks against llama.cpp.

Any Apple Silicon, NVIDIA, or AMD GPU, or nothing but a CPU. There is no fixed minimum. Smaller machines run smaller models, and additional recollection lets you run larger ones.

What functioning systems does it support?

macOS, Linux, and Windows.

Which models does it support?

See the complete catalog at magnitude.dev/models. We compose optimized kernels for the most famous open-weight families, which is how we attack generalist engines.

Which agents activity alongside it?

One click connects Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. Anything alternatively plant through the OpenAI-compatible API.

Yes. Prompts, files, and models remain on your machine. No net needed formerly a example is downloaded.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads