Pair it pinch a section coding agent.
Run llama serve, instal the pi-llama plugin and motorboat Pi. It will automatically observe your section model. No config, nary API keys. Files enactment connected your machine, requests ne'er time off it.
# 1. Serve a model llama serve # 2. Install the pi-llama plugin pi install git:github.com/huggingface/pi-llama # 3. Run Pi, everything is set pi
Optimized for immoderate hardware.
From your laptop to a cluster, llama.cpp runs connected immoderate you have. Same binary, aforesaid models, aforesaid hand-tuned kernels for each GPU and CPU.
Apple Silicon
M Ultra
RTX 5090
CPU
Jetson
H100
MI300
RTX 4090
A100
M Pro
M Max
DGX Spark
T4
Radeon RX
B200
Intel Arc
RTX 3090