llama.cpp

Hacker News by 1 min read 501x views
llama.cpp

Share Post

Pair it pinch a section coding agent.

Run llama serve, instal the pi-llama plugin and motorboat Pi. It will automatically observe your section model. No config, nary API keys. Files enactment connected your machine, requests ne'er time off it.

# 1. Serve a model llama serve # 2. Install the pi-llama plugin pi install git:github.com/huggingface/pi-llama # 3. Run Pi, everything is set pi

Pi

Optimized for immoderate hardware.

From your laptop to a cluster, llama.cpp runs connected immoderate you have. Same binary, aforesaid models, aforesaid hand-tuned kernels for each GPU and CPU.

Apple Silicon

M Ultra

RTX 5090

CPU

Jetson

H100

MI300

RTX 4090

A100

M Pro

M Max

DGX Spark

T4

Radeon RX

B200

Intel Arc

RTX 3090

Other Article Hacker News
↑
Close Right Ads
Close Left Ads