Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia

Hacker News by 10 min read 26x views
Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia

Share Post

Janus is a single Go binary that runs .gguf models on your device (GPU or CPU) and exposes an OpenAI-compatible API affirmative a built-in web UI. No Python, no Docker, no Ollama required — although Ollama is supported as a backend if you prefer.

Use it your way: call it from the command line (curl, PowerShell, scripts), cable it into Cursor / Cline / any OpenAI client, or use the Web UI — identical local models, any workflow fits you. Janus is the runner and router; you choose the forefront end.

Design idea: the example decides what to do; Go runs inference, routes requests, executes tools, and keeps everything local.


  • Local inference — llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback
  • OpenAI-compatible API — /v1/chat/completions, /v1/models, tool listing/calling
  • Web UI — Assistant, Chat, Kernel (tool loop), Config, Memory, Skills
  • Built-in tools — read/write files, run commands, math, docx in/out, PDF output, OCR (with Tesseract), and more
  • Hot-swap models — alter .gguf in the UI without restarting
  • Optional auth — Basic Auth for admin endpoints whenever JANUS_AUTH=true

Platform What you need
Windows (primary) Windows 10/11, Go 1.22+, Vulkan-capable GPU recommended
Linux Go 1.22+, Vulkan or CPU
macOS Go 1.22+, CPU backend (Vulkan varies by hardware)

Disk: scheme for the example size (often 2–8 GB per model) affirmative ~50 MB for Janus + llama.dll.

Optional: Tesseract OCR if you desire scanned-document OCR tools.


git copy https://github.com/Vibra-Ingenn/Janus.git cd janus .\build.ps1

build.ps1 downloads pre-built llama.cpp Vulkan DLLs and compiles dist\janus.exe.

If you already have DLLs in lib\windows\:

.\build.ps1 -SkipDownload

Put a .gguf document in the models\ folder. Easiest way — use the included downloader:

go build -o dist\modelget.exe .\cmd\modelget .\dist\modelget.exe -repo meta-llama/Llama-3.2-3B-Instruct -file Llama-3.2-3B-Instruct-Q8_0.gguf -out .\models\

Or download any compatible GGUF from Hugging Face manually.

Edit .env — you must set the example path:

INFERENCE_BACKEND=vulkan JANUS_MODEL_PATH=./models/Llama-3.2-3B-Instruct-Q8_0.gguf JANUS_MAX_TOKENS=4096 JANUS_AUTH=false
Variable Default Meaning
INFERENCE_BACKEND vulkan vulkan, cpu, or ollama
JANUS_MODEL_PATH (required) Path to your .gguf file
JANUS_GPU_LAYERS -1 -1 = all layers on GPU, 0 = CPU only
JANUS_VRAM_CEILING_MB 9216 VRAM prosperity clue (MiB)
JANUS_MAX_TOKENS 4096 Max tokens per reply
JANUS_LISTEN_ADDR 127.0.0.1:8990 Bind address
JANUS_AUTH false Set true to necessitate login on admin routes
JANUS_SAFE_MODE false Set true to obstacle casing commands in tools

Or rebuild + initiate in one step:

Janus opens http://127.0.0.1:8990 in your browser (disable alongside JANUS_NO_BROWSER=1).

First startup loads the example into VRAM — expect 10–60 seconds depending on example size and disk speed.

curl http://127.0.0.1:8990/health

You should see {"status":"ok",...}.


Quick commencement (Linux / macOS)

git copy https://github.com/Vibra-Ingenn/Janus.git cd janus go mod tidy go build -o dist/janus ./cmd/janus cp .env.example .env # edit .env — set JANUS_MODEL_PATH and INFERENCE_BACKEND=cpu if no Vulkan ./dist/janus

On Linux you need libllama.so next to the binary or on LD_LIBRARY_PATH. See docs/AIR_GAPPED_INSTALL.md for offline setup.


After .\dist\janus.exe starts, your browser should open http://127.0.0.1:8990 (or open that URL yourself). Wait for the position pill to display your example — archetypal burden can obtain 10–60 seconds.

Everyday use (Assistant or Chat)

Assistant — finest for “get this done” tasks:

  1. Type what you desire in the big content box
  2. Optional: drag a document onto the upload area or click attach
  3. Click Get Started
  4. Read the result; use Copy or commencement a New Request

Chat — finest for back-and-forth conversation:

  1. Open the Chat tab
  2. Type a communication and media Enter (Shift+Enter for a new line)
  3. Pick a example from the dropdown if you have additional than one

Click ⚙ Advanced in the top nav to reveal:

Tab What to do
Kernel Type a project → Run Kernel → observe all tool call in the live log
Config Pick a .gguf from the dropdown → Save & Load Model; edit the conversation scheme immediate here
Memory Browse former chats; preserve key/value facts the AI can reuse
Skills Install community tools from JSON (optional power-user feature)
  1. Advanced → Config
  2. Choose a example from the catalog (or paste a way to ./models/your-model.gguf)
  3. Click Save & Load Model — no restart needed

Focus the terminal anywhere Janus is operating and media Ctrl+C.

More walkthroughs (uploads, troubleshooting, tools): docs/USER_MANUAL.md


Janus is meant to remain out of your way — run prompts from a terminal, a script, or a third-party app by pointing it at Janus akin any another OpenAI endpoint:

curl http://127.0.0.1:8990/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{  "model": "local",  "messages": [{"role": "user", "content": "Hello!"}]  }'

Base URL: http://127.0.0.1:8990/v1
API key: not required whenever JANUS_AUTH=false

Method Path Description
GET /health Liveness (?deep=true for component checks)
GET /v1/models Model list
POST /v1/chat/completions Chat (streaming supported)
GET /v1/tools/list Built-in + community tools
POST /v1/tools/call Call a tool directly
POST /kernel/run Run the ReAct kernel on a task
POST /upload Upload a document (multipart, 50 MB max)

Full tool reference: docs/TOOLS_REFERENCE.md
Operator guide: docs/USER_MANUAL.md


Ollama backend (optional)

If you already use Ollama alternatively of local GGUF:

INFERENCE_BACKEND=ollama OLLAMA_BASE_URL=http://127.0.0.1:11434 OLLAMA_MODEL=mistral:7b

Janus proxies conversation to Ollama; tools and the web UI motionless work.


Pitfalls (learned the difficult way)

If you're construction or hacking on Janus, these are the gotchas that burned us repeatedly. None of this is apparent the archetypal time.

Problem What’s going on Fix
“It built but my changes aren’t there” On Windows, Go can’t overwrite a operating .exe. Old janus.exe keeps operating on the port. End all janus.exe in Task Manager, or .\run.ps1, afterward rebuild.
“Address already in use” / two Januses A leftover procedure (sometimes elevated) motionless holds harbor 8990. Janus tries to spotless this up, but an admin case may survive. Task Manager → end all janus.exe. Run your terminal as the identical person that started the old one.
Server starts, conversation fails / no model JANUS_MODEL_PATH wrong, or no .gguf in models/. Copy .env.example → .env, set the path, put the document in models/. Use Config → Save & Load Model.
“local motor unsuccessful to start” Missing llama.dll, bad GPU drivers, or example way typo. Run .\build.ps1 archetypal (copies DLLs into dist/). Update GPU drivers, or set INFERENCE_BACKEND=cpu.
Built alongside lone go build go build solitary doesn’t fetch llama.cpp DLLs. Use .\build.ps1 for a complete Windows build, or copy DLLs from lib\windows\ into dist\ yourself.
Wrong URL Default is http://127.0.0.1:8990, not 8080. Bookmark 8990. API basis is http://127.0.0.1:8990/v1.
“.env ignored” Janus searches upward from current directory and next to the exe. Running from the incorrect directory loads a distinct .env (or none). Run from project root, or keep .env beside dist\janus.exe.
Paths to models break ./models/foo.gguf is related to where you launch Janus, not anywhere the origin code lives. Launch from repo root, or use an complete way in .env.
First answer takes forever Model is loading into RAM/VRAM — normal. Wait 10–60s; smaller quants (Q4) burden faster than Q8.
401 on Config / example load Auth is on (JANUS_AUTH=true). Set JANUS_AUTH=false for local dev, or use the admin password printed on archetypal run.
OCR / PDF tools fail Tesseract not installed; PDF input is constricted in this build. Install Tesseract for scans; use Word/text uploads or ocr_extract.
Committed secrets Never commit .env — it holds keys and passwords. Only commit .env.example.

If item motionless feels possessed, inspect logs/janus.log in the project directory and the terminal output from startup.


“Model not found” / server starts but conversation fails

  • Check JANUS_MODEL_PATH in .env matches a genuine document under models/
  • Use Config → Save & Load Model in the web UI to choice from detected .gguf files
  • Paths are related to anywhere you run janus.exe (usually project base or dist/)

Normal — the example loads on archetypal petition or at startup. Smaller quantizations (Q4, Q5) burden faster than Q8.

  1. Update GPU drivers
  2. Try CPU: INFERENCE_BACKEND=cpu and JANUS_GPU_LAYERS=0
  3. Or use Ollama backend (above)

Port already in use / rebuilt but nothing changed

See the pitfalls table — nearly continually a old janus.exe on Windows, or the incorrect harbor (default 8990, not 8080). If another app owns the port, set JANUS_LISTEN_ADDR=127.0.0.1:8991 in .env.

When JANUS_AUTH=true, Janus prints an admin password on archetypal run. Use it for Config/Model admin routes, or set JANUS_ADMIN_PASSWORD in .env before archetypal launch.

Tool errors (OCR, PDF input)

  • OCR needs Tesseract installed and on PATH
  • PDF content extraction in OSS is constricted — use ocr_extract for scans, or nexus plain content / Word files
  • Output PDFs via render_pdf activity without additional installs

cmd/janus/ Main server cmd/modelget/ Hugging Face example downloader internal/kernel/ ReAct iteration (model + tools) internal/tools/ Built-in and community tools internal/engine/ llama.cpp Vulkan/CPU backend internal/bridge/ DLL loader models/ Put .gguf records current (not committed) dist/ janus.exe + llama.dll following build docs/ Manuals and references 

Hacking on Janus is welcome. Read Pitfalls (learned the difficult way) archetypal — particularly the Go-on-Windows rule: stop all janus.exe processes before go build, or you’ll think your fix didn’t work.

Typical loop:

Get-Process -Name "janus" -ErrorAction SilentlyContinue | Stop-Process -Force go test ./... go build -o dist\janus.exe .\cmd\janus .\dist\janus.exe

Or fair .\run.ps1 (stop → build → run in one step). For a complete Windows build including llama DLLs, use .\build.ps1.

Contributions greeted — see CONTRIBUTING.md.


Also from the identical team: Vibe Engine PRO

Janus is complete on its own — nothing current is missing or locked rearward a paywall. Use it from the CLI, scripts, Cursor, the Web UI, or any OpenAI-compatible client.

If you bask operating local models and desire to see what alternatively is out there, the identical squad makes Vibe Engine PRO — a distinct merchandise focused on recipe-based automation (multi-step workflows, MCP, pipelines that run in Go without calling the example all step). It is one option, not a requirement. You can additionally broaden Janus yourself via the API.

Janus (this repo) Vibe Engine PRO
What it is Free example runner & API router Paid automation / formula engine
Price Free (MIT) 7-day trial, afterward $8.99/month

Vibe Engine connects to a local OpenAI-compatible endpoint — Janus plant for that:

Base URL: http://127.0.0.1:8990/v1 API key: (leave blank whenever JANUS_AUTH=false) 

Early subscribers at $8.99/month remain at that charge if the listed cost goes up later.

Take a appearance if it appears interesting. If Janus is all you need, that's fine too.


MIT — see LICENSE.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads