Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

Hacker News by 8 min read 53x views
Jeff – Jev-compatible 0.8B decision models, trained at home, ~30 ms

Share Post

Fine-tunes of Qwen3.5 and Gemma 4 for zero-shot classification: small, accelerated decision models you slot into your code, alongside the identical petition format as Jev. You depict a circumstance and catalog the options in plain words; Jeff returns a calibrated probability for all choice from a sole onward pass. No generated text, no parsing: concerning 22 ms per decision on an RTX PRO 6000 and 28 ms on an Apple M4 Max (MLX).

Zero-shot method the options can be anything: assistance queues, person intents, moderation labels, sound commands, game moves. Your categories don't need to appear in the training data; you depict them, and Jeff picks.

What it is, and what it isn't. These are extremely small models. They create extremely fast, well-calibrated judgement calls between options, and they slot effortlessly into your local code. On benchmarks they approach, and sometimes beat, Jev; but at this size their reasoning won't equivalent Jev's, which runs on a much larger model. If zero-shot accuracy isn't good adequate for your purposes, a abbreviated fine-tune on your own examples takes you much further: our voice-navigation fine-tune moved held-out accuracy from 31.7% to 95.8% in under fractional an hr on one GPU.

Built entirely on local hardware. Training on one RTX PRO 6000 workstation GPU (the 0.8B trains in concerning 2 hours, the 2B in concerning 3.5), all synthetic training data written by an open example (Qwen3.8-Flash-Next) on two DGX Sparks, testing on a MacBook. No haze GPUs, and no closed-model output in the training data; a closed example was used lone to spot-check the norm of a example of the synthetic data.

Independent project. Jeff uses the identical petition format as Jev, but it is not affiliated alongside or endorsed by TypeSafe, the makers of Jev. Our training code starts from the open-source AutoJev recipe.

Models on Hugging Face: Jeff-Qwen3.5-0.8B · Jeff-Qwen3.5-2B · Jeff-Gemma4-E2B

uv sync uv run hf download mstrasser/Jeff-Qwen3.5-0.8B --local-dir checkpoints/jeff-0.8b # NVIDIA GPU or CPU (PyTorch) JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve # Apple silicon (MLX, much faster on a Mac; Qwen models only) uv sync --extra mac JEFF_BACKEND=mlx JEFF_CHECKPOINT=checkpoints/jeff-0.8b PORT=8765 uv run jeff-serve
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{  "model": "jeff-latest",  "state": "Refund request: the client says the parcel arrived shattered and wants their prosperity back.",  "questions": {  "route": {"type": "choice", "instructions": "Which squad should grip this?",  "criteria": {"1": "Refunds and payments", "2": "Damaged or misplaced parcels", "3": "Account and login problems"}},  "angry": {"type": "noul", "instructions": "Is the client angry?"}  } }'

Each answer has a probability per option, the chosen choice and a confidence. Three inquiry types: choice (pick one of up to 255 options), noul (yes/no, returned as a probability) and mark (a item on a measure you describe). Several autonomous questions in one petition are answered together.

4,599 questions from five community benchmarks, affirmative JevBench's community difficult tier (105 items, scored separately):

Accuracy of Jeff-Qwen3.5-0.8B, Jeff-Qwen3.5-2B and Jeff-Gemma4-E2B against Jev's published figures, per benchmark

Benchmark Qwen3.5-0.8B untrained Jeff-Qwen3.5-0.8B Qwen3.5-2B untrained Jeff-Qwen3.5-2B Gemma 4 E2B untrained Jeff-Gemma4-E2B Jev (published) AutoJev-27B (published)
Overall (5 benchmarks) 45.3 79.1 46.5 83.1 62.5 81.6 83.0 84.9
BBH 39.5 64.0 46.0 68.0 51.3 66.4 94.3 82.8
Financial PhraseBank 36.0 96.4 53.4 96.3 86.0 96.1 77.0 84.2
JudgeBench 56.6 62.6 57.4 64.6 46.9 60.6 78.6 78.9
RAGTruth 49.1 86.1 35.9 88.9 63.8 87.4 77.3 88.9
WinoGrande 49.2 68.6 52.2 79.0 51.0 77.4 90.7 83.3
JevBench difficult (separate) 36.2 47.6 45.7 53.3 41.0 48.6 73.3 70.3

Bold: the victor of Jeff against Jev in all row. Bold italic: AutoJev-27B anywhere it is the finest of all models in the row (on RAGTruth, tied alongside Jeff-Qwen3.5-2B); it is shown for reference, since the head-to-head difference is with Jev. The published Jev and AutoJev figures were measured on a distinct example of the identical benchmarks. Jeff's overall mark comes from classification and grounding, anywhere it matches or strikes the ample models; on the reasoning-heavy benchmarks (BBH, JudgeBench, JevBench) it stays fine below them, as you would anticipate at this size.

To test zero-shot achievement on tasks dissimilar item in the benchmarks, we had Jeff perform three games. Games aren't the ideal zero-shot test, since a game's province isn't representative unstructured data; but they are a common, and fun, way to test a System 1 model. Each turn, the code describes the circumstance and the lawful moves in words, and the example picks one. The options province what all move leads to (Frogger: "you would be hit by a car and endure a life"; Doom: "the nearest monster is a small to your left"), but never which move is right. Each outcome is 20 episodes, kernel 1234; ▶ opens a video of the run's archetypal episode.

Jeff-Qwen3.5-0.8B playing, zero-shot (the bold row in the array below; click a clip for the complete video):

Model Doom, kills (monster's direction in words) Frogger, crossings (consequences) Pac-Man, pellets of 98 (consequences)
Random moves −0.05 0 11.2
Hand-coded regulation bot 6.55 ▶ 10.25 ▶ 94.1 ▶
Qwen3.5-0.8B, untrained 5.0 ▶ 1.0 ▶ 25.8 ▶
Jeff-Qwen3.5-0.8B 6.55 ▶ 10.3 ▶ 57.0 ▶
Qwen3.5-2B, untrained 0.55 ▶ 0.05 ▶ 72.1 ▶
Jeff-Qwen3.5-2B −0.9 ▶ 6.0 ▶ 41.2 ▶
Gemma 4 E2B, untrained −0.55 ▶ 0 ▶ 3.2 ▶
Jeff-Gemma4-E2B 0.55 ▶ 0.15 ▶ 53.2 ▶
Jev (published, Doom) 6.55, told the aiming rule; −0.60 without it — —

Jeff-0.8B decides in 29–49 ms per move on an M4 Max; Jev's published Doom run took 212 ms per call complete its API. The two times were not measured on the identical hardware. To perform them yourself:

uv sync --extra games uv run python -m jeff.games --game destiny --player jeff --criteria circumstance --url http://127.0.0.1:8765 --video --out runs/games/doom.json uv run python -m jeff.games --game frogger --player jeff --criteria outcomes --url http://127.0.0.1:8765 --out runs/games/frogger.json uv run python -m jeff.games --game pacman --player regulation --out runs/games/pacman-rule.json

Median period per decision complete the identical 200 benchmark questions (about 200 input tokens each), one inquiry at a time, from raw content to probabilities:

Model Parameters Weights (16-bit) NVIDIA RTX PRO 6000 Apple M4 Max (MLX) CPU (32 threads)
Jeff-Qwen3.5-0.8B 0.8B 1.7 GB 22 ms 28 ms 463 ms
Jeff-Qwen3.5-2B 2B 4.2 GB 24 ms 60 ms 708 ms
Jeff-Gemma4-E2B 2B productive (4.6B stored) 9.3 GB 29 ms — (MLX runs Qwen only) 1.0 s
AutoJev-27B 27B ~54 GB not published — —
Jev not disclosed API only 114–212 ms per call in published Doom runs, including the network
  • Reason in code, decide alongside Jeff. It's a classifier, not a planner. State what all choice leads to ("this move gets you hit by a car"); asked to forecast ("a car arrives in 2 turns"), it does no improved than random.
  • Wording matters enormously. Describe options consistently: giving Frogger's goal choice the identical words as every other onward choice took one event from 15 crossings to 23.
  • Use abbreviated choice keys and descriptive text: {"1": "Engagement letter"}, not lengthy IDs, which disbursal period and add nothing.
  • Ask autonomous questions together in one request.
  • Fine-tune it if zero-shot isn't enough. A voice-navigation fine-tune on ~11k app-specific examples took concerning half an hr on one GPU and moved held-out accuracy from 31.7% to 95.8%, at concerning 40 ms per decision on an M4 Max: autojev-train --initial-checkpoint <jeff> --epochs 1 ....
  • Pick the size for the job. For accelerated choice picking the 0.8B is the saccharine spot: the 2B is additional careful and plays the games worse, notwithstanding scoring higher on the benchmarks.
uv run autojev-mix ... # build the training set (public data, synthetic data, leak filter) scripts/train.sh RUN data/mix/public.jsonl data/mix 5e-6 40 Qwen/Qwen3.5-0.8B <revision> --epochs 1 uv run autojev-evaluate --data data/panel.jsonl --local --checkpoint checkpoints/RUN/selected --output runs/eval/RUN.json

The complete pipeline (synthetic data from a local teacher, leak filter, learning-rate sweeps, dashboard) is described in scripts/train_all.sh, and all training origin alongside its licence in docs/data-sources.md. Training recipe: full-weight fine-tuning, one epoch, batches of 256, cross-entropy complete the option letters, afterward one fitted heat for calibration; checkpoints are chosen on a betterment set, never on the benchmark panel. At smallest fractional of all training family follows the panel's layout conventions (formats only; no panel item is always trained on).

  • Small models don't reason. Expect fast, calibrated choices between the options you describe, not multi-step reasoning. At 0.8B–2B parameters this holds for all model, not fair Jeff.
  • Jeff-2B is a weaker equivalent participant than Jeff-0.8B. The untrained 2B already appears additional risk-averse than the untrained 0.8B, and our training seems to have made that worse. This needs additional investigation.
  • Benchmark scores don't foretell equivalent play. The untrained Gemma 4 E2B strikes the untrained Qwen models on the benchmarks yet plays the games worst: correct most of the time, but not reliably. Training fixed its Pac-Man (3.2 → 53.2 pellets) but not its Doom or Frogger.
  • Prompts matter. Jev's own Doom immediate (a raw direction figure affirmative an aiming rule) does not activity for any of our models; options that province consequences in words do.
  • English and content only.

Jeff began as a fork of AutoJev by Denis Yarats (MIT licence), an open recipe that fine-tunes Qwen3.8-27B to come back Jev-style decisions. We kept its center scheme (one onward continue per decision, a trained answer readout, a fitted heat for calibration) and built on it: small students (0.8B and 2B Qwen, Gemma 4 E2B), a local synthetic-data pipeline alongside a leak filter, immediate layouts for domain fine-tunes, MLX serving on Apple silicon, equivalent tests and a training dashboard. The first copyright notice is kept in LICENSE.

Code: MIT (including AutoJev's). Model weights: Apache 2.0. Doom harness adapted from jev-plays-doom (MIT). Training data: see the dataset card; each source keeps its licence and is listed in docs/data-sources.md. We publish the weights and code, not the training data; several sources are share-alike (CC BY-SA).

Other Article Hacker News
↑
Close Right Ads
Close Left Ads