Jeeves. Reasoning improves Jev-like decision models

Hacker News by 8 min read 518x views
Jeeves. Reasoning improves Jev-like decision models

Share Post

A reasoning Jev-style classifier alongside a diffusion drafter, trained alongside SFT and CISPO.

Jeeves

 9B  MIT

Inspired by Kev.

  • A 9B Jev-like example (Qwen3.5-9B, LoRA, pointer head) that thinks before it decides, alongside a block-4 diffusion drafter and the complete training code and train/dev/test data.
  • Beats Kev-9B and Jev on test data it was never trained on (0.889 vs 0.822 and 0.857) and on JevBench's community tiers (0.935 vs 0.866 for Jev).
  • Supports yes/no (noul), multiple-choice (choice), and ranking (score) questions in the identical request, through a Jev-compatible API.
  • About 0.3 s per petition without thinking and a 3.3 s median alongside it on one H100. Can be sped up by truncating sequence length.
  • Runs on CUDA (Hopper for the FP8 kernel).

Jev-like models provision calibrated decision probabilities, but at low accuracy. A lot of pipelines hence depend on a reasoning example as a fallback. Jeeves trains a Jev-like Qwen3.5-9B (LoRA and a pointer head) using CISPO to logic before it decides.

This results in improved achievement on out of domain tasks, and outperforms Jev in JevBench difficult (public).

Accuracy alongside thinking, greedy, 2,560-token cap. The Kev-9B and Jev columns are the numbers Kev publishes.

bench Kev-9B Jev Jeeves
Test overall (out-of-domain and held-out, item-weighted) 0.822 0.857 0.889
Transfer overall (MMLU-Pro and buried state) 0.579 0.800 0.746
JevBench overall (231 community items) 0.715* 0.866 0.935
QNLI 0.925 0.925 0.913
SciQ 0.963 0.988 0.991
TweetEval offensive 0.775 0.813 0.813
PAWS 0.763 0.788 0.875
MMLU 0.738 0.900 0.793
Emotion 0.600 0.588 0.647
Held-out regulation structures 0.896 0.885 1.000
Contrastive policies 0.900 0.963 1.000
MMLU-Pro (10-way) 0.515 0.840 0.739
Buried state 0.740 0.700 0.759
Unknowable answered at p ≥ 0.9 (lower is better) 0.000 0.090 0.055
JevBench difficult (111 community items) 0.451* 0.730 0.865
JevBench ECE (public items) 0.049 0.037

* No Kev-9B JevBench outcome is published. These are Kev-8B (Qwen3).

All JevBench numbers are on the community easy, norm and difficult tiers (231 items). The sealed fairness tier is not included, and the Jev and Kev numbers are restricted to the identical community items.

Without thinking the identical checkpoint scores 0.804 on our test divided (2,962 items), against 0.840 alongside it.

Requirements: Python 3.12 and a CUDA GPU.

pip instal -r requirements.txt

Download the released weights and assist them:

hf download PostHog/jeeves --local-dir jeeves-weights python -m inference.serve --model jeeves-weights --drafter jeeves-weights/drafter_k4.safetensors --port 8009

Or fuse your own trained checkpoint into a standalone example and assist it alongside a drafter:

python export.py runs/cispo/final --out runs/fused python -m inference.serve --model runs/fused --drafter runs/drafter_k4/drafter.safetensors --port 8009

Then dispatch a petition in Jev's format:

curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{  "state": "Shoes arrived two weeks delayed and in the incorrect size. Also I see two charges on my card.",  "questions": {  "department": {"type": "choice", "instructions": "Which squad should grip this?",  "criteria": {"returns": "Exchanges, refunds, incorrect or damaged items",  "shipping": "Delivery status, delays, misplaced packages",  "billing": "Charges, invoices, fee problems"}},  "escalate": {"type": "noul", "instructions": "Does this need urgent individual attention?"},  "frustration": {"type": "score", "instructions": "How disappointed is the customer?",  "criteria": ["Calm", "Frustrated", "Very angry"]}  },  "options": {"max_think": 512}}'

Response on one H100 (FP8), alongside the three questions thinking in parallel:

{ "model": "jeeves-latest", "answers": { "department": { "type": "choice", "choice": "billing", "confidence": 0.19, "probabilities": { "returns": 0.4, "shipping": 0.14, "billing": 0.46 } }, "escalate": { "type": "noul", "noul": 0.72 }, "frustration": { "type": "score", "score": 1.5, "legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" }, "probabilities": { "0": 0.04, "1": 0.43, "2": 0.54 }, "confidence": 0.75 } }, "usage": { "input_tokens": 129, "output_tokens": 160, "reasoning_tokens": 1536 }, "latency_ms": 8141.6 }

sdk/ is a drop-in replacement for Jev's Python SDK (typesafe-sdk):

from jeeves_sdk import Choice, Noul, Score, TypeSafeClient with TypeSafeClient() as client: result = client.system_one( state="I was charged twice. Please help.", questions={ "billing": Noul(instructions="Is this concerning billing?"), "tone": Choice(instructions="What is the tone?", criteria={"calm": None, "angry": None}), "urgency": Score(instructions="How urgent is this?", criteria=["can wait", "this week", "today"]), }, max_think=768, return_reasoning=True, ) print(result.nouls["billing"].noul, result.choices["tone"].choice, result.scores["urgency"].score) print(result.reasoning["tone"].text)

The client connects to http://127.0.0.1:8009 by default (or JEEVES_BASE_URL), needs no API key, and waits up to 120s.

options is optional and ignored by Jev clients that don't dispatch it. Server-wide defaults are set alongside the matching assist flags.

option default effect
think true false answers from the immediate solitary (about 0.3 s)
max_think 2560 truncates all reasoning sequence at this many tokens, afterward answers
nothink_threshold null answers without thinking whenever the no-think confidence is at smallest this value
return_reasoning false adds all question's reasoning content to the response

On 325 dev questions:

setting accuracy mean reasoning tokens median / p90 latency
full thinking 0.825 1,138 3.3 s / 17.1 s
max_think 768, nothink_threshold 0.9 0.806 344 2.0 s / 5.6 s
no thinking 0.775 0 about 0.3 s

Questions, states and answers are loaded into the Qwen conversation template like

<state> …state… <q> instructions <opt> choice 1 </opt> <opt> choice 2 </opt> … <think> 

The example afterward rolls out its reasoning chain, and following the </think> token we append

</think> <q> instructions <opt> choice 1 </opt> <opt> choice 2 </opt> … <decide> 

A pointer caput scores all choice alongside a scaled dot merchandise between a query projection of the hidden province at <decide> and a key projection of the hidden province at that option's </opt>, where

<state>, <q>, <opt>, </opt>, <decide> = "<|fim_prefix|>", "<|fim_middle|>", "<|box_start|>", "<|box_end|>", "<|fim_suffix|>" 

These are rare, mostly unused tokens in the Qwen tokenizer. Ablations established that using plain content akin "State" in the immediate alternatively worsened performance. Likewise, not repeating the questions following the reasoning obstacle additionally decreases performance. The final probabilities are a softmax complete the choice scores, divided by a heat fitted on the dev set.

  1. SFT (2 epochs, 596 steps on 8 GPUs). LoRA r=16 on all projections of Qwen3.5-9B affirmative the pointer head, trained on 19,126 questions from 12 community datasets and synthetic guideline data. Half the questions transport a reasoning sequence sampled from the basis model.
  2. CISPO (a 624-step agenda stopped at stage 402). 9,992 RL questions, 8 rollouts all at heat 1, capped at 2,560 thinking tokens.
  3. Calibration. A sole heat fitted on dev, stored alongside the checkpoint.

Stopping at stage 402 keeps the finest calibration and dev score. Past it, the caput over-sharpens on the soaked RL pool.

A diffusion perspective of the icy example (drafter/), inspired by Orthrus.

Unlike Orthrus, which supports attention-only models, it supports Qwen3.5's Gated DeltaNet layers by letting disguise tokens cross-attend to those layers' post-convolution keys and values.

chain tokens per second
plain graphed greedy decoding, one question 109
block 4, one question 176 (1.6×)
block 8, one question 193 (1.76×)
block 4, eight questions batched about 960 in total

Block 4 is the default since it stays cheap whenever multiple questions are batched.

You can build the datasets locally using the prep scripts. This downloads the community datasets from Hugging Face at the revisions pinned in prep/public.py:

Each community dataset stays under its own license.

On 8 GPUs, alongside the data in data/, attack run.sh runs the entire pipeline:

torchrun --nproc_per_node 8 train.py sft --run-dir runs/sft torchrun --nproc_per_node 8 train.py cispo --run-dir runs/cispo --init runs/sft/final torchrun --nproc_per_node 8 test.py runs/cispo/final torchrun --nproc_per_node 8 jevbench.py runs/cispo/final python export.py runs/cispo/final --out runs/fused torchrun --nproc_per_node 8 -m drafter.gen --model runs/fused torchrun --nproc_per_node 8 train.py drafter --model runs/fused --block 4 --run-dir runs/drafter_k4
path contents
model/ Qwen3.5 (Gated DeltaNet + gated attention), LoRA, pointer head
loader/ prompt format, tokenisation and batching
prep/ dataset building (prep.py) and synthetic generators
trainer.py, train.py SFT, CISPO and drafter training
test.py, jevbench.py, calibrate.py evaluation, JevBench, heat fitting
export.py fuses LoRA into a standalone example alongside the caput and temperature
drafter/ drafter model, sequence sampling, fused speculative decoder
inference/ FP8 kernel, batched speculative engine, Jev-compatible server and benchmark
sdk/ jeeves_sdk, a drop-in replacement for Jev's Python SDK alongside the reasoning options
  • Knowledge questions trail Jev (MMLU 0.793 vs 0.900, MMLU-Pro 0.739 vs 0.840).
  • Thinking is dilatory at the tail: 17 s at p90 alongside complete chains. Use max_think and nothink_threshold whenever latency matters.
  • The Kev and Jev comparisons exterior JevBench use distinct items from the identical sources.
  • No tongue consistency reward was included so thinking chains are not fine interpretable.

If you use Jeeves, its training formula or its drafter, delight cite:

@software{waltz2026jeeves, author = {Waltz, Nicholas P.}, title = {Jeeves: Reasoning Improves Jev-like Decisions}, year = {2026}, url = {https://github.com/PostHog/jeeves}, note = {Qwen3.5-9B decision example trained alongside SFT and CISPO, alongside a block-4 diffusion drafter} }
Other Article Hacker News
↑
Close Right Ads
Close Left Ads