A reasoning Jev-style classifier alongside a diffusion drafter, trained alongside SFT and CISPO.
Inspired by Kev.
- A 9B Jev-like example (Qwen3.5-9B, LoRA, pointer head) that thinks before it decides, alongside a block-4 diffusion drafter and the complete training code and train/dev/test data.
- Beats Kev-9B and Jev on test data it was never trained on (0.889 vs 0.822 and 0.857) and on JevBench's community tiers (0.935 vs 0.866 for Jev).
- Supports yes/no (noul), multiple-choice (choice), and ranking (score) questions in the identical request, through a Jev-compatible API.
- About 0.3 s per petition without thinking and a 3.3 s median alongside it on one H100. Can be sped up by truncating sequence length.
- Runs on CUDA (Hopper for the FP8 kernel).
Jev-like models provision calibrated decision probabilities, but at low accuracy. A lot of pipelines hence depend on a reasoning example as a fallback. Jeeves trains a Jev-like Qwen3.5-9B (LoRA and a pointer head) using CISPO to logic before it decides.
This results in improved achievement on out of domain tasks, and outperforms Jev in JevBench difficult (public).
Accuracy alongside thinking, greedy, 2,560-token cap. The Kev-9B and Jev columns are the numbers Kev publishes.
| bench | Kev-9B | Jev | Jeeves |
|---|---|---|---|
| Test overall (out-of-domain and held-out, item-weighted) | 0.822 | 0.857 | 0.889 |
| Transfer overall (MMLU-Pro and buried state) | 0.579 | 0.800 | 0.746 |
| JevBench overall (231 community items) | 0.715* | 0.866 | 0.935 |
| QNLI | 0.925 | 0.925 | 0.913 |
| SciQ | 0.963 | 0.988 | 0.991 |
| TweetEval offensive | 0.775 | 0.813 | 0.813 |
| PAWS | 0.763 | 0.788 | 0.875 |
| MMLU | 0.738 | 0.900 | 0.793 |
| Emotion | 0.600 | 0.588 | 0.647 |
| Held-out regulation structures | 0.896 | 0.885 | 1.000 |
| Contrastive policies | 0.900 | 0.963 | 1.000 |
| MMLU-Pro (10-way) | 0.515 | 0.840 | 0.739 |
| Buried state | 0.740 | 0.700 | 0.759 |
| Unknowable answered at p ≥ 0.9 (lower is better) | 0.000 | 0.090 | 0.055 |
| JevBench difficult (111 community items) | 0.451* | 0.730 | 0.865 |
| JevBench ECE (public items) | 0.049 | 0.037 |
* No Kev-9B JevBench outcome is published. These are Kev-8B (Qwen3).
All JevBench numbers are on the community easy, norm and difficult tiers (231 items). The sealed fairness tier is not included, and the Jev and Kev numbers are restricted to the identical community items.
Without thinking the identical checkpoint scores 0.804 on our test divided (2,962 items), against 0.840 alongside it.
Requirements: Python 3.12 and a CUDA GPU.
pip instal -r requirements.txt
Download the released weights and assist them:
hf download PostHog/jeeves --local-dir jeeves-weights python -m inference.serve --model jeeves-weights --drafter jeeves-weights/drafter_k4.safetensors --port 8009
Or fuse your own trained checkpoint into a standalone example and assist it alongside a drafter:
python export.py runs/cispo/final --out runs/fused python -m inference.serve --model runs/fused --drafter runs/drafter_k4/drafter.safetensors --port 8009
Then dispatch a petition in Jev's format:
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{ "state": "Shoes arrived two weeks delayed and in the incorrect size. Also I see two charges on my card.", "questions": { "department": {"type": "choice", "instructions": "Which squad should grip this?", "criteria": {"returns": "Exchanges, refunds, incorrect or damaged items", "shipping": "Delivery status, delays, misplaced packages", "billing": "Charges, invoices, fee problems"}}, "escalate": {"type": "noul", "instructions": "Does this need urgent individual attention?"}, "frustration": {"type": "score", "instructions": "How disappointed is the customer?", "criteria": ["Calm", "Frustrated", "Very angry"]} }, "options": {"max_think": 512}}'
Response on one H100 (FP8), alongside the three questions thinking in parallel:
{ "model": "jeeves-latest", "answers": { "department": { "type": "choice", "choice": "billing", "confidence": 0.19, "probabilities": { "returns": 0.4, "shipping": 0.14, "billing": 0.46 } }, "escalate": { "type": "noul", "noul": 0.72 }, "frustration": { "type": "score", "score": 1.5, "legend": { "0": "Calm", "1": "Frustrated", "2": "Very angry" }, "probabilities": { "0": 0.04, "1": 0.43, "2": 0.54 }, "confidence": 0.75 } }, "usage": { "input_tokens": 129, "output_tokens": 160, "reasoning_tokens": 1536 }, "latency_ms": 8141.6 }sdk/ is a drop-in replacement for Jev's Python SDK (typesafe-sdk):
from jeeves_sdk import Choice, Noul, Score, TypeSafeClient with TypeSafeClient() as client: result = client.system_one( state="I was charged twice. Please help.", questions={ "billing": Noul(instructions="Is this concerning billing?"), "tone": Choice(instructions="What is the tone?", criteria={"calm": None, "angry": None}), "urgency": Score(instructions="How urgent is this?", criteria=["can wait", "this week", "today"]), }, max_think=768, return_reasoning=True, ) print(result.nouls["billing"].noul, result.choices["tone"].choice, result.scores["urgency"].score) print(result.reasoning["tone"].text)
The client connects to http://127.0.0.1:8009 by default (or JEEVES_BASE_URL), needs no API key, and waits up to 120s.
options is optional and ignored by Jev clients that don't dispatch it. Server-wide defaults are set alongside the matching assist flags.
| option | default | effect |
|---|---|---|
| think | true | false answers from the immediate solitary (about 0.3 s) |
| max_think | 2560 | truncates all reasoning sequence at this many tokens, afterward answers |
| nothink_threshold | null | answers without thinking whenever the no-think confidence is at smallest this value |
| return_reasoning | false | adds all question's reasoning content to the response |
On 325 dev questions:
| setting | accuracy | mean reasoning tokens | median / p90 latency |
|---|---|---|---|
| full thinking | 0.825 | 1,138 | 3.3 s / 17.1 s |
| max_think 768, nothink_threshold 0.9 | 0.806 | 344 | 2.0 s / 5.6 s |
| no thinking | 0.775 | 0 | about 0.3 s |
Questions, states and answers are loaded into the Qwen conversation template like
<state> …state… <q> instructions <opt> choice 1 </opt> <opt> choice 2 </opt> … <think> The example afterward rolls out its reasoning chain, and following the </think> token we append
</think> <q> instructions <opt> choice 1 </opt> <opt> choice 2 </opt> … <decide> A pointer caput scores all choice alongside a scaled dot merchandise between a query projection of the hidden province at <decide> and a key projection of the hidden province at that option's </opt>, where
<state>, <q>, <opt>, </opt>, <decide> = "<|fim_prefix|>", "<|fim_middle|>", "<|box_start|>", "<|box_end|>", "<|fim_suffix|>" These are rare, mostly unused tokens in the Qwen tokenizer. Ablations established that using plain content akin "State" in the immediate alternatively worsened performance. Likewise, not repeating the questions following the reasoning obstacle additionally decreases performance. The final probabilities are a softmax complete the choice scores, divided by a heat fitted on the dev set.
- SFT (2 epochs, 596 steps on 8 GPUs). LoRA r=16 on all projections of Qwen3.5-9B affirmative the pointer head, trained on 19,126 questions from 12 community datasets and synthetic guideline data. Half the questions transport a reasoning sequence sampled from the basis model.
- CISPO (a 624-step agenda stopped at stage 402). 9,992 RL questions, 8 rollouts all at heat 1, capped at 2,560 thinking tokens.
- Calibration. A sole heat fitted on dev, stored alongside the checkpoint.
Stopping at stage 402 keeps the finest calibration and dev score. Past it, the caput over-sharpens on the soaked RL pool.
A diffusion perspective of the icy example (drafter/), inspired by Orthrus.
Unlike Orthrus, which supports attention-only models, it supports Qwen3.5's Gated DeltaNet layers by letting disguise tokens cross-attend to those layers' post-convolution keys and values.
| chain tokens per second | |
|---|---|
| plain graphed greedy decoding, one question | 109 |
| block 4, one question | 176 (1.6×) |
| block 8, one question | 193 (1.76×) |
| block 4, eight questions batched | about 960 in total |
Block 4 is the default since it stays cheap whenever multiple questions are batched.
You can build the datasets locally using the prep scripts. This downloads the community datasets from Hugging Face at the revisions pinned in prep/public.py:
Each community dataset stays under its own license.
On 8 GPUs, alongside the data in data/, attack run.sh runs the entire pipeline:
torchrun --nproc_per_node 8 train.py sft --run-dir runs/sft torchrun --nproc_per_node 8 train.py cispo --run-dir runs/cispo --init runs/sft/final torchrun --nproc_per_node 8 test.py runs/cispo/final torchrun --nproc_per_node 8 jevbench.py runs/cispo/final python export.py runs/cispo/final --out runs/fused torchrun --nproc_per_node 8 -m drafter.gen --model runs/fused torchrun --nproc_per_node 8 train.py drafter --model runs/fused --block 4 --run-dir runs/drafter_k4
| path | contents |
|---|---|
| model/ | Qwen3.5 (Gated DeltaNet + gated attention), LoRA, pointer head |
| loader/ | prompt format, tokenisation and batching |
| prep/ | dataset building (prep.py) and synthetic generators |
| trainer.py, train.py | SFT, CISPO and drafter training |
| test.py, jevbench.py, calibrate.py | evaluation, JevBench, heat fitting |
| export.py | fuses LoRA into a standalone example alongside the caput and temperature |
| drafter/ | drafter model, sequence sampling, fused speculative decoder |
| inference/ | FP8 kernel, batched speculative engine, Jev-compatible server and benchmark |
| sdk/ | jeeves_sdk, a drop-in replacement for Jev's Python SDK alongside the reasoning options |
- Knowledge questions trail Jev (MMLU 0.793 vs 0.900, MMLU-Pro 0.739 vs 0.840).
- Thinking is dilatory at the tail: 17 s at p90 alongside complete chains. Use max_think and nothink_threshold whenever latency matters.
- The Kev and Jev comparisons exterior JevBench use distinct items from the identical sources.
- No tongue consistency reward was included so thinking chains are not fine interpretable.
If you use Jeeves, its training formula or its drafter, delight cite:
@software{waltz2026jeeves, author = {Waltz, Nicholas P.}, title = {Jeeves: Reasoning Improves Jev-like Decisions}, year = {2026}, url = {https://github.com/PostHog/jeeves}, note = {Qwen3.5-9B decision example trained alongside SFT and CISPO, alongside a block-4 diffusion drafter} }
- Jev's Architecture Unmasked, the Jev scheme that Kev and Jeeves follow.
- Qwen Team. Qwen3.5-9B, the basis model.
- Yang, Kautz, Hatamizadeh. Gated Delta Networks: Improving Mamba2 alongside Delta Rule. ICLR 2025.
- Hu et al. LoRA: Low-Rank Adaptation of Large Language Models. ICLR 2022.
- MiniMax. MiniMax-M1: Scaling Test-Time Compute Efficiently alongside Lightning Attention. 2025. Introduces CISPO.
- Guo, Pleiss, Sun, Weinberger. On Calibration of Modern Neural Networks. ICML 2017. Temperature scaling.
- Orthrus, arXiv 2605.12825. The diffusion drafter ours adapts to Gated DeltaNet.
- Leviathan, Kalman, Matias. Fast Inference from Transformers via Speculative Decoding. ICML 2023.
