Build your own decision model

Hacker News by 5 min read 18x views
Build your own decision model

Share Post

"System one" decision models are models that infer and react alongside calibrated probabilities or all allowed answer.

Consider your mundane tongue model, to get typed output from it (JSON), you may use Structured Output to constrain the output to guaranteed valid JSON. While example prefills the input in one pass, it motionless has to go execute a continue for all token in command to create a valid response.

In this example, 11 passes are required to create the final output. (We're not accounting for speculative decoding and another conclusion optimization techniques.)

prompt (prefill) generated (predicted) being predicted

Decision models specified as Jev, create the assumption that there are fixed options we can choose from and we can do so quickly by making a sole pass. In this example we constrain the set of imaginable outputs to the options A, B, C, D, E. By masking another items in the vocabulary, the example can lone emit those tokens. By selecting the highest probability output, we get our answer.

prompt (prefill) generated (predicted) being predicted

Since the outputs are constrained to lone a fixed set of options, the example can't choose item exterior of those. This nevertheless does not justify that the output volition be correct. It's additionally average to treat the output token probabilities a confidence scores in this context, but without additional training, those scores apt indicate its confidence in what the next token volition be fairly than the true probability of the reply being the accurate answer.

We can emulate this behavior by constraining output tokens using an LLM. Here I'm using Qwen/Qwen3-1.7B

import argparse import json import torch from transformers import AutoModelForCausalLM, AutoTokenizer model_name = "Qwen/Qwen3-1.7B" options = ["A", "B", "C", "D", "E"] parser = argparse.ArgumentParser() parser.add_argument("--input", default="question.json") args = parser.parse_args() # burden the tokenizer and the model tokenizer = AutoTokenizer.from_pretrained(model_name) model = AutoModelForCausalLM.from_pretrained( model_name, torch_dtype="auto", device_map="auto" ) # the token the example would emit for all choice as the archetypal aide token option_token_ids = [tokenizer.encode(opt, add_special_tokens=False)[0] for opt in options] def format_prompt(item): immediate = item["question"] + "\n" for opt in options: immediate += f"{opt}. {item[opt]}\n" immediate += "Answer:" messages = [ {"role": "user", "content": prompt} ] come back tokenizer.apply_chat_template( messages, tokenize=False, add_generation_prompt=True, enable_thinking=False ) with open(args.input) as f: item = json.load(f) model_inputs = tokenizer(format_prompt(item), return_tensors="pt").to(model.device) with torch.no_grad(): logits = model(**model_inputs).logits[0, -1] # constrained decoding: lone the choice tokens are allowed probs = torch.softmax(logits[option_token_ids].float(), dim=-1) print(f"prediction: {options[probs.argmax().item()]}") for opt, prob in zip(options, probs.tolist()): print(f"{opt}: {prob:.4f} {item[opt]}")

Running it alongside an uncomplicated inquiry to test it yeilds the following output

// input { "question": "What color is the sky?", "A": "Red", "B": "Blue", "C": "Green", "D": "Purple", "E": "I don't know" } // output prediction: B A: 0.0000 Red B: 0.9988 Blue C: 0.0000 Green D: 0.0000 Purple E: 0.0012 I don't know

The example was capable to create awareness of our input and create a prediction that fairly corresponds to the accurate answer.

We can test the accuracy of the example by operating it against community datasets. I ran this against a random example holdout of CommonsenseQA

 precision recall f1 support A 0.5733 0.7197 0.6382 239 B 0.5506 0.7686 0.6416 255 C 0.5372 0.6598 0.5922 241 D 0.7206 0.3904 0.5065 251 E 0.7519 0.4255 0.5435 235 accuracy: 725/1221 = 0.5938 macro f1: 0.5844

Not bad for a 1.7B model, Running a quick finetune on the dataset gives us slightly improved performance

 precision recall f1 support A 0.6475 0.6611 0.6542 239 B 0.6113 0.6784 0.6431 255 C 0.6234 0.5975 0.6102 241 D 0.6700 0.5418 0.5991 251 E 0.5808 0.6426 0.6101 235 accuracy: 762/1221 = 0.6241 macro f1: 0.6234

Testing the example against a extremely ambiguous issue demonstrates an engaging problem.

// input { "question": "Where would you most apt discover a bat?", "A": "Cave", "B": "Baseball game", "C": "Attic", "D": "Zoo", "E": "Sporting products store" } // output prediction: A A: 0.9978 Cave B: 0.0004 Baseball game C: 0.0017 Attic D: 0.0000 Zoo E: 0.0001 Sporting products store

There have to be no apparent answer here, but treating the output probabilities as a pseudo "confidence" score, shows that the example is extremely overconfident in this answer.

If we bin the confidence mark ranges in the eval I ran earlier, we can see that the model's confidence does not equivalent its accuracy. This method that the example is not calibrated.

 bin figure confidence accuracy (0.00, 0.10] 0 0.0000 0.0000 (0.10, 0.20] 0 0.0000 0.0000 (0.20, 0.30] 3 0.2834 0.0000 (0.30, 0.40] 26 0.3761 0.2692 (0.40, 0.50] 41 0.4538 0.2683 (0.50, 0.60] 70 0.5490 0.3286 (0.60, 0.70] 74 0.6476 0.3649 (0.70, 0.80] 77 0.7495 0.4286 (0.80, 0.90] 121 0.8555 0.4711 (0.90, 1.00] 809 0.9855 0.7009

We can notice that the example tends to be extremely overconfident in the 0.9 - 1.0 bin but it's lone accurate 70% of the time. When it makes a prediction alongside 0.8 - 0.9 confidence it's lone exact ~40% of the time. This method that the example is mostly overconfident in its predictions.

Since our goal is to have the example output scores that is reflective of its accuracy, one method we can use to callibrate it is through heat scaling. By modifying the heat value, we can flatten its output probability allocation curve and measure it to approximate its accuracy.

Curve fitting fit the heat indicator to the model's accuracy, I established 3.797280788421631 as a temp value.

 bin figure confidence accuracy (0.00, 0.10] 0 0.0000 0.0000 (0.10, 0.20] 0 0.0000 0.0000 (0.20, 0.30] 82 0.2712 0.2317 (0.30, 0.40] 217 0.3507 0.3917 (0.40, 0.50] 199 0.4472 0.5126 (0.50, 0.60] 166 0.5475 0.5482 (0.60, 0.70] 139 0.6562 0.5827 (0.70, 0.80] 140 0.7492 0.7714 (0.80, 0.90] 169 0.8507 0.7988 (0.90, 1.00] 109 0.9333 0.9541

This gets us a much improved calibration. If you desire to perform about alongside this, I made a GitHub repo alongside scripts that stroll you through construction a dataset, evaluating, finetuning and calibrating your own model. I advance pulling it and trying it on another bigger models.

Other Article Hacker News
↑
Close Right Ads
Close Left Ads