Most business decisions do not need an AI system to write an essay. They need it to choose one valid action from a controlled list: route a ticket, approve a workflow step, assign a document type, or escalate a case. That distinction is the foundation of a reliable AI decision system.
In this tutorial, you will build your own decision model with Python, Hugging Face Transformers, and Qwen3 1.7B. Instead of allowing unrestricted generation, the model will score predefined answers using language model next-token probabilities. You will then evaluate its accuracy on CommonsenseQA, fine-tune the small language model, and calibrate its confidence scores.
This pattern reflects a broader shift toward compact open-weight models, constrained language model outputs, adapter-based fine-tuning, and measurable AI model reliability. It is useful, but it does not turn an LLM into an infallible decision-maker. High-impact decisions still require domain validation, safeguards, monitoring, and human oversight.
Decision Model vs Language Model: What Are You Building?
A conventional decision model compares options against explicit criteria. A loan policy might evaluate income, debt, risk limits, and documentation before selecting approve, review, or decline. Its inputs, rules, and outputs are intentionally bounded.
A language model works differently. It estimates the probability of the next token given the tokens already present. It can produce an answer that resembles a decision, but its underlying operation is sequential text prediction. An LLM does not inherently know that only four workflow actions are valid unless you enforce that constraint.
An LLM decision model combines these ideas. The language model interprets unstructured input, while the surrounding decision framework:
- Defines the permitted choices.
- Scores each candidate under the same prompt.
- Selects the highest-scoring valid answer.
- Measures accuracy on labeled examples.
- Calibrates confidence using separate validation data.
- Applies thresholds, abstention rules, and human review.
This is usually more dependable than prompting a model to generate an answer and attempting to parse whatever text it returns.
Why Use Qwen3 1.7B for a Decision Model?
Qwen3 1.7B is a compact model that is practical for experimentation, local inference, and fine-tuning with limited hardware. Its smaller footprint makes repeated evaluation less expensive than testing a much larger model. The model and usage notes are available on the Qwen3 1.7B Hugging Face page.
A small model is not automatically accurate, fast, or calibrated. Performance depends on hardware, precision, prompt length, batching, quantization, and the task itself. The advantage is control: developers can inspect the complete scoring pipeline, adapt the model, and measure whether changes genuinely help.
Set Up Python and Hugging Face Transformers
Create an isolated environment and install the required packages. A CUDA-capable GPU is recommended, although CPU inference is possible at a slower rate.
python -m venv .venv
source .venv/bin/activate
pip install 'transformers>=4.51' torch datasets accelerate peftFor a reproducible experiment, save the resolved package versions with pip freeze, record the model revision, set random seeds, and retain the exact evaluation IDs. Loading an unpinned repository later may retrieve changed files.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
MODEL_ID = 'Qwen/Qwen3-1.7B'
torch.manual_seed(42)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
torch_dtype='auto',
device_map='auto'
)
model.eval()Score Fixed Answers with Next-Token Probabilities
Suppose a question has five choices labeled A through E. Rather than asking the model to generate arbitrary text, calculate the conditional log probability of every label and select the largest score.
The prompt should end at the point where the answer begins. Qwen3 supports disabling its reasoning mode through the chat template, which is helpful when the desired output is a single classification label.
def build_prompt(question, labels, texts):
options = 'n'.join(
f'{label}. {text}' for label, text in zip(labels, texts)
)
message = {
'role': 'user',
'content': (
'Choose the best answer. Return only its label.nn'
f'Question: {question}n{options}nAnswer:'
)
}
return tokenizer.apply_chat_template(
[message],
tokenize=False,
add_generation_prompt=True,
enable_thinking=False
)
@torch.inference_mode()
def candidate_log_probability(prompt, candidate):
prompt_ids = tokenizer(
prompt, add_special_tokens=False
).input_ids
candidate_ids = tokenizer(
candidate, add_special_tokens=False
).input_ids
ids = torch.tensor(
[prompt_ids + candidate_ids], device=model.device
)
logits = model(ids).logits[0]
log_probs = torch.log_softmax(logits, dim=-1)
score = 0.0
start = len(prompt_ids)
for offset, token_id in enumerate(candidate_ids):
score += log_probs[start + offset - 1, token_id].item()
return score
def predict_choice(question, labels, texts):
prompt = build_prompt(question, labels, texts)
scores = torch.tensor([
candidate_log_probability(prompt, label)
for label in labels
])
probabilities = torch.softmax(scores, dim=0)
best = int(torch.argmax(probabilities))
return labels[best], probabilities.tolist(), scores.tolist()The softmax values are normalized only across the offered choices. They are not universal probabilities that the answer is true. If the correct answer is missing, the system will still assign all probability mass among the wrong candidates.
Tokenization Can Change the Comparison
Before trusting LLM probability scores, inspect how every candidate tokenizes:
for label in ['A', 'B', 'C', 'D', 'E']:
print(label, tokenizer.encode(label, add_special_tokens=False))Single-token labels make comparison straightforward. Multi-token candidates require summing token log probabilities. Raw sums tend to favor shorter strings, while length-normalized scores can favor longer ones, so neither method is neutral. Labels also may tokenize differently after a space or newline. Use a fixed prompt delimiter, inspect token IDs, and test formatting changes. For classification, scoring short labels is generally safer than scoring full answer sentences.
This candidate-scoring approach is a form of constrained decoding: only approved outputs can win. Production systems can further enforce valid tokens with a logits processor, grammar, or finite-state constraint.
Evaluate LLM Accuracy with CommonsenseQA
CommonsenseQA provides multiple-choice questions with labeled answers and is suitable for demonstrating a complete evaluation pipeline. Review its fields and documentation on the CommonsenseQA dataset page.
from datasets import load_dataset
validation = load_dataset(
'tau/commonsense_qa', split='validation'
)
def evaluate(dataset):
correct = 0
rows = []
for item in dataset:
labels = item['choices']['label']
texts = item['choices']['text']
prediction, probabilities, scores = predict_choice(
item['question'], labels, texts
)
correct += int(prediction == item['answerKey'])
rows.append({
'prediction': prediction,
'target': item['answerKey'],
'confidence': max(probabilities),
'scores': scores
})
return correct / len(dataset), rows
baseline_accuracy, baseline_rows = evaluate(validation)
print({'baseline_accuracy': baseline_accuracy})Run the baseline before changing prompts or weights. Save per-example predictions, not just aggregate accuracy. Error records reveal label bias, malformed prompts, truncation, and categories the model consistently misunderstands.
Prompt format, model revision, precision, chat template, sample selection, and scoring rules can materially change the result. Claims sometimes cited for this experiment—59.38% baseline accuracy and 62.41% after fine-tuning—should not be treated as verified benchmarks without the original checkpoint, data split, code, and evaluation logs. Generate and report results from your own pinned setup instead.
Fine-Tuning a Small Language Model
Fine-tuning teaches the model to map the prompt format to the correct label. Parameter-efficient LoRA training is often preferable to updating every weight. Train only on the CommonsenseQA training split, preserve a calibration split, and evaluate once on untouched examples.
from peft import LoraConfig, get_peft_model
from transformers import Trainer, TrainingArguments
lora = LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
target_modules=['q_proj', 'k_proj', 'v_proj', 'o_proj'],
task_type='CAUSAL_LM'
)
model = get_peft_model(model, lora)
args = TrainingArguments(
output_dir='qwen3-decision-model',
learning_rate=2e-4,
num_train_epochs=1,
per_device_train_batch_size=4,
gradient_accumulation_steps=8,
bf16=True,
logging_steps=25,
save_strategy='epoch',
report_to='none'
)
trainer = Trainer(
model=model,
args=args,
train_dataset=tokenized_training_data,
data_collator=decision_data_collator
)
trainer.train()The preprocessing step represented by tokenized_training_data should concatenate each prompt with its correct label. Set prompt positions in the training labels to -100 so loss is calculated only on the answer tokens. The collator should pad input IDs and use -100 for padded label positions.
After training, run the identical evaluator against the same untouched examples. Compare baseline and fine-tuned accuracy, per-class recall, latency, and failure cases. A gain is meaningful only when the evaluation conditions remain constant. Repeated seeds or bootstrap confidence intervals help determine whether a small difference is stable or sampling noise.
Why Model Confidence Does Not Equal Accuracy
A model can assign 95% of the candidate probability mass to an incorrect answer. This is the central distinction in LLM confidence vs accuracy. Accuracy asks how often predictions are correct; calibration asks whether predictions made with confidence near 0.8 are correct about 80% of the time.
Useful calibration metrics include expected calibration error, which compares confidence and observed accuracy across bins, and the multiclass Brier score, which measures squared error between predicted probabilities and one-hot targets. Reliability diagrams make overconfidence and underconfidence visible.
Temperature scaling is a simple post-training calibration method. It learns one positive scalar on a dedicated calibration set:
# candidate_scores: shape [examples, choices]
# targets: index of the correct choice
log_temperature = torch.nn.Parameter(torch.zeros(()))
optimizer = torch.optim.LBFGS([log_temperature], max_iter=50)
def closure():
optimizer.zero_grad()
temperature = log_temperature.exp()
loss = torch.nn.functional.cross_entropy(
candidate_scores / temperature, targets
)
loss.backward()
return loss
optimizer.step(closure)
temperature = log_temperature.exp().detach()
calibrated = torch.softmax(candidate_scores / temperature, dim=1)Use one data partition to fit the temperature and another to report calibration. Temperature scaling changes confidence values but normally does not change the highest-scoring class, so it can improve calibration without improving accuracy. It also cannot repair missing choices, distribution shifts, biased labels, or fundamentally wrong reasoning.
Turning Scores Into a Reliable AI Decision System
A production AI decision-making framework needs more than an argmax operation. Define a minimum calibrated confidence, abstain when the top choices are too close, validate inputs, log model and prompt versions, and monitor accuracy by class and user segment. Route uncertain or sensitive cases to a person.
Practical applications include:
- Routing support requests to billing, technical, account, or security teams.
- Classifying contracts, invoices, resumes, and compliance documents.
- Selecting the next permitted action in an automated workflow.
- Allowing structured AI agents to choose from approved tools.
- Prioritizing records for review without automatically making final decisions.
For faster LLM inference, batch candidates or compute shared prompt states once and reuse the key-value cache. Quantization can reduce memory requirements, but evaluate accuracy and calibration again after any optimization. Even mathematically equivalent-looking changes can alter numerical scores.
Frequently Asked Questions
Is a decision model the same as an AI classification model?
They overlap. A classifier maps an input to predefined classes, while a decision model may also include criteria, thresholds, abstention, policies, and downstream actions. Using an LLM as the scoring component does not eliminate the need for that surrounding logic.
Why not ask the LLM to return JSON?
Structured JSON reduces parsing problems, but the model may still generate an unsupported label or extra content. Candidate scoring guarantees that the selected result belongs to the allowed set. JSON remains useful for transporting the final validated decision.
Does the highest probability identify the correct answer?
No. It identifies the candidate preferred by the model under a specific prompt and scoring method. Calibration estimates how confidence relates to observed correctness, but it does not guarantee that an individual prediction is right.
Should full answer text or letter labels be scored?
Short labels usually make comparisons cleaner because they minimize token-length effects. Full text can be appropriate when labels have undesirable tokenization, but the scoring policy must account for different token counts and be validated empirically.
When is human oversight required?
Use meaningful human review whenever errors could affect safety, rights, employment, healthcare, credit, legal status, security, or substantial financial outcomes. A calibrated confidence score is decision support—not proof, accountability, or authorization.
Build the Model, Then Test the Decision Process
To build a decision model with Python, start by narrowing the output space. Score every permitted answer consistently, evaluate on untouched labeled data, fine-tune only when the baseline justifies it, and calibrate confidence on a separate split. The result is far more testable than unrestricted generation.
The most important work happens around the LLM: defining valid choices, controlling tokenization, preserving reproducible evaluation, detecting uncertainty, and deciding when the system must abstain. That is what turns next-token probabilities into a practical decision tool rather than an unverified guess.