Home / Chapter 8 · Choosing a model
Last edited · 6 min read
When not to use an LLM: classifiers and System One
When a system needs a decision rather than text (a label, a score, yes or no), an LLM still writes it token by token, and its probabilities are often missing or poorly calibrated. A non-generative model is often faster, cheaper and better calibrated; an LLM wins when the answer space is open or the task needs reasoning. TypeSafe’s System One serves here as the worked example of a new model type.
In plain wordsThe difference between writing half a page of justification and ticking boxes on a form. A form needs no writer, but it only has the boxes someone thought of in advance.
Run both lanes: the LLM writes a three-field JSON token by token, Jev (TypeSafe’s System One model) gets the same three questions in a single call
State: the ticket “You charged me twice, I want a refund!”. Questions: which team, is it a refund request, how frustrated is the customer.
LLM
writes JSON, one model pass per token
Passes: 0
Jev
state and three questions in one call
Calls: 0
An illustrative animation, not a benchmark. Each LLM step is one decode pass after the prompt has been processed; a model that reasons before the JSON writes hundreds of tokens more. The Jev side shows one API call; the vendor does not describe how much work the model does inside it.
The options for a decision
- A fine-tuned encoder (a BERT-type model) or embeddings with logistic regression: one forward pass, milliseconds, a fraction of a cent, running next to the service, with probabilities you can calibrate on a validation set. The price is labelled data for every task and retraining for every new class.
- Zero-shot, with no training data: an NLI model scores each label, written in natural language, as an entailment of the text (Yin et al., 2019), and a cross-encoder scores text–label pairs. A cheap baseline to measure everything else against.
- An LLM as a classifier: a prompt with the list of labels and a single-token answer read from logprobs. That is one decode step after prefill, so it is fast too, but each question is a separate request or a longer output. Logprobs are often unavailable (never in the Claude API; in OpenAI’s current models only with reasoning effort set to none, as of September 2026), and post-training degrades calibration: in the GPT-4 report the base model was well calibrated and the post-trained one noticeably worse. Measure calibration and correct it, e.g. with temperature scaling.
- An LLM wins when you need text, code, an explanation, multi-step reasoning, arithmetic or tool calls, when the answers cannot be closed into a list up front, and at low volume, where a prompt is cheaper than a labelled dataset. Strict structured output gives an LLM typed answers too (see “Enforcing output format”), so the real differences are cost, latency and the quality of the probabilities. Which LLM to pick is covered in “Choosing a model”.
Worked example: System One (vendor claims)
- According to TypeSafe’s documentation (September 2026, version jev-1.13), its model Jev takes a text “state” (a string, a JSON object or an array of texts) and typed questions: Choice (pick from a list), Score (a level on a described scale) and Noul (the probability of “yes”). Choice and Score return a distribution and a
confidencefield. It generates no text. The vendor positions it as needing no task-specific training, unlike a fine-tuned encoder. - All questions in a request are evaluated in parallel and independently against the same state. The vendor says adding questions barely increases response time but publishes no latency figures. Pricing: $0.042 per million input tokens, output free, up to 64k tokens per request, of which the state plus the longest question may take at most 32k. English is the primary language.
- The probabilities were trained for calibration with a method the vendor calls RLCD (reinforcement learning for calibrated decisions), and the vendor cautions that calibration holds across groups of decisions, not for a single answer. The documentation describes neither the architecture nor any independent comparisons, so this whole section is claims to test on your own data.
- Limitations the vendor lists for jev-1.13: it does not count reliably, compares dates and numbers poorly, reads questions literally, gets lost with indirect questions and in a large state full of irrelevant data, and the content of the state can steer it the way prompt injection steers an LLM. Answers to a question and to its negation need not sum to 1.
Move the threshold: decisions above it run automatically, those below go to an LLM or a human. Watch how many mistakes the automation lets through
The dots are the probability of the chosen answer for 12 illustrative tickets. With good calibration, decisions at probability 0.9 are wrong 10% of the time on average. That does not tell you whether this particular decision is right.
Fitting it into a system
- The pattern: a fast decision model settles the easy cases, and on low confidence the case goes to a reasoning LLM or a human (see “LLMs in production”). Automation pays off when (1 − p) × cost of a mistake < cost of escalation. With a mistake costing $100 and an escalation $5, the threshold is p > 0.95.
- Check calibration yourself, whatever the model: group decisions by their stated probability and compare it with the actual accuracy in each group (a reliability diagram, ECE). Do it separately for each question type and language, because calibration in English does not guarantee it in any other language.
- Arithmetic, dates and hard rules stay in code. Once thresholds are tuned, pin the model version (for Jev
jev-1.13.0, not thejev-latestalias): an alias moves with each release, and the probability distribution moves with it.
Check yourself
When does a classifier or a decision model such as System One beat an LLM, and when doesn’t it?
A non-generative model wins when the answer is a label, a score or yes/no from a closed list and volume, latency or calibrated probabilities matter. A fine-tuned encoder answers in milliseconds but needs labelled data for each task; zero-shot NLI encoders need none. An LLM needs no data but writes the decision token by token, and its logprobs are often unavailable or poorly calibrated after post-training. Decision models such as TypeSafe’s Jev claim typed answers with calibrated probabilities in one call, which has to be measured on your own data. An LLM wins for text, reasoning, tools and open answer spaces. The pattern: the fast model settles confident cases and escalates the rest to an LLM or a human.
Po polsku
Model niegeneratywny wygrywa, gdy odpowiedzią jest etykieta, ocena albo tak/nie z zamkniętej listy, a liczą się wolumen, opóźnienie albo skalibrowane prawdopodobieństwa. Dostrojony enkoder odpowiada w milisekundach, ale wymaga oznaczonych danych dla każdego zadania; enkodery zero-shot oparte na NLI ich nie potrzebują. LLM obywa się bez danych, ale pisze decyzję token po tokenie, a jego logprobs często są niedostępne albo po post-trainingu źle skalibrowane. Modele decyzyjne, takie jak Jev od TypeSafe, obiecują typowane odpowiedzi ze skalibrowanymi prawdopodobieństwami w jednym wywołaniu, co trzeba zmierzyć na własnych danych. LLM wygrywa przy tekście, rozumowaniu, narzędziach i otwartej przestrzeni odpowiedzi. Wzorzec: szybki model rozstrzyga pewne przypadki, a resztę eskaluje do LLM-a albo człowieka.
Follow-up questions (4)
- When is an LLM still the better choice?
- When you need text, code, an explanation of the decision, multi-step reasoning, arithmetic or tool calls. Also when the space of answers cannot be closed into a list of options up front, or the volume is too low to justify collecting labelled data.
- How do you check calibration?
- On your own labelled data: split the decisions into bins by stated probability and in each bin compare it with the actual accuracy. Report the weighted average gap as ECE. Do it separately for each question type and language.
- How does a decision model differ from an LLM classifier with logprobs?
- A classifier that emits a single label token is one decode step after prefill, so it is fast too. But each question is a separate request or a longer output, logprobs are often unavailable (never in the Claude API; in OpenAI’s current models only with reasoning off), and post-training does not optimise them for calibration and usually spoils what pretraining gave. Measuring accuracy, calibration, latency and cost on the same dataset settles it.
- How would you set the escalation threshold?
- From costs: automate when (1 − p) × cost of a mistake is lower than the cost of escalation. Then check on data that the actual error rate at that threshold matches the expected one, and repeat that after every model version change.