Build your own decision model
This article explains how to build a 'system one' decision model using large language models by constraining output tokens to a fixed set of options. It provides a technical walkthrough on using LLMs to perform classification tasks efficiently by masking vocabulary.
Why it matters
Understanding how to constrain LLM outputs is critical for developers building reliable, deterministic AI applications that require structured data.
"System one" decision models are models that infer and respond with calibrated probabilities or every allowed answer.
Consider your everyday language model, to get typed output from it (JSON), you may use Structured Output to constrain the output to guaranteed valid JSON. While model prefills the input in one pass, it still has to go perform a pass for every token in order to generate a valid response.
In this example, 11 passes are required to generate the final output. (We're not accounting for speculative decoding and other inference optimization techniques.)
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in