AI Model Temperature: What it is and What it means for Judgment

Every AI agent runs at a temperature. Here is what that means, why it matters, and how each Vantage agent gets its own.



AI Model Temperature: What it is and What it means for  Judgment

Every AI system you will ever evaluate has a dial inside it that its makers rarely discuss, and the setting of that dial changes the character of everything the system produces. It is called temperature. It is a single number, usually between zero and one, and it governs one thing: how much randomness the model uses when it writes.

Most platforms set it once, globally, and never speak of it again. We think that is a mistake, and this essay explains why: what temperature actually is, why a financial analysis and a brainstorm should never run at the same setting, how every agent in the Vantage fleet carries a temperature chosen for its task, and how those settings become inputs to a learning loop rather than constants in a config file.

The dial no one talks about and what temperature actually is

At every step of generating text, an AI model does not decide on one next word. It produces a probability distribution over its entire vocabulary. After "the capital of France is," the token for "Paris" might carry 95 percent of the probability, "located" 2 percent, "a" 1 percent, with thousands of others splitting the remainder. Temperature is a dial applied to that distribution before the model picks.

Low temperature sharpens the distribution. Likely tokens become more dominant, and at zero the model simply takes the single most probable token every time. The output becomes deterministic, or close to it: ask the same question twice, get essentially the same answer. High temperature flattens the distribution. The gap between the likely and the unlikely narrows, so the model samples from a wider spread of plausible continuations. Same question twice, noticeably different answers.

The name comes from physics, and the borrowing is nearly literal. In thermodynamics, temperature governs how much particles jitter around their most stable state: cold systems settle, hot systems explore. Token sampling uses the same equation.

One misconception is worth retiring immediately. Temperature does not make a model smarter or dumber, and it does not control creativity in any deep sense. The model's knowledge and reasoning are identical at every setting. Temperature only governs how the final dice are rolled over conclusions the model has already reached.

A field guide to the scale

The numbers themselves deserve context, because the scale is not linear in feel and the useful range is narrower than the possible one.

At 0, the model takes the single most probable token every time. Ask it to describe a company's quarter and you get the same sentences on every run. This is the setting for arithmetic, extraction, and anything an auditor might re-run.

At 0.1 to 0.2, the model is nearly deterministic but no longer frozen. Two runs agree on every fact and conclusion; a word choice or sentence boundary may differ. The output still behaves like a measurement.

At 0.3 to 0.5, the model writes like a careful professional with latitude. The substance is stable across runs, but the framing, emphasis, and connective tissue vary. Ask twice and you get two well-formed memos that agree with each other while reading differently. Most conversational AI products you have used run somewhere in this band.

At 0.7 to 1.0, variety becomes the product. The model reaches for less probable phrasings and less obvious angles, which is what you want when the task is imagining genuinely different futures and what you do not want when the task is extracting the numbers from a filing. At 1.0 the model samples from its raw, unmodified distribution, every token drawn at exactly the probability the model assigned it.

Above 1.0, the dial keeps turning but the returns invert. The distribution flattens past the model's own judgment, tokens it rated unlikely for good reason start getting picked, and coherence decays: first the prose loosens, then the logic, then the grammar. By 2.0 the output is word salad. No serious analytical system runs here; the territory exists mostly to demonstrate why the dial matters.

The full dial runs from frozen to incoherent, and the entire analytical range lives in its bottom half. That is worth holding onto when reading what follows: the difference between 0.1 and 0.7 is not a nuance. It is the difference between a measurement and a brainstorm.

Why one setting cannot serve every task

What temperature buys, and what it costs, depends entirely on the task.

For work with a verifiable right answer, extraction, classification, computation, reading a financial statement, randomness is pure liability. You want the model's best single judgment, reproducibly. Run the same board deck through the same parser twice and the extracted figures should match, byte for byte, because reproducibility is what makes machine work auditable. A fund's auditor does not want to hear that the pipeline produces different numbers on different days.

For work whose value lies in range, framing scenarios, surfacing angles, drafting narrative, some randomness is the point. The second-most-likely phrasing is often the more interesting one, and a model held at zero produces the same safe synthesis every time, which is a quiet way of narrowing what a human reader gets to consider.

So the question is never "what temperature should the AI run at." It is "what temperature does this task deserve." A platform that runs everything at one setting has answered a question it never asked.

How the Vantage fleet is tuned

Every reactive agent in the Vantage fleet, the agents that read, analyze, and report on each portfolio company, carries a temperature chosen for the kind of judgment it renders. The doctrine has three bands.

Verdict reads run near zero (0.1). Financial Analysis, the parsing pass of Deck Storyteller, and the Product and Technology Analyst all render verdicts on evidence: what the numbers are, what the deck claims, what the technology is. These are reads where two runs should agree, so they run nearly cold.

Reconstruction runs low (0.2). Company Timeline and House View reassemble what happened and what the firm concluded. The facts are fixed; the assembly requires slight judgment in ordering and emphasis, but not invention. Cold, with a degree of freedom.

Synthesis runs moderate (0.3 to 0.4). Company Profile, Peer Benchmarks, Predictive Signals, News Briefing, Valuation Insights, Talent Assessment, and the synthesis pass of Deck Storyteller weave many sources into narrative. Here phrasing, framing, and connection carry real value, so these agents run warm enough to write well while staying anchored to cited evidence. The oracles sit in this band too: Pythia converses at 0.4, and the fund-level Narrative agent, whose whole job is prose, runs at 0.5.

Generation runs hot (0.7). One agent sits far above the rest of the fleet, and the gap is the doctrine working, not failing. The Liquidity agent runs at 0.7 because it performs the most generative task in the fleet: where every other agent produces a verdict, a classification, or a reconstruction of facts, the Liquidity agent writes. Its output is three distinct scenario narratives, downside, base, and upside, each a full prose paragraph painting a genuinely different future, delivered in partner voice: direct, specific, conviction-driven. Run that task cold and the result is flat, templated scenario prose, three futures that read like one future with the adjectives swapped, which defeats the entire purpose of scenario work. The heat is what buys three paths that sound as different as they would actually feel. The numbers underneath the scenarios are not generated at 0.7; they arrive from the cold agents upstream. The temperature governs only the writing, and the writing is the deliverable.

Notice what the bands encode: the closer an agent's output sits to a number an auditor might rely on, the colder it runs, and the more an agent's value lies in humans genuinely considering alternatives, the warmer. Temperature is not a style preference. It is a statement about what kind of claim the agent is licensed to make.

What the spread proves

Publishing a tuning doctrine invites a fair question: is this real engineering or a tidy story written after the fact? The spread itself is the evidence. A platform that never asked the temperature question runs everything at one vendor default, and its settings, if you could see them, would be a flat line. The Vantage fleet runs from 0.1 to 0.7, a sevenfold range, and every position on that range traces to an argument about the task: verdicts cold because auditors re-run them, reconstruction low because facts are fixed, synthesis moderate because framing carries value, scenario generation hot because three futures that read alike are not three futures. When one agent sits far from the others, there is a reason you can read, and disagree with, and hold us to. That is the same principle that runs through the whole platform, the same reason every claim in Engram carries provenance: systems earn trust not by being uniform but by being inspectable.

Temperature as a learned parameter

Everything above describes settings chosen by people, informed by the task and checked by evals. The next step, already underway, is to close the loop.

Vantage agents write their outputs back into the record: predictions with dates, assessments with citations, forecasts that are later joined by their resolutions. That means every agent accumulates a track record, and a track record is exactly what you need to tune with. If the predictive agents' resolved forecasts show that a warmer setting surfaces signals a colder one misses, the setting should move. If a synthesis agent's citation-verification rate degrades above a threshold, it should cool. The same evidence decides larger questions than temperature: which underlying model each agent should pull from, as new model generations arrive and are benchmarked against each agent's actual task.

This is the learning loop: settings that begin as engineering judgment become parameters adjusted by measured performance, per agent, per task, on the client's own work. The fleet does not just produce analysis. It produces the evidence for its own recalibration.

The point beneath the dial

None of this is really about a number between zero and one. It is about a discipline: every degree of freedom a machine is given should be a decision someone made, stated plainly enough to be audited, and instrumented well enough to be improved. Temperature happens to be the cleanest place to show that discipline, because it is one dial with visible consequences. But the same discipline governs model selection, grounding, and citation, all the way down.

The machines run at the temperature their task deserves, and leave the record that proves it. Your team keeps the conviction.


Conviction Made Citable.

← All Perspectives