typia.llm.evaluation โ typed questions for evaluation models
An evaluation model generates no text. You give it a shared state and a map of typed questions, and it answers every question with probabilities: a yes/no probability, one option of a closed set, or a position on ordered levels. TypeSafeโs Jevย is the native one, and Vercel AI SDKโs experimental_evaluate() runs the same provider-neutral question shape on OpenAI, Anthropic, and Google as well.
typia.llm.evaluation<T>() turns a TypeScript decision type into those questions, and folds the answers back into T.
export namespace llm {
function evaluation<
T extends Record<string, any>,
Config extends Partial<ILlmEvaluation.IConfig> = {},
>(): ILlmEvaluation<T>;
}interface ILlmEvaluation<T> {
questions: Record<string, ILlmEvaluation.IQuestion>; // hand this to the model
config: ILlmEvaluation.IConfig; // what it was generated with, defaults filled
decode: (answers: unknown) => IValidation<T>; // answers in, T out
}This feature is experimental. It follows Vercel AI SDKโs evaluation model specification, which is itself experimental and may change in patch releases.
First example
import { experimental_evaluate } from "ai"; // AI SDK >= 7.0.103
import typia from "typia";
enum Department {
/**
* Payments, invoicing, refunds.
*
* @probability 0.8
*/
billing = "billing",
/**
* Bugs, outages, integrations.
*
* @probability 0.5
*/
technical = "technical",
/**
* Pricing, upgrades, new accounts.
*
* @probability 0.5
*/
sales = "sales",
}
enum Frustration {
/** Calm or neutral */
calm = 0,
/** Annoyed but cooperative */
annoyed = 1,
/** Angry or threatening to leave */
angry = 2,
}
interface ITicketTriage {
/** Does the customer convey urgency? */
urgent: boolean;
/** Which team should handle this ticket? */
department: Department;
/** How frustrated is the customer? */
frustration: Frustration;
/** Which products does the customer mention? */
products: Array<"card" | "loan" | "deposit">;
/**
* Does the customer ask for a refund?
*
* @probability 0.8
*/
refund: boolean
}
const ticket = "I was charged twice. Refund it now, or I leave.";
const triage = typia.llm.evaluation<ITicketTriage>();
const result = await experimental_evaluate({
model: "typesafe-ai/jev",
state: ticket,
questions: triage.questions,
});
const decoded = triage.decode(result.answers);
if (decoded.success) decoded.data; // ITicketTriage
// raw probabilities stay in the answer map, keyed by readable paths
result.answers["refund.requested"]; // { type: "boolean", probability: 0.83 }undefined
export namespace llm {
export function evaluation<T extends Record<string, any>>(): ILlmEvaluation<T>;
}The transform replaces the call with one compile-time plan. Calling TypeSafeโs API directly looks like this:
TypeScript Source
import { TypeSafeClient } from "@typesafe-ai/sdk";
import { toJevQuestions } from "@typia/jev";
import typia, { tags } from "typia";
enum Department {
/**
* Payments, invoicing, refunds
*
* @probability 0.5
*/
billing = "billing",
/**
* Bugs, outages, integrations
*
* @probability 0.75
*/
technical = "technical",
/**
* Pricing, upgrades, new accounts
*
* @probability 0.5
*/
sales = "sales",
}
enum Frustration {
/** Calm or neutral */
calm = 0,
/** Annoyed but cooperative */
annoyed = 1,
/** Angry or threatening to leave */
angry = 2,
}
interface ITicketTriage {
/** Does the customer convey urgency? */
urgent: boolean;
/** Which team should handle this ticket? */
department: Department;
/** How frustrated is the customer? */
frustration: Frustration;
/** Which products does the customer mention? */
products: Array<"card" | "loan" | "deposit">;
/** Does the customer ask for a refund? */
refund: boolean & tags.Probability<0.8>;
}
const main = async (): Promise<void> => {
// Generate the questions and checked answer decoder.
const triage = typia.llm.evaluation<ITicketTriage>();
// Ask TypeSafe's Jev in its own wire format
const client = new TypeSafeClient(); // reads TYPESAFE_API_KEY
const { answers } = await client.systemOne({
model: "jev-1.13.0", // pinned: thresholds are tuned per model version
state: "I was charged twice this morning. Refund it now, or I leave.",
questions: toJevQuestions(triage.questions),
});
// Check and decode the answers into ITicketTriage.
const result = triage.decode(answers);
if (result.success === false) {
console.error("Evaluation failed:", result.errors);
return;
}
console.log("Triage:", result.data);
console.log("Department probabilities:", answers.department);
};
main().catch(console.error);When to use this
| If you needโฆ | Use |
|---|---|
| Decisions over closed sets, with probabilities, from an evaluation model | evaluation<T>() |
| Generated data of any JSON shape, from a language model | structuredOutput<T>() |
| Function calling, where the LLM picks functions | application<Class>() |
Type mapping
Every leaf property of T becomes one question. Its JSDoc description is the question text, because the question key is not sent to the model.
| Property type | Question | Value in T |
|---|---|---|
boolean | boolean | true when P(true) reaches the threshold, 0.5 by default |
| string enum, or string literal union | choice | the returned option |
| numeric enum, or numeric literal union | score, levels in ascending value order | the level (see below) |
Array<U> of a string literal union or string enum | one boolean per member | the members decided true |
| nested object | flattened, one question per leaf | the object rebuilt |
Anything an evaluation model cannot answer is a compile error that names the property: string, number, and other open types; a single literal; a union mixing question kinds; optional, nullable, @hidden, or @internal properties; arrays of anything but a string literal union or string enum, and array type tags; tuples, dynamic keys, Map, Set, and recursive types; a leaf without a JSDoc description; a choice above 255 options or a score above 10 levels; enum members or literals that share a value, since the question could keep only one description and requirement; a boolean whose true and false carry different tags.Probability; and validation refinements such as tags.MinLength that the answer decoder cannot enforce. Documentation-only tags may remain.
Public class and interface getters or setters are decision properties, with their JSDoc supplying the question and optional @probability. Methods, private or protected members, symbol-keyed members, and call or construct signatures cannot be decoded into T and are compile errors.
criteria tells the model how to distinguish the available answers. For a choice it maps each string value to its enum memberโs JSDoc description. For a score it is an array of numeric enum member descriptions, sorted by numeric value. The providerโs score is a possibly fractional position on the arrayโs index scale, from 0 to the last index. decode() selects a level by the rules below and maps it back to the enumโs numeric value. A missing choice description becomes null, while a missing score description becomes the numeric value written as text.
For the Frustration enum above, triage.questions.frustration is:
{
type: "score",
instructions: "How frustrated is the customer?",
criteria: ["Calm or neutral", "Annoyed but cooperative", "Angry or threatening to leave"],
}A set memberโs question is the propertyโs description followed by Does the option "card" apply? and the memberโs description, if any. That sentence is written by typia in English, so keep it in mind when the rest of your JSDoc is in another language.
A nested objectโs own JSDoc is not sent anywhere; only the leavesโ descriptions become questions.
Question keys are the property paths in typiaโs accessor notation, such as refund.requested, products.card, or ["with space"].
Decode answers
decode() checks the providerโs answer map and converts it into T, returning IValidation<T>. It is not typia.validate<T>(): that function checks an already-formed T, while decode() reads the different answer-map shape. For a supported decision type, a successful decode needs no second general validation. It accepts both the neutral boolean answer { type: "boolean", probability } and TypeSafeโs native { type: "noul", noul }, and ignores TypeSafeโs extra confidence and legend fields.
- Boolean:
truewhen P(true) reaches the threshold. - Choice: the returned option.
- Score: the most probable level when the answer has
probabilities, where a tie picks the lower level; otherwise the level nearest to the fractionalscore, where a half rounds up. The distribution wins because a bimodal answerโs rounded mean can be its least likely level. - Set: the members whose P(true) reaches their threshold.
A missing or extra answer, a wrong answer type, an undeclared option, or a probability or score out of range fails with typiaโs usual error paths, such as $input.refund.requested. A choice or score may omit probabilities; when it includes them, the map must contain every option exactly once and sum to 1 within the rounding of decimals. The selected choice must have maximum probability, and a score must agree with the distributionโs weighted mean within that precision, following AI SDKโs evaluation answer contract.
Config
The second generic argument is ILlmEvaluation.IConfig.
| Option | Default | Meaning |
|---|---|---|
decimals | 2 | Decimal places of the modelโs probabilities and scores. A @probability finer than this is a compile error, and decode() tolerates that much rounding. |
const triage = typia.llm.evaluation<ITicketTriage, { decimals: 3 }>();
triage.config.decimals; // 3In practice Jevโs answers come rounded to two decimals, though TypeSafe does not document the precision. decimals is an integer from 0 to 15, like AI SDKโs rounding declaration. The default tolerates two-decimal rounding for every provider, so pass 15 to check a full-precision provider almost exactly.
Probability requirements
tags.Probability<N> and its JSDoc spelling @probability N attach a probability requirement to a decision. They carry metadata only: is(), validate(), and JSON schemas ignore them.
| Spelling | Meaning |
|---|---|
boolean & tags.Probability<N> | boolean threshold: true only when P(true) โฅ N |
"a" & tags.Probability<N> in a union | that optionโs acceptance minimum; in an array set, its inclusion threshold |
@probability N on every enum member | each memberโs acceptance minimum; in an enum array set, its inclusion threshold |
@probability N on a property | the boolean threshold, or the default minimum (for a set, the default threshold) of members without an override |
A memberโs own requirement wins over the propertyโs. Probability requirements are all-or-none for a choice, score, or set: once one member declares one, every member must declare one or the property must provide the default. The compiler rejects a partially covered type. If no member and no property declares a requirement, a choice or score has no minimum and every set member uses the 0.5 threshold.
Put @probability on the decision property or enum member, not on a type, interface, class, or enum declaration. When the program uses typia.llm.evaluation(), the compiler rejects declaration-level tags even on unused declarations: TypeScript erases the identity of primitive aliases, so use-site detection cannot reliably distinguish an unused alias from a used one.
@probability belongs to the property where it is written. An indexed-access type such as Source["threshold"] extracts the propertyโs type, not its JSDoc; put another @probability on the resulting evaluation property if it needs the same requirement. Passing Source itself to evaluation<Source>() does read the original propertyโs comment. For a requirement that must travel through indexed access or generic aliases, put tags.Probability<N> on the boolean or individual literal member type; type tags survive that extraction, including members of an array set.
enum Action {
/**
* Page the on-call.
*
* @probability 0.8
*/
escalate = "escalate",
/**
* Answer the customer.
*
* @probability 0.5
*/
reply = "reply",
}
interface IDecision {
/** What should happen next? */
action: Action;
}When the selected option has a minimum and its probability is below it, decode() fails for that path. It never falls back to a less likely option, because that would invert the modelโs judgment. A required option also fails when the answer has no probabilities to prove it, which is always the case on the OpenAI, Anthropic, and Google adapters. If the type declares no probability requirement at all, an answer without probabilities passes.
A member requirement lives on the enum, so it applies at every property that uses that enum.
Thresholds are tuned against one modelโs calibration. TypeSafeโs aliases such
as jev-latest move when a new release ships, so pin the versioned model ID
when you tune thresholds.
Calling Jev directly
Jevโs own wire format, shared by TypeSafeโs API and OpenRouterโs Decisions API, spells the boolean question type "noul". @typia/jev converts the questions; the answers need no conversion.
import { TypeSafeClient } from "@typesafe-ai/sdk";
import { toJevQuestions } from "@typia/jev";
import typia from "typia";
enum Urgency {
/** Can wait */
low = 0,
/** Respond today */
medium = 1,
/** Respond immediately */
high = 2,
}
interface ITicketTriage {
/** Does the customer need an urgent response? */
urgent: boolean;
/** How soon should the team respond? */
urgency: Urgency;
}
const ticket = "The payment failed and today's deadline is approaching.";
const triage = typia.llm.evaluation<ITicketTriage>();
const client = new TypeSafeClient();
const { answers } = await client.systemOne({
state: ticket,
questions: toJevQuestions(triage.questions),
});
const decoded = triage.decode(answers);
if (decoded.success) decoded.data.urgency; // Urgency member
answers.urgency; // { type: "score", score: 1.6, probabilities: { "0": 0, "1": 0.4, "2": 0.6 }, ... }Jevโs answers come rounded to two decimals in practice, the default of decimals, so decode() accepts them as they are.
Providers
- Calibration: probabilities are calibrated only on native evaluation models such as Jev. The AI SDK adapters for OpenAI, Anthropic, and Google ask a language model to write each number itself in one structured-output request, and return no choice or score distribution.
- Independence: Jev evaluates each question independently, and the AI SDK adapters for language models instruct the model to do the same, so
refund.requestedat0.1next to a confident refund reason is a valid result. Ask dependent questions in a second request. - Limits: TypeSafe accepts at most 255 choice options and 10 score levels, so a larger choice or score is a compile error, even for a provider without that limit. Other providers may have limits of their own, which typia does not check.
- Language: Jev documents lower accuracy for non-English state and instructions, which includes non-English JSDoc.