A model interface built around decisions
TypeSafe AI presents Jev as its first public System One model, intended for decisions inside software. Its interface accepts state and typed questions and returns results that application code can consume. This article covers that design, rather than treating vendor performance claims as independent benchmarks. Source: TypeSafe AI.
The documentation defines three primitives:
- Choice: Select an option from a supplied list, with probabilities and confidence.
- Score: Evaluate state against a rubric, with probabilities and confidence.
- Noul: Estimate whether a statement is true, returning a value from zero to one.
Questions are evaluated independently against the supplied state. TypeSafe recommends asking narrow questions and combining their answers in code when a decision has several dimensions. Source: Jev introduction.
A practical routing example
Imagine a support inbox for a portfolio site. An application could ask a Choice question with categories such as technical issue, collaboration request, and other. A separate Score question could assess urgency using a clearly defined rubric.
My suggested workflow would first validate the incoming message, request those judgments, and propose a queue. It would keep ambiguous messages available for manual review rather than silently discarding them. Moving a message into a suggested queue is also a different action from sending a response.
This example describes an application architecture; it is not an SDK code sample or a measured result.
Confidence is a signal to interpret
For Choice and Score, Jev's confidence summarizes the shape of the returned probability distribution. It is not the same field as the selected option's probability. Noul does not return a separate confidence field. TypeSafe's guidance recommends testing thresholds on your own data and adjusting them to the consequences of an error. Source: TypeSafe confidence documentation.
A useful evaluation set would include short messages, mixed languages, missing context, and requests that fit several categories. Check actual routing accuracy and the cases sent for review. A typed response makes software integration easier; the application still needs evidence that its decisions are useful.
Keep policy in your application
Use explicit code for allowed actions, validation, logging, and fallback behavior. A model's uncertainty signal can inform a policy, but it should not define account permissions or replace business rules.
For an initial integration, I would prefer suggestions that can be reviewed and reversed. Expand automation only after the results support that change.