Jev Doesn’t Write, It Decides: Games Today, Company Data with Care, Computer Vision Next
Last Updated on September 25, 2026 by Editorial Team Author(s): luisacsfreitas Originally published on Towards AI. Jev Doesn’t Write, It Decides: Games Today, Company Data with Care, Computer Vision Next Most of what we build with LLMs is not writing. It is deciding. Which team gets this email. Is this message spam. Which move should the bot make. Is this action allowed. We use a text generator for these decisions because it is the tool we have, and then we spend time parsing its prose, fixing its JSON and wondering whether its “90% sure” means anything. In September 2026, TypeSafe released Jev, a model built for exactly that gap. It does not generate text at all. You give it a state and a set of typed questions, and it returns answers with probabilities, in one parallel pass, in a fraction of a second. I have spent the last weeks looking at where a model like this fits in real systems. My conclusion so far has three parts, and they are the structure of this post: Games are its natural home today. Bounded choices, code that owns the rules, tight latency and no sensitive data. Company data needs care. Jev is API-only and closed. That is not a deal-breaker, but it changes what you can send it. Computer vision is where it gets interesting. Vision models have been “System One” for a decade. Jev brings the same idea to language, and the two fit together naturally. 1. What Jev is, in one picture Figure 1. A generative LLM writes its answer token by token. Jev reads the state once and returns typed decisions with probabilities. The public facts, from TypeSafe’s documentation and early coverage: Input: a state (a string, a JSON object or a list of text values) and a set of questions. Three question types: choice (pick one option from a list), score (a value on a scale) and noul (yes/no). Output: for each question, the answer and a probability distribution. No text, no rationale. How it runs: the state is read once and every question is evaluated against it in parallel. There is no token-by-token decoding, which is why calls take roughly 70–500 ms. Limits: text only (no images, audio or video), up to 64k tokens per request. Price: about $0.04 per million input tokens; output tokens are free. Training: a post-training method TypeSafe calls RLCD (Reinforcement Learning for Calibrated Decisions), aimed at making the probabilities mean what they say. The recipe is not public. On TypeSafe’s own four-workflow evaluation, Jev scored about the same accuracy as GPT-5.6 Terra (67.8% vs 67.9%) at roughly 1/75 of the cost per case and about 25 times faster, while larger reasoning models scored higher. Those numbers are the vendor’s. The most useful independent look I found (an analysis on archerhume.com) supports the core claim that probabilities are read out directly, measured a calibration error of about 0.03 on 1,200 MMLU items, and also found two things worth remembering: calibration was weaker on harder, freshly generated problems, and the same option’s probability moved between 0.84 and 0.96 depending on where it sat in the option list. So: fast, cheap, typed and reasonably calibrated, with no explanation of why. 2. Why games are its natural home A game is almost the ideal environment for a model like this: The choices are bounded. The engine knows the legal moves, the NPCs in the room and the actions each one can take. Code owns the rules. The model picks; the engine validates and executes. A bad pick costs a strange NPC reaction, not a lost customer. Latency matters. A decision loop runs every few seconds or every frame. Ten seconds for a reasoning model is unplayable; 200 ms is fine. The data is not sensitive. Game state is yours, synthetic and public. Nothing personal leaves your perimeter. Figure 2. The game loop with Jev as the decision step, and a concrete example: working out which NPC the player is talking to. An easy example. A player with a microphone says: “hey, you with the sword, how much for a room?”. The speech-to-text transcript and the NPCs in range go into the state, and Jev gets one yes/no question per NPC: is the player talking to this character? It comes back with innkeeper 0.81, guard 0.12, bard 0.03. The guard is the one carrying a sword, but the request is for a room, and Jev weighs both. The code decides what happens next: above 0.6 the innkeeper answers, below it the NPC asks “talking to me?”. This is not hypothetical. A community benchmark (jev-benchmark, on jev-1.13.0) tested exactly this task on 79 hand-labelled utterances designed to be tricky, and reported an F1 of 0.96 with precision 1.0, against 0.82 for a fuzzy name-matching heuristic, and 0.93 when names were phonetically misspelled. The same repository used Jev as a chess engine: about 37% best-move accuracy, an estimated ~950 Elo and a median of 166 ms per move, provided the code first computes the facts about the position (Jev does not calculate ahead). Another project, WorldKit, packages the pattern as an NPC runtime: the engine enforces the rules and Jev only ever sees the actions that are currently valid. What about game engines? Jev is an HTTP API with official Python and JavaScript SDKs, so any engine can call it. The most complete integration I found is jev-unreal-statetree, a C++ plugin (public alpha) that adds a “Jev Decision” task to Unreal Engine 5.8’s StateTree. It sends one request when a state is entered, not every frame, and tags each request with a world “revision” number so that answers arriving after the game state has changed are simply discarded. For Unity, Godot or browser engines like Three.js the pattern is the same over HTTP. One rule applies everywhere: never ship the API key inside the game build; put a small server or gateway between the game and Jev. The samples are small, and these are community projects, not peer-reviewed […]
