Build an AI Agent Evaluation with JEV
Author(s): Quan Huynh Originally published on Towards AI. Build an AI Agent Evaluation with JEV Build a small eval harness for a tool-using AI agent: code checks the work it did, and JEV judges the words it wrote. One run of my incident agent told me a checkout slowdown was caused by a config deploy that shrank the database pool from 50 connections to 5. It was right. The explanation was clear; it cited four tools, and it even ruled out a payment-provider warning that showed up later in the logs. Then I ran it with a shorter system prompt and got an answer that read almost the same. That one was wrong in two ways. It cited a tool it never called, and it never saw the warning it was supposed to rule out. If I had only read the paragraph, I would have shipped it. That is the problem with grading an agent by reading its answer. A good paragraph and a good investigation are two different things. This post builds a small harness that checks both. Plain Python checks the work. Jev, a fast structured evaluation model, judges the explanation. You can clone it and run it in a few minutes. Run it first, read about it later I think the fastest way to understand an eval harness is to watch it grade something. So let’s start there. You need Python 3.10 or newer, an OpenAI API key, and (for the judge step) a Cloudflare account. The code is in the devops-ai-guidelines repo: git clone https://github.com/VersusControl/devops-ai-guidelines.gitcd devops-ai-guidelines/07-evaluating-ai-agents/code/chapter-08python -m pip install -r requirements.txtcp .env.example .env Open .env and fill in OPENAI_API_KEY. Leave the Cloudflare fields empty for now. OPENAI_MODEL defaults to gpt-4.1-mini; I used gpt-5.4-mini. Now run the agent once: python run_agent.py Before you look at the output, here is what that command does. It takes about ten seconds, and nothing in it touches a real system. It loads a recorded incident and checks that the case hasn’t expired. It hands the agent two things: the alert text and four read-only tools (get_metrics, get_logs, get_deploys, get_db_status). Each tool returns a response recorded from the incident, not live data. The model picks one tool per turn. The runner calls it, writes down the call, and passes the result back. This repeats for up to five turns. When the model thinks it knows the cause, it calls submit_diagnosis with a paragraph plus a few structured fields. The script prints that. The answer key is in the same JSON file, but the agent never gets it. That matters later: it’s the only reason a grade means anything. Here is what one run printed for me on September 24, 2026: Scenario: checkout-latency-after-pool-changeAgent model: gpt-5.4-mini-2026-03-17Alert: checkout-service p95 latency > 2sTool calls: get_metrics, get_deploys, get_db_status, get_logsRoot cause: A deploy at 14:02 changed DB_MAX_CONNECTIONS 50 -> 5, which immediatelyreduced database pool capacity and caused checkout requests to queue and time out.Category: deployCause change: DB_MAX_CONNECTIONS 50 -> 5Cause effect: pool_exhaustedCited evidence: get_metrics, get_logs, get_deploys, get_db_statusRejected signals: payment_provider_latency, checkout_image_v1_9_2, general_capacity_limitSteps: 5 Your run will not match this word for word. The tool data is frozen, but the model can pick a different tool order or different wording every time. That’s normal, and it’s exactly why we need a grader instead of eyeballs. What that output actually tells you It’s worth slowing down here, because every line in that output is either something the agent did or something the agent said. The whole harness is built on keeping those two apart. Read the figure from top to bottom: Alert is the input. The agent gets only this and four read-only tools. Tool calls are recorded by the runner, not reported by the agent. The agent can’t edit it after the fact. This is the most trustworthy line. Root cause is the agent’s own paragraph. It’s free text, so no == check can grade it. This is the line Jev reads. Category, Cause change, and Cause effect are short structured fields the agent must fill in. Code can compare them exactly with a hidden answer key. Cited evidence is a claim. The agent says it relied on these tools. The grader checks that claim against the Tool calls line. Rejected signals is also a claim: “I saw these and ruled them out.” The recording has a planted misleading signal (a payment-provider latency line that appears after the alert). The grader checks that the agent actually saw it and named it. Steps are counted by the runner. This case allows five: four tool calls and one answer. So in this run, the agent did the work: it called all four tools, including get_logs, so it really could have seen the payment line. The paragraph also sounds right. But "sounds right" is the part we can't check with code, and that is where Jev comes in. The incident behind the example The example is an incident agent for a checkout service. The alert says p95 latency crossed two seconds. The real cause is a config deploy one minute earlier that cut DB_MAX_CONNECTIONS from 50 to 5. The pool fills up, requests wait for a connection, and checkout times out. The order on that timeline is the whole puzzle. The deploy comes before the alert, so it can be the cause. The payment-provider line comes after, so it can’t be, even though “payment provider slow” sounds like a checkout problem. A good agent notices the order. A lucky one just picks the scariest line. Every tool response comes from a recorded JSON file, not from production. The same file holds a hidden answer key: the true cause, the evidence that proves it, the planted distraction, and the step budget. The runner never gives the answer key to the agent. It loads it only after the agent has committed to an answer. Incidents are just my example. The pattern works for any agent that uses tools and then explains itself: a support agent, a SQL agent, a code-review agent. Two checks, one […]
