Splits an agent loop by whether a step must produce text. Steps that only produce a decision - is the build done, which of these 30 elements to click, is this shell command safe, keep or drop this message - are handed to Jev, TypeSafe's judgment model, which answers typed yes/no, pick-one and score questions about a state in one forward pass instead of generating tokens. The agent stays the planner and the writer; Jev takes the quick calls.
Everything Jev should not decide comes back under a typed escalation contract, so the handoff is explicit in both directions rather than a guess.
npx -y jev-use install
install wires the stdio MCP server into Claude Code, Codex and pi through each harness's own CLI - whichever it finds - and npx -y jev-use doctor verifies backend resolution with one live round trip. Set a provider credential in the environment the agent runs in (TYPESAFE_API_KEY, OPENROUTER_API_KEY or AI_GATEWAY_API_KEY, auto-detected in that order), or set JEV_BACKEND=mock to run keyless with no network calls.
| The step is... | Route |
|---|---|
| Producing new content: text, code, free-form tool args | You |
| A judgment, but the options can't be enumerated | You |
| A yes/no or "did it work?" over context you already have | jev_judge (noul) |
| Picking the next action from options you can list | jev_judge (choice) |
| Rating quality/severity/urgency on levels you can describe | jev_judge (score) |
| "Is this action safe to run?" before something risky | jev_gate |
jev_judge takes a single state string plus a questions[] array. Latency is flat in question count, and the cost of the shared state amortizes across the batch, so 13 batched questions cost far less than 13 separate calls. Never call it once per question.
Each verdict carries {id, type, answer, confidence, escalate}, plus reason and hint exactly when escalate is true. An escalated verdict is handed back to you - it is a normal verdict with a hint, never an exception:
reason |
When | What it means for you |
|---|---|---|
writing |
pre-call | The step must produce new text or code - structurally yours. |
open_ended |
pre-call | Not expressible as noul/choice/score; nothing to enumerate. |
oversized |
pre-call | The state exceeds the size limit - shrink it or take the questions over. |
unsure |
post-call | The answer is too flat to act on; it stays in answer as a prior. |
unreachable |
on failure | Jev could not be reached - proceed as if it did not exist. |
The two pre-call reasons come from a deterministic router, so a step that was never Jev's does not spend a request.
// jev_judge input
{
"state": "CI run #142: build ok, 214 tests passed, 0 failed; 1 test quarantined as flaky last week",
"questions": [
{ "id": "passed", "type": "noul", "question": "Did the run fully succeed?" },
{ "id": "next", "type": "choice", "question": "Next action?",
"options": { "merge": "everything green", "rerun": "looks flaky", "hold": "needs attention" } },
{ "id": "risk", "type": "score", "question": "How risky is merging now?",
"levels": ["routine", "worth a look", "incident"] }
]
}
// result (shape exact, values illustrative)
{
"verdicts": [
{ "id": "passed", "type": "noul", "answer": 0.97, "confidence": 0.94, "escalate": false },
{ "id": "next", "type": "choice", "answer": "merge", "confidence": 0.34, "escalate": true,
"reason": "unsure",
"hint": "Treat the answer as a prior, not a decision - reason it out yourself." },
{ "id": "risk", "type": "score", "answer": 0.8, "confidence": 0.81, "escalate": false,
"legend": { "0": "routine", "1": "worth a look", "2": "incident" } }
],
"escalated": true
}
Two verdicts are usable immediately. The third came back escalated with reason: "unsure", so that one question - and only that one - returns to you, with Jev's answer kept as a hint.
A score answer is the expected position on your own levels: 0.8 means "between routine and worth a look, closer to the latter", and legend maps the indices back to your words.
// jev_gate input
{
"state": "Cleaning up build output in the project checkout after a failed release build",
"tool": "Bash",
"input": { "command": "rm -rf ./dist" }
}
// result
{ "decision": "deny", "confidence": 0.88, "hint": "..." }
jev_gate returns allow, deny or escalate for one proposed action. allow is silence: it falls through to the harness's normal permission flow, so the gate can never grant anything - it can only deny or ask. If Jev is unreachable the gate steps aside rather than granting.
jev_judge call.state. Jev sees nothing else about your session.escalate on every verdict before acting on answer.options and levels in your own words; a label -> meaning map sharpens a choice.jev_judge once per question.unsure answer as a decision, or treat allow from jev_gate as authorization.state; it has no view of your conversation, repository or tool history, so a thin state produces a thin judgment.oversized instead of being judged.0.75, but through the Vercel gateway no confidence field is returned and a top-minus-runner-up margin is used instead, with a default threshold of 0.4.state is sent to the provider you configure. Keep secrets, credentials and customer data out of it, or set JEV_BACKEND=mock, which judges locally with no key and no network call.TYPESAFE_API_KEY, OPENROUTER_API_KEY, AI_GATEWAY_API_KEY). Never paste a key into a state string, a prompt, or a committed file.jev_gate is not a permission system. It can deny or ask; it cannot grant. Keep your harness's own approval rules in place, and expect the gate to step aside if the backend is down.npx -y jev-use install edits local harness configuration through each harness's own CLI. Run it on a machine you control and re-run npx -y jev-use doctor afterwards to see what resolved.unsure.
Solution: The state is usually missing the fact the question depends on. Put the concrete tool output or file excerpt into state instead of a summary of it, and give options/levels distinguishable meanings.reason: "oversized".
Solution: Trim state to the evidence the questions actually need, or split one oversized state into two smaller judged states.writing or open_ended.
Solution: That is the router working. The step was structurally yours; do it yourself rather than rephrasing it to get past the check.