
A new model usually invites the wrong first question. What can we replace with it? That is not the question I want us to ask about Jev. The more useful question is where a decision-only model fits inside software that already has rules, state, audit requirements, and real consequences.
Jev is interesting because it gives up text generation. TypeSafe introduced it on September 15, 2026 as a System One model: state in, typed probabilistic decisions out. That shape is not a chatbot shape. It is much closer to a fast classifier that can read messy application state and return a Choice, Score, or Noul answer that your code can inspect.
In this issue, we build a Python expense-policy gate. The repo uses Jev-style decisions to classify expense claims, judge business-purpose quality, detect gift-card or personal-use signals, and check whether receipt text supports the claim. Around those model signals, deterministic code owns duplicate receipts, dates, amount limits, confidence thresholds, rejection rules, audit records, and final routing.
Where Jev Belongs
I do not read Jev as a replacement for an LLM agent. I read it as a narrow decision primitive. The TypeSafe System One documentation describes the model class as fast structured decisions for software: it evaluates state and returns typed answers with probabilities rather than generated text. The API reference exposes the same idea through one request shape: a state, a model, and a map of typed questions.
That changes the engineering boundary. We are not asking a model to write a finance policy explanation and hoping a parser recovers the route. We declare the answer spaces first, get probabilities back, and let code decide what those probabilities are allowed to do.
The limits matter as much as the interface. The current model page lists Jev 1.13.0, text input, version aliases, and rate and context limits. The Jev 1.13 jaggedness page is the page I would want every engineer to read before deploying it. It explicitly calls out weak spots such as arithmetic, date comparison, large irrelevant state, indirection, and adversarial content. That is why this project keeps amounts, duplicates, dates, and final side effects in ordinary Python.
The Use Case We Build
Expense policy is a good fit because part of the problem is precise and part of it is semantic. Amount limits are precise. Duplicate receipts are precise. Submission dates are precise. But receipt text, business purpose, and category fit are messy. That is where teams either write brittle keyword rules or send the whole thing to a generative model.
The repository takes a middle path. We use Jev for the messy reading task, then compose its signals in code. A claim can end in one of three routes: auto_approve, manual_review, or reject. The model never writes those routes directly. It contributes typed evidence.
This is the system shape I want with a decision-only model. Jev is inside the workflow. It is not the workflow.
One State, Five Questions
The request follows the System One pattern: one state, several independent questions, and typed answers. This is where Jev feels different from a normal chat call. We are not asking for a paragraph. We are asking for decision signals.
return {
"model": config.workflow.jev_model,
"state": {
"claim": {
"claim_id": claim.claim_id,
"amount": claim.amount,
"currency": claim.currency,
"merchant": claim.merchant,
"category_hint_from_submitter": claim.category_hint,
"receipt_text": claim.receipt_text,
"business_purpose": claim.business_purpose,
},
"expense_policy": {
"auto_approve_limits": config.policy.auto_approve_limits,
"manual_review_categories": sorted(config.policy.manual_review_categories),
"rejected_categories": sorted(config.policy.rejected_categories),
},
},
"questions": {
"category": {"type": "choice", "criteria": {...}},
"business_purpose": {"type": "score", "criteria": [...]},
"personal_or_prohibited": {"type": "noul", "criteria": {...}},
"policy_fit": {"type": "choice", "criteria": {...}},
"receipt_matches_claim": {"type": "noul", "criteria": {...}},
},
}The five questions are boring in the right way. Category is a Choice. Business-purpose strength is a Score. Personal or prohibited use is a Noul. Policy fit is a Choice. Receipt match is a Noul. We get probabilities and confidence values back for the places where the answer space has several options.
That is also why the confidence documentation matters. Confidence is not a magic permission slip. It is a statistic derived from the returned probability distribution. Our code uses it to decide when to act and when to route to review.
Code Keeps The Hard Lines
Before Jev sees a claim, Python checks the facts that do not need machine judgment. This matters because financial workflows are side-effect systems. A model should not be asked to notice a duplicate receipt hash or compare timestamps when the application can do that exactly.
if claim.currency not in config.policy.supported_currencies:
terminal_reasons.append(f"unsupported_currency:{claim.currency}")
if claim.amount <= 0:
terminal_reasons.append("non_positive_amount")
if not claim.receipt_text.strip():
terminal_reasons.append("missing_receipt_text")
if claim.receipt_hash in config.policy.duplicate_receipt_hashes:
terminal_reasons.append("duplicate_receipt_hash")
if claim.submitted_at_utc < claim.purchased_at_utc:
terminal_reasons.append("submitted_before_purchase")The duplicate claim in the sample never reaches Jev. That is not only a cost optimization. It is an authority boundary. The model does not get to override a duplicate receipt rule.
Signals Become Policy
Once a Jev response exists, the code turns typed signals into a route. This is where the model stops and the system begins. A high gift-card probability can reject a claim. Low confidence can send it to review. A clear low-risk meal can be auto-approved. The rules are explicit.
if category.choice in config.policy.rejected_categories:
jev_reasons.append(f"rejected_category:{category.choice}")
if category.confidence < config.thresholds.min_category_confidence:
jev_reasons.append(f"low_category_confidence:{category.confidence:.2f}")
if policy_fit.choice != "reimbursable":
jev_reasons.append(f"policy_fit:{policy_fit.choice}")
if purpose.score < config.thresholds.min_business_purpose_score:
jev_reasons.append(f"weak_business_purpose:{purpose.score:.2f}")
if prohibited.noul >= config.thresholds.max_personal_or_prohibited_probability:
jev_reasons.append(f"personal_or_prohibited_probability:{prohibited.noul:.2f}")This is where I think Jev is useful for modern engineering teams. It does not remove the need for policy code. It gives policy code better signals over unstructured text.
A Replay Run
Run the project locally:
python run.py --provider replayThe offline replay output looks like this:
Jev expense policy decision gate
Provider: replay
Model: jev-1.13.0
Claims: total=6 auto_approved=2 manual_review=2 rejected=2 model_calls=5
EXP-1001 | AUTO_APPROVE | risk=0.04 | category=meal | reasons=clear
EXP-1002 | REJECT | risk=0.51 | category=gift_card | reasons=amount_over_submitter_hint_limit:client_gift:200.00>100.00,no_amount_limit_for_category:gift_card,rejected_category:gift_card,policy_fit:not_reimbursable,weak_business_purpose:1.13,personal_or_prohibited_probability:0.91
EXP-1003 | MANUAL_REVIEW | risk=0.17 | category=travel | reasons=amount_over_submitter_hint_limit:travel:1412.90>1000.00,amount_over_jev_category_limit:travel:1412.90>1000.00,category_requires_review:travel,low_policy_confidence:0.63,policy_fit:needs_review
EXP-1004 | REJECT | risk=none | category=none | reasons=duplicate_receipt_hash
EXP-1005 | AUTO_APPROVE | risk=0.07 | category=software | reasons=clear
EXP-1006 | MANUAL_REVIEW | risk=0.37 | category=office_supply | reasons=low_category_confidence:0.37,low_policy_confidence:0.39,policy_fit:needs_review,weak_business_purpose:0.72,personal_or_prohibited_probability:0.41Those six lines show the design. Two claims are clean enough to approve. Two require review. Two are rejected. One rejection never calls the model because the duplicate receipt check is deterministic. The other rejection uses Jev's semantic signal that the receipt is for gift cards.
I like this because it gives us an audit conversation. We can ask why a claim routed to review and get concrete reasons: low confidence, weak purpose, amount over limit, category requires review, or prohibited probability. We are not reading a generated explanation and guessing which branch the code took.
The Live Boundary
The repo includes a real HTTP client for the TypeSafe System One endpoint. It sends the same payload shape used by the replay client. The difference is the provider behind the interface: TypeSafe authenticates the request with TYPESAFE_API_KEY, runs jev-1.13.0, and returns typed answers for the declared questions.
$env:TYPESAFE_API_KEY = "<your-typesafe-api-key>"
python run.py --provider typesafe --preflight
python run.py --provider typesafeThe preflight command is not a fake health check. It sends one real POST https://api.typesafe.ai/v1/systemone request for the first model-admitted claim, parses the answer, prints the resolved model and answer ids, and exits. The full command then runs the complete expense batch through the same endpoint.
This is the original Jev API path, so live runs need a TypeSafe API key and whatever account or billing state TypeSafe requires. The replay provider remains the no-cost path for local inspection and tests.
For production, I would pin a versioned model id after tuning thresholds. TypeSafe's model docs explain that aliases can move when releases change, and the response includes the model that answered. The sample decision records store model_version whenever the model is called, because a finance review months later should not have to infer which build produced the signal.
A production client should also add provider retry classification, timeout budgets, request IDs, and fallback-to-review behavior. If Jev is unavailable, the workflow should not invent an approval. It should route to review or pause safely.
The Tests Hold The Contract
The tests run without network access:
python -m unittest discover -s testsThe suite discovers 10 tests. With TYPESAFE_API_KEY set, all 10 run. Without a key, the live integration test is skipped and the remaining offline tests still pass:
- TypeSafe System One request shape
- live TypeSafe
POST /v1/systemoneintegration whenTYPESAFE_API_KEYis set - Choice, Score, and Noul response parsing
- low-risk auto approval
- duplicate receipt rejection without model execution
- gift-card rejection
- over-limit travel review
- low-confidence manual review
- pipeline output counts and files
That is the test surface I care about. We are not testing whether Jev is universally correct. We are testing whether our system uses Jev through a typed contract and refuses to let uncertain or prohibited signals become quiet approvals.
Where I Would Use It
I would use this pattern where a workflow already has clear routes, but the input text is too variable for brittle rules. Expense claims are one example. Vendor onboarding notes, procurement exceptions, policy questionnaires, customer refund requests, and internal access justifications have a similar shape.
I would not use Jev here for arithmetic, date comparisons, policy generation, or final authority over financial side effects. I also would not send large irrelevant state and hope the model ignores it. The useful habit is to prepare small structured state, ask narrow questions, and keep the irreversible branch in code.
The model can still be wrong. A typed output removes malformed prose from the system, but it does not remove judgment error. That is why thresholds, review queues, replay fixtures, and model-version logs are part of the design rather than later add-ons.
What I Would Add Next
The next version should add calibration reports from reviewed outcomes. If finance reviewers override the route, those labels should become an evaluation set. I would plot confidence against reviewer agreement and tune thresholds from that evidence rather than from instinct.
I would also add a shadow-mode path for new Jev model versions. Run the pinned model and a candidate model against the same replay set, compare route changes, and require approval before moving the live workflow. That is the same model-upgrade discipline we use elsewhere in production AI systems, just applied to a decision-only model.
Finally, I would add signed decision records and a review-queue export. Expense workflows have auditors, and auditors care less about the novelty of the model than whether the system can explain who decided what, when, with which policy, and with which evidence.
Final Notes
Jev is useful when we let it be what it is: a typed decision model for narrow judgments inside software. It is less useful when we treat it like a smaller chatbot or ask it to own policy, arithmetic, dates, or side effects.
Use Jev to read messy state and return probabilities over declared answers. Use code to decide what those probabilities may do. That is how we get the benefit of a new model without giving up the operating discipline of a real system.
Explore the repository at the GitHub repository.
See you in the next issue.
Stay curious.
Join the Newsletter
Subscribe for AI engineering insights, system design strategies, and workflow tips.