Using Jev to Reduce Our Reasoner Cost by 30%
Using Jev Inside Altis, The Insurance Claims AgentAt Altis, we're building an insurance claims agen 2026-10-6 06:30:36 Author: hackernoon.com(查看原文) 阅读量:1 收藏

Using Jev Inside Altis, The Insurance Claims Agent

At Altis, we're building an insurance claims agent for property and casualty service providers such as auto repair shops, roofing companies, home restoration vendors. A core part of the product is deep research across thousands of pages of policy documents, OEM repair procedures, and insurer guidelines to produce a well-substantiated claim.

We used to walk these documents with an army of parallel agents on frontier models to maximize recall. But this was prohibitively expensive, roughly ~1M input and ~100K output tokens per run, with no smart way to filter documents before processing. Until Jev.

Our reasoner searches a large technical corpus before making recommendations. The corpus holds manufacturer procedures, estimating guidance, and reference material; the input describes the vehicle, visible damage, estimate lines, and repair context. Most documents have nothing to do with a given repair, yet the original pipeline read all of them, because excluding the wrong one could mean missing a valid recommendation.

That worked, but cost scaled with the size of the corpus rather than the difficulty of the repair. As the corpus grew, so did the requests, input tokens, and intermediate candidates. We first treated this as a model-replacement problem. It wasn't. We needed the best models, run only on the relevant subset.

Our First Model Swap Made the Workflow Worse

The first experiment was simple, we swapped the frontier model doing the corpus walk for a cheaper open-weight one (DeepSeek V4.1, Kimi K3, GLM 5.3), kept the prompt and output shape mostly fixed, and compared against historical runs.

Raw recall looked promising. The cheaper model surfaced most previously-approved procedures at a fraction of the model cost. But downstream stages fell apart given it ran 4-10x more verbose, represented source conditions inconsistently, and sometimes attached the wrong source ID to plausible operations among other issues. We'd cut cost at the one stage we were watching and raised it everywhere else.

Prompt tuning didn't help. Aggressive semantic dedup could strip prerequisites, warnings, or single-use part requirements the reasoner needed. Applicability filtering just duplicated the audit stage under another name. A split extractor-and-judge design was technically workable but economically pointless, the volume sent to the frontier judge erased the savings.

The failed migration reframed the question. Not "can a cheaper model reproduce the whole walk?" but "can a specialized model keep obviously unrelated documents from reaching it?"

We Built the Evaluation Set Before Adding the Router

By now we had completed runs from several frontier models with source citations, grounded operations, intermediate candidates, and final recommendations. Reviewers had approved or rejected many of the resulting line items, giving us signal on which parts of those traces mattered.

We converted these runs into replayable evaluation cases. Each froze the input evidence, corpus, stage output, source IDs, token usage, and reviewer outcomes, so we could rerun one stage without quietly changing the rest.

The frontier traces were a high-quality baseline, not ground truth. Reviewer decisions marked the recommendations that survived the real workflow, and we traced each back through synthesis, grounded operations, and citations. That produced several evaluation layers:

  • Source recall: did the replacement keep the documents supporting approved recommendations?
  • Extraction recall: did the corpus walk recover the relevant operation from them?
  • Candidate recall: did synthesis produce the same recommendation, even under different wording or estimate line?
  • Reviewed precision: among comparable recommendations, how many matched approved rather than rejected work?
  • Cost and latency: did stage-level savings survive the downstream work they created?

A source could survive routing while its operation vanished in extraction; an operation could survive extraction while synthesis merged or dropped it. Comparing only final recommendations would have blamed all three on whichever model we were testing.

The harness also let us hot-swap models behind stable stage setups where we test a new router against the same corpus, replace the corpus-walk model without rerunning damage understanding, run a new synthesis pass against recorded Phase 2 output. A model no longer had to become the default before we knew if it worked.

The leverage came from using frontier models to build the benchmark rather than to repeatedly do work a narrower model might handle. They stayed in the quality loop as teachers, controls, and evaluators, even when they were no longer the best choice for every production decision.

Jev Fit the Decision We Actually Needed

Jev is TypeSafe's first System One model, a class built to make typed probabilistic decisions instead of generating text. TypeSafe calls the training approach Reinforcement Learning for Calibrated Decisions (RLCD), meaning the model is trained around constrained decisions and their probabilities rather than preferred prose. Its primitives are Choice, Score, and Noul, where Noul answers whether a statement is true on a zero-to-one scale. TypeSafe's docs cover the interface and primitives.

That fit our routing problem better than a generative model. For every technical source, we needed exactly one decision. Could this source contain a procedure, prerequisite, inspection, calibration, part requirement, or technical condition relevant to the repair in this case? Jev did not decide whether the procedure was missing from the estimate, whether a carrier would pay for it, or whether it should become a recommendation. Those stayed downstream, where the reasoner has the estimate, repair context, and supporting evidence.

We then tuned the threshold to minimize the volume of sources analyzed while retaining the source material behind more than 90% of previously reviewed suggestions. This took some experimenting but gave us a cleaner test, since every other agent in the pipeline kept its behavior constant. Those numbers measured source coverage, not final recommendation recall, but that was exactly the behaviour we wanted the router to satisfy, shrinking the search area without hiding the evidence the existing reasoner needs.

Wiring in models from different providers has been an exercise in resilient infrastructure, anticipating failure modes we're unfamiliar with and allowing for generous layers of fallback. Given TypeSafe's API is still in beta, we had to account for provider timeouts, malformed answers, missing probabilities, and exhausted retries. All of these fail open, cutting savings during a degraded run but avoiding silently dropped recall. Routing decisions were checkpointed in bounded batches, so an interrupted job could resume without repeating completed calls, and off, shadow, and enforce modes let us watch exclusions before letting them touch the corpus walk.

What This Changes About Model Selection

Model comparisons usually start with capability, token price, latency, and context length. Agentic systems add another variable, the downstream computation each model creates. A cheaper generative model may bring longer outputs, duplicate candidates, weaker citations, or extra verification work; a decision model may look weaker in isolation while removing most of the inputs to an expensive stage. The only useful unit of comparison is the whole workflow under the same inputs and evals.

Jev's value here comes from moving away from using generation to achieve high quality classification. The same pattern is potentially useful in other expensive decision points, like selecting tools, choosing which cases need a frontier model, or routing uncertain ones to a person, while synthesis and technical interpretation stay with generative reasoning.

Jev has broadened how we think about using AI in real work. We're taking a hard look at which parts of our reasoner are truly classification versus truly generation. Excited for the future of highly capable zero-shot classification models like Jev.


文章来源: https://hackernoon.com/using-jev-to-reduce-our-reasoner-cost-by-30percent?source=rss
如有侵权请联系:admin#unsafe.sh