Typed Decisions Lab
GitHubOpen the lab
An open, reproducible evaluation · September 2026

Fast and local, or smart and hosted? What we learned putting @ruvector/typesafe against Jev

Six days after TypeSafe AI launched Jev, its "System One" model for typed decisions, an open-source clone of its API appeared in RuVector. It promised 5 ms decisions, no network and no token bill. We rebuilt it, reproduced its claims, tested both systems live on fresh data, and then combined them. This is what we found, and what we still don't know.

Investigated and orchestrated by Mondweep Chakravorty · DxSure / Agentics Consulting · built with Claude
01 · The question

Two ways to make software decide

Jev is a hosted model that returns typed answers: pick one of these options, place this on a scale, or say how true this statement is. You describe the options in words, and it decides.

@ruvector/typesafe copies Jev's request and response format exactly, but runs in-process. It compares text embeddings and learns from your labelled examples.

Could the free, local one stand in for the hosted one?

The lab's start page comparing Jev and typesafe
02 · The approach

Reproduce, challenge, go live, go fresh, then build

  1. Reproduce. We built the package from source and re-ran its author's own benchmarks. The headline numbers matched exactly.
  2. Challenge. Was the published comparison like-for-like?
  3. Go live. We ran Jev on the same tickets with a real API key.
  4. Go fresh. We tested on 150 new messages that neither system had seen.
  5. Build. Could the two work better together?

Every step is live in a public lab, so anyone can re-run it.

Headline findings
03 · Like-for-like

The published comparison wasn't apples to apples

Jev scored 85% routing support tickets to the right team using only a one-line description of each team. typesafe's matching numbers used around 120 labelled tickets. Given the same descriptions and no examples:

84.7% vs 44.0%

The ticket explorer shows Jev's recorded answer, Jev live, and typesafe untrained and trained, side by side for all 150 tickets.

Ticket explorer with four answers
04 · Does it read the question?

Change one word and see who notices

We changed "Which team should own this message" to "avoid". typesafe gave the same answer, because it never reads the instruction of a choice question. Jev noticed the question had become odd, and its confidence dropped from 100% to 31%.

With a negation, "needs a response soon" vs "does NOT need one", typesafe scored 0.84 vs 0.81. Jev scored 0.97 vs 0.04.

typesafe needs meaning in labelled examples, not in wording. With 12 examples, the same negated question works.

Thanks to Chris Poulter for pointing this out and prompting further testing and evaluation.

Negation experiment
05 · Off-topic messages

Neither system knows when to say "none of these"

Asked "What is the capital of Australia?", Jev routed it to feedback at 96% confidence. With an explicit other option, which Jev's own guidance recommends, Jev caught 77% of off-topic messages with 1% false alarms. typesafe caught 17%, with 24% false alarms.

Out-of-scope experiment
06 · Confidence you can trust

typesafe earns calibration once it has labels

When a system says "90% sure", is it right 90% of the time? With about 100 labels per question, typesafe's confidence was better calibrated than Jev's: an error of 0.035–0.039, against 0.056–0.066. Untrained, it was badly off at 0.31.

We also corrected one claim. The package's README said Jev's urgency detection was worse than guessing. That was an artefact of a 0.5 cut-off. With the cut-off tuned on separate data, Jev scored 91%.

Reliability diagrams
07 · Fresh data

150 new messages, three unrelated tasks

The tasks were council enquiry routing, construction-site safety severity, and "does this message set a deadline?". Training and test messages were kept strictly apart, and each test message is tagged by what makes it tricky: slang, negation, misleading keywords, unstated intent or sarcasm.

93–98% vs 61–75%

The trained typesafe model rated 9 of 15 immediate dangers as merely "Medium". Jev caught all 15.

Fresh benchmark results
08 · Where it breaks

Negation and misleading keywords

"No council tax question here, I just want to pay my parking fine" was routed to council tax. Across the tricky categories, typesafe scored 50–71% and Jev 100%.

typesafe's real strengths are elsewhere: 10–40× lower latency, answers that never leave your servers, and no per-token cost.

Accuracy by difficulty tag
09 · Build

A cascade: cheapest first, Jev last

Each message goes to typesafe first. If typesafe isn't confident enough, it passes to a small re-reader, a 22M-parameter model fine-tuned in 6 minutes on CPU that reads the message and each option together. Next comes a local LLM, Qwen2.5-1.5B with the training examples in its prompt, and finally Jev.

We tried RuVector's own LLM runtime, ruvllm, first. It returned random characters, or silently gave canned "mock" replies, so the local stage runs llama.cpp for now. That finding is shared with the author.

Cascade diagram
10 · Result

Jev-level accuracy with 41% of the traffic sent to Jev

93.3% with only 41% sent to Jev

At 31% sent to Jev the cascade scored 92.7%; with everything kept local, about 81%. One caveat: these cut-offs were tuned on the same 150 test messages, so treat them as an upper bound. The lab lets you move the sliders yourself.

Cascade tuner
11 · Honest failure

A confidently wrong local stage stops escalation

The local LLM was 98% sure the parking-fine message was about council tax, so it never reached Jev, which would have been right. Deciding when to trust a cheap stage is the real engineering problem, and where we'd most value your ideas.

Try a message in the cascade
12 · Verdict

Not substitutes: two different tools

Jev understands a decision you describe in words, with few or no examples. typesafe is a fast, private classifier that is well calibrated once you have labels. The most promising pattern combines them: Jev labels and adjudicates, typesafe serves the known, high-volume cases locally, and anything uncertain escalates.

Verdict table

Review it, break it, build on it

This is one investigator's evaluation, on one hand-written dataset, over two days. It needs more eyes. We'd especially welcome:

Peer review

Challenge the method, the dataset or the numbers. Every result has a runnable script and a downloadable receipt.

Real use cases

Bring a real routing, triage or guardrail problem, ideally with labelled data, and let's test whether a local-first cascade holds up.

Better stages

Better out-of-scope guards, calibrated escalation rules, a working ruvllm path, or RuVector's coherence gates as the "should I escalate?" signal.

Lab

Open the Typed Decisions Lab →

Run the experiments live

GitHub

mondweep/typesafe-lab →

Source, dataset, scripts, results

Contact

Mondweep Chakravorty →

Get involved or share a use case

Tested: @ruvector/typesafe (ruvnet/RuVector, MIT) and Jev by TypeSafe AI. Not affiliated with either. Independent evaluation by DxSure Ltd (Agentics Consulting), built with Claude.