Six days after TypeSafe AI launched Jev, its "System One" model for typed decisions, an open-source clone of its API appeared in RuVector. It promised 5 ms decisions, no network and no token bill. We rebuilt it, reproduced its claims, tested both systems live on fresh data, and then combined them. This is what we found, and what we still don't know.
Jev is a hosted model that returns typed answers: pick one of these options, place this on a scale, or say how true this statement is. You describe the options in words, and it decides.
@ruvector/typesafe copies Jev's request and response format exactly, but runs in-process. It compares text embeddings and learns from your labelled examples.
Could the free, local one stand in for the hosted one?

Every step is live in a public lab, so anyone can re-run it.

Jev scored 85% routing support tickets to the right team using only a one-line description of each team. typesafe's matching numbers used around 120 labelled tickets. Given the same descriptions and no examples:
The ticket explorer shows Jev's recorded answer, Jev live, and typesafe untrained and trained, side by side for all 150 tickets.

We changed "Which team should own this message" to "avoid". typesafe gave the same answer, because it never reads the instruction of a choice question. Jev noticed the question had become odd, and its confidence dropped from 100% to 31%.
With a negation, "needs a response soon" vs "does NOT need one", typesafe scored 0.84 vs 0.81. Jev scored 0.97 vs 0.04.
typesafe needs meaning in labelled examples, not in wording. With 12 examples, the same negated question works.
Thanks to Chris Poulter for pointing this out and prompting further testing and evaluation.

Asked "What is the capital of Australia?", Jev routed it to feedback at 96% confidence. With an explicit other option, which Jev's own guidance recommends, Jev caught 77% of off-topic messages with 1% false alarms. typesafe caught 17%, with 24% false alarms.

When a system says "90% sure", is it right 90% of the time? With about 100 labels per question, typesafe's confidence was better calibrated than Jev's: an error of 0.035–0.039, against 0.056–0.066. Untrained, it was badly off at 0.31.
We also corrected one claim. The package's README said Jev's urgency detection was worse than guessing. That was an artefact of a 0.5 cut-off. With the cut-off tuned on separate data, Jev scored 91%.

The tasks were council enquiry routing, construction-site safety severity, and "does this message set a deadline?". Training and test messages were kept strictly apart, and each test message is tagged by what makes it tricky: slang, negation, misleading keywords, unstated intent or sarcasm.
The trained typesafe model rated 9 of 15 immediate dangers as merely "Medium". Jev caught all 15.

"No council tax question here, I just want to pay my parking fine" was routed to council tax. Across the tricky categories, typesafe scored 50–71% and Jev 100%.
typesafe's real strengths are elsewhere: 10–40× lower latency, answers that never leave your servers, and no per-token cost.

Each message goes to typesafe first. If typesafe isn't confident enough, it passes to a small re-reader, a 22M-parameter model fine-tuned in 6 minutes on CPU that reads the message and each option together. Next comes a local LLM, Qwen2.5-1.5B with the training examples in its prompt, and finally Jev.
We tried RuVector's own LLM runtime, ruvllm, first. It returned random characters, or silently gave canned "mock" replies, so the local stage runs llama.cpp for now. That finding is shared with the author.

At 31% sent to Jev the cascade scored 92.7%; with everything kept local, about 81%. One caveat: these cut-offs were tuned on the same 150 test messages, so treat them as an upper bound. The lab lets you move the sliders yourself.

The local LLM was 98% sure the parking-fine message was about council tax, so it never reached Jev, which would have been right. Deciding when to trust a cheap stage is the real engineering problem, and where we'd most value your ideas.

Jev understands a decision you describe in words, with few or no examples. typesafe is a fast, private classifier that is well calibrated once you have labels. The most promising pattern combines them: Jev labels and adjudicates, typesafe serves the known, high-volume cases locally, and anything uncertain escalates.

This is one investigator's evaluation, on one hand-written dataset, over two days. It needs more eyes. We'd especially welcome:
Challenge the method, the dataset or the numbers. Every result has a runnable script and a downloadable receipt.
Bring a real routing, triage or guardrail problem, ideally with labelled data, and let's test whether a local-first cascade holds up.
Better out-of-scope guards, calibrated escalation rules, a working ruvllm path, or RuVector's coherence gates as the "should I escalate?" signal.
Run the experiments live
GitHubSource, dataset, scripts, results
ContactGet involved or share a use case
Tested: @ruvector/typesafe (ruvnet/RuVector, MIT) and Jev by TypeSafe AI. Not affiliated with either. Independent evaluation by DxSure Ltd (Agentics Consulting), built with Claude.