Typed Decisions Lab@ruvector/typesafe × Jev

Two ways to make software decide

This lab compares Jev, typesafe.ai's hosted "System One" model, with @ruvector/typesafe, an open-source local engine that copies Jev's API. Both run live from this server: typesafe in-process, and Jev through its API (jev-1.13.0). Every number is labelled with where it came from.

Jev · typesafe.ai

A foundation model that returns typed values

Jev is a new model class trained for "fast, structured decisions that software can use directly". It uses parallel sampling and a method typesafe.ai calls RLCD (reinforcement learning for calibrated decisions). It gives up free-text generation. You send a state and a set of questions. It returns typed answers with probabilities.

Runs
Hosted API, early access
Learns from
Criteria text you write (zero-shot)
Measured here
111 ms p50 / 160 ms p95 from London
Price
$0.042 / M input tokens; output free
Scope claimed
Workflows, guardrails, game state (Doom), Wikipedia racing
@ruvector/typesafe 0.1.0

A local classifier that uses Jev's request and response format

A Rust core (napi-rs native, with a WASM fallback) embeds the text with a small sentence model (bge-small, 384 dimensions). It then answers with a decision head: nearest-prototype with no labels, then a linear probe or logistic head once you add examples. Temperature scaling calibrates the confidence. A governed optimisation loop writes a receipt for every change.

Runs
In-process, offline, MIT licence
Learns from
Criteria text + your labelled examples
Measured here
13–26 ms p95 on 2 vCPU
Price
Your CPU; no per-token cost
Scope
Bounded classification over embeddings

The shared contract

Both take the same request body. Three question types cover most in-app decisions.

choice

Pick one of N (≤255)

Options are described in words. The answer is the chosen key, a probability per key, and a confidence.

"department": { "type": "choice",
  "criteria": {
    "billing": "Funds already collected…",
    "shipping": "Physical parcels not yet…"
  } }
score

Place on an ordinal scale

A legend of buckets. Jev returns a continuous score. typesafe returns the expected bucket plus a probability per bucket.

"frustration": { "type": "score",
  "criteria": ["Neutral or calm",
    "Irritated but polite",
    "Openly angry"] }
noul

A 0–1 predicate

"Is this true of the state?" typesafe marks an untrained noul calibrated:false, because it is a raw similarity, not a probability.

"urgent": { "type": "noul",
  "instructions":
    "The sender needs a response very soon" }

Headline findings independent run, 25 Sep 2026

How to use this lab

1
Ticket explorer. Open any frozen test ticket. See the ground truth, Jev's recorded answer, and two live typesafe engines (zero-shot and trained) side by side. Filter for disagreements.
2
Playground. Write your own state and questions in Jev's JSON format. The local engine answers in Jev's exact shape, or with typesafe's extra fields.
3
Experiments. Five guided exercises: zero-shot gap, out-of-scope inputs, teaching with examples, calibration, and the governed loop.
4
Evidence → Verdict. Every number with its provenance, how it fits the wider RuVector stack, and when to use which.

Ticket explorer

These are 150 synthetic support tickets from the frozen test split (seed 1337). When you pick a ticket you see four answers: Jev's recorded response from 21 Sep, Jev live, and two live typesafe engines (zero-shot and trained). Jev and typesafe zero-shot see only the criteria text.

Pick a ticket on the left.

Playground

The request body is Jev's POST /v1/systemone. The server exposes the same endpoint, so an existing Jev client can point here unchanged. Question IDs department, urgent and frustration hit the heads trained at boot. Any other ID is answered zero-shot.

Run a decision to see the zero-shot and trained engines side by side.
Call it yourself

Jev (needs an early-access key):


        

This lab (same body, same response shape):


        

In your own Node app:

npm i @ruvector/typesafe
node node_modules/@ruvector/typesafe/scripts/fetch-models.mjs

import { createTypesafe, choice, noul } from '@ruvector/typesafe';
const ts = createTypesafe({ embedder: { kind: 'onnx',
  modelDir: 'node_modules/@ruvector/typesafe/models',
  manifest: 'node_modules/@ruvector/typesafe/models/manifest.json',
  model: 'bge-small-en-v1.5-int8' } });
// ⚠ createTypesafe() with no options uses the "hash" test double:
//    word-overlap only, answers are near-meaningless.
await ts.train('dept', [{ text: '…', label: 'billing' }, …]);
const r = await ts.decide(state, { dept: choice({ billing: '…', fraud: '…' }) });
r.dept.choice; // typed "billing" | "fraud"

Experiments

Five hands-on exercises. Each one tests a specific claim. Results are computed live on this server.

E1

Is the comparison like-for-like?

Jev scores ~85% from criteria text alone. The README's local numbers use about 119–137 labelled examples. Draw a random sample of test tickets and score every arm on the same items. The Jev live arm calls the real API.

E2

Does it know when it doesn't know?

Real traffic includes things no option fits. ADR-003 says typesafe should answer those with low confidence and high abstain. Jev's own guidance says to add an explicit other option. This runs both systems with and without an other option on the states below. Edit them freely.

about 4 Jev calls per state
E3

Teach it something new

A fresh engine is built for your request. It trains on your examples (one text → label per line) and answers your test states. Watch which head answers and whether it reports calibrated. The engine needs about 5 examples per option before a trained head answers, and at least 20 held back (the default is 20% of examples) before it reports calibrated:true.

E4

Calibration: does 90% confident mean 90% right?

Reliability diagrams on the 150 test tickets. The diagonal is perfect calibration. Bar height shows accuracy within each confidence bin, and the label shows how many tickets fell in it.

E5

The governed optimisation loop, receipt by receipt

This campaign was re-run here with bge-small INT8. Each proposal is a configuration change. It is promoted only if a paired, anytime-valid sequential test on the frozen validation split rejects "no improvement" (wealth ≥ 20). One test is on accuracy and one on per-item log-loss (calibration). Receipts are hash-chained.

E6

Does it read the question?

Change one word of the instruction and see who notices. typesafe never reads the instructions of a choice question; it matches the message against the option descriptions. For noul it embeds the instruction, and embeddings largely ignore negation. Jev reads the whole request.

Fresh benchmark

These are 150 test messages neither system has seen, across three unrelated tasks, with a separate training set. It was written for this lab to avoid the package author's synthetic tickets. Four setups are compared: typesafe with criteria only, typesafe trained on the training split, Jev with criteria only, and Jev with the whole training split passed in the state.

How it was built. The evaluator (Claude) hand-wrote the messages and labels, so there is one annotator. No sentence appears in both train and test, and the closest train/test pair share only 36% of their words. Each test item carries a tag for the kind of difficulty it tests: plain, slang, negation, misleading (a keyword points to the wrong answer), implicit (the intent is never named) or sarcasm. For the deadline (noul) task, each setup's decision threshold is tuned on the training split, never on test. The "Jev + data in state" setup sends {message, labelled_examples} as the state; when scoring a training item it leaves that item out of the examples.

Reference run

Where each setup breaks: accuracy by difficulty tag

Reproduce it live

This runs the same four setups on this server, item by item: two local typesafe engines and two live Jev calls per item. Deadline thresholds come from the reference run's training split. Live results should match the reference to within a point or two. Jev is not guaranteed to be deterministic.

Browse the items

Show the training split

Cascade: local first, Jev last

Each message goes through a chain of stages, cheapest first. A stage answers only when its confidence clears the threshold you set. Otherwise the message moves to the next, more capable stage. All three local stages run on our own servers, so nothing leaves them until the final Jev step.

1 · typesafebge-small embeddings + trained head~5 ms · local 2 · re-readercross-encoder, reads message + option~30–70 ms · local 3 · local LLMQwen2.5-1.5B + training examples~0.3–0.6 s compute · own server 4 · Jevhosted System One model~150 ms · leaves your boundary unsureunsureunsure confident → answerconfident → answerconfident → answeralways answers

Each stage on its own

Tune the cascade

Stages switched on
Confidence needed to answer
typesafe
re-reader
local LLM
Where each message was answered

How little Jev do you need?

For each cap on the share sent to Jev, this is the best mix of thresholds found on the 150 test items. The thresholds were chosen on the test data itself, so treat these figures as an upper bound. Real traffic needs thresholds picked on held-out data.

Try a message

This runs every stage live, then applies your thresholds above.
About ruvllm. Stage 3 was meant to run on RuVector's own LLM runtime.
  • The npm package (2.6.2) returns random characters.
  • The Rust server loads only Llama/Mistral-style models. For anything else, including Qwen, Phi and Gemma, it silently switches to a canned-reply "mock mode".
  • With SmolLM2-1.7B it gave real answers, but took 8–13 s per call on 2 CPUs.
Stage 3 therefore runs llama.cpp (MIT licence) with Qwen2.5-1.5B-Instruct (Apache-2.0), hosted in us-central1, Cloud Run's cheapest region. It uses the same model format and could switch back to ruvllm once those issues are fixed. The re-reader is DeBERTa-v3-xsmall (MIT), fine-tuned in 6 minutes on CPU on the 124 training examples only.

Graph + cascade: does a knowledge graph help?

Real routing traffic usually comes with history: which company a complaint is about, which council a report goes to, and how past cases were routed. We added that history to the cascade as an extra signal, in two ways. One is plain counts ("how often did this company's past complaints fall under each product?"). The other is @ruvector/kge, which learns embeddings from the same history as a knowledge graph. Both were tested on real public data.

typesafe reads the textprobability for each option · local history priorcounts or KGE, from past cases · local combinetext × prior^weight (weight tuned on val) Jevonly the least confident share unsureconfident → answer

Each signal on its own

How little Jev do you need?

Target share sent to Jev

The confidence threshold for each target is picked on the validation month, then applied to the test month. So the share actually sent can differ a little from the target.

Full frontier table

What did the graph learn about a company?

Share of past complaints (counts)
KGE prior (softmax over link scores)

Why the two differ: the graph as built here records that a company has complaints about a product, not how often. Counts keep the proportions; the KGE link scores for the 11 products sit close together, so its prior comes out flatter. Giving the graph a way to express frequency is the natural next step, and the section below tries two ways to do that.

Walk through real cases

Teaching the graph about frequency

We tried two ways to give @ruvector/kge the "how often" information, in the small-history setting (1% of the history) where a graph should have the most to offer.

What's next

Method and caveats

Download the results file (all arms, frontiers, sample items).

Evidence

All runs use the package's own frozen fixtures and harness (commit 5356a84), plus a few new experiments. Provenance tags: reproduced = re-run here and matches the README, new = experiment added for this lab, recorded = the author's capture and not re-runnable here.

Tickets (8 departments, 150 test items)

Read this before quoting the headline. The like-for-like comparison is criteria text with no examples. On that basis Jev scores 84.7% live and typesafe 44.0%. The README's "close to Jev" rows use about 119–137 labelled tickets from the same synthetic generator as the test set. Given just 3 examples per option in its criteria, Jev reaches 89.3%.

Out-of-scope detection (CLINC150)

This is the package's own release gate (OOS AUROC ≥ 0.85), and abstain passes it: 0.90 with 150 intents after 8-shot training, 0.90 zero-shot, and 0.94–0.97 on 8-intent subsets. What doesn't carry over is the scale. abstain is one slot of a 151-way softmax, and training fits a sharp temperature that is applied to it too, so values sit near 1e-8 and a fixed threshold can't move between questions. A proposed opt-in abstainMode: 'sigmoid' reports the same signal on a fixed 0–1 scale. With it, one threshold tuned on the 150-intent validation split caught 78% of out-of-scope test items at 14% false alarms, and 92–97% at 11–19% on the 8-intent subsets.

Correction (27 Sep). An earlier version of this page reported OOS AUROC 0.50 at 150 options. That figure came from the benchmark harness at commit 5356a84, which left in-scope items unflagged, so the AUROC fell back to 0.5. The package author had already fixed the harness upstream (8825933a). Re-measured with the published package, as above.

Speed and cost

Code-level observations

AreaObservationImpact
Default embeddercreateTypesafe() defaults to hash, a word-overlap test double. The README quick-start answers every example with 50/50 probabilities.high First-run experience is misleading
AbstentionOn the 8-option ticket question, abstain ranks off-topic states well (AUROC 0.97), but its values are tiny: about 1% off-topic vs 0.1% on-topic, not the "high abstain" the ADR describes. The ranking holds at 150 options (CLINC150 AUROC 0.90), but the values shrink with the option count and after training (median ~1e-8).medium Needs a threshold set per question; opt-in fixed-scale mode proposed
README claim about Jev"Jev's noul urgency sits below a constant not-urgent baseline" comes from thresholding at 0.5. Live, Jev's urgency scores reach AUROC 0.94, and 91.3% accuracy with a threshold tuned on the val split. The author's harness recorded booleans, so it could not see this.medium Understates Jev
Untrained noul/scoreZero-shot urgent 28.7% (AUROC 0.51, chance level). Frustration 36.7%. These rows are missing from the README.medium Needs labels for these question types
Train reporttrain() infers the head from label strings. A choice whose option keys look boolean (pos/neg, yes/no) is reported as logistic, but a prototype or probe head actually answers.low Cosmetic edge case
ReceiptsThe hash chain uses 128-bit FNV-1a. It catches accidental corruption, but it is not cryptographic tamper-evidence.low RVF witness signing would close the gap
Dependenciesruvector-router-core is declared but never imported. Cargo comments say ONNX is "TODO", yet the code path is wired and works.info Stale docs
Distributionnpm 0.1.0 (published 21 Sep) ships a linux-x64 native binary with ONNX, plus a 16.5 MB WASM. macOS and Windows fall back to WASM. Model weights come from a hash-pinned fetch script.info Young release, one author
Tests295 JS tests here: 294 pass, 1 skipped, 0 failing, and a security test forbids child_process, network clients and state logging.good Disciplined engineering

Where it sits in RuVector

RuVector describes itself as a Rust substrate for "high speed decisions and persistent memory for AI agents". typesafe is its decision layer. These are the integrations that already exist and the ones that look realistic, tagged by maturity.

A reference architecture worth prototyping

state + questions(Jev JSON) @ruvector/typesafetrained heads · ~10 msconfidence + abstain act (high confidence)receipt → graph / RVF escalate (low confidence)Jev / ruvllm / human label → bank (tier A/B)governed loop promotes retrain

The live results favour Jev on accuracy, including off-topic handling with an other option. typesafe is 10–15× faster, runs locally, and is better calibrated once trained. So combine them. Use Jev (or a local LLM) to label and adjudicate. Serve the high-volume, known-distribution decisions locally from trained typesafe heads. Escalate low-confidence or low-abstain-margin cases to Jev. Keep a cheap out-of-scope guard in front, such as Jev with an other option on a sample, a Cognitum coherence gate, or a tuned abstain threshold.

Verdict

These are not substitutes. Jev is a foundation model that makes typed decisions from a description, and on accuracy it wins almost everywhere we measured, live. @ruvector/typesafe is a well-engineered, API-compatible toolkit for training small, fast, calibrated classifiers on your own labels. Its case rests on latency, locality and calibration, not accuracy.

DimensionJev@ruvector/typesafe
Zero-shot (criteria only)strong 84.7% live (85.3% recorded); urgency AUROC 0.94weak 44.0%; urgency ≈ chance
With labelled examplesstrong 89.3% with 3 examples per option in criteria; urgency 91.3% with a val-tuned threshold79–84% department with ~120–320 labels; urgency 83–85%; frustration 71–72% (Jev 67%)
Calibration (ECE, dept.)0.066 live (0.056 with examples)better 0.035–0.039 when all heads are trained; 0.31 untrained
Out-of-scope inputsForces an answer without an escape option. With other: 77% caught, 1% false alarmsWith other: 17% caught, 24% false alarms. abstain ranks well (AUROC 0.90–0.97) but its values are tiny, so thresholds need tuning per question
Latency111 ms p50 / 160 ms p95 live from London (185 / 231 ms recorded)10–15× faster 13–26 ms p95 on 2 vCPU, in-process
Cost at 1M decisions≈ $23 (≈546 input tokens each; ≈$52 with 3 examples per option)Compute only: ≈ 3 hours of wall time on 2 vCPU at ~10 ms/decision
Data residencyText leaves your boundaryNever leaves the process; no input logging
Open-endednessReasons over game state, navigation, guardrailsBounded to embedding similarity; no reasoning
Fresh 150-item benchmark (3 new tasks)strong 93% criteria only, 98% with training data in state; caught every High-severity hazard61% criteria only, 75% trained; missed 9 of 15 High-severity hazards (rated Medium)
Maturity / verifiabilityEarly access, closed, vendor benchmarksv0.1.0 (published to npm 21 Sep), one author, open and reproducible

Recommendations

Use typesafe when…

  • Latency must be tens of milliseconds, or data must not leave your boundary (regulated data, edge or offline devices)
  • You have (or can bootstrap, for example from Jev) 30–200 labels per question
  • You need confidence scores you can calibrate and audit, with gated, receipted changes
  • The input distribution is known and closed, or guarded upstream

Use Jev when…

  • You only have a description of the decision, or just a handful of examples
  • Inputs are open-world: add an other option and it handles off-topic states well
  • The decision needs reasoning over structured state, not topic similarity
  • A ~100–160 ms network hop and sending text externally are acceptable
  • You use noul thresholds tuned on data, not 0.5

Before relying on it in production

  1. Re-run both on your own, non-synthetic data. These tickets are templated and come from the package author's generator.
  2. Always give choices an other option, and test against your real off-topic traffic. For typesafe, set an abstain threshold empirically.
  3. Always pass an ONNX embedder explicitly. Never ship the default hash.
  4. Measure on a non-templated dataset. Banking77 and HWU64 fetchers exist in bench/datasets.
  5. Tune noul thresholds on a labelled validation slice. Jev's urgency jumps from 58.7% at 0.5 to 91.3% at 0.91.
  6. Pin the package version. It is 0.1.x and the API may still change.