RESEARCH
How does Jev compare to the alternatives?
A routing study. Accuracy, latency, and cost.
I wanted a small decision before a large reply.
I wasn’t building another chatbot or an agent with seventeen tools. I just wanted a router that could look at the last messages in a roleplay conversation and decide whether we should keep going, think harder, or look up an older fact.
That sounds like exactly the job for a tiny decision model. Run it locally, keep the expensive model for the actual writing, stop treating every wink like an archaeological expedition.
So I tried Laya. Then Luna. Then more reasoning. Then Jev. Then an English-only Laya checkpoint and Sol, because apparently I can’t leave a small experiment small.
Luna high had the highest overall accuracy: 81 of 90 decisions. Luna low missed none of the 30 decisions that required recall. High missed three.
For the application I am building, I care a lot about that second number.
If the reply needs an old fact, skipping the lookup doesn’t help.
01 / The decision, not the conversation
Pocket RP is a work-in-progress roleplay application. A reply can be cheap improvisation inside an established scene. Or it can depend on something the character promised a long time ago.
A serious conversation might need care even when no historical lookup is needed. The router has to tell these workloads apart.
I gave the router three options:
- Flow: answer from the available context. Stay in the scene; no extra processing.
- Deliberate: prepare a more careful response. The problem is interpretation or an important conversational turn, not necessarily missing history.
- Recall: fetch older evidence before answering. The reply depends on something outside the visible context.
“Think longer” and “remember more” are different operations. “I don’t trust you anymore” may need deliberation. “What was the name of that bartender?” may need recall. A fast answer that invents the bartender’s name is not a useful fast path.
The router does not get to see the hidden memory store. It sees the visible context packet, including the recent-message window and the relevant checkpoint or corrections. It chooses the next operation. A shared dispatcher handles what happens after that choice.
I am testing the decision policy here. Whether a model can write a convincing character is outside this experiment.
One context. Three different kinds of work.
Recent messages
Checkpoint + corrections
02 / What we actually tested
The frozen corpus contains 45 synthetic scenarios: 15 scenario families in English, Ukrainian and mixed language. Each configuration answered each scenario twice, giving 90 attempts per configuration.
Some scenarios form present/absent pairs. In one, the relevant fact is still visible. In the other, it sits beyond the recent-message window or only in the synthetic memory haystack. The correct next operation changes even when the latest message looks similar.
The corpus also includes negation, implicit callbacks, conflict and facts superseded by newer information. These are useful places to catch a router that reacts to keywords rather than the state of the conversation.
Expected routes and evidence IDs are recorded in an oracle. The oracle and the hidden haystack never go to the decision model. The comparisons check the visible packet hashes: no mismatches across the matched case/repetition pairs.
This is not 90 independent user conversations. It is 45 synthetic cases, repeated, with correlated translations. That limits what the numbers can establish.
The comparison covers:
- Jev:
jev-1.13.0, through TypeSafe’s decision API. - Laya multilingual: a local, CPU-only resident service.
- Laya English: a separate service using the English checkpoint.
- Luna: the exact model ID
gpt-6-luna, at low, medium and high reasoning effort through the GPT subscription transport. - Sol:
gpt-6.1-sol, low effort, through that same subscription transport.
Medium has two usable runs. They are kept separately, not averaged into a more flattering result. An older medium run cut off streams because of a client adapter defect; it is not part of the model-quality comparison. An early OpenRouter Luna run is also kept out of the main table, because I changed transport afterward.
There is no gpt-6.1-luna result. That model ID was unavailable on the tested endpoint. We did not silently replace it and keep the label.
03 / A leaderboard that hides the important mistake
| Configuration | Correct / attempted | Accuracy | Missed required Recall / 30 | Unnecessary Recall | Non-decisions |
|---|---|---|---|---|---|
| Jev 1.13.0 | 69/90 | 76.7% | 12 | 9 | 0 |
| Laya multilingual | 36/90 | 40.0% | 30 | 0 | 0 |
| Laya English | 29/90 | 32.2% | 26 | 12 | 1 |
| Luna low | 78/90 | 86.7% | 0 | 12 | 0 |
| Luna medium — first usable run | 78/90 | 86.7% | 4 | 6 | 1 |
| Luna medium — repeat | 75/90 | 83.3% | 5 | 10 | 0 |
| Luna high | 81/90 | 90.0% | 3 | 6 | 0 |
| Sol low | 74/90 | 82.2% | 10 | 5 | 1 |
Accuracy uses all attempted decisions. A timeout remains in the denominator; it does not become a convenient missing row. Non-decisions are shown separately from wrong route choices.
Which mistake are we optimizing?
The interactive chart needs JavaScript. The table above has the same overall counts.
Accuracy includes every attempt, including non-decisions. Recall misses are among expected-Recall attempts. Medium first/repeat are separate runs on the same cases, not independent datasets.
High wins on overall accuracy. Low was the measured run least willing to answer without required historical evidence, which is what I want from this router.
A false-positive recall adds work. A missed recall can produce an answer without the evidence needed to keep the character’s history consistent. In this application, an extra lookup may be an acceptable price for not forgetting the important thing.
Luna low made 12 unnecessary recalls. On this sample, it was conservative in the direction I wanted.
Zero observed misses doesn’t promise zero future misses. Low has one complete measured run. Before letting it skip memory retrieval in production, I want a larger, independently authored holdout and shadow-mode traces.
04 / More reasoning wasn’t a monotonic upgrade
Luna medium scored 78/90 in its first usable run, with four missed recalls and one real timeout at the 30-second deadline.
We repeated it on the same inputs. The timeout disappeared. The accuracy did not improve: 75/90, with five missed recalls and ten unnecessary recalls.
The two runs chose different routes on 10 of the 90 attempts; nine differences were between attempts that produced valid decisions in both runs.
That matters when the gap between the headline scores is only a few decisions. A three-decision advantage is not a stable ranking merely because I can draw it as a taller bar.
High scored 81/90, but still missed three required recalls. The higher reasoning setting did not remove the failure mode that mattered most to me.
Same medium. Different answers.
10 of 90 route outcomes changed; 9 changes among 89 jointly valid attempts. The timeout disappeared. The routing errors did not.
For now, I’d try low: its observed trade-off fits this router’s job. That remains a hypothesis to test. I haven’t made a release decision.
05 / Jev was fast. English-only accuracy wasn’t the whole story
Jev had the lowest measured decision latency: approximately 220 ms at the median, versus 1,358 ms for Luna low. These were sequential runs over different transports, not simultaneous measurements of raw inference speed.
On the English slice, both got 26 of 30 decisions right, but Jev missed two of the ten English Recall cases. Luna low missed none. On the complete multilingual sample, Jev missed 12 of 30 required recalls.
A fast, inexpensive Jev decision could be useful as an advisory signal or in a future cascade. This test did not validate a cascade, though, and I would not make Jev the sole gatekeeper for memory based on these results.
An accuracy tie hides a memory miss.
Jev / English
26/30Correct routes
Luna low / English
26/30Correct routes
Same score. Different failure cost. These are repeated synthetic cases, not 30 independent conversations.
06 / Laya returned valid answers to the wrong question
The multilingual Laya run chose Flow on all 90 attempts.
Its 36 correct answers were exactly the cases whose expected route was Flow. It missed all 30 required recalls and all 24 required deliberations. A 40% score can sound like partial competence; here it was also the score of the constant-Flow baseline.
This was not a parser guessing Flow when it failed. The recorded responses really chose that route. The checkpoint was checked, and no input was truncated.
The English checkpoint did not fix it. It stopped choosing Flow everywhere, but its score fell to 29/90. Even restricting the comparison to English gives 11/30, versus multilingual’s 12/30. It missed all ten English cases requiring recall, though not all of those misses were Flow.
The English checkpoint was also trained at a shorter context length than our evaluation budget. The largest input was 1,045 tokens; the documented English training length was 512. Its encoder can accept the input, but “accepted” does not mean “within the training distribution.”
These checkpoints, with this packet and these route definitions, did not work as this router. That gives us little basis for judging Laya on other tasks. A shorter, more natural decision packet could be a different experiment. We have not run that controlled comparison yet.
07 / Valid JSON isn’t a correct decision
The GPT arms use a strict route schema. Jev and Laya expose typed decisions through their own API contracts. The adapters are different; the visible context is matched.
The completed decisions were structurally valid: Luna low and high each produced 90 valid outputs; medium’s first usable run produced 89, its repeat 90; Sol and English Laya each produced 89. The missing attempts are non-decisions, not a pile of repaired JSON.
There was no hidden retry loop, JSON repair or fallback model rescuing those scores.
A perfectly parseable {"route":"flow"} can still tell the writer to answer without a fact it needs. The schema won’t catch that mistake.
Confidence doesn’t solve this either. On the multilingual Laya run, the report’s mean confidence was slightly higher on wrong decisions than on correct ones. The score is descriptive here, not a calibrated permission to skip recall. We did not tune a confidence threshold on the test corpus and then pretend it generalized.
08 / Price: cheap is not the same as free
The GPT tests ran through a subscription. There was no observed per-request dollar charge. To compare economics, the report computes an API-equivalent estimate from returned token usage and the official Standard rates frozen for the experiment.
| Configuration | Estimated USD per 1,000 decisions with reported usage |
|---|---|
| Jev 1.13.0 | $0.0423 |
| Luna low | $0.0679 |
| Luna medium — first usable run | $0.0972 |
| Luna medium — repeat | $0.1002 |
| Luna high | $0.1114 |
| Sol low | $1.4556 |
| Laya | Not priced: local hardware and electricity |
Jev’s number is also a usage-based estimate, not a confirmed charge. Laya has no provider fee, but that doesn’t turn the CPU, machine and power bill into zero.
For the two configurations with a GPT timeout, the report has usage for 89 of 90 attempts. Their per-1,000 estimates normalize the attempts with known usage. The missing usage is unknown, not free.
Cached input would replace the ordinary input rate; reasoning tokens are already inside output. Neither should be charged twice. These runs reported no cached or cache-write tokens.
On this workload, Sol low’s estimate is about 21.4 times Luna low’s, with no observed accuracy advantage. I haven’t found a reason to pay its equivalent price for this job. Other tasks would need their own comparison.
Paying more did not buy the route we needed.
The interactive chart needs JavaScript. The cost table above has the same estimates.
X = estimated USD / 1,000 decisions with known usage, log scale. Y = correct routes / all 90 attempts. Sol and first medium have usage for 89/90 attempts; unknown usage is not zero. Laya is excluded because local hardware/power cost is unpriced.
Roundtrip latency, with the incident left in.
The interactive chart needs JavaScript. The latency table below has the same measurements, with the incident-affected rows labeled.
Dot = median; line end = p95 of successful decisions. Fixed 0–10 seconds scale. Runs were sequential, with different transports. Hatched rows include client memory pressure and must not be used as a clean speed comparison. Timeout durations are excluded from successful-decision percentiles.
09 / The latency chart needs an incident label
| Configuration | Measured median / p95, seconds | Interpretation |
|---|---|---|
| Jev | 0.220 / 0.359 | Sequential remote-API run |
| Laya multilingual | 2.095 / 2.871 | Warm local service roundtrip |
| Luna low | 1.358 / 2.918 | Subscription transport |
| Luna medium — first usable run | 2.154 / 5.506 | One timeout; successful-decision timings |
| Luna medium — repeat | 2.106 / 6.042 | No timeout |
| Luna high | 2.756 / 7.119 | Subscription transport |
| Laya English | 5.192 / 8.985 | Client memory-pressure incident; contaminated |
| Sol low | 3.356 / 6.557 | Client memory-pressure incident; contaminated |
All medians and p95 values above describe successful decisions. They do not include the missing replies as zero milliseconds.
The client VPS was swapping during the English Laya and Sol tests. The kernel OOM-killed an unrelated service. English Laya’s client timeout was recorded after about 117.6 seconds even though the intended deadline was 30 seconds; the model server logged successful responses for every main request. Client starvation is not proof of a slow or broken model.
Sol had a genuine 30-second stream stall. Its measured run still shares the client’s resource confound.
The chart hatches those two rows because the incident prevents a clean speed ranking. They need a rerun under controlled conditions before I can make speed claims from them.
The other runs were also sequential, not a controlled head-to-head throughput study. And we have not measured full reply latency or net application savings. Adding a router before the writer costs time; the pipeline only benefits if the work avoided is worth more than that extra decision.
10 / What I would actually ship next
I would start Luna low in shadow mode and keep the normal reply workflow. I’d record the suggested route and compare it with the evidence the reply actually needs. Then I’d test a narrowly scoped active path with fallbacks on timeout or uncertainty. I don’t want a tiny model permanently closing the door to memory.
Before calling the experiment production-ready, I want:
- More independently authored cases, especially implicit references and facts superseded by newer events.
- More repeats for low and high, not only medium.
- A clean latency rerun without client memory pressure, with controlled warm/cold state.
- Actual Hindsight retrieval and end-to-end replies, rather than a synthetic lookup and a deliberation flag.
- A failure-cost policy: how much extra retrieval we tolerate to avoid losing a necessary memory.
- A shorter, training-length-aware Laya packet as a separate controlled test.
The experiments have not changed production routing. They measure the router’s choices, not roleplay quality or a real memory system’s retrieval performance.
Jev is fast. Fucking fast. In our tests, its median decision took about 220 ms, compared with 1,358 ms for gpt-6-luna at low effort. For English-first routing where speed matters, that is compelling: both scored 26/30 on the English slice. These were sequential runs over different transports, so this is an observed workload result, not a universal speed claim.
The trade-offs are language coverage and decision accuracy. Jev scored 20/30 in Ukrainian and 23/30 in mixed language; good old gpt-6-luna at low effort scored 26/30 in each. If non-English inputs or higher accuracy matter more, Luna looks better on this sample: low scored 78/90 overall versus Jev’s 69/90, and high reached 81/90. Even in English, equal overall scores hid Jev’s two missed required recalls versus low’s zero.
I’d keep Jev as a candidate for fast English-first advisory routing and try Luna low in the memory-sensitive shadow-mode trial above. Neither result validates a production cascade. On this sample, more reasoning and a larger general-purpose model still missed required recalls.
Evidence and scope
Measurements frozen on 3 October 2026. This is a preliminary synthetic experiment, not a finished study. Frozen corpus and append-only attempts are stored in the Pocket RP router lab; implementation revision da3479f. The primary table uses Laya multilingual, Laya English, Jev 1.13.0, Luna subscription low/high, both usable medium replicates, and Sol subscription low. It excludes the historical OpenRouter arm and the adapter-defective medium run.
Download the sanitized results (JSON) ↓ Per-configuration counts for each language slice, route confusion matrices, latency percentiles and cost estimates. It contains no deployment configs, credentials, private hostnames, account identifiers or raw transport headers.
Design reference: Anthropic’s interactive effort/cost charts. This reference informs the visual treatment, not any claim about our results.