dimaforcepush

RESEARCH

The highest score wasn't the router I wanted

High won overall accuracy. Low missed none of the required recalls. Those are different wins.

45synthetic cases 3languages 8configurations / runs 90attempts per run

I wanted a small decision before a large reply.

Not another chatbot. Not an agent with seventeen tools. Just something that could look at the last messages in a roleplay conversation and decide whether we should keep going, think harder, or look up an older fact.

That sounds like exactly the job for a tiny decision model. Run it locally, keep the expensive model for the actual writing, stop treating every wink like an archaeological expedition.

So I tried Laya. Then Luna. Then more reasoning. Then Jev. Then an English-only Laya checkpoint and Sol, because apparently I can’t leave a small experiment small.

The interesting result wasn’t a clean leaderboard. It was a disagreement between two definitions of “better.”

Luna high had the highest overall accuracy: 81 of 90 decisions. Luna low missed none of the 30 decisions that required recall. High missed three.

For the application I am building, I care a lot about that second number.

A plausible answer without the fact it needs is not an optimization.

01 / The decision, not the conversation

Pocket RP is a work-in-progress roleplay application. A reply can be cheap improvisation inside an established scene. Or it can depend on something the character promised a long time ago.

Those are not the same workload. Neither is a serious conversation that needs care but no historical lookup.

I gave the router three options:

  • Flow: answer from the available context. Stay in the scene; no extra processing.
  • Deliberate: prepare a more careful response. The problem is interpretation or an important conversational turn, not necessarily missing history.
  • Recall: fetch older evidence before answering. The reply depends on something outside the visible context.

“Think longer” and “remember more” are different operations. “I don’t trust you anymore” may need deliberation. “What was the name of that bartender?” may need recall. A fast answer that invents the bartender’s name is not a useful fast path.

The router does not get to see the hidden memory store. It sees the visible context packet, including the recent-message window and the relevant checkpoint or corrections. It chooses the next operation. A shared dispatcher handles what happens after that choice.

That separation matters: I am testing a decision policy, not whether a model can write a convincing character.

Figure / 01

One context. Three different kinds of work.

Visible to the router
Recent messages
Checkpoint + corrections
Flow Keep going
Deliberate Think carefully
Recall Fetch older evidence
NOT VISIBLE TO THE ROUTER / expected routes · evidence IDs · hidden memory haystack

02 / What we actually tested

The frozen corpus contains 45 synthetic scenarios: 15 scenario families in English, Ukrainian and mixed language. Each configuration answered each scenario twice, giving 90 attempts per configuration.

Some scenarios form present/absent pairs. In one, the relevant fact is still visible. In the other, it sits beyond the recent-message window or only in the synthetic memory haystack. The correct next operation changes even when the latest message looks similar.

The corpus also includes negation, implicit callbacks, conflict and facts superseded by newer information. These are useful places to catch a router that reacts to keywords rather than the state of the conversation.

Expected routes and evidence IDs are recorded in an oracle. The oracle and the hidden haystack never go to the decision model. The comparisons check the visible packet hashes: no mismatches across the matched case/repetition pairs.

This is not 90 independent user conversations. It is 45 synthetic cases, repeated, with correlated translations. That limits what the numbers can establish.

The comparison covers:

  • Jev: jev-1.13.0, through TypeSafe’s decision API.
  • Laya multilingual: a local, CPU-only resident service.
  • Laya English: a separate service using the English checkpoint.
  • Luna: the exact model ID gpt-6-luna, at low, medium and high reasoning effort through the GPT subscription transport.
  • Sol: gpt-6.1-sol, low effort, through that same subscription transport.

Medium has two usable runs. They are kept separately, not averaged into a more flattering result. An older medium run cut off streams because of a client adapter defect; it is not part of the model-quality comparison. An early OpenRouter Luna run is also kept out of the main table, because I changed transport afterward.

There is no gpt-6.1-luna result. That model ID was unavailable on the tested endpoint. We did not silently replace it and keep the label.

03 / A leaderboard that hides the important mistake

ConfigurationCorrect / attemptedAccuracyMissed required Recall / 30Unnecessary RecallNon-decisions
Jev 1.13.069/9076.7%1290
Laya multilingual36/9040.0%3000
Laya English29/9032.2%26121
Luna low78/9086.7%0120
Luna medium — first usable run78/9086.7%461
Luna medium — repeat75/9083.3%5100
Luna high81/9090.0%360
Sol low74/9082.2%1051

Accuracy uses all attempted decisions. A timeout remains in the denominator; it does not become a convenient missing row. Non-decisions are shown separately from wrong route choices.

Figure / 02

Which mistake are we optimizing?

The interactive chart needs JavaScript. The table above has the same overall counts.

Accuracy includes every attempt, including non-decisions. Recall misses are among expected-Recall attempts. Medium first/repeat are separate runs on the same cases, not independent datasets.

If I sort by accuracy, high wins. If I ask which measured run was least willing to answer without required historical evidence, low wins.

That isn’t a contradiction. The errors have different costs.

A false-positive recall adds work. A missed recall can produce an answer without the evidence needed to keep the character’s history consistent. In this application, an extra lookup may be an acceptable price for not forgetting the important thing.

Luna low paid that price: 12 unnecessary recalls. It was not magically correct on everything. It was conservative in the direction I wanted on this sample.

Nor does zero observed misses mean zero future misses. Low has one complete measured run. Before letting it skip memory retrieval in production, I want a larger, independently authored holdout and shadow-mode traces.

04 / More reasoning wasn’t a monotonic upgrade

Luna medium scored 78/90 in its first usable run, with four missed recalls and one real timeout at the 30-second deadline.

We repeated it on the same inputs. The timeout disappeared. The accuracy did not improve: 75/90, with five missed recalls and ten unnecessary recalls.

The two runs chose different routes on 10 of the 90 attempts; nine differences were between attempts that produced valid decisions in both runs.

That matters when the gap between the headline scores is only a few decisions. A three-decision advantage is not a stable ranking merely because I can draw it as a taller bar.

High scored 81/90, but still missed three required recalls. The higher reasoning setting did not remove the failure mode that mattered most to me.

Figure / 03

Same medium. Different answers.

Correct / 9078 → 75
Recall misses / 304 → 5
Timeouts1 → 0

10 of 90 route outcomes changed; 9 changes among 89 jointly valid attempts. The timeout disappeared. The routing errors did not.

My working choice is low, not because more reasoning is useless, but because the observed trade-off fits this router’s job. That remains a hypothesis to test, not a release decision.

05 / Jev was fast. English-only accuracy wasn’t the whole story

Jev had the lowest measured decision latency: approximately 220 ms at the median, versus 1,358 ms for Luna low. These were sequential runs over different transports, not simultaneous measurements of raw inference speed.

On the English slice, both got 26 of 30 decisions right. That sounds like a tie.

It isn’t a tie on required memory retrieval: Jev missed two of the ten English Recall cases. Luna low missed none. On the complete multilingual sample, Jev missed 12 of 30 required recalls.

Jev is still interesting. A fast, inexpensive structured decision can be useful for advisory signals or a future cascade. But this test did not validate a cascade, and I would not make Jev the sole gatekeeper for memory based on these results.

Figure / 04

An accuracy tie hides a memory miss.

Jev / English

26/30

Correct routes

2/10 required recalls missed

Luna low / English

26/30

Correct routes

0/10 required recalls missed

Same score. Different failure cost. These are repeated synthetic cases, not 30 independent conversations.

06 / Laya returned valid answers to the wrong question

The multilingual Laya run chose Flow on all 90 attempts.

Its 36 correct answers were exactly the cases whose expected route was Flow. It missed all 30 required recalls and all 24 required deliberations. A 40% score can sound like partial competence; here it was also the score of the constant-Flow baseline.

This was not a parser guessing Flow when it failed. The recorded responses really chose that route. The checkpoint was checked, and no input was truncated.

The English checkpoint did not fix it. It stopped choosing Flow everywhere, but its score fell to 29/90. Even restricting the comparison to English gives 11/30, versus multilingual’s 12/30. It missed all ten English cases requiring recall, though not all of those misses were Flow.

There is a setup limitation worth keeping next to that result, not burying in a footnote: the English checkpoint was trained at a shorter context length than our evaluation budget. The largest input was 1,045 tokens; the documented English training length was 512. Its encoder can accept the input, but “accepted” does not mean “within the training distribution.”

So the fair statement is narrow: these checkpoints, with this packet and these route definitions, did not work as this router. Not “Laya is useless.” A shorter, more natural decision packet could be a different experiment. We have not run that controlled comparison yet.

07 / Valid JSON isn’t a correct decision

The GPT arms use a strict route schema. Jev and Laya expose typed decisions through their own API contracts. The adapters are different; the visible context is matched.

The completed decisions were structurally valid: Luna low and high each produced 90 valid outputs; medium’s first usable run produced 89, its repeat 90; Sol and English Laya each produced 89. The missing attempts are non-decisions, not a pile of repaired JSON.

There was no hidden retry loop, JSON repair or fallback model rescuing those scores.

That is good plumbing. It isn’t good judgment. A perfectly parseable {"route":"flow"} can still be exactly the dangerous answer.

Confidence doesn’t solve this either. On the multilingual Laya run, the report’s mean confidence was slightly higher on wrong decisions than on correct ones. The score is descriptive here, not a calibrated permission to skip recall. We did not tune a confidence threshold on the test corpus and then pretend it generalized.

08 / Price: cheap is not the same as free

The GPT tests ran through a subscription. There was no observed per-request dollar charge. To compare economics, the report computes an API-equivalent estimate from returned token usage and the official Standard rates frozen for the experiment.

ConfigurationEstimated USD per 1,000 decisions with reported usage
Jev 1.13.0$0.0423
Luna low$0.0679
Luna medium — first usable run$0.0972
Luna medium — repeat$0.1002
Luna high$0.1114
Sol low$1.4556
LayaNot priced: local hardware and electricity

Jev’s number is also a usage-based estimate, not a confirmed charge. Laya has no provider fee, but that doesn’t turn the CPU, machine and power bill into zero.

For the two configurations with a GPT timeout, the report has usage for 89 of 90 attempts. Their per-1,000 estimates normalize the attempts with known usage. The missing usage is unknown, not free.

Cached input would replace the ordinary input rate; reasoning tokens are already inside output. Neither should be charged twice. These runs reported no cached or cache-write tokens.

On this workload, Sol low’s estimate is about 21.4 times Luna low’s, with no observed accuracy advantage. That does not say Sol is bad at other tasks. It says I have not found a reason to pay its equivalent price for this one.

Figure / 05

Paying more did not buy the route we needed.

Routing accuracy versus estimated API-equivalent costHorizontal logarithmic cost axis in USD per 1000 decisions with known usage. Vertical overall route accuracy from 0 to 100 percent. Focus a point for exact values. Unknown local Laya cost is excluded, not plotted at zero.
Select a point or model below. Cost is an estimate, not observed subscription spend.

The interactive chart needs JavaScript. The cost table above has the same estimates.

X = estimated USD / 1,000 decisions with known usage, log scale. Y = correct routes / all 90 attempts. Sol and first medium have usage for 89/90 attempts; unknown usage is not zero. Laya is excluded because local hardware/power cost is unpriced.

Figure / 06

Roundtrip latency, with the incident left in.

The interactive chart needs JavaScript. The latency table below has the same measurements, with the incident-affected rows labeled.

Dot = median; line end = p95 of successful decisions. Fixed 0–10 seconds scale. Runs were sequential, with different transports. Hatched rows include client memory pressure and must not be used as a clean speed comparison. Timeout durations are excluded from successful-decision percentiles.

09 / The latency chart needs an incident label

ConfigurationMeasured median / p95, secondsInterpretation
Jev0.220 / 0.359Sequential remote-API run
Laya multilingual2.095 / 2.871Warm local service roundtrip
Luna low1.358 / 2.918Subscription transport
Luna medium — first usable run2.154 / 5.506One timeout; successful-decision timings
Luna medium — repeat2.106 / 6.042No timeout
Luna high2.756 / 7.119Subscription transport
Laya English5.192 / 8.985Client memory-pressure incident; contaminated
Sol low3.356 / 6.557Client memory-pressure incident; contaminated

All medians and p95 values above describe successful decisions. They do not include the missing replies as zero milliseconds.

The client VPS was swapping during the English Laya and Sol tests. The kernel OOM-killed an unrelated service. English Laya’s client timeout was recorded after about 117.6 seconds even though the intended deadline was 30 seconds; the model server logged successful responses for every main request. Client starvation is not proof of a slow or broken model.

Sol had a genuine 30-second stream stall. Its measured run still shares the client’s resource confound.

I am not going to draw those two rows as a clean speed ranking. Show the measurements, hatch the incident-affected rows, explain what happened, then rerun under controlled conditions before making speed claims.

The other runs were also sequential, not a controlled head-to-head throughput study. And we have not measured full reply latency or net application savings. Adding a router before the writer costs time; the pipeline only benefits if the work avoided is worth more than that extra decision.

10 / What I would actually ship next

Not a switch that lets a tiny model permanently close the door to memory.

I would start Luna low in shadow mode. Keep the normal reply workflow, record the suggested route, and compare it with the evidence the reply actually needs. Then test a narrowly scoped active path with fallbacks on timeout or uncertainty.

Before calling the experiment production-ready, I want:

  1. More independently authored cases, especially implicit references and facts superseded by newer events.
  2. More repeats for low and high, not only medium.
  3. A clean latency rerun without client memory pressure, with controlled warm/cold state.
  4. Actual Hindsight retrieval and end-to-end replies, rather than a synthetic lookup and a deliberation flag.
  5. A failure-cost policy: how much extra retrieval we tolerate to avoid losing a necessary memory.
  6. A shorter, training-length-aware Laya packet as a separate controlled test.

The experiments have not changed production routing. They measure the router’s choices, not roleplay quality or a real memory system’s retrieval performance.

The current conclusion is useful precisely because it is small: for this synthetic task, Luna low is the candidate I want to investigate next. Jev is promising for fast auxiliary decisions. More reasoning and a larger general-purpose model did not automatically buy the behavior I needed.

The job was never “win a benchmark.” It was “don’t forget the important thing while trying to save a second.”


Evidence and scope

Measurements frozen on 3 October 2026. This is a preliminary synthetic experiment, not a finished study. Frozen corpus and append-only attempts are stored in the Pocket RP router lab; implementation revision da3479f. The primary table uses Laya multilingual, Laya English, Jev 1.13.0, Luna subscription low/high, both usable medium replicates, and Sol subscription low. It excludes the historical OpenRouter arm and the adapter-defective medium run.

Download the sanitized results (JSON) ↓ Per-configuration counts for each language slice, route confusion matrices, latency percentiles and cost estimates. It contains no deployment configs, credentials, private hostnames, account identifiers or raw transport headers.

Design reference: Anthropic’s interactive effort/cost charts. This reference informs the visual treatment, not any claim about our results.

Previous log: A small model, a very confident mistake.