In August we switched on two narrator features together and measured a large gain. This month we pulled them apart in a 2x2 ablation, and the gain didn't survive. Both features against neither now come out within noise on both of our test journals; the no-feature baseline has itself risen a long way on the same playthrough; and taken separately, the sensory palette shows small gains while the story critic adds nothing we can measure on top of it. The most useful thing we learned isn't about either feature: re-run your baseline whenever the model underneath you changes.
This is the write-up we promised when we reported the August result. It's mostly a null result, which is the kind worth publishing.
What were the two features?
The first is a sensory palette. Each turn a weighted random table rolls two sensory channels (touch, weight and temperature are favoured; smell is weighted lowest) and offers them to the narrator as an invitation rather than an instruction.
The second is a story critic. After each turn a small, fast model reads the recent passages and writes up to two private notes that steer the next passage.
In August both were switched on in a single arm, and the critic spoke on only 2 of the 11 turns in that run. So whatever the gain was, we couldn't credit it to either feature. That's the question the ablation was built to answer.
How did we run the ablation?
Four arms: neither feature, palette only, critic only, both. Two journals: the same real playthrough as before, 14 turns, and a 42-turn probe set. Two samples per turn, every arm re-run on the same day.
Every arm had to be re-run, including the baseline. Since August the model provider had shipped new versions of both the narrator model and the model we use as judge, so the August reports couldn't be compared with anything generated today.
We fixed the harness first. It had been losing passages to a formatting failure: on scenes where nobody speaks, the narrator would write an empty cast list and stop. Live play now rescues those by handing the model its own header back as the start of its reply, and the harness was changed to do exactly the same before this run, so no arm lost turns unevenly. The August harness had the same failure and retried those passages from scratch instead.
This time the critic spoke on 11 of 14 turns on the real playthrough and 26 of 42 on the probe set, so it was at least present to be measured.
Did the August gain replicate?
No. On the real playthrough, both features against neither moved the overall deterministic score by 3.2 points (t = 1.5, 9 wins, 2 ties, 3 losses), which leans better but isn't a finding by our bar. Rule compliance leaned the same way, up 14.3 (t = 1.34, 3/10/1), as did the judge's freshness score, up 2.3 (t = 1.45, 9/1/4). Everything else was noise, including the judge's momentum and character scores.
On the 42-turn probe set, both against neither was flatter still: overall up 0.3 (t = 0.44, 14/15/13). Opening variety leaned better (up 1.1, t = 1.63, 24/5/13) and judge continuity leaned better (up 1.8, t = 1.31, 19/8/15). Momentum, freshness and character were all noise.
Meanwhile the no-feature baseline had risen a long way since August on the same playthrough. We can't separate how much of that is the newer narrator model, the newer judge, or the August harness flaw that retried silent scenes from scratch.
Which feature was doing the work?
Palette alone against neither. On the real playthrough the judge's character score rose 7.1 points (t = 2.29, 10 wins, 0 ties, 4 losses), from 56.3 to 63.4, which is a finding by our bar. Everything else on that journal was noise. On the probe set, lexical variety rose 1.5 (t = 2.11, 24/4/14), also a finding; overall leaned better (up 1.4, t = 1.98, 16/19/7) and judge continuity leaned better (up 2.5, t = 1.62, 26/2/14).
Critic alone against neither. On the real playthrough lexical variety rose 1.3 (t = 2.39, 9/3/2), a finding, and nothing else moved past noise. On the probe set opening variety rose 1.4 (t = 2.17, 21/2/19), a finding with a nearly even win/loss split, while judge momentum leaned worse, down 3.6 (t = -1.86, 17/3/22).
Both against palette alone, which is the cleanest question about the critic. On the real playthrough every metric was noise; the judge's character score fell 4.5 (t = -1.16, 4 wins, 0 ties, 10 losses), which isn't a finding but doesn't point the way we'd hoped. On the probe set, overall leaned worse (down 1.1, t = -1.39, 13/10/19) and judge freshness leaned worse (down 1.9, t = -1.71, 17/5/20). Nothing leaned better.
So the critic didn't add anything measurable on top of the palette, and where the numbers lean at all, they lean slightly the wrong way.
What can't these numbers tell you?
Quite a lot. The judge is a language model, and a newer one than in August; a shift in its taste is indistinguishable from a shift in the narrator's craft. Its per-passage scores are noisy: on the real playthrough the paired differences in momentum had a standard deviation of 16.8, against a mean difference of 2.1.
The turn counts are small. Fourteen turns on the real playthrough means a single strong passage can carry a metric; 42 on the probe set is better but not by an order of magnitude. Both journals are from one game.
And there were about eighty paired comparisons across the four arms and two journals. At that count, a few results crossing |t| >= 2 is roughly what chance alone would produce. The palette's character finding and the critic's lexical and opening-variety findings each sit in that shadow; read them modestly.
The most frequent weakness the judge flagged, in every arm, was the scene not advancing: between 23 and 26 times across the four real-playthrough arms, and between 48 and 55 on the probe set. Neither feature moved that count much. If you want the longer thought on why that matters, we've written about why stories need to feel like they're going somewhere.
What did we decide?
We switched the critic off in live play and kept the palette. The critic's code stays behind a switch so it can be re-tested later.
The result we're taking most seriously isn't about either feature. A baseline is only valid for the model it was measured on. When the provider ships a new narrator or a new judge, the August numbers become history rather than a control, and a gain measured against them can't be trusted until it's measured again.
Wordfarer is in development for Android and iOS. There's a one-email waitlist at wordfarer.app.