Everyone who has tried an AI-narrated text adventure has an opinion about why they go wrong. The world forgets. Nothing happens. The prose is fine and somehow nothing lands. We had all the same opinions about the one we're building, so this summer we stopped arguing with ourselves and measured it: 90 judged passages across 15 tales, a second model scoring each one, and a harness that could replay a real playthrough under a changed prompt and compare like with like.
The short version: across 90 judged passages, the two dominant failures were flat endings (41% of passages) and scenes in which nothing actually changed (29%). Both traced to a single omission in the prompt, and fixing it moved continuity, momentum and freshness by statistically significant margins on a real playthrough. The surprises were elsewhere: a shorter prompt wrote better prose but broke something we cared about more; the biggest improvement we measured, from a second model whispering two sentences to the first, didn't survive a re-run a month later; and the judge's own diagnostic labels got worse as the writing got better.
How do you measure whether an AI narrator is any good?
The trap with prompt work is that every change is judged by eye, on a handful of turns, by the person who made the change. So the first thing we built wasn't a fix; it was a way to stop fooling ourselves.
The narrator keeps a structured world state per tale: characters with wants and fears, an object ledger with locations and conditions, a quest graph, a ledger of established facts. With journaling on, the server snapshots that entire state before every player turn, alongside the action the player typed. A replay tool can then restore the exact pre-turn state, regenerate the passage with a different set of rules, and score it. Run it twice, once per prompt, and the only variable between the two runs is the prompt. Generation is stochastic, so every turn is sampled twice and averaged.
Scoring is a separate, cheaper model acting as judge. It sees the player's action, the passage, and the facts that were standing when it was written, and returns 0-100 on four axes: continuity (does it honour the facts and the action?), momentum (does the story move: a reveal, a change, a cost, a decision?), freshness (specific, world-rooted images, or stock phrasing?), and character (do people feel like themselves?). It is told to be strict: 50 is competent but flat, 80 is genuinely good. It also names one weakness per passage from a fixed list, which turned out to matter later.
The comparison is paired: same state and same action in both arms, differenced turn by turn, with a t statistic. We hold the line at |t| ≥ 2 for a finding and treat 1.3-2 as a lean. An earlier version of the tool called anything with t > 1 "BETTER", which dressed up p ≈ 0.25 as a result. That bug produced one of the more useful lessons in this whole exercise: the harness is only as honest as its verdict labels.
What actually goes wrong in AI-narrated stories?
The baseline, on the strict scale: continuity 70, momentum 63, freshness 67, character 63. Momentum lowest, which matched how the game felt to play. The weakness labels across the 90 passages: flat ending 37, scene not advanced 26, unearned resolution 8, thin character 6, then passive protagonist, ignoring the action, and no consequence at 3 each. Two failures accounted for 70% of everything the judge flagged.
Reading the passages made the pattern obvious in a way the counts didn't. When the player probed a character, asked, confronted, demanded, the character deflected in atmosphere and yielded nothing concrete. And nearly every passage closed on a settling image. The question settles into your bones as cold as the cellar air. The flame steadies, and it is only damp again. The quiet is worse. Individually these are decent sentences. Collectively they are a narrator rounding off every turn instead of leaving it leaning forward, and a player who has just typed something now has nothing to push against.
Why do AI stories end flat?
Because we asked them to. Not deliberately, but the only instruction about endings anywhere in an 18,000-character system prompt was "end on a complete sentence", and nothing in that prompt demanded that anything be different by a passage's last line. The model did the reasonable thing with that: it finished neatly. Every rule about momentum operated at the level of the whole story, the quest, the drive toward the next objective, and none at the level of the single passage the player was about to read.
The fix was two rules, adding 6.5% to the prompt and leaving everything else byte-identical so the comparison isolated the change. The first says that by a passage's last line something concrete must be different: a fact learned, a decision forced, ground gained or lost, a cost paid; and that when the player probes, the world must yield something real, and even a character's refusal must reveal what they protect or what refusing costs. The second bans the settling coda by name, silence returning, night pressing in, a feeling settling into your bones, tells the model to cut such a line and end on the live beat before it, and carves out the story's true final ending as the one place a passage may come to rest.
On the 42-turn probe set, every judged axis moved positive and one deterministic metric, whether the model marked how each line of dialogue should be delivered, jumped 19 points at t = 3.06. On a real 11-turn playthrough, continuity rose 11.8 points at t = 2.14, with freshness and character leaning up behind it. Nothing regressed beyond noise. The solo scenes that used to end "and it is only damp again" now surface something and end mid-pull: The current changes. It tugs, faint but definite, toward the door.
Does a shorter prompt write better?
This was the experiment we expected to win, and it is the one we didn't ship. We compacted the rules by 25% by merging redundancies, and separately cut them by 49% into telegraphic fragments, and ran both against the original across the same 15 tales.
Both compactions left craft leaning better on all four judged axes, by 5 to 7 points. Constraint overload is real; the model writes more freely with fewer rules. But both also cost 4 to 5 points of voice attribution, the mechanism by which the narrator declares who is speaking each line so the app can give every character their own voice, and that regression sat at t ≈ 2.1, better evidenced than the craft gain. So the shorter prompt writes better and breaks the feature that makes the game sound like a cast rather than one narrator. Neither version went live.
Two side findings we'd not have guessed. The number of dialogue lines rose as the rules shrank, 257 to 323 to 381: fewer constraints change what the model writes, not just how many tokens it costs. And the telegraphic version was no better on craft than the merely compacted one, and worse on voice. The gain came from removing redundant rules, not from clipped grammar. If you're tempted to write your system prompt in caveman, don't; write it once, well.
Can a second model steer the first?
Static rules can say "don't end flat". They can't say "Mother Gruul has stalled the player twice now; this turn she yields something or leaves". So we added a critic: at the end of each turn, a small fast model reads the last three passages and writes zero to two private notes for the next one, tendency corrections specific to this story, injected as private direction the narrator sees and the player never does. It ran on a background thread alongside image generation and cost $0.00014 a turn, under 1% of the turn's spend.
It took three rounds to calibrate, and each round taught us something about instructing a model. The first prompt let the critic write notes on 34 of 42 turns. On the real playthrough momentum rose 17 points at t = 2.10, but lexical variety fell in both datasets at t ≈ -2.0, because the notes kept re-anchoring on the same motifs, the brass wire, the amber thread, and the prose dutifully reused them. So we tightened the prompt: never anchor on a motif your previous notes used, and "an empty list is the expected answer". The critic went mute: notes on 4 of 42 turns and 0 of 11, and the gains vanished with them. The third version replaced the mood words with a base rate, "roughly a third of turns deserve a note, the rest do not", and landed at 12 of 42 and 2 of 11.
That version, paired with a small weighted table that rolls two sensory channels per turn and deliberately under-weights smell (where the stock words live: tallow, petrichor, ozone), swept the real playthrough: momentum +16.2 (t = 2.25), continuity +19.8 (t = 2.18), freshness +16.8 (t = 2.12), character +13.7, and lexical variety flipped from a regression to +1.0 at t = 2.67, nine wins in eleven turns. Measured in August, across the two changes together, the real-play scores went from continuity 72 to 92, momentum 56 to 72, and freshness 63 to 80.
One caveat we owe you, and one that a later write-up caught before we did: that run tested the critic and the sensory table together, and in it the critic spoke on only 2 of the 11 turns, so most of those passages were written with the sensory roll and no note at all. So in September we separated them: all four combinations, critic on and off, sensory table on and off, on the same playthrough and on a 42-turn probe set, every arm re-run on the same day. Every arm had to be, because the model provider had shipped new versions of both the narrator and the judge in between.
The August result did not replicate. With both features against neither, the gains shrank to noise: momentum +2.1 (t = 0.48), continuity +1.1, freshness +2.3 (t = 1.45). Part of the reason is that the baseline itself had moved: with neither feature on, the same playthrough now scored continuity 85.5 and freshness 75.3, against 72 and 63 a month earlier. We can't separate how much of that came from the newer narrator, the newer judge, or a flaw in the August harness, which had been losing some passages to a formatting failure and quietly retrying them.
Taken apart, the sensory table showed small gains (character +7.1 at t = 2.29 on the playthrough, vocabulary variety +1.5 at t = 2.11 on the probe set), and the critic added nothing measurable on top of it, leaning slightly the wrong way. Four significant results out of eighty comparisons is about what chance alone would produce, so we read the ablation modestly: the table probably helps a little, and the critic doesn't earn its place. We switched the critic off and kept the table.
The calibration lesson is the one we'd pass on to anyone: telling a model that an empty answer is "expected" muzzles it; stating a base rate lands it where you want. "Most turns need none" is a mood. "About a third" is an instruction.
Can you trust the judge's labels?
Not the way we first did. When we compared the fixed rules against the baseline, the count of passages labelled "flat ending" went up, 35 to 39, while every score went up too. That looked like a contradiction until we remembered the judge names exactly one weakness per passage, whatever the passage's quality. As scenes started advancing, the label migrated to the next-worst thing, which was still, often, the ending. The weakness counts are a diagnostic of what to work on next; they are not a paired metric of whether the last change worked. Trust the scores, and read the passages, because the numbers miss things the eye catches in thirty seconds.
And read the raw output, not just the parsed one. The single worst failure we found all summer wasn't a craft failure: on certain player phrasings the narrator model emitted its speaker header, an empty cast list, and then simply stopped, no prose at all, and the app committed the empty turn. The eval harness had been quietly retrying exactly this and clearing it; live play gave up after one salvage attempt and shipped a blank. That is now fixed by doing what the harness did, but only because the harness dumps let us see it.
What does a turn cost?
Since the question behind all of this is whether such a game can exist commercially: the marginal cost of a text-only turn is about $0.018, with essentially no fixed cost, about five cents to bootstrap a whole world. Of that per-turn figure, images are 25%, the premium model used for pivotal beats and chapter openings 37%, the body prose 20%, and everything else, state extraction, the critic, replotting, the rest. Narration audio is not in that number and roughly doubles it. The chapters are budgeted, 8 turns for the first, 12 thereafter, with a valve that forces a close after four turns of overrun, so a tale is bounded at around 76 turns and a chapter's cost has a ceiling. That last design decision turned out to matter more for pricing than anything about the prose.
What the numbers can't tell you
The judge is a language model, and a cheap one, scoring the output of another language model; it may share the generator's tastes and blind spots, and per-passage scores have a standard deviation around 30, which is why paired t statistics on 42 or 11 turns are the right tool and why anything under |t| = 2 is noise here. The real playthrough is one tale, played by one person, and the probe set uses three deliberately provocative actions per tale rather than natural play. This is one game, one engine, one summer. And we are the people who built the thing being measured. What the method gives us isn't certainty; it's the ability to be wrong in a way we can detect, which is more than we had before.
What we'd tell anyone building one
Instrument before you tune; a harness that restores real pre-turn state and compares paired is worth more than any single prompt insight. Look for the rule you never wrote: the dominant failures came from an absence, not a bad instruction. Expect every change to have a cost somewhere else, and measure the thing you're not looking at; the compaction only "won" until we checked voice. Calibrate models with base rates, not mood words. Re-run your baseline whenever the model underneath you changes; ours moved a long way in a month. And keep the humans in the loop reading passages, because the judge will tell you the ending got worse when the whole passage got better.
This is the engine behind Wordfarer, a narrated text adventure for phones that we're still building; it isn't out yet, and there's a one-email waitlist at wordfarer.app if you'd like to know when it is. The replay tool, the rules, and the experiments described here all live in the project, and we'll keep publishing what the measurements say, including when they say we were wrong.