Entrepreneur · CTO · Principal AI architect · Consulting

Will Frasier

CTO and founder (Story Stream, Family OS), 18-year Microsoft veteran, and dad of three.

Essay ·

A Crowd of Truths

In January 2024 I made the first commit to what became Story Stream, software that reads a novel manuscript and gives the writer editorial feedback. Almost every writer who used it wanted to know one thing: is the book good?

Software isn't built for that question. Nearly everything we know about evaluating a system assumes a ground truth somewhere: the right label, the expected output, the test that passes or fails. A novel has no such thing. It has a crowd of truths that disagree with each other, and they're all correct. This essay is about building against that, and what it taught me about quality.

Many ground truths

Machine translation met a version of this first: BLEU scores against several human reference translations, because two excellent translators will render a sentence differently and both are right. A novel is that problem with the dimensions turned all the way up. Readers judge on style, tone, pace, word choice, humor, subtext, the cultural moment, whether the book meets or subverts what its genre promised, whether the voice sounds like a person. These dimensions trade against each other, and the trades are the art. A slow chapter can be a flaw on the pacing axis and the whole point on the emotional one.

A flaw on one axis is often a choice on another. The grader has to know which.

I studied classical composition before I wrote software. Every theory student learns that parallel fifths are wrong. Then you hear Debussy stacking them on purpose and it's gorgeous. The rule isn't false. It's a default, and the skill is knowing when you've left it deliberately.

Two lenses: what we measure, and what we think

Story Stream assesses a manuscript through two lenses at once: measurement, meaning what can be observed in the text, and a subjective read, meaning an agent with a point of view deciding whether what it observed works. Every piece of feedback is a synthesis of the two.

The measurements come from editors, not from us. We distilled the books editors learn from (Browne and King, McKee, Truby, Save the Cat!, The Bestseller Code, and a shelf of others) into core insights, and asked of each whether it could be checked against a manuscript. That prompt has my favorite line in the codebase:

"Not all insights will be measurable in a novel… 'Drink lots of water before writing' is not measurable… 'Action scenes should use shorter sentences' is measurable."

What survived became the rubric: 791 metrics across 16 areas of craft, plus 19 genre-specific sets, each traceable to a sentence an editor once wrote in a book.

An interesting test is the seemingly illusive concept of theme. Sometimes this feels like the most abstract thing a writer works with. So how do you even go about measuring theme? I learned that it's actually easier than you think, editors know all about this. Let's say the book's central theme is corruption. A theme feels well realized with the reader sees corruption from many perspectives: the victims, the people in power, the people on the sidelines, the people quietly benefiting. Readers don't know why, but when a theme is explored through many different characters, it lands. I don't pretend to know why, it just does. Our theme guidance puts it this way:

"Themes feel well-explored when examined from multiple different angles: personal, professional, governmental, local, global, through the villain's perspective, the hero's perspective, supporting characters, and different social strata. Each angle adds depth and nuance."

That is observable. Which characters touch the theme? From which side? Is anyone profiting from the rot, or is everyone either a victim or a villain?

Is it universal? No. Nothing in craft is. It's generally true, not exclusively true, and we measure it because the editorial community has learned over a great many books that it tends to work. A low score doesn't mean the theme is badly executed. It just means the manuscript isn't following a practice that other novels have collectively found reliable. That might be exactly the intent (a novel sealed inside one man's corruption is a legitimate book), or it might be an omission. The measurement can't tell which and shouldn't try. When we measure, we don't pass judgment. We observe.

Judgment is the second lens, and it lives in the agents.

Two editors

The Developmental Editor reads the whole manuscript. More than a dozen specialists (pacing, theme, character, dialogue, plot, emotional impact, structure, market) run in parallel alongside sixty-odd genre specialists, because good pacing in a thriller is a flaw in a family saga. Each grade is split into execution (how well is this done?) and impact (how much would fixing it matter?). A clunky hallway and a clunky climax are both clunky. Only one is why a reader puts the book down.

The Manuscript Editor lives on the page: chapter reviews, line edits, a panel of sixteen beta reader personas, and a conversation where the writer can argue back. Same two lenses, at the scale of a sentence.

Both end in the same place. Something has to take many voices and produce one answer. That's the judge.

The judge

This is the part I find richest. I've done countless experiments where I've let hoards of different literary editors loose on a chapter. What is the collective noun for literary editors? A revision of editors? I've sat back and watched various conversations play out, from long rambling ruminations about historical context and precedent, stylistic relevance, arguments over 'cultural moment' to one hilarious interaction where the editors all ended up in a round robin thanking each other profusely. But at the end of it all, when a dozen agents disagree about a chapter, how you you decide who "wins"? Here are the handful of popular patterns for judging that I considered for Story Stream, each carrying a hidden theory of quality.

  • Voting. Sample many answers and take the majority. Self-consistency (Wang et al., 2022) made this standard for reasoning, premised on problems with a "unique correct answer." This is where I naievly started. This method suffers from the answers all tending toward the average. In a novel, this means tending toward predictability, safeness and the 'median' reader. Who is, let's be honest, the worse dinner date possible.
  • Most similar wins. For free-form output, where you can't count identical answers, Universal Self-Consistency (Chen et al., 2023) picks the candidate most consistent with the rest. Agreement becomes a proxy for truth, which leans toward the note anyone could have written.
  • Strongest voice wins. One capable judge reads everything and decides. Zheng et al. (2023) found strong LLM judges agree with humans over 80% of the time, about as often as humans agree with each other. The ceiling is human disagreement. They also documented position and verbosity biases: the strongest voice is sometimes just the longest. But what if you have two voices to are loud and equally shout-ey?
  • Juries. Verga et al. (2024) replaced one big judge with a panel of smaller models from different families: better agreement with humans, less bias, a seventh of the cost. Diversity beat size.
  • Debate. Du et al. (2023) have agents argue until they converge, which improves factuality. For fiction, convergence can mean the minority reading got argued out of existence.
  • Keep the disagreement. The data perspectivism work (Cabitza, Campagner and Basile, 2023) argues majority-vote gold labels throw away real signal on subjective tasks, and that the spread of perspectives is the data.

So which did we choose? A jury with a chair, where the chair has to publish the minority report. It takes one piece from four of the patterns above and rejects the other two.

From juries we took diversity. Our panelists differ by perspective rather than by model family. An acquisitions editor, an adaptation scout, a first-time reader and a sensitivity reader want different things from the same chapter, and that difference is the reason to ask all of them.

From similarity we took confidence. The aggregator clusters feedback by underlying issue, but agreement doesn't decide which note survives. It decides how sure we're allowed to be. The strongest signal is a cluster where a beta reader's reaction and the editor's craft diagnosis land on the same passage, what the prompt calls "reader reaction plus craft cause." The Developmental Editor's Senior Editor works the same way across a whole book: it looks for convergence, where it can be confident, and divergence, where it has to say "it depends what you're going for," and mean it.

From the strongest voice we took a decision, because a writer needs an answer: move on, or revise. One aggregator has final say, and the prompt gives it permission to overrule: "It is reasonable to choose 'ready to move on' even when one reader or the editor gave high-priority feedback, if you judge that feedback is an outlier."

From perspectivism we took disclosure. The same instruction goes on: "your readiness_summary MUST acknowledge the discrepancy." The chair can overrule a dissent. It can't hide one.

We left out majority voting because our voters aren't interchangeable. Persona groups (craft, publishing, diverse voices) are scored separately and then averaged with equal weight, so seven craft readers can't outvote the diverse-voices group on headcount, and sixteen readers flagging one problem counts as one problem. We left out debate because we never wanted the panel to converge. We wanted the writer to see where it didn't.

None of these patterns, ours included, solves the question at their center: how do you tell an outlier from the only reader who noticed? In a vote, fifteen to one is a clear result. In a novel, the one might be the sensitivity reader who caught what everyone missed. That's where I think the most interesting technical work in subjective evaluation sits.

When it got it right

The Manuscript Editor is where writers lived, and where I learned what "right" means when there's no right answer. A few moments from real sessions, names removed:

A thriller writer asked for a chapter review of a scene where a hardened soldier finally explains herself. The review flagged her big speech:

"[It] reads as [the narrator's] internal monologue transplanted into her dialogue. It's articulate, thematic, and too clean. Real people, especially people trained to hide, don't deliver speeches about their trauma. They hint, deflect, and let others infer."

The writer: "OK, you're right. The Sergeant speech is a bit on the nose and long." That's not a metric. It's a reading of character.

A fantasy novelist was tuning one line in which a father makes a mildly exasperated remark. She'd landed on "remarks in a slightly acid tone." The Manuscript Editor pushed on the one word left:

"The adverb 'acidly' is doing the work now. And here's the craft question: Does Father's tone actually need to be labeled at all? … That said: if you feel the adverb earns its place by adding a specific shade of tone that the dialogue alone doesn't quite capture, that's a legitimate judgment call."

"That's actually a really good catch," she replied, and settled on "dryly," which wasn't what the model argued for. That's fine. Its job was to make the choice visible. She made it.

Later the same novelist got flagged for weak qualifiers ("very," "rather," "a bit," "probably," "almost") and pushed back: "As a general rule, I agree. But these are terms an eleven year old boy would use, so are they still wrong in his narration?" The answer is the best demonstration of many ground truths the system ever produced:

"Yes, these qualifiers are still 'wrong' by craft standards, but they're right for [the narrator's] voice… The graduated certainty ('almost as sure' vs. 'sure') reflects what [the narrator] actually knows versus what [the narrator is] inferring."

It held both truths at once: the rule is real, and this narrator is an exception for a reason you can point to in the text. It also noticed the hedges are calibrated. An eleven-year-old who's "almost as sure" about one thing and "sure" about another is a character with an epistemology. The guidance behind it ends with a line I'd put on the wall: "Sterile is a failure mode, not a target."

The wisdom is the work

The clearest evidence that the editorial layer matters came from a beta build where it was missing. The Manuscript Editor is fed curated craft wisdom, the distilled principles from those editing books, selected for the task at hand. In one beta configuration that wisdom was left out of the response step, and the results were demonstrably worse. When a writer said "I like it the way it is," the model simply dropped a sound note. Our write-up: "without wisdom, the response model has no principled craft basis to defend a flagged issue, so it defaults to capitulation."

I read that as good news: an accidental, clean ablation. Remove the editorial context and quality collapses; restore it and the model can hold a position for a reason. The good answers above aren't the base model being clever. They come from accumulated craft knowledge plus the engineering around it: near-deterministic settings, a stable identity for every flagged issue, an "author intentional" flag that retires a note for good, and a per-book knowledge graph. The model supplies fluency. The wisdom supplies judgment.

What do you show when there is no right?

Partly, a refusal. The Manuscript Editor has an explicit rule against scores: "Numerical scores invite unproductive questions like 'why not an 8?' and shift focus away from craft principles toward arbitrary numbers." A 7 out of 10 for a novel is a precise answer to a question nobody can ask.

Instead we show reasoning: what needs work, why it matters, the craft principle behind it, the exact text, and how you might fix it. The principle is what lets a writer argue back and the model hold its ground. When a disagreement is purely stylistic, the model has a script:

"This is a stylistic call. My case for X was [reason]; your case for Y is also valid. It's your book."

Alongside it: "The anchor is the craft principle, not the author's agreement," and "Intentional rule-breaking is a legitimate craft tool… Their freedom to choose does not require your agreement." That's the closest thing Story Stream has to a theory of quality. Observe without judging. Judge for a stated reason. Change your mind only for a new reason. Name the trade-off, and hand the decision to the one person whose ground truth counts most.

The README says it in one line: "Story Stream is a mirror, not a ghostwriter." I used to read that as a product boundary. Now I think it's the only honest answer to the technical problem. When there are many ground truths, the system's job isn't to pick one. It's to show the writer what's on the page clearly enough that they can pick for themselves.

The questions I'm left with

This isn't only about novels. More of what we ask models to judge looks like this: essays, designs, code review, hiring, anything where experts disagree and all of them have a point. What I'd put to anyone building there, including myself:

  • Can you measure something without the measurement turning into a verdict?
  • If your evaluators disagree, is that noise to be averaged away, or the most valuable thing you have?
  • How would your judge tell an outlier from the only reader who noticed?
  • When your judge overrules a dissent, can the user see that it did?
  • If a style guide is applied perfectly and the result is lifeless, which one was wrong: the output, or the guide?

I spent almost three years trying to get software to tell writers whether their book was good. The most useful version it became was the one that stopped pretending it could, and learned to say instead: here's what I observe, here's what I think it means, and here's the one thing only you can decide.