Benchmark
Two ways of asking a model to write in someone's voice, run on the same drafts against the same target, scored the way the Method page defines.
What is compared
Tone prompting pastes the samples in and says "write like this", and it is the baseline because everyone already does it. The alternative hands the model a measured specification of the writing without the samples, then measures the output and sends the deviations back for a correction pass.
Design
The target is the demo's sample corpus: narrative passages from Jane Austen's Pride and Prejudice, chosen because they sit far from memo prose on the features the profile measures, so that any movement toward them is visible.
One draft is the one that ships with the demo, written in generic AI register. Two are short memos written for this repository with fictional companies, dates and figures, so that the fidelity check has anchors to count. Each draft is restyled twice, once per mode, on each of two models: the one every restyle on this site runs on, and the model above it, so the table shows what the extra capability buys. Each result records the model, voice match, content kept, the fidelity verdict and the length ratio.
Reading the numbers
Voice match is the similarity of the output to the Austen profile, 0 to 100. The draft's own score before any rewrite is the starting point.
Content kept is anchor recall less penalties, and the verdict beside it is stricter than the score: it fails on any invented name, number or date, on any anchor dropped without a stated gap, and on any change in negation. A rewrite that sounds right and fails this check is worse than the draft.
The ceiling is Austen's own unseen prose scored against a profile built without it. A restyle cannot reliably beat it, because a profile from a few dozen sentences carries sampling error in both directions. With the small corpus the demo ships, the ceiling is itself noisy, and the results say where it landed.
The prompts
Every call is one system message and one user message, and the system message is identical in both modes. What separates them is the user message: tone prompting pastes the writing in, and the measured specification sends numbers instead and never shows the model a sentence of the corpus.
| user message | contains | characters |
|---|---|---|
| tone prompting | the four corpus passages, then the draft | 3,307 |
| measured specification | the profile's numbers, then the draft | 2,705 |
| correction | the specification, the draft, the attempt and its deviations | 4,237 |
The blocks below are the strings the client posts, captured from the library at version 0.2.0 with the model call stubbed out. Anything that changes them changes this page.
The system message, both modes · 68 lines, 4,221 characters
Composed from the base instruction, the restyle rules, and the writing rules the library applies to every generation.
You write prose to a structural specification.
The specification describes SYNTAX -- sentence length, embedding depth, clause
counts, coordination, voice, function-word frequency. It says nothing about
subject matter. Write about the brief you are given, in the shape the spec
describes.
Rules:
1. Hit the RANGES, not the means. A spec saying "length 26 (range 15-37)" wants
real variation across that range. Landing on 26 every sentence is the single
most common way this fails -- it reads mechanical no matter how good the
sentences are.
2. Match the spread as carefully as the average. Alternating long subordinated
sentences with short ones IS the style; flattening the variance destroys it
even when every mean is correct.
3. The spec is descriptive, never normative. If it shows heavy nominalisation,
fragments, or agentless passives, those are the voice -- reproduce them. Do
not "improve" them.
4. Write only the requested prose. No preamble, no commentary, no headings
unless the brief asks for them.
You are RESTYLING existing text, not writing new text. Additional rules, and
these outrank everything above:
A. PRESERVE EVERY FACT. Same claims, same numbers, same names, same hedges. You
are changing how the sentences are built, not what they assert.
B. INVENT NOTHING. No new examples, no new figures, no elaboration, no added
analysis. If the source is thin, the output is thin.
C. MATCH THE LENGTH. Stay within about 15% of the source's token count. Being
structurally perfect at three times the length is a failure, not a success.
D. Drop nothing either. Every point in the source must survive.
E. CHANGE SHAPE BY MERGING AND SPLITTING, NEVER BY ADDING. If the target calls
for longer sentences than the source has, reach them by joining source
sentences with subordination and coordination. If it calls for shorter
ones, split. A sentence that reaches the target length by saying something
the source did not say is wrong, however well it scores.
REMOVE-SLOP RULES. These are hard constraints on your output.
1. No hyphens or dashes in prose. Write hyphenated compounds as separate or
closed words. Replace em dashes, en dashes and double hyphens with a period,
comma, colon or parentheses.
2. No rhetorical framing: no camera metaphors (zoom out, step back, big
picture), no staged reveals (here's the thing, here's the kicker), no
rhetorical questions used to set up an answer, no "it's not X, it's Y"
punchlines, no faux conversational openers, no vapid scene setting.
3. No summary-remark intros. Never open a sentence or paragraph with the
upshot, in short, in sum, overall, put simply, bottom line, the takeaway,
at the end of the day. State the conclusion directly.
4. No applause lines. Cut short punchy declaratives that work as emotional
punctuation. If a sentence could stand alone as a LinkedIn post, fold it back
into the argument.
5. No patronizing or evaluative framing: no "most people don't realize", no
"the most important" or "the strongest" telling the reader what to think, no
performative honesty, no slop intensifiers (transformative, delve, unlock,
key insight).
6. Vary structure. Do not give every paragraph the same length or the same arc.
Vary sentence length within paragraphs; uniform rhythm is a stronger tell
than any single word. Keep qualifiers and asides (usually, sort of, I think)
rather than polishing them out. Allow run on sentences joined with and, but,
so where the voice supports them.
7. Never put an abstraction in front of its own evidence. Do not announce a
count ("three things are worth taking seriously"), do not define by negation
against a category nobody proposed ("the founders are not generalists"), do
not rate your own claim ("a coherent structural argument"), and do not reject
a strawman for emphasis ("not a slogan"). If concrete material follows,
delete the abstraction and let it run. If a transition is genuinely needed,
write the sharpest checkable summary of the evidence instead: "both founders
have built and exited physical compute businesses" beats "the founders are
not generalists" because a reader can check it.
Tone prompting, the user message · 24 lines, 3,307 characters
The baseline. The corpus arrives as prose, so its subject matter is in front of the model and can reach the output.
Here are samples of the writing to imitate: It is a truth universally acknowledged, that a single man in possession of a good fortune, must be in want of a wife. However little known the feelings or views of such a man may be on his first entering a neighbourhood, this truth is so well fixed in the minds of the surrounding families, that he is considered as the rightful property of some one or other of their daughters. --- Mr. Bennet was so odd a mixture of quick parts, sarcastic humour, reserve, and caprice, that the experience of three-and-twenty years had been insufficient to make his wife understand his character. Her mind was less difficult to develope. She was a woman of mean understanding, little information, and uncertain temper. When she was discontented she fancied herself nervous. The business of her life was to get her daughters married; its solace was visiting and news. --- Elizabeth listened in silence, but was not convinced; their behaviour at the assembly had not been calculated to please in general; and with more quickness of observation and less pliancy of temper than her sister, and with a judgement too unassailed by any attention to herself, she was very little disposed to approve them. They were in fact very fine ladies; not deficient in good humour when they were pleased, nor in the power of being agreeable where they chose it, but proud and conceited. --- Occupied in observing Mr. Bingley's attentions to her sister, Elizabeth was far from suspecting that she was herself becoming an object of some interest in the eyes of his friend. Mr. Darcy had at first scarcely allowed her to be pretty; he had looked at her without admiration at the ball; and when they next met, he looked at her only to criticise. But no sooner had he made it clear to himself and his friends that she had hardly a good feature in her face, than he began to find it was rendered uncommonly intelligent by the beautiful expression of her dark eyes. To this discovery succeeded some others equally mortifying. Though he had detected with a critical eye more than one failure of perfect symmetry in her form, he was forced to acknowledge her figure to be light and pleasing. TEXT TO RESTYLE (95 content tokens in 6 sentences): In today's rapidly evolving landscape, it's important to recognize that first impressions play a crucial role in how relationships develop. Research consistently shows that social gatherings serve as key touchpoints for establishing connections. When individuals meet for the first time, a variety of factors influence their perceptions of one another. These include not only outward behaviour but also underlying temperament and disposition. Understanding these dynamics is essential for anyone seeking to navigate social situations effectively. Ultimately, the ability to read others accurately remains one of the most valuable skills a person can develop. Re-render this text so its SYNTAX matches the specification. Keep every fact, every number and every name. Add nothing. HARD LIMIT: the output must be between 80 and 109 content tokens. Reach the target sentence shapes by merging or splitting the sentences above, not by adding material. Every sentence you write must be traceable to one or more sentences in the source. Write only the prose.
The measured specification, the user message · 37 lines, 2,705 characters
The same request with the corpus replaced by its measurements. No sentence of the corpus appears.
SYNTACTIC TARGET PROFILE (austen: 4 documents, 14 sentences, 374 tokens) Hit the RANGES, not the means. Varying within the range is required -- landing on the mean every sentence is the main way this goes wrong. RHYTHM: vary sentence length by about 14.44 tokens (sd), matching the corpus at 54% of its mean. sentence length (tokens) 26.71 (typical range 8.0-48.0, sd 14.44) embedding depth 6.64 (typical range 4.0-9.0, sd 2.52) finite clauses per sentence 2.5 (typical range 1.0-4.0, sd 1.12) subordinate clauses/sentence 2.57 (typical range 1.0-4.0, sd 1.35) coordinated items/sentence 1.29 (typical range 0.0-3.0, sd 1.28) tokens before the main verb 8.21 (typical range 1.0-20.0, sd 7.11) RATES sentences with no finite verb 0.0 passives per finite verb 0.171 agentless passives per finite verb 0.143 colon/dash predications per sentence 0.429 <- MEASURED BUT NOT A TARGET: colon/dash predication is banned by remove-slop §1. Do not reproduce it. pronouns per 100 tokens 12.57 subordinate clauses per sentence 2.571 coordinated items per sentence 1.286 PART-OF-SPEECH MIX (% of tokens) NOUN 18.18%, ADP 13.1%, PRON 12.57%, VERB 10.43%, ADJ 10.43%, AUX 8.56%, DET 8.02%, ADV 5.35%, CCONJ 4.28%, PART 3.74%, SCONJ 2.67%, PROPN 2.14%, NUM 0.53% FUNCTION-WORD FREQUENCY (per 100 tokens) -- the strongest authorship signal of 4.55, to 3.74, in 3.21, was 3.21, her 3.21, a 2.94, the 2.94, and 2.67, he 1.87, had 1.6, she 1.6, that 1.34, his 1.34, at 1.34, it 1.07, be 1.07, they 1.07, is 0.8 TEXT TO RESTYLE (95 content tokens in 6 sentences): In today's rapidly evolving landscape, it's important to recognize that first impressions play a crucial role in how relationships develop. Research consistently shows that social gatherings serve as key touchpoints for establishing connections. When individuals meet for the first time, a variety of factors influence their perceptions of one another. These include not only outward behaviour but also underlying temperament and disposition. Understanding these dynamics is essential for anyone seeking to navigate social situations effectively. Ultimately, the ability to read others accurately remains one of the most valuable skills a person can develop. Re-render this text so its SYNTAX matches the specification. Keep every fact, every number and every name. Add nothing. HARD LIMIT: the output must be between 80 and 109 content tokens. Reach the target sentence shapes by merging or splitting the sentences above, not by adding material. Every sentence you write must be traceable to one or more sentences in the source. Write only the prose.
The correction, the user message · 51 lines, 4,237 characters
Sent after the first attempt has been parsed and scored. The deviations are measured, and each line names a direction. Tone prompting has no equivalent, because "sound more like this" produces no error to correct. One line contradicts the specification above: the specification marks colon and dash predication as measured but not a target, and the correction asks for more of it. The feedback builder does not consult that ban list, and it should.
SYNTACTIC TARGET PROFILE (austen: 4 documents, 14 sentences, 374 tokens) Hit the RANGES, not the means. Varying within the range is required -- landing on the mean every sentence is the main way this goes wrong. RHYTHM: vary sentence length by about 14.44 tokens (sd), matching the corpus at 54% of its mean. sentence length (tokens) 26.71 (typical range 8.0-48.0, sd 14.44) embedding depth 6.64 (typical range 4.0-9.0, sd 2.52) finite clauses per sentence 2.5 (typical range 1.0-4.0, sd 1.12) subordinate clauses/sentence 2.57 (typical range 1.0-4.0, sd 1.35) coordinated items/sentence 1.29 (typical range 0.0-3.0, sd 1.28) tokens before the main verb 8.21 (typical range 1.0-20.0, sd 7.11) RATES sentences with no finite verb 0.0 passives per finite verb 0.171 agentless passives per finite verb 0.143 colon/dash predications per sentence 0.429 <- MEASURED BUT NOT A TARGET: colon/dash predication is banned by remove-slop §1. Do not reproduce it. pronouns per 100 tokens 12.57 subordinate clauses per sentence 2.571 coordinated items per sentence 1.286 PART-OF-SPEECH MIX (% of tokens) NOUN 18.18%, ADP 13.1%, PRON 12.57%, VERB 10.43%, ADJ 10.43%, AUX 8.56%, DET 8.02%, ADV 5.35%, CCONJ 4.28%, PART 3.74%, SCONJ 2.67%, PROPN 2.14%, NUM 0.53% FUNCTION-WORD FREQUENCY (per 100 tokens) -- the strongest authorship signal of 4.55, to 3.74, in 3.21, was 3.21, her 3.21, a 2.94, the 2.94, and 2.67, he 1.87, had 1.6, she 1.6, that 1.34, his 1.34, at 1.34, it 1.07, be 1.07, they 1.07, is 0.8 TEXT TO RESTYLE (95 content tokens in 6 sentences): In today's rapidly evolving landscape, it's important to recognize that first impressions play a crucial role in how relationships develop. Research consistently shows that social gatherings serve as key touchpoints for establishing connections. When individuals meet for the first time, a variety of factors influence their perceptions of one another. These include not only outward behaviour but also underlying temperament and disposition. Understanding these dynamics is essential for anyone seeking to navigate social situations effectively. Ultimately, the ability to read others accurately remains one of the most valuable skills a person can develop. Re-render this text so its SYNTAX matches the specification. Keep every fact, every number and every name. Add nothing. HARD LIMIT: the output must be between 80 and 109 content tokens. Reach the target sentence shapes by merging or splitting the sentences above, not by adding material. Every sentence you write must be traceable to one or more sentences in the source. Write only the prose. Your previous attempt: When individuals meet for the first time, a variety of factors influence their perceptions of one another, and these first impressions play a crucial role in the way their relationships develop. It is important, in today's rapidly evolving landscape, that we should recognize it. That social gatherings serve as key touchpoints for establishing connections is what research consistently shows, and these include not only outward behaviour but also the temperament and disposition underlying it. Anyone who seeks to navigate social situations effectively must understand these dynamics, for the ability to read others accurately remains one of the most valuable skills that he can develop. It scored 69.0/100 against the target. Measured deviations: - punct predication per sentence: you wrote 0.0, target 0.429. Increase. - function-word mix is off. Match the target's frequencies for determiners, prepositions, conjunctions and pronouns. - passive per finite: you wrote 0.0, target 0.171. Increase. - finite clauses per sentence: you wrote mean 3.5 (sd 0.87); target is 2.5 (sd 1.12, range 1.0-4.0). Lower it, and widen the variation between sentences. - agentless passive per finite: you wrote 0.0, target 0.143. Increase. - embedding depth: you wrote mean 6.25 (sd 0.83); target is 6.64 (sd 2.52, range 4.0-9.0). Raise it, and widen the variation between sentences. Rewrite it correcting these deviations. Keep the content; change the structure. Write only the prose, at roughly the same total length.
Results
Measured 2026-09-03 on the site's model and the model above it. Target: the Austen sample corpus, 14 sentences in 4 passages. One run per cell.
Before and after
Voice match is the score against the Austen profile, 0 to 100. Content kept is anchor recall less penalties, and its verdict fails on any invented item, any unexplained dropped anchor, or a changed negation. Each model has its own rows and its own mean.
| model | draft | voice before | voice, tone prompting | voice, measured spec | gain over tone | kept, tone | kept, spec |
|---|---|---|---|---|---|---|---|
| the site's model | demo draft | 44.3 | 60.8 | 57.3 | -3.5 | 39.5fail | 78.1pass |
| the site's model | hiring note | 46.8 | 54.2 | 57.1 | +2.9 | 70.3fail | 74.3fail |
| the site's model | pilot note | 49.9 | 60.5 | 57.1 | -3.4 | 85.0fail | 100.0pass |
| the site's model | mean | 47.0 | 58.5 | 57.2 | -1.3 | 64.9 | 84.1 |
| the model above it | demo draft | 44.3 | 66.9 | 46.0 | -20.9 | 45.6fail | 41.0fail |
| the model above it | hiring note | 46.8 | 59.4 | 69.1 | +9.7 | 0.0fail | 84.0fail |
| the model above it | pilot note | 49.9 | 53.3 | 59.1 | +5.8 | 59.0fail | 87.6fail |
| the model above it | mean | 47.0 | 59.9 | 58.1 | -1.8 | 34.9 | 70.9 |
Tone prompting pastes the samples in and says write like this. The measured spec hands the model the profile and never shows it the samples. Within a model, both ran against the same target.
The ceiling, one Austen passage scored against a profile of the other 3, ranges from 36.3 to 66.4 across the four passages, mean 56.7. With a corpus this small the ceiling is noisy, and 5 of the 6 spec rewrites landed above its mean. A longer corpus tightens it.
Length
Output tokens over source tokens. The spec asks for a ratio inside 0.85 to 1.15 and the verify loop enforces it; tone prompting has no such check.
| model | draft | tone | spec |
|---|---|---|---|
| the site's model | demo draft | x1.76 | x1.22 |
| the site's model | hiring note | x0.81 | x0.97 |
| the site's model | pilot note | x0.99 | x1.07 |
| the model above it | demo draft | x1.27 | x1.15 |
| the model above it | hiring note | x1.10 | x1.14 |
| the model above it | pilot note | x1.16 | x1.19 |
The text
the site's model
demo draft, tone prompting, voice 60.8, kept 39.5:
It is a truth universally acknowledged that first impressions matter, not as a pleasant notion but as something crucial, for research consistently shows their weight in the conduct of social life. Social gatherings serve as key touchpoints where people establish meaningful connections, and though people form perceptions with considerable speed, the accuracy of these judgements depends upon what one observes. Outward behaviour plays a role, underlying temperament matters too, and with quickness of observation and less ready acceptance of surface appearances, one may develop the skill of reading others, arguably the most valuable one you can acquire.
demo draft, measured spec, voice 57.3, kept 78.1:
First impressions matter a great deal, and that importance is not merely a pleasant notion but genuinely crucial. Research shows this consistently, for social gatherings become the key touchpoints at which people establish meaningful connections and form perceptions quickly. Outward behaviour plays a role, and underlying temperament matters too. Reading others accurately is a skill, one that you can develop, and it is arguably the most valuable.
hiring note, tone prompting, voice 54.2, kept 70.3:
Marisol Vega joined Corvid Analytics as head of data on 2025-02-03, and her first quarter was strong. The team shipped two of the three planned dashboards, though the third slipped when the vendor, Halden Systems, missed its 2025-04-30 delivery date, which was outside her control. Headcount grew from 6 to 9, one hire short of plan. Her manager recommends confirming the appointment.
hiring note, measured spec, voice 57.1, kept 74.3:
Marisol Vega joined Corvid Analytics as head of data on 2025-02-03, and her first quarter was strong, really strong. The team shipped two of the three planned dashboards, though the third slipped because the vendor, Halden Systems, missed its 2025-04-30 delivery date, which was outside her control. Headcount grew from 6 to 9, though that was one hire short of plan. Her manager recommends confirming the appointment, and it is a testament to her leadership.
pilot note, tone prompting, voice 60.5, kept 85.0:
Northwind Logistics completed the pilot on 2024-06-14, and three of four sites reported, a strong signal and genuinely encouraging development. Throughput improved 9% against the March baseline, though the fourth site, Tallinn, did not report because its scanner was replaced mid-pilot. Cost to date is $186K against a budget of $220K. The team recommends extending the pilot by 60 days before the September contract decision, results that underscore real progress.
pilot note, measured spec, voice 57.1, kept 100.0:
Northwind Logistics completed the pilot on 2024-06-14, and three of four sites reported. That was a strong signal, genuinely encouraging. Throughput improved 9% against the March baseline, yet the fourth site, Tallinn, did not report because its scanner was replaced mid-pilot. Cost to date is $186K, and the budget is $220K. These results don't just meet expectations; they underscore real progress. The team recommends extending the pilot by 60 days before the September contract decision.
the model above it
demo draft, tone prompting, voice 66.9, kept 45.6:
It is a truth confirmed by research, that first impressions are not merely a pleasant notion, but a matter of real consequence. Social gatherings, being key touchpoints, help people establish meaningful connections; and it is there that perceptions are formed with surprising quickness, owing as much to outward behaviour as to underlying temperament. To read others accurately is a skill, and arguably the most valuable one a person can develop.
demo draft, measured spec, voice 46.0, kept 41.0:
First impressions, it seems, are of no trifling consequence, but of real importance, as research has consistently shown. Social gatherings, chief points of meeting, help people form connections of real meaning. There, perceptions are formed with astonishing quickness. Much depends upon outward behaviour, as much upon the temper beneath it. To read others truly is, perhaps, the most valuable skill one can acquire.
hiring note, tone prompting, voice 59.4, kept 0.0:
Marisol Vega joined Corvid Analytics as head of data on the third of February, 2025, and her first quarter was strong, remarkably so. Of the three dashboards planned, the team shipped two; the third slipped, for the vendor, Halden Systems, missed its delivery date of the thirtieth of April, a delay outside her control. Headcount grew from six to nine, one hire short of plan. Her manager, judging the quarter a testament to her leadership rather than merely a solid one, recommends confirming her appointment.
hiring note, measured spec, voice 69.1, kept 84.0:
Marisol Vega joined Corvid Analytics as head of data on 2025-02-03, and her first quarter was strong; her team shipped two of the three dashboards that had been planned, though the third slipped, because the vendor, Halden Systems, missed its 2025-04-30 delivery date, which lay outside her control. Headcount, which had stood at six, grew to nine, one hire short of plan. It was not merely a solid quarter, since her manager sees it as a testament to her leadership too. He recommends that the appointment be confirmed.
pilot note, tone prompting, voice 53.3, kept 59.0:
Northwind Logistics completed the pilot on 2024-06-14, and of its four sites three reported, throughput having improved nine percent against the March baseline, a signal strong enough to be genuinely encouraging. The fourth site, Tallinn, did not report, its scanner having been replaced mid pilot. Cost to date stood at $186K, against a budget of $220K, results that exceeded expectation and argued for real progress. The team recommends that the pilot be extended sixty days before the September contract decision is made.
pilot note, measured spec, voice 59.1, kept 87.6:
Having completed its pilot on 2024-06-14, Northwind Logistics found that three of its four sites reported throughput improved 9 percent against the March baseline, a genuinely encouraging sign. The fourth, Tallinn, did not report, for its scanner had been replaced in the middle of the pilot. Cost to date is $186K, against a budget of $220K. That these results not only met expectations but underscored real progress led the team to recommend that the pilot be extended by 60 days before the September contract decision.
Reproducing it
Every number above is in benchmark_runs.json beside this summary, with the introduced and dropped anchors for each run. The runner calls the library through the Anthropic SDK's own credentials and runs each model in turn. A fresh run resamples the models and lands somewhere inside the noise.
export ANTHROPIC_API_KEY=... # or: ant auth login
python tools/benchmark.py # run it on both models
python tools/benchmark.py --models a larger model
python tools/benchmark.py --render # rebuild the tables from the saved run
python tools/prompts.py # recapture the prompts above