One set of writing principles does not serve all genres.
Pick what you are writing, and this says which technique fits it — and which will quietly damage it. An incident notification and a research proposal are not the same job. STE100 makes prose unambiguous and strips the hedging that honest uncertainty requires. The Amazon memo forces a decision and resists being skimmed. No technique wins everywhere, so the useful question is never what is good writing but good for which reader, doing what.
Then it helps you do the writing. The review prompt packages these same judgements into questions you can take to any model: it works out what you are writing, asks only what changes the answer, and interrogates your draft against the technique that fits it. It does not write the prose. That is the point — a method you can apply, not an output you have to trust.
How to read it. Purpose and technique never meet directly. One grid scores what each genre demands; the other scores what each technique does to a draft; both hang off the same ten considerations in the middle. A recommendation is those two joined — so you can disagree with half of it without discarding the other, and every cell stays one claim about one thing.
Three models scored every cell independently and the median is what is published. Nothing was rounded toward agreement: where they genuinely split, the cell says so rather than averaging the disagreement away.
The genre is the thing you are writing. It stays chosen while you move between the three views, and it is what the selection rule needs. Leave it unset to read the matrices on their own.
What each genre requires, before any tool is chosen.
Scale 0–3: 0 irrelevant, 1 minor, 2 important, 3 critical. The chosen genre's column is outlined.
What each technique does to that quality.
Scale −2 to +2: +2 serves it strongly, 0 neutral, −2 actively damages it. Click any cell marked ! to see the three votes behind it.
! the three models differ by two points or more — the median is not a consensus and should not be read as one. † one model declined to score the cell. hatched no value is published: the recorded votes permit none, and any number printed would be one no model assigned.
Match a genre's column in A against the technique columns in B.
That is the published rule, and the point of it is to argue with the cells rather than trust a recommendation.
Pick a genre above.
Nothing is ranked until you say what you are writing — the demand column is half of the calculation.
This ranking is arithmetic added here, not a finding of the research. It weights each technique's supply score by how much the chosen genre demands that consideration, and sums. The source document prescribes matching the columns; it does not prescribe this or any other formula. Read it as a way into the two matrices, and argue with the cells rather than the total.
Every technique, ranked in every genre.
The same arithmetic run ten times. Read along a row to see where a technique earns its place and where it does not. 1 is the best fit for that genre, 10 the worst — a rank is a position, not a score, so a 1 in a weak field is still only first among ten. Rows are ordered by mean rank across all ten genres, shown in the last column, and techniques that tie on score share a rank rather than being separated by an arbitrary tiebreak.
What each technique is, and where to read it.
Ten techniques, in the order they rank on the composite. The description says what the technique covers; the reference says where it is defined. A link appears only where the URL could be checked. Everything else carries its citation in full, which is what a bibliography is for — a reference you can look up beats a link nobody verified.
Every technique was examined in dhk/alexandria ↗; the row links through to the investigation that covered it.
The same standard, as something you can hand to a model.
The matrices are a judgement about mechanism. This turns that judgement into a review procedure: what to hold the draft to, in what order to ask, and — the part that matters — what the reviewer is forbidden to do. Copy it into whatever model you already use. Nothing here is sent anywhere; the prompt is built in your browser and goes no further unless you take it.
The refusals are the point. A reviewer that rewrites your draft has answered a question you did not ask, and one that scores it has replaced a judgement with a number. Both are in the prompt as prohibitions rather than preferences, because that is the difference between a review and a wash of plausible edits.
The study, how it was tested, and what it does not establish.
Everything above is a judgement about mechanism, produced by a method you can inspect. Here is that method, the control that was run against it, and the limits it does not cross — in that order, because a reader deciding how much weight to give the grid needs all three.
The research behind the grid.
In dhk/alexandria ↗, the open research corpus. The tables above are 05-analysis/matrices.md ↗; the per-model scores are scores.csv ↗ and the reasoning is analysis.md ↗.
Canonical run
r-2026-0813-03, 13 Aug 2026 — 26 claims, $0.97.
Control run
r-2026-0814-01 ↗
— same question, framing removed, $0.85. Assurance level: silver.openai/gpt-5.4,
anthropic/claude-opus-4.7 and x-ai/grok-4.5.
Graded by anthropic/claude-sonnet-4.6. Every cell above is the
median of the three; no model's number was overridden.claude-writing-skills measured against them ↗
— literature synthesis, then tool evaluation.Is ASD-STE100 the right tool for clear, concise, persuasive writing? ↗ — run
r-2026-0813-02, 34 claims, $2.16. Found STE is not a general-purpose writing standard and is
actively harmful for persuasion, because it removes by design the devices
persuasion depends on. That finding is carried into Matrix B rather than
re-derived.What is solid here, and what to check yourself.
Solid: the scores. Three models reasoned over the same material independently, a fourth graded them, every cell is a median that was never rounded toward agreement, disagreements are published as disagreements, and a control run tested how much the framing was doing. That is more scrutiny than most opinion in this space gets.
Check yourself: the citations. Web search was on during the run, so references to standards and papers came from the models' own retrieval and were not audited against primary sources afterwards. For several techniques the models worked from publicly described principles rather than the specification itself — which is fine for judging what a technique optimises for, and not good enough to quote. If you are citing ASD-STE100 or ISO 24495-1, go to the standard.
Deliberately out of scope: externally-facing communication — customer notices, status pages, regulatory filings — where legal and disclosure constraints need their own treatment rather than a row in a table; and performance and promotion writing, where norms are too institution-specific to transfer.
Read the grid as an informed lens with its provenance attached, not as a settled result. Everything it rests on is public in dhk/alexandria ↗ under MIT, including the per-model votes for every contested cell and both readings of the calibration row.
The same question, asked twice, to find out how much the asking mattered.
A brief that tells models what to look for tends to get told it back. The grader noticed the three models agreed more than it expected, and said so: the source materials may have been constraining the conclusions. So the whole thing was run a second time as a control — same ten genres, same ten considerations, same matrix — with the brief's hints stripped out. No mention of the tensions it had flagged, no statement of what the earlier investigations had found.
If a finding appears in both runs, the models found it in the material. If it appears only in the first, it came from the question.
Most of it held. Across 200 cells, 133 came out identical and 57 moved by a single point. Ten moved by two or more — and three of those ten are the same row, the row the brief had argued about hardest.
Told that commitment versus calibration was a live tension, three models scored these techniques as damaging calibration. Not told, the same three scored them as helping it. Both readings are published here because neither has a claim to be the true one — if you cite a calibration score from this work, cite the pair.
What survived the strip-out is the more important half.
Without being handed it, all three models independently read the
humanize skill's instruction to force claims that commit as damaging
to calibrated writing — off the skill's own text — and all three landed again on
incident notification as the genre no technique serves.
The defect is in the material. The severity was in the brief.
That is a distinction most research never gets to make about itself.
Outcome evidence.
No technique here has been shown to improve organizational decision quality, trust calibration, incident recurrence, or the durability of reasoning. Every score is a judgement about mechanism, and the strongest claim in Matrix B — Amazon's prose-over-bullets — remains unproven outside the institution that practises it.
Nor is it established that these ten considerations are the right or complete set. They come from the brief's framing, and no run tested whether different ones would change the rankings. The matrices are built on proxies. That is the same Goodhart problem all three models raised about readability scores, turned on the instrument itself.