Alexandria · organizational writing by genre

One set of writing principles does not serve all genres.

Pick what you are writing, and this says which technique fits it — and which will quietly damage it. An incident notification and a research proposal are not the same job. STE100 makes prose unambiguous and strips the hedging that honest uncertainty requires. The Amazon memo forces a decision and resists being skimmed. No technique wins everywhere, so the useful question is never what is good writing but good for which reader, doing what.

Then it helps you do the writing. The review prompt packages these same judgements into questions you can take to any model: it works out what you are writing, asks only what changes the answer, and interrogates your draft against the technique that fits it. It does not write the prose. That is the point — a method you can apply, not an output you have to trust.

How to read it. Purpose and technique never meet directly. One grid scores what each genre demands; the other scores what each technique does to a draft; both hang off the same ten considerations in the middle. A recommendation is those two joined — so you can disagree with half of it without discarding the other, and every cell stays one claim about one thing.

Three models scored every cell independently and the median is what is published. Nothing was rounded toward agreement: where they genuinely split, the cell says so rather than averaging the disagreement away.

Choose a genre, then a view

The genre is the thing you are writing. It stays chosen while you move between the three views, and it is what the selection rule needs. Leave it unset to read the matrices on their own.

What each genre requires, before any tool is chosen.

Scale 0–3: 0 irrelevant, 1 minor, 2 important, 3 critical. The chosen genre's column is outlined.

About this work

The study, how it was tested, and what it does not establish.

Everything above is a judgement about mechanism, produced by a method you can inspect. Here is that method, the control that was run against it, and the limits it does not cross — in that order, because a reader deciding how much weight to give the grid needs all three.

The study

The research behind the grid.

This investigation
Which writing technique for which organizational genre ↗
In dhk/alexandria ↗, the open research corpus. The tables above are 05-analysis/matrices.md ↗; the per-model scores are scores.csv ↗ and the reasoning is analysis.md ↗.

Canonical run r-2026-0813-03, 13 Aug 2026 — 26 claims, $0.97. Control run r-2026-0814-01 ↗ — same question, framing removed, $0.85. Assurance level: silver.
Who scored it
Answered independently by openai/gpt-5.4, anthropic/claude-opus-4.7 and x-ai/grok-4.5. Graded by anthropic/claude-sonnet-4.6. Every cell above is the median of the three; no model's number was overridden.
The two before it
Written-communication best practices, and claude-writing-skills measured against them ↗ — literature synthesis, then tool evaluation.

Is ASD-STE100 the right tool for clear, concise, persuasive writing? ↗ — run r-2026-0813-02, 34 claims, $2.16. Found STE is not a general-purpose writing standard and is actively harmful for persuasion, because it removes by design the devices persuasion depends on. That finding is carried into Matrix B rather than re-derived.
The series in prose
In Defense of Mechanical Writing ↗ and More on Mechanical Writing ↗, DHK On Data and AI. The instrument built on this research is dhk/praxis ↗.
The ten techniques
ASD-STE100 · BLUF and the US Army writing standard · Minto's Pyramid Principle · the Amazon six-page narrative memo and PR-FAQ · ISO 24495-1:2023 plain language · blameless postmortem convention as practised in SRE culture · classical and Toulmin argumentation · readability optimisation as embodied by the Hemingway Editor · Hotaling, S. (2020), "Simple rules for concise scientific writing", Limnology and Oceanography Letters 5:379–383 · and the three skills in github.com/nonatofabio/claude-writing-skills ↗.

What is solid here, and what to check yourself.

Solid: the scores. Three models reasoned over the same material independently, a fourth graded them, every cell is a median that was never rounded toward agreement, disagreements are published as disagreements, and a control run tested how much the framing was doing. That is more scrutiny than most opinion in this space gets.

Check yourself: the citations. Web search was on during the run, so references to standards and papers came from the models' own retrieval and were not audited against primary sources afterwards. For several techniques the models worked from publicly described principles rather than the specification itself — which is fine for judging what a technique optimises for, and not good enough to quote. If you are citing ASD-STE100 or ISO 24495-1, go to the standard.

Deliberately out of scope: externally-facing communication — customer notices, status pages, regulatory filings — where legal and disclosure constraints need their own treatment rather than a row in a table; and performance and promotion writing, where norms are too institution-specific to transfer.

Read the grid as an informed lens with its provenance attached, not as a settled result. Everything it rests on is public in dhk/alexandria ↗ under MIT, including the per-model votes for every contested cell and both readings of the calibration row.

The method, and its control

The same question, asked twice, to find out how much the asking mattered.

A brief that tells models what to look for tends to get told it back. The grader noticed the three models agreed more than it expected, and said so: the source materials may have been constraining the conclusions. So the whole thing was run a second time as a control — same ten genres, same ten considerations, same matrix — with the brief's hints stripped out. No mention of the tensions it had flagged, no statement of what the earlier investigations had found.

If a finding appears in both runs, the models found it in the material. If it appears only in the first, it came from the question.

Most of it held. Across 200 cells, 133 came out identical and 57 moved by a single point. Ten moved by two or more — and three of those ten are the same row, the row the brief had argued about hardest.

Confidence calibrationTension namedHints removed
ASD-STE100−2+1
BLUF / Army−1+1
The repo's skills−1+1

Told that commitment versus calibration was a live tension, three models scored these techniques as damaging calibration. Not told, the same three scored them as helping it. Both readings are published here because neither has a claim to be the true one — if you cite a calibration score from this work, cite the pair.

What survived the strip-out is the more important half. Without being handed it, all three models independently read the humanize skill's instruction to force claims that commit as damaging to calibrated writing — off the skill's own text — and all three landed again on incident notification as the genre no technique serves. The defect is in the material. The severity was in the brief. That is a distinction most research never gets to make about itself.

The evidence, and its limits

Outcome evidence.

No technique here has been shown to improve organizational decision quality, trust calibration, incident recurrence, or the durability of reasoning. Every score is a judgement about mechanism, and the strongest claim in Matrix B — Amazon's prose-over-bullets — remains unproven outside the institution that practises it.

Nor is it established that these ten considerations are the right or complete set. They come from the brief's framing, and no run tested whether different ones would change the rankings. The matrices are built on proxies. That is the same Goodhart problem all three models raised about readability scores, turned on the instrument itself.