How many measurable properties of a brand's writing move when nothing about the brand has changed.
This first release is the method and the metric, plus one dataset of 30 drafts (15 matched pairs). That corpus is written by us, not collected from anybody, so nothing here is a claim about brands in general and the page does not make one.
What this sample cannot tell you
These come first rather than last. A study page that puts its sample size under its finding is choosing which half gets quoted, and this is the half that decides whether the other half means anything.
- The corpus is written by Repic. It is a controlled pair, not a sample of anybody's real publishing, and it cannot support any claim about brands in general.
- n = 30 drafts across 15 pairs. That is enough to describe this corpus and not enough to generalise from.
- The "before" drafts are composites written to match a register observed in ranking pages. They quote no company, which means they are representative by construction rather than by sampling, and that is a real weakness of this release.
- Every measurement is arithmetic over text. It cannot tell you whether a piece was worth publishing, and that is the more common failure.
The metric, defined once
The Brand Drift Index counts, across a set of texts published under one identity, how many of ten measured properties vary beyond a stated tolerance. It is a count out of ten, not a score out of a hundred, because the underlying instrument produces ten independent readings and averaging them into one number would invent a precision the measurement does not have.
A single 0-100 number would be more quotable and less true. Ten properties do not share a unit, so weighting them into one figure is an editorial decision, and the weights would then be doing work the reader cannot see. A count is checkable: you can disagree with it by pointing at a property.
Choosing a subset would be choosing the result, so the list is the engine's and not an editor's. The instrument is free_tools.voice_tool.measure(), which is pure arithmetic over the text: no model, no randomness, same input gives the same reading.
- Sentence count
- Word count
- Average sentence length
- Longest sentence
- Contraction rate
- Whether it addresses the reader
- First-person mentions
- Second-person mentions
- Vocabulary richness
- Presence of stock phrases
What this corpus showed
Announcement pairs: Fifteen announcement moments, each written twice: once in the register the ranking "25 examples" listicles publish, and once obeying a rule specific to that moment. The pair is controlled: same message, same moment, one variable.
Mean 5.07 before, 1.67 after, and it fell in 15 of 15. No rule in the set said "use fewer first-person words"; each said something about the reader.
0 of 15. The instrument flags reader-address when second person is aimed at the person reading. Every standard version failed it, including the moments whose whole purpose is telling a customer something that affects them.
Average sentence length moved 13.71 to 13.83 words and got shorter in only 8 of 15; total length rose, 48.93 to 50.27 words. The two most common pieces of advice for copy that sounds corporate, shorten it and say less, are not what separates these thirty texts.
Each of the fifteen pairs is published with its own numbers on the announcement pages, so every figure above can be checked against a page that shows its working.
Method, in full
- Fifteen announcement moments. For each, a draft in the register the ranking "25 examples" pages publish, and a second obeying one rule specific to that moment. Nothing else differs between the pair.
- Both drafts read by
free_tools.voice_tool.measure(), which returns the ten properties listed above. - A property counts as moved when the two readings differ at all. No threshold, no weighting, no rounding.
- Reproduce it with
seo-audit/capture/announce-measurements.py, which contains every draft and re-runs the measurement.
The aggregates
Every number this page shows, as JSON, including the limitations and the sample counter. It is generated from the same record the page renders, so it cannot disagree with what you have just read.
/data/brand-drift-index-r1.json →Cite it as: Brand Drift Index 0.1, Release 1, Repic, 2026-08-17.
How the visitor sample is collected
Samples come from one place, the free consistency checker, and from nowhere else. The other tools return text built from the visitor's own answers and this schema does not fit them, so they are out of scope rather than included by implication.
One row is kept per completed check, and this is all of it: which tool ran, the ISO week rather than the date, the band the overall score fell into, a band for each of the six checks, how many posts were pasted and roughly how long they were as ranges rather than counts, whether the two optional boxes were filled in at all, and a one-way hash of the removal code. Not the text, not the brand name, not the palette, and no user, account, session or address that could join this row to anything else.
Why bands and not scores
The checker is deterministic arithmetic, so an exact tuple of scores is a fingerprint: anybody who knew which posts a brand had pasted could reproduce it and find the row. Bands make many different inputs land on one row. This reduces re-identification, it does not remove it, and we do not say otherwise. Nothing this study publishes needs more precision than a band, and precision you publish nothing from is precision only an attacker can use.
Why there is no timestamp
A date to the second is close to unique, so storing one beside a field deliberately blurred to the week would make the blurring decorative. The week is the only temporal value held, and it is what the retention window is measured in as well. For the same reason rows carry a random identifier rather than a sequential one: a counter records the order things arrived, which is precision we never publish.
Retention, and the floor
Rows are deleted after 26 weeks. It is stated in weeks and not in days because the week is the only date this store holds, and a promise to delete on a particular day is one a week-granular store cannot keep. The reason for the length is data minimisation and only that: an earlier draft justified a longer window with a year on year comparison the arithmetic did not support, and that claim was withdrawn rather than rounded to fit. No published cell rests on fewer than 30 records, and the collector enforces it rather than the writer remembering it.
What we call this
Pseudonymised, not anonymous. Singling out a row is re-identification even with no name attached, and the honest reading is that a visitor who kept their posts could find their own. So the data is handled as personal data, it is not shared, and there is a way to delete it, immediately below.
Method last read against the code on 18 August 2026 by Saffi. This page describes how software behaves, and the software is changed by people who are not reading this page, so the review is a re-read against the implementation rather than a re-dating. The window and the reason are on our editorial standard, which fails our build when this goes stale.
Remove your row
If you used the consistency checker you were given a removal code with your result. Paste it here and the one row it opens is deleted. You do not need to send us the posts you pasted, and you should not: we never had them.
We store a one-way hash of your code rather than the code, so we can check one you give us and we cannot work yours out. If you have lost it we genuinely cannot find the row for you, and saying otherwise would be the kind of reassurance this page exists to avoid. Everything is deleted after 26 weeks regardless. Releases we have already published are not recalculated when a row goes, because a number somebody has cited should not change underneath them; the row is left out of every release after the removal.
The visitor corpus, and what is in it today
Our free tools score real writing that visitors paste in. Until this release every one of those results was returned and thrown away, so there was no corpus to publish and there never had been one. Collection now exists, and it holds 0 samples.
Related reading: the thirty drafts in full on how to announce it, what drift is on the glossary, the standard this page is held to on our editorial standard, and the checker itself at the brand consistency checker.
This is the first release. The metric definition above is versioned at 0.1 and every future release cites it rather than restating it, so a change to the definition is visible as a version bump rather than as a quietly different number.