The Index / Methodology
Version 1.0. Published unchanged alongside the results. Anything decided after scoring begins appears in Amendments at the bottom, with a reason.
Universe. US-headquartered staffing firms, $100M–$1B in annual US staffing revenue.
Source. Staffing Industry Analysts, Largest Staffing Firms in the United States, most recent edition, plus its segment reports. All 150 revenue figures come from this one source. No substitutions and no estimates from any other database — the moment revenue comes from two places, every comparison here is arguable.
Selection. Rank all qualifying firms by revenue within each of six segments. Take the top 25.
An excluded company is replaced by the next firm down the revenue rank in the same segment. Every exclusion and its reason is logged and published.
Scoring "above the fold" is not reproducible. The fold moves with the viewport and two scorers see different amounts of page.
The unit is the first message unit: the h1 — or the visually dominant headline where no h1 exists — plus the first subordinate line of copy directly beneath it, whichever renders first in the DOM. Navigation, eyebrow tags, utility bars and buttons are excluded.
Applied in order, stopping at the first rule that fires:
The generosity rule. Where a reading exists under which the unit names a buyer situation, code it Negative Present — even if a stricter reading would code it otherwise. This runs against the Index's own thesis deliberately. A rule that was strict about what qualifies would manufacture the finding.
Does not count as Negative Present: "challenges" or "complexity" with no situation specified · a question that names no condition · a problem stated as the company's expertise rather than the reader's experience.
Informational / Aspirational / Transformational. Where two apply, code the one carried by the headline, not the subhead.
Don't-Hurt-Me if nothing in the unit could be disputed in good faith. Assertive if at least one claim could be contested on the facts. Aggressive if it attacks, presumes failure, or disparages an alternative. The court test governs: confidence of tone is not assertion.
Scored conservatively — a message earns a level only if it clearly clears it.
| Score | Standard | Anchor |
|---|---|---|
| 0 | Cannot be repeated from memory after one reading | "Partnering with you across the talent lifecycle" |
| 1 | Repeatable, concrete enough to restate to a peer | "We place clinical staff in rural hospitals" |
| 2 | Carries a reason for a VP to spend time | "We fill rural clinical roles in under 30 days" |
| 3 | Contains a number, stake or cost reaching a P&L | "Rural hospitals lose $1.2M a year to unfilled clinical roles" |
| 4 | Distinguishable from three competitors described secondhand | A claim no competitor in the sample could publish unchanged |
Score 4 is assigned last, after the segment is complete, because it can't be judged without the other 24.
Substitute the nearest competitor's name. Pass if the unit contains something only this company could say. The competitor used is recorded.
Single-scorer categorical judgment is the weakest point in a study like this, so it gets tested rather than asserted.
A random 20% subsample — 30 of 150 — is scored independently by a second scorer who has read the coding rules and not the book. Published: percent agreement per measure, Cohen's κ for Position, Level and Band, and mean absolute difference for Upstream.
If a second scorer is unavailable, the subsample is scored twice, blind, at least fourteen days apart, and reported as intra-rater agreement — which is weaker evidence and will be labelled as such rather than quietly presented as reliability.
Threshold. If κ falls below 0.60 on any categorical measure, the coding rule is not clear enough to publish. It gets revised, that measure is rescored across all 150, and the change is recorded as an amendment.
A language model produces a first-pass score and a one-sentence rationale per measure, from the captured text only. Every one of the 150 is then reviewed by a human who accepts or overrides. The override rate is published whatever it is.
The reliability subsample is scored by humans from scratch, with model output withheld.
Using a model for the first pass is efficient. Not disclosing it would be the fastest available way to lose the argument.
This scores homepages. A company may lead with the Negative Present in outbound, paid or sales conversations and not on the homepage. The Index does not measure that and should not be described as measuring a company's messaging overall.
Scores are a snapshot. Homepages change. Every score carries a capture date for this reason.
Position and Level are judgments applied under stated rules, not measurements. The reliability statistics exist so a reader can see how much that judgment moved.
It models a real process — hand-offs up an org chart — but it is scored from text rather than observed. It is a structured estimate of portability, not a measurement of what happened inside any company.
Dean Waye sells advisory work premised on the central finding. That is why past clients are excluded, why the coding rule is generous against the thesis, why the raw data is published, and why this document is dated before scoring begins. None of that removes the interest. It makes the work checkable by someone who assumes the worst.
Anything changed after scoring begins is logged here with a date and a reason. An empty section is a good sign. A section with entries and reasons is still honest. A silently edited methodology is neither.
No amendments recorded.