A walk through one night of research — what runs on its own, what gets checked, and how a tier actually changes. The live figures below — last night's scan, tonight's rotation slot, and this cycle's tier moves — are rebuilt every time the site redeploys, from the same files that build the rest of the index.
21 practices rescanned · 124 evidence items added · 0 tier changes
A night rescans the practices that are due for another look, up to 54.
Every entry in this index is the output of a nightly research run: software that searches, reads, scores, and writes, checked by a second layer before anything reaches the site. This page walks through what that run does, start to finish. For what the six tiers and five trends mean, see Methodology — this page is about the machinery that decides them.
Clean, or approved, lands automatically. If review flags something, a further check can land anyway, pull one item, ask for one redo, or — rarely — escalate to a person.
A run starts at 00:05 London time and moves through eleven checkpointed steps in order — interrupted partway, it resumes from the last completed step instead of starting over. Open picks the batch: up to 54 practices, stalest-scanned first, from whichever domain the 14-day rotation has due. Prep and scan then run once per practice, independently, before history bullets land on every practice that gained evidence, and summary, picks, and header write up the domain as a whole. Review checks the entire night's work; land merges it into the live site, publish sends that night's brief out, and close writes the run report and tears down after it.
Two different kinds of automation are doing this work. Some steps are ordinary code — arithmetic, deduplication, schema validation, the parts with one right answer. Others are a model making a judgment call: is this page evidence, does this evidence clear the bar for promotion, does this summary hold up. Nothing pauses for a person by default; a person only enters the picture if the review layer specifically asks for one, which is rare (see "Checking the work," below).
What runs each step:
| STAGE | MODEL / VENDOR | WHAT IT DOES |
|---|---|---|
| Query writing | DeepSeek | Writes each practice's six search queries a night — three chasing adoption, three chasing trouble. Claude Haiku steps in if DeepSeek doesn't answer. |
| Search | Perplexity + Firecrawl | Runs every query and returns candidate pages. Run together, not as a fallback pair: Perplexity's date filter finds what's in the scan window; Firecrawl finds relevant pages alongside it that Perplexity's own search missed. |
| Page fetching | Direct HTTP, then Firecrawl | Retrieves the full text of every candidate page. PDFs are read locally, with OCR for scanned pages. |
| First-pass triage | TypeSafe's Jev | Decides whether a fetched page is worth keeping, and why — for most pages, this is the only judgment a page gets. A typed-question model, not a Claude or DeepSeek model: it answers a fixed set of yes/no and multiple-choice questions about one page and returns probabilities, never prose. |
| Careful escalation read | DeepSeek | Takes a second, full look at the pages the first pass couldn't confidently place, and writes the explanation and summary attached to every surviving page. Claude Haiku steps in if DeepSeek doesn't answer. The same call also writes up every page the first pass kept, whether or not it was escalated. |
| Reader: evidence and practice text | Claude Haiku | Chooses which written-up candidates become permanent evidence on the practice, and rewrites the practice's Overview and Current Landscape sections to match. |
| Tier/trend decision | Claude Sonnet | Builds the qualitative case for or against a tier promotion, steel-mans its own conclusion, and sets the practice's trend. |
| History writer | Claude Sonnet | Appends this month's dated bullet to the practice's tier history. |
| Domain summary | Claude Opus | Writes the domain's recap for this scan. |
| Picks | Claude Sonnet | Chooses the ten evidence items the domain page and brief lead with; every link is checked before it ships. |
| Brief | Claude Opus | Writes the newsletter brief for the domain whose turn it is to publish. |
| Header art: concept | Claude Opus | Reads the finished brief and names one visual subject for the header image. |
| Header art: image | OpenAI | Renders the header PNG from that subject, wrapped in a fixed style block that comes from editorial policy, never from the model. |
| Review | Claude Opus | Reads a sample of that night's tier and trend decisions, looking for a call that doesn't hold up against its own cited evidence. |
| Approver | Claude Fable | Runs only when the review flags something: decides whether that night publishes anyway, is corrected, gets one item pulled, or goes to a person. If its call fails, Claude Opus decides instead. |
Each practice gets six search queries a night — three hunting for adoption, three for trouble. Every result is de-duplicated by URL in code and dated from what the page itself says, not guessed by a model — a model is good at reading a claim and bad at date arithmetic, so dates stay in code before and after a page is judged. A page nobody can date is still kept and judged, not discarded.
Not every fetched page becomes evidence. A typed-question model decides most fetched pages outright, kept or dropped, from a fixed set of yes/no and multiple-choice answers. A second model writes the explanation and summary for every page that survives, and makes the call itself on the pages the first model leaves undecided.
Every kept item is scored for source credibility, not just relevance. Peer-reviewed research, standards documentation, and engineering blogs with real deployment detail score highest. Vendor material — even a strong case study — is capped below independent sources: a vendor describing its own product isn't a neutral witness, and self-reporting is useful context, not proof. When evidence points two ways, both directions are kept on file — at least one negative or critical signal survives selection whenever one exists, specifically so an all-positive set can't make a promotion look inevitable.
The confidence floor for each kind of source:
| BAND | SOURCE KIND | FLOOR |
|---|---|---|
| HIGH | engineering blog of a well-known company, with deployment detail shopify.engineering, blog.google, engineering.atspotify.com, netflixtechblog.com | 0.8 |
| HIGH | peer-reviewed paper or preprint arxiv.org, aclanthology.org | 0.8 |
| HIGH | official product documentation docs.anthropic.com, docs.github.com | 0.85 |
| HIGH | government or standards body nist.gov, ico.org.uk | 0.9 |
| MEDIUM | industry analyst report gartner.com, forrester.com, mckinsey.com | 0.65 |
| MEDIUM | reputable tech journalism techcrunch.com, theverge.com, arstechnica.com | 0.6 |
| MEDIUM | conference talk (named speaker, named venue) youtube.com (conference channels), infoq.com | 0.65 |
| MEDIUM | company or vendor blog (a business writing about its field or its own product) datadoghq.com/blog, cast.ai/blog, any vendor's /blog or /resources | 0.55 |
| MEDIUM | practitioner blog with specific technical detail dev.to, a named engineer's own site | 0.5 |
| LOW | personal blog or newsletter medium.com/@*, substack.com | 0.4 |
| LOW | social media twitter.com, linkedin.com | 0.3 |
| LOW | press release or marketing prnewswire.com, businesswire.com | 0.35 |
| LOW | unknown, regional or content-farm site a site you cannot place, a regional reseller, an SEO article page | 0.45 |
CURRENT CYCLE · Sep 28–Oct 11 · tonight is day 5 of 14
gold edge = a paired night (5 of the 14); the newsletter's lead alternates, so each publishes every four weeks.
Scanning all 19 domains every night would mean a shallow once-over instead of a real dig into a handful of practices. Instead, the rotation works through fourteen slots on a two-week cycle: most nights cover one domain, five nights cover two related ones. There's no memory of where the cycle last stopped — the slot for any date is just that date's position in the 14-day loop, fully reproducible from the calendar alone.
On a night that covers two domains, only one gets that night's newsletter brief; the other's turn comes four weeks later. A domain covered alone publishes every time its slot comes up. Either way, every domain is rescanned every fourteen days.
qualitative case first, numeric gate second — every step below has to say yes
The default outcome, reached whenever new evidence is missing, the defining question isn't answered yes, or the steel-man doesn't come out weak — a false promotion costs more than a delayed one. When a strong qualitative case exists but the citation-and-gate check itself comes back no, that's recorded as a gap for next cycle rather than promoted anyway.
A tier promotion isn't a number crossing a threshold on its own. The model first builds a qualitative case: state the next tier's defining question, answer it citing specific evidence by title, then argue the strongest case against its own conclusion — a steel-man. Only if that steel-man comes out weak does the promotion reach a numeric gate at all: at least two citations that match a real evidence item's title, a minimum amount of evidence for that tier, and the right mix of evidence types. A practice moves at most one tier per scan, never skipping one.
When the case doesn't hold up, nothing happens, and that's deliberate: the default is not to promote. A wrongly-early promotion is treated as worse than a delayed one.
A concrete example of where the bar sits: evidence from a single vendor can support a Leading Edge classification, but not Good Practice or above — a single-vendor case that technically clears a higher tier's numeric gate gets flagged for a person, not promoted automatically.
Demotion works differently on purpose: it's never automatic. If the evidence behind a practice's current tier stops holding up, that's an editorial call, not a script. Trend — a practice's momentum toward its next tier — is a separate, related question, judged from the same cited evidence: a change is written only once two scans running have agreed on it (a promotion sets Accelerating at once). See Methodology for the five values.
The numeric gate for each tier. "X OR Y" means at least one of the two is present; "N+ X" means at least that many items of that type; "any" means no specific type is required. Where a tier lists more than one clause, every clause must be met.
| TIER | MIN. EVIDENCE | REQUIRES (ALL CLAUSES) |
|---|---|---|
| Invisible | 5 | adoption-metric OR 3+ case-study |
| Established | 5 | adoption-metric OR 3+ case-study |
| Good Practice | 3 | case-study OR product-gaproduct-ga OR industry-report |
| Leading Edge | 2 | case-study OR significant-repo |
| Bleeding Edge | 1 | any |
| Research | 1 | any |
Before any of a night's work reaches the live site, it passes a review layer with two parts. First, deterministic checks that need no judgment at all: is every file valid against its schema, does every date make sense, has an already-published evidence item been quietly edited (never allowed), does a promotion actually clear its numeric gate. A single failed check decides that run's outcome regardless of what any model says about it.
Second, a sampled model review reads a handful of that night's decisions — practices with no tier change, and practices with a reconfirmed or newly-pending trend — looking for the kind of problem a checklist can't catch: does this tier decision actually hold up against its cited evidence, does this trend match the evidence's arc, does anything look invented. The sample is picked by a formula seeded on the date, reproducible after the fact but not predictable in advance.
If that review comes back clean, the night publishes on its own. If it flags something, a further, independent check — reading the validator report, the review, and the finished brief together — decides whether to publish anyway, correct one item, ask for a narrow rewrite of the write-up (never the underlying research), or put the decision in front of a person. That last option is rare, reserved for things like a validator failure nothing already explains, a specific brief claim resting on one thin source, or a pattern across many practices that looks systematic. Every decision is logged, whichever way it goes.
A one-off audit on September 27, 2026 checked every specific claim in the index, meaning a named organization together with a figure or a dated event, against the evidence on file. Of 26,092 such claims, 25,301 matched an evidence item automatically. The other 791, plus 7 flagged on a close read, were each checked against their original source. 637 were accurate and already backed by evidence on file. 83 misstated their source and were corrected to match it. 38 were accurate but had never been given a source; one was found and filed for each. 16 had lost their source in an earlier rewrite, and it was restored from the page's history (one of these was also corrected). 25 had no source anywhere, even after a second search, and were removed.
If you're deciding whether to trust a number on this site: every tier and trend claim traces back to named, dated, linked evidence on that practice's own page, and the mechanism above is what produced it — not editorial opinion applied after the fact.
The same facts are also available as structured data: a machine-readable summary at /how-it-works.json, a plain-text mirror at /how-it-works.md, and citation metadata embedded directly in this page. /llms.txt points here and at the rest of the site. The State of Play's data and text are CC BY 4.0 — reuse them, including for training, with attribution and a link back to thestateofplay.ai.
See Methodology for the full tier and trend definitions, the license, and who's behind this.