HomeEvidence Protocol

How We Decide What to Publish

By Fırat Mıhcı · Computational linguist · NLP researcher. Protocol HEP-1.0, adopted 3 August 2026.

TL;DR

Every number we publish comes from one written method: fixed queries, logged candidates, live-refetched sources, two independent labelers, and a denominator and date on every count. We built it after auditing our own pages and finding nine kinds of sourcing error. That is the standard behind the tool, which you can try free.

If you are reading this, you are probably deciding whether a number on one of our pages is worth anything. That is the right instinct, and the question is harder here than on most sites, because we sell a product in the category we write about. So the useful thing is not an assurance. It is the method itself, stated precisely enough that you could repeat a sweep against us and see whether the same sources come back.

Where the protocol stands today

Protocol version
HEP-1.0
Registry entries
80
Completed runs
6

Screening flow, first completed run: 210 identified → 210 screened → 200 excluded → 10 included. Run from 26 preregistered queries, captured 3 August 2026.

Exclusions by reason code for the 3 August 2026 run
Exclusion reasonWhat it meansCount
astroturfCoordinated or paid promotional posting78
off-topicMatched a query but made no claim about the subject61
no-methodology-affiliateA ranking or test with no method disclosed behind it27
unverifiable-this-runWould not resolve when fetched on the capture date24
duplicateThe same item reached through more than one query10

Independent coding: two coders agreed on 10 of 10 categorical fields (100%), with 0 items left unresolved.

These figures are read directly from the evidence registry when this page builds, so they move when the registry moves. Pages published before 3 August 2026 predate the protocol and are labelled as such.

Why This Protocol Exists

On 31 July 2026 we audited the sourcing across this site and found nine distinct classes of error spread over roughly 34 pages. That audit is the reason this page exists, and the protocol set out below is what came out of it.

The most damaging class was citation laundering. A figure would appear carrying an institution name, a date and a sample size, and the figure would not be in the source it was attributed to. Those three details are not decoration on a fabricated number. They are the thing that lets it survive a sceptical reading. A bare number gets challenged; a number wearing an institution and a sample size gets repeated.

Repetition is where the real damage happens. One page states a figure, a second cites the first, a third cites the second, and by the fourth the number reads as established with nobody in the chain having opened a primary source. We found that pattern inside our own pages. We also found it in the pages a reader meets first on this subject, including one that cited three community threads whose identifiers turned out to be placeholders. Resolved live, one returned an error and the other two led to unrelated forums on unrelated subjects. Its headline engagement figures had no source at all, and a search engine’s AI summary was already handing them back as though they were the community consensus.

Catching that afterwards is not good enough, because by the time anyone checks, the number has been indexed, quoted and summarised. Everything below is built to stop it before publication rather than to detect it after.

What Happens to the Sources We Do Not Use

Every candidate a documented query returns is logged before any decision is taken about it. That log is what the identified count refers to, and it is the number that makes the rest of the flow mean anything: a page that reports only what it kept is reporting a selection, not a sweep, and the two are indistinguishable from outside unless the discarded pile is counted too.

From there each item is either included or excluded against a recorded reason. Inclusion needs all four of these to hold: the item is dated or reliably datable, it is publicly linkable, it is first-person experience or a test with a disclosed method or a vendor speaking about its own conduct, and it carries enough detail to label at all, which at minimum means naming a tool and an outcome.

Exclusions use a fixed vocabulary rather than free text: undated, unlinkable, astroturf, unverifiable-this-run, duplicate, off-topic and no-methodology-affiliate. Fixed codes are what let the discards be counted and compared between runs instead of being explained away one at a time. Vendor material can be included, because a vendor describing its own conduct is primary evidence about that conduct, but it is labelled self-interested everywhere it is cited.

In that first run, 210 candidates were logged and 200 were excluded. An exclusion rate that high is less a sign of a strict filter than a description of what a commercial search returns in this category. The largest single bucket, 78 items, was coordinated promotional posting, which is never cited, linked or paraphrased regardless of what it says.

How We Verify a Source

Every included item is fetched again, live, during the run that publishes it, and the quote is captured word for word with the date it was captured. If that fetch fails for any reason, the item moves to unverifiable-this-run and does not appear on the page. It is not published on the strength of what we remember it saying six weeks ago.

Two things that look like verification and are not:

A source repeating a claim is not verification of that claim. Ten pages carrying the same figure are ten copies of one claim. If none of the ten opened the primary, ten is worth precisely what one is worth, and the only thing the repetition establishes is that the figure travels well.

Our own earlier work is not verification either. Not an older page on this site, not an internal brief, not something we posted on another platform. Citing our own archive is how a mistake becomes permanent, because from the second citation onward it starts to look externally supported. Our July audit found this happening in both directions at once, which is what convinced us the rule had to be absolute rather than a matter of judgement.

Nothing is registered from a search result snippet, either. A snippet is a search engine’s compression of a page, and it is very often compressing a claim that the page itself attributes to somebody else.

Why Two People Label the Same Evidence

Every included item is labelled twice, independently: which tool it concerns, which detectors it names, what the outcome was, the plan or tier if one is stated, the item’s own date, and its stance. Two coders do that without seeing each other’s labels. Disagreements go to a third resolver, and anything still unsettled after that publishes as coding-ambiguous rather than being quietly dropped, because a disagreement that disappears from the record is just an undisclosed editorial decision.

We publish how often the two agreed, since a labelling scheme nobody checks is one person’s reading with a procedure wrapped around it.

In that first run the two coders agreed on all 10 items, and no resolver was needed. The figure is published so it can be watched as the registry grows, and so that a fall in it later is visible instead of quietly absorbed.

How We Report Numbers

A count means nothing without a denominator and a date window, so both are mandatory and both appear next to the figure rather than in a note at the bottom. The shape we are allowed to publish looks like this: of the 210 candidates that 26 documented queries returned on 3 August 2026, 10 passed screening. Somebody else can re-run that sentence.

What is banned is the population percentage. We do not publish “X% of users” about anything, ours or anybody else’s. There is no sample of users behind a number like that, only whoever happened to post, and a reader has no way to tell whether it came from a survey, an estimate, or nowhere at all. The same rule rules out invented rates and the quiet promotion of a handful of anecdotes into a prevalence.

A single account of something stays a single account of something, and is presented as one. Three people saying the same thing in one thread is three people in one thread, which is occasionally interesting and is never a rate.

What We Publish About Our Own Tool

Community evidence never supports a claim about our own product. Not a review, not a testimonial, not a star rating, not a thread where somebody says it worked for them. That material is outside the protocol entirely when the subject is us, and the reason is obvious enough: it is the easiest evidence in the world for an interested party to accumulate.

Any number about our humanizer or our detector comes from a dated measured run and from nothing else. It is a run we can name, on inputs we can describe, on a date printed beside the figure. When the tooling on either side of a measurement changes, the figure is stale until it is measured again, and we say so and re-run it rather than leave an old number standing.

What those dated runs show is a detector we can stand behind. In our largest false-positive test it stayed clean on 99.8% of 15,542 human-written passages (a 0.2% false-positive rate) and flagged none of the non-native TOEFL essays or native student essays most likely to be wrongly accused. You can run the detector yourself and check what it does with writing you know the origin of.

The same measured-run discipline produces the lab's public research: every study in the Computational Linguistics Lab ships with its method, its code, and a DOI, so any figure we cite from our own work can be traced to a document anyone can open. Independent press coverage of the tool and this work is collected on the press page, each piece linked in full on the outlet’s own site rather than quoted back at you here.

The Other Site That Runs This Protocol

The same owner runs a second site, metegpt.com, on this same protocol. That is where the method was developed and first published, in July 2026. It is adopted here under this site’s own name because it is applied here, by the same author, to this site’s own claims. You should know that before you weigh anything either site says.

Two rules follow from the shared ownership, and neither is optional.

Neither site may cite the other as evidence, corroboration or independent support. Not in prose, not as a registry entry, not as a link positioned so it reads as third-party backing. A claim appearing on both sites has one source behind it and not two. The other site has already deleted a registry entry for precisely this, and the same standard binds here.

First-party measurements do not cross. A figure measured on one site’s tool is that site’s figure and says nothing whatever about the other’s, even though the same person ran both measurements.

The two methodology pages are written separately rather than copied, and this one carries no link to the other. Publishing near-identical method pages on two domains the same person owns would invite a reader to count one method as two, and it is also the pattern that cost this site most of its search visibility in June 2026. Disclosing the arrangement is the point of this section. Running one protocol across two sites while keeping quiet about the shared owner would be the dishonest version of the same setup.

What This Protocol Does Not Fix

Most of this site predates it. The protocol was adopted on 3 August 2026. Pages published before that date were written with claims dated and linked, but with no census sweep, no flow counter and no independent second labelling behind them. They carry a pre-protocol label rather than being quietly upgraded, and they are re-run as they come up for refresh. The registry stands at 80 entries from 6 completed runs, and it grows with every refresh.

Some places cannot be swept cleanly. Parts of the web block automated fetching, and a run that hits that wall can produce a census of a platform containing nothing from that platform without ever announcing the gap. Where it happened, the run records how the material was actually reached, so you can judge the coverage rather than assume it.

Who posts is not who uses. People write about a tool when it failed them far more often than when it worked. The demographics of any given platform are not the demographics of anything else, and a sample of posts is not a sample of people. A count produced by this method describes what was said in a defined place across a defined window, which is a narrower claim than that kind of number usually gets read as.

Everything here goes stale. A capture date sits on every item because sources change, get edited and disappear after we have read them. A verified quote is verified as of a date and not in perpetuity.

What a method does is make a number traceable. You can see where it came from, what was discarded on the way to it, and when somebody last checked. Where we get something wrong, the correction is logged and published rather than the page being silently edited, because a site that quietly fixes its own record is asking you to trust the version you happen to be reading.

This page and the protocol are versioned together. The current version is HEP-1.0; any change to the method is a version bump, and this page is updated in the same change. Fırat Mıhcı is a computational linguist and NLP researcher; his published work on AI-text detection and second-language writing is on ResearchGate. If a claim on this site does not survive the checks described here, write to us at hello@humanizemy.ai and it will be corrected in public.