Verification asks you to say a phrase the server picks, right now. Not a password, not a recording you uploaded earlier — a phrase you couldn't have known a minute ago. Here's exactly what the check does, why it's built this way, and — just as important — what it does not prove.
A verification session is four rounds. In each one the server hands you a random five-word phrase — ordinary words mixed with spoken digits — and you read it aloud. The phrase is single-use and expires after 90 seconds. Nothing you can prepare in advance is any use, because the thing you have to say didn't exist until you asked for it.
A fresh 5-word phrase — words + spoken digits
Within 90 seconds. The phrase is single-use.
Timing, the words, the voice
Then one final check across all four
Each round is judged on three independent things. They fail differently, which is the point — an attack that beats one usually trips another.
A human has to hear the phrase, read it, and speak it. An answer that arrives faster than that is machinery, not a person.
Speech recognition transcribes what you said and requires at least 4 of the 5 words. The tolerance is there because microphones and accents are real; it is not there to let you say something else.
The audio is compared against the voiceprint held for this account. Right words in the wrong voice fails.
The four recordings are checked against each other. This is what catches handing the microphone to someone else halfway through, and stitching a session together out of clips from different sources.
Every practical attack on voice is an attack you prepare. Someone records you off a livestream. Someone trains a clone on your podcast. Both take time, and both produce audio of you saying things the attacker chose.
A server-chosen phrase deletes that advantage. A recording of you cannot know the phrase. A clone built offline cannot know the phrase. To pass, the attacker has to synthesize new speech, in your voice, saying five specific words, inside 90 seconds — live. That's the expensive case, and forcing every attacker into it is the entire purpose of the challenge.
| Attack | What it needs | What the challenge does to it |
|---|---|---|
| Replay a recording | Any audio of you | Dead on arrival — the recording says the wrong words |
| Offline voice clone | Samples of your voice + time | Useless unless it can generate the new phrase live |
| Splice clips together | Words harvested from many recordings | Has to beat the timing check and the same-speaker check across four rounds |
| Pass the mic to a real person | A cooperating human | Voice won't match the account; the cross-round check catches a swap mid-session |
The intuition that "more tests = harder to fake" is wrong here, and it's worth being blunt about why. Repeating the same test doesn't compound. A bot that can solve one round can solve fifteen — it's the same problem, handed to it again. What you'd be buying with round eleven is not security; it's the same wall, one more time.
What rounds do buy is two real things: a more stable score, because averaging across several samples smooths out a cough, a bad microphone moment, or one unlucky clip; and the same-speaker check across rounds, which needs more than one sample to exist at all. Both of those saturate quickly — around four. Past that, the curve is flat and the only thing still going up is how long a real person has to sit there talking to their phone.
The badge attests to a moment, so it's dated like one. After 90 days it lapses and has to be re-earned by running the check again.
This matters more than it looks. Cloning and synthesis get better every year; a check that was hard to beat in one season may not be in the next. If badges were permanent, a single successful break — one clone that got through once — would buy a permanent mark of trust. Expiry means every account has to keep proving it against whatever the check looks like today.
This is the part people are right to interrogate, so here it is plainly.
A biometric you can't rotate is a serious thing to hand anyone, which is why the design keeps the smallest possible artifact and gives you a way to destroy it.
The badge means one specific thing: a live human spoke a phrase we chose, at a moment in time, in a voice matching this account. That's it. It is worth stating what it isn't:
Verification isn't a gate on speech. Nothing is blocked, nothing is removed, and no one is banned for not doing it. What it changes is ranking: verified accounts are boosted in the feed and in comments, and unverified accounts rank below them.
The reasoning is ordinary: attention is finite, and the cheapest way to flood a network is with accounts that cost nothing to make. Making a live-voice check the thing that earns priority puts a real, repeated, human cost in front of anyone who wants scale. It doesn't make abuse impossible — it makes it expensive, which is the honest goal.
Boosted in the feed and in comment threads. Re-earned every 90 days.
Ranks below verified. Still posts, still comments, still fully present — just not boosted.
No ranking effect at all — it's live voice end to end, so the thing verification measures is already happening.
To be precise about what we are not saying: an unverified account is not a bot, and we don't treat it as one. Plenty of real people won't verify, and that's a legitimate choice. This is a ranking preference, stated openly, not an accusation.