A difficulty label is a promise. When a question is tagged Hard, you're being told something specific: most people who see this will get it wrong, and getting it right is worth feeling good about. The fastest way to break trust in a quiz game is to break that promise — to call something Hard because the author found it tricky, when really half the world knows it cold. So in Quizzer, nobody decides a question is hard. The players do, one answer at a time, and the game just keeps score. Here's how.
01“Hard” is a promise, not a sticker
The difficulty you see on a question card is not a tag someone typed in at 1 a.m. It's a measurement — a question is Hard because a lot of capable players have actually gotten it wrong, not because it looked intimidating to the person who wrote it. That distinction is the whole philosophy of how we keep the game fair.
A question doesn't get to claim it's hard. It has to prove it, against real players, over and over.
02Every question has its own rating
You already carry a hidden rating — an ELO number that climbs when you win and drops when you lose. The twist is that every question has one too. A question isn't a fixed obstacle; it's a competitor. When you answer it, the two of you play a tiny one-on-one match, and one of you wins. Every new question walks in at 1600 — the same starting line as a new player — then plays thousands of little matches and drifts to wherever it belongs.
Crucially, the math accounts for who is answering. When a 1900-rated player misses a question, that's a far louder signal than when a beginner misses it, and the question's rating jumps accordingly. A question that even strong players keep failing must really be hard; one that beginners keep acing must really be easy. The math sorts this out on its own.
And there's one deliberate asymmetry worth stating plainly: only the question's rating moves. Your personal ELO is decided by the match against your human opponent — never by individual questions. A brutal question can't quietly tank your rating, and an easy one can't pad it. You're an unpaid, anonymous, very large panel of difficulty judges, and you never even feel it happening.
03Watch a question find its level
Drag the slider to set how often players get a question right. As the share rises, the question is clearly easier, its rating slides down, and the badge changes to match. This is exactly how the live system settles every card.
A coin-flip question sits right in the middle of the pack.
04“Hard” is a reading off the dial
Once a question has a rating, its label falls out automatically. There's no separate “difficulty” field anyone fills in — we just read the live ELO and translate it into one of three fixed bands:
“Hard” isn't a hunch; it's a measurement, the same way “92°F” is a measurement. And it's why the label can be trusted to change. A question an author was sure was fiendish goes live at 1600 — Medium — and then a surprising number of people get it right; its rating sinks past 1450 and the badge quietly flips to Easy, because that's the truth now. The reverse happens too: a gentle-looking question with a sneaky trap climbs above 1750 and earns its Hard badge honestly.
There's no “provisional” gate where a question is benched until it's been answered a thousand times. Every answer counts from the very first one. A new question starts at the most honest guess we have — dead average — and the crowd corrects it from there, quickly at first and then more gently as the rating settles. The author's opinion got exactly one vote; everyone else's answers cast the rest.
05The part most quizzes get wrong: difficulty never picks your questions
Here's the line we hold carefully. A question's rating is used to describe difficulty. It is never used to choose your questions. The game doesn't notice you're on a hot streak and start feeding you 1800-rated monsters to cut you down. It doesn't hand the weaker player softballs.
Question ELO is a thermometer, not a thermostat. It reads the temperature — it doesn't crank the heat.
That separation is the whole fairness argument. Both players in a match get the same questions, in the same order, and difficulty is something we observe afterward, not a lever we pull during the game. If difficulty secretly steered selection, “Hard” would stop meaning “hard for everyone” and start meaning “hard for you, right now, because you were winning” — and the label would be a lie again. So we keep the measurement and the matchmaking strictly apart.
06Before a question ever earns a rating: the gate
All of this assumes the question is worth answering in the first place. Self-correcting difficulty is wonderful, but it can't fix a question that's flat-out wrong, ambiguous, or has two defensible answers — ELO would just dutifully rate a broken question. So there's a human gate in front of the whole system.
- Player submissions land as pending. A question sent from inside the app is never auto-published, no matter who sent it. It waits in a review queue.
- A reviewer approves or rejects — with a reason. Only a signed-in admin can move a question through the gate, and a rejection comes back with a note rather than a silent drop.
- Team-authored questions go in pre-approved, because they've already been written to standard — but they meet the same bar.
We also lean on AI to help draft questions — but never to wave them through. Generated candidates run a verification pass before a human ever sees them: current-events questions get fact-checked against live web search so a stale answer can't slip in, evergreen ones get a second model-driven review, and obvious duplicates are caught before they're created. Anything ambiguous, outdated, miscategorized, or lacking exactly one clearly-correct answer is thrown out. The AI fills the funnel; a person guards the gate.
07Why we bother
Stack it up and the fairness comes from three plain ideas working together:
- A human gate keeps broken questions out.
- A per-question rating lets the crowd, not the author, decide how hard each one truly is.
- A firewall between difficulty and selection means a label describes reality instead of manipulating it.
The payoff is small but it's the entire point: when Quizzer tells you a question is Hard and you get it right anyway, the badge means exactly what it says. Thousands of people genuinely struggled with that one, the rating proves it, and the win is yours to keep.
Quizzer never hand-labels difficulty. Every question carries its own ELO-style rating — start 1600, K-factor 32 — and plays a tiny match against each person who answers it: a correct answer pushes its rating down, a wrong one pushes it up, weighted by how strong the answerer is. Only the question's rating moves, never yours. That live number reads straight into three fixed bands — Easy below 1450, Medium 1450–1750, Hard above 1750 — so the label is a measurement that can change as the crowd corrects it. Difficulty only ever describes a question; it never selects one, so both players always face the same questions in the same order. And before any of that, a human gate (with AI only drafting, never approving) keeps broken questions out. “Hard” is measured, not invented.