QuizCraft / Item analysis
Item analysis calculator
Paste your students' answers and the key. Get the difficulty index, discrimination index and a distractor breakdown for every question, plus the test's KR-20 reliability.
Your data
What is item analysis?
Item analysis is a set of statistics that tells you how each question on a test behaved: how many students got it right, and whether the students who did well overall were the ones who got it right. It comes from classical test theory and is the standard way teachers, exam boards and textbook publishers decide which questions to keep, fix or throw out.
Difficulty index (p)
The proportion of students who answered correctly. A p of 0.85 means 85% got it right, so despite the name, a higher number means an easier question. For a classroom test that ranks students, most items between 0.30 and 0.90 work well; mastery checks can reasonably run higher.
Discrimination index (D)
Take the top 27% and bottom 27% of students by total score. D is the proportion correct in the top group minus the proportion correct in the bottom group. It runs from −1 to +1. Robert Ebel's widely used guideline reads it like this:
| D | Verdict | What to do |
|---|---|---|
| 0.40 and up | Very good | Keep it. |
| 0.30 – 0.39 | Good | Keep; small wording tweaks at most. |
| 0.20 – 0.29 | Fair | Review the distractors. |
| 0.00 – 0.19 | Poor | Rewrite or replace. |
| Below 0 | Negative | Check the key first, then look for a trick or ambiguity. |
See it move: flip answers and watch D change
Ten students, five questions. Students are sorted by total score. Click any cell to flip a right answer (filled) to a wrong one (empty) and watch each question's numbers move.
| Student | Q1 | Q2 | Q3 | Q4 | Q5 | Total |
|---|---|---|---|---|---|---|
| Ava top | 4 | |||||
| Ben top | 4 | |||||
| Chloe top | 3 | |||||
| Dev | 3 | |||||
| Emma | 3 | |||||
| Grace | 3 | |||||
| Finn | 2 | |||||
| Hugo bottom | 2 | |||||
| Isla bottom | 2 | |||||
| Jonah bottom | 1 | |||||
| Difficulty p | 0.9 | 0.5 | 0.4 | 0.4 | 0.5 | |
| Discrimination D | 0.33 | 1.00 | 0.67 | 0.33 | -0.33 |
- Q1: Very easy, good. Almost everyone got it, so it can’t tell students apart. Fine as a warm-up.
- Q2: Moderate, very good. It separates students who know the material from those who don’t.
- Q3: Moderate, very good. It separates students who know the material from those who don’t.
- Q4: Moderate, good. It separates students who know the material from those who don’t.
- Q5: Moderate, negative — check key. Weaker students did better than stronger ones: check the answer key or look for a misleading word.
Distractor analysis
For multiple-choice items, count how many students chose each option. A good distractor attracts some students, mostly from the lower group. A distractor nobody picks is doing no work; replace it with a common misconception. A distractor that pulls lots of top students usually means it is also defensibly correct, or the stem is ambiguous.
Point-biserial (rpb) and KR-20
The point-biserial correlation is a more precise version of D that uses every student instead of just the top and bottom groups; above 0.20 is generally acceptable. KR-20 (Kuder–Richardson Formula 20) estimates the reliability of the whole test from right/wrong items. Classroom tests commonly land between 0.50 and 0.80; short quizzes score lower simply because they have fewer items.
Small classes
With 20–30 students the numbers wobble: a single student can shift D by 0.1 or more. Treat the verdicts as prompts to look at a question, not as rulings, and compare across a couple of classes before retiring an item.