SUS Calculator
System Usability Scale scores, with the confidence interval.
What the System Usability Scale is
The System Usability Scale (SUS) is a ten-item questionnaire, each item answered on a five-point agreement scale from strongly disagree to strongly agree. John Brooke published it in 1996 as a quick, technology-agnostic way to measure perceived usability, and it has been in continuous use since — largely because it is short enough that participants actually finish it, and because there is now a huge body of published scores to compare a result against. The items alternate positive and negative wording on purpose, to stop a participant from agreeing their way down the page without reading each one.
How it is scored
Odd-numbered items are positively worded, so they score (response − 1). Even-numbered items are negatively worded, so they are reversed and score (5 − response). Summing all ten gives a number from 0 to 40, which is multiplied by 2.5 to land on a familiar 0–100 scale.
A SUS score is not a percentage, even though it looks like one. It is the result of that reversal-and-sum arithmetic, benchmarked against other studies — not a proportion of questions answered correctly. 68 is the average score across a large body of published research, not a bare pass mark. A system scoring 68 is performing exactly as well as a typical system, no better and no worse.
Why the confidence interval matters more than the mean
A mean of 72 from eight participants looks precise. It is not. With that few responses, the true average opinion of the wider user population this sample was drawn from might reasonably sit anywhere from the mid-sixties to the high seventies — and the mean alone gives no hint of that range. Reporting “we scored 72” states a single number with a confidence the data does not support and invites everyone downstream to treat it as exact. Reporting the interval instead — “72, likely somewhere between 64 and 80” — carries the same evidence honestly. That is why this tool puts the interval at the centre of the results, not tucked away as a footnote to the mean.
How to use the benchmarks
The percentile and letter grade shown alongside the mean come from the Sauro–Lewis curved grading scale, built from several hundred published usability studies. They tell you how a score compares to that wider body of research, not against some fixed notion of perfection. A grade of C is genuinely average — the middle of the distribution — rather than a failing mark, and a D or an F is a real signal that the system is underperforming most of what has been studied, worth investigating rather than dismissing.
Frequently asked questions
How many participants do I need?
Twelve to fourteen participants typically narrows the 95% confidence interval to roughly ±10 points, tight enough to report a mean with some confidence. Five participants is enough to find usability problems in a qualitative session — the rule of thumb behind most moderated usability testing — but not enough to measure satisfaction with any precision. Those are different jobs, and this tool is built for the second one.
Can I change the wording of the items?
You can, but the benchmarks no longer apply if you do. The percentile and grade are calibrated on the original ten items, in their original wording; changing wording, adding items or dropping them produces a score that is no longer comparable to the published distribution, even though the arithmetic still runs.
What if someone skips an item?
Discard that response. There is no defensible way to impute a missing SUS item — the scale was validated on complete responses, and guessing a value one way or another will bias the score in a direction you cannot know in advance.