Single Ease Question (SEQ)

The Single Ease Question (SEQ), also known as the Task Ease Question, is a post-task attitudinal survey looking at the ease of completing a task. It’s particularly useful as a diagnostic measure, complementing larger post-scenario surveys. It’s simple to administer and score, and easy for participants to answer, so it can be a useful tool for comparing perceived usability before and after a design change, and for diagnosing UX problems in larger workflows.

This article explains how to use the Single Ease Question in the context of UX research.

seq question

How to formulate the Single Ease Question?

The SEQ is usually presented as a statement about task difficulty, such as “Overall, this task was…”, or “Overall, how difficult or easy was this task to complete?”, or just simply “How easy or difficult was this task?”; and then asking participants to select the answer on a 5-point or a 7-point scale ranging from Very Difficult to Very Easy.

A Single Ease Question (SEQ) with a 7-point scale.

1. Overall, this task was...

1 2 3 4 5 6 7
Very Difficult Very Easy

In the 2022 article Evaluation of Three SEQ Variants, Sauro and Lewis reported on an experiment to check if respondents will be biased to provide more optimistic answers if “Very Easy” was on the left side of the scales. Although top-box scores were somewhat affected for difficult tasks, the means were not statistically different, so the “data from either format are likely comparable”. In the same article, they argue for adding numbers on each item of the scale, and not just labelling the ends, as this helps respondents differentiate between options more easily, and the participants strongly preferred that option (more than two to one).

In 2023, Lewis and Sauro also proposed a variant where options are labelled with adjectives. They asked participants to rate the same activities on both the SEQ and a 6-point adjective scale labelled Most difficult imaginable, Very difficult, Difficult, Easy, Very Easy and Easiest Imaginable, then used the results to map SEQ score ranges onto those adjectives. For a quick reference, see the section on interpreting results.

seq question meaning

SEQ name and meaning

The SEQ acronym is a registered trademark in the US. The registration covers downloadable software for measuring user experience, not the research method, so it does not restrict asking the question or publishing SEQ scores. Single Ease Question and Task Ease Question are the generic names for the same type of survey.

To add to the confusion, the acronym SEQ often has other common usages in research, including “Social Experience Questionnaire”, “Student Engagement Questionnaire”, “Self-Efficacy Questionnaire” and others.

Origins of the Single Ease Question method

Donna Tedesco and Tom Tullis compared five ways of collecting post-task ratings at Fidelity Investments in 2006, including two variants single-ease questions, though they never used that name. In A Comparison of Methods for Eliciting Post-Task Subjective Ratings in Usability Testing they report that the most consistent one at small sample sizes is “Overall, this task was…” with a 5-point scale, and ends labelled “Very Difficult” to “Very Easy”.

Jeff Sauro and Joseph Dumas published a paper in 2009 comparing three one-question post-task usability questionnaires, concluding that the 7-point scale version was statistically as sensitive as the more elaborate SMEQ, and that both were more sensitive than Usability Magnitude Estimation. It correlated strongly with task time and moderately with a post-test SUS questionnaire. They did not use the SEQ name in the paper, but instead presented it as a “variant of the Likert scale found most reliable” by Tedesco and Tullis, directly establishing the lineage. Notably, in the variant, Sauro and Dumas switch the labels so that “Very Easy” is the start of the scale, and “Very Difficult” is at the end.

The label “Single Ease Question” was first coined by Jeff Sauro in the 2010 MeasuringU blog post If You Could Only Ask One Question, Use This One., referring to the 7-point scale and crediting it to joint work with Joe Dumas from the prior two years, so the name likely emerged somewhere between 2008 and 2010. In the 2012 web article 10 Things To Know About The Single Ease Question (SEQ) Sauro suggests switching the endpoint labels back to the order used by Tedesco and Tullis, with “Very Difficult” first.

Sauro’s company Measuring Usability LLC holds the US trademark for the acronym SEQ, and Sauro is the main driver behind the industry adoption of this questionnaire, including compiling benchmarks and normative tables.

Tracing the intellectual lineage from Tedesco and Tullis, and Sauro’s references, SEQ was influenced by Jim Lewis’s three-item post-scenario questionnaire at IBM (The After-Scenario Questionnaire ASQ, 1991, Psychometric Evaluation of an After-Scenario Questionnaire for Computer Usability Studies: The ASQ). ASQ similarly uses a 7 point scale but it has 3 questions instead of one. One of the three is “Overall, I am satisfied with the ease of completing the tasks in this scenario”, which is a direct inspiration for SEQ. Fred Zijlstra’s scale, known in usability research as the Subjective Mental Effort Questionnaire (SMEQ), presented this measurement in a single question, but on a 150 point scale. Zijlstra himself called it the Rating Scale Mental Effort, and describes how he built it in Efficiency in Work Behaviour: A Design Approach for Modern Tools: rather than picking labels by intuition, he had respondents rate a set of effort descriptions and placed each one on the scale at the geometric mean of their ratings, which is why the labels sit at uneven intervals. Mick McGee published a similar single-question method at Oracle in 2004 (Master Usability Scaling: Magnitude Estimation and Master Scaling Applied to Usability Measurement), called Usability Magnitude Estimation, where respondents would invent their own scale. Sauro worked for Oracle at the time when the 2009 paper with Dumas was published, and that paper tested the 7-point question head to head against both SMEQ and UME.

seq survey

How many people are needed for a Single Ease Question Survey?

Since SEQ is an example of a Likert Scale, the same general principles for calculating sample sizes apply.

Based on Comparison of Three One-Question, Post-Task Usability Questionnaires, A Comparison of Methods for Eliciting Post-Task Subjective Ratings in Usability Testing and Nielsen Norman Group’s practitioner guidance, 10–12 is the floor for detecting any difference at all, and 20–30 is a minimum for relevant results. Below 10 participants, Sauro and Dumas found that post-task ratings “may be unable to reliably” separate two products even when the difference in usability is large. For a margin of error that can be used to compare against benchmarks, surveys require more than 100 participants.

Two independent studies of similar research methods agree on the same lower threshold. Albert and Dixon concluded in Is This What You Expected? The Use of Expectation Measures in Usability Testing that “It is unlikely that fewer than 10 or 12 participants will produce statistically reliable findings”, and Tullis and Stetson A Comparison of Questionnaires for Assessing Website Usability sub-sampled a 123-person study concluding that “sample sizes of at least 12-14 participants are needed to get reasonably reliable results”.

seq score

Interpreting Single Ease Scores

Although the scores mathematically range from 1 to 7, the score of 5.5 should be interpreted as average, not good, and a small gain at the top end is worth far more than the same gain at the bottom. Raw scores can be easily misinterpreted, and it’s better to rank against benchmark percentiles.

Across more than 400 tasks and 10,000 users, Sauro reports that the average SEQ score on a 7-point scale hovers between 5.3 and 5.6. In separate research on that database, he finds a strong correlation between a SEQ score and the likelihood that a user will successfully complete a task.

The SEQ and completion rates have a strong correlation (r = .66). This means that perception of ease (SEQ scores) can explain about 44% of task completion rates. … A raw SEQ score of 4.7 will correspond to a completion rate of 58% and task time of 2.8 minutes. A raw SEQ score of 5.9 will correspond to a completion rate of 86% and task time of about 2 minutes.

– Jeff Sauro, Using Task Ease (SEQ) to Predict Completion Rates and Times

The relationship between task completion and SEQ is not linear, as the same research shows that completion rates level off at about 90% once SEQ scores pass the 70th percentile (around 5.9), and at about 50% once they drop below the 30th (around 4.7). The relationship is only close to linear between those two points.

In Describing SEQ Scores with Adjectives, Lewis and Sauro present a mapping from SEQ scores to adjective interpretation:

AdjectiveLowMeanHigh
Most difficult imaginable1.001.001.49
Very difficult1.501.932.69
Difficult2.703.534.29
Easy4.305.095.59
Very easy5.606.146.49
Easiest imaginable6.506.827.00

In the historical percentiles published in the same source, a score of about 5.5 (“Easy”) sits around the 50th percentile, 5.9 at the 70th, and 4.7 at the 30th. So the scale compresses significantly at the top end. Note that French-Lazovik and Gibson point to the possibility that commonly used labels on SEQ could be part of the reason for this clustering. Their research includes an experiment where changing labels on two performance rating questionnaires impacted the means by a quarter to a half of a scale point. French-Lazovik and Gibson concluded that commonly used labels are more negative than what most people running SEQ queries assume, which causes the answers to group towards the high end.

There are no academically evaluated methods specific to interpreting SEQ results, so it’s best to follow the general advice on interpreting rating scales. Report the mean and the top-box percentage (the share of participants who answered 7), each with a confidence interval. Show the distribution as well if you can. The intervals keep a single number from being read as more precise than it is, and the box score provides context to resolve the clustering.

In Quantifying the User Experience, Sauro and Lewis suggest computing the mean and standard deviation of the responses, then use the t-distribution. Use the interval to report the precision. A mean of 5.68 from twelve participants, with a typical spread, spans everything from below average to beyond the 99th percentile.

Compute the standard deviation from your own data. The publicly available research figures differ significantly, so it’s best not to rely on them. Sauro and Lewis aggregated 465 seven-point items and got an average of 1.46, from items ranging between 0.68 and 2.02. Christophersen and Konradt measured 1.90 on a single seven-point usability item in Reliability, Validity, and Sensitivity of a Single-Item Measure of Online Store Usability, while Del Grande and Kaczorowski assumed 1 when powering the trial in Rating versus ranking in a Delphi survey: a randomized controlled trial. Gregor Burger, Jože Guna and Matevž Pogačnik report on thirteen tasks in Suitability of Inexpensive Eye-Tracking Device for User Experience Evaluations with SEQ standard deviations 0.62 to 1.99.

For a seven-point ease item, Sauro and Lewis prefer the top-box over the top-two box reporting, since “measurements of extreme responses tend to be better predictors of future behavior than tepid responses”. Bottom-box scores are usually not relevant for SEQ as almost nobody picks 1. With small samples, the recommendation for the interval for the top-box percentage from Quantifying the User Experience is adjusted-Wald binomial interval rather than the t-distribution.

Check the distribution of the data first before applying this general advice, because the mean is the wrong thing to report for two types of distributions. The first is where answers cluster at both ends (some people complete the task easily and some struggle) and nobody actually votes around the average value. Sullivan and Artino warn about that in Analyzing and Interpreting Data From Likert-Type Scales: “if responses are clustered at the high and low extremes, the mean may appear to be the neutral or middle response, but this may not fairly characterize the data”. The second is where most responses pick 7, so a mean above 6 already produces an interval running past the end of the scale. In such cases it’s best to let the distribution and the top-box percentage carry the result, and treat the mean as supporting detail.

seq ux

Using SEQ in UX Research

In Comparison of Three One-Question, Post-Task Usability Questionnaires, Sauro and Dumas compared the seven-point variant (which still was not called SEQ) to SMEQ and Usability Magnitude Estimation, and concluded that it was easy for participants to use, easy for researchers to set up and administer in electronic form, and that the survey was easy to score. They also note that the “Likert question” was highly correlated with the other measures they investigated, and that despite being a single question requiring no explanation, it was statistically “as sensitive as SMEQ”. In addition, Sauro suggests that the comparative benefit of SEQ is that it’s technology agnostic (“We use the SEQ on mobile devices, websites, consumer and business software and even tasks on paper prototypes”).

Pierce and colleagues used the SEQ while testing new features for clinical software, and followed what happened to those features in production. For one feature, the SEQ scores “improved from 2.7 to 4.2 or greater”, while task time fell from 287.2 seconds to 140.8 and 59.8 seconds and “Total errors dropped from 9 to 3.” A different feature was delivered regardless of SEQ scores of 2.5 and 2.8, and “it was subsequently withdrawn after firing 2,473 times in” six months, which the team estimate cost users around 50 hours on an “ineffective feature with poor usability.” A third feature went the other way: testing found “generally low completion rates, relatively high error rates, but good SEQ scores”, traced to a calendar icon that was invisible without scrolling at lower screen resolutions. The researchers conclude that the SEQ score was simple and easily obtained, and that in the withdrawn feature’s case it “appears to have portended the withdrawal of that feature from production.” Two important notes for this research: there were six participants per test event, well below what’s recommended in practice, and the paper reports no standard deviations, confidence intervals or significance tests for any SEQ score. These are illustrations of the method in use, not evidence about its reliability.

Page Laubheimer suggests SEQ for testing usability of individual UI components, but pairing it with qualitative research:

If you’re specifically interested in the usability of the individual components of the UI, use the Single Ease Question after each task and ask users to explain their score.

Page Laubheimer, Beyond the NPS: Measure Perceived Usability with the SUS, NASA-TLX, and the Single Ease Question

Reviewing seven different usability questionnaires in the context of health IT evaluation, Alissa L. Russ-Jara, Jason J. Saleem and Jennifer Herout in A practical guide to usability questionnaires that evaluate clinicians’ perceptions of health information technology list the key strength of SEQ as “Very rapid”, but note that because it consists of a single item with limited scope, it is “likely more appropriate as a formative, rather than summative, usability satisfaction score.”

Single Ease Question is best used as a post-task question during each important step of a UX test, to provide additional diagnostic information that post-test questionnaires would not be able to identify. Post-task questions can help pinpoint individual steps or tasks that cause frustration. To avoid distracting the users during a longer test session, post-task surveys need to be quick and easy, both to answer and to score, and Single Ease Question seems to fit those criteria nicely.

When diagnosing issues in a longer workflow, ask the participants to fill in SEQ surveys immediately after each task, not once at the end of a group. Schuessler, Fischer and Walpuski found in Investigating Construct Validity of Cognitive Load Measurement Using Single-Item Subjective Rating Scales that a single difficulty rating after a set of items tracks the hardest item in the set rather than the average, and they “advise to precisely examine at which point and how frequently cognitive load is measured”.

All the general advice that applies to post-task questions is also relevant for SEQ, and is covered in the separate article on post-task surveys. What is specific to the SEQ is the budget it leaves you: because it asks only one question, researchers have a bit more space for follow-up questions than more complex post-task surveys allow, and it’s very useful to ask participants to briefly explain their score. This can provide additional contextual information and a very narrow focus, particularly when working with a smaller sample. Sauro suggests setting a threshold to diagnose tasks that are more difficult than average:

Ask Why? : When users rate a task difficult, it’s good to know why they did. When a user provides a rating of less than 5 we ask them to briefly describe why they found the task difficult. This provides immediate diagnostics information right when the user is cognizant of what is driving the poor rating.

Jeff Sauro, 10 Things To Know About The Single Ease Question (SEQ)

The Votito task ease template does exactly that. It is a two-question survey. Participants provide the rating in the first question, and an optional second question shows for people who rate the task below 5, asking to explain the score.

SEQ can be useful to compare the difficulty of completing the same task before and after a change. If you do not have a baseline already, then use it as part of the Expectation Ratings method, which follows a similar approach but also compares actual outcomes to perceived difficulty. Some tasks are inherently harder than others, and SEQ cannot differentiate between those. An Expectation Ratings survey can provide the additional context.

Applicability and limitations

Because Sauro and his company Measuring Usability LLC are the main driving force behind industry usage of SEQ, most practical advice about it traces back to them. The benchmarks, percentile tables, adjective mappings and format experiments are almost entirely the work of one company, MeasuringU, which sells a research platform with the SEQ built in and holds the trademark on the acronym. None of that work is peer reviewed.

Most of the limitations that apply to post-task questionnaires apply to SEQ: a rating shows how hard a task was, without explaining why it was hard (Mapping the Landscape of Standardized Usability Questionnaires in Healthcare: A 26-Year Scoping Review). Ratings from a small group of participants will be far less reliable than behavioural measures collected from the same people (An Empirical Comparison of Lab and Remote Usability Testing of Web Sites).

Since SEQ is a single question, it provides a single score that reflects a combination of effectiveness, efficiency and satisfaction, without being able to track how those dimensions move independently of each other A single question has to collapse all three into one number, and Cairns argues in A Commentary on Short Questionnaires for Assessing Usability that any such collapse “must make simplifying compromises” while nobody publishing a short scale ever says what those compromises are. Russ-Jara and colleagues evaluated seven questionnaires against a list of usability attributes, they found the SEQ assesses only one attribute of usability (ease of use), against eight of ten attributes for the CSUQ/PSSUQ and three for the SUS. They also point out that it’s not possible to calculate Cronbach alpha for SEQ, because internal consistency cannot be computed for a one-item instrument. More complex instruments allow for measuring consistency.

A potential problem with asking just a single question is that people sometimes answer something different. Cairns suggests that a participant asked how usable something was may in fact be answering how much they enjoyed the task (A Commentary on Short Questionnaires for Assessing Usability).

For SEQ in particular, it’s worth noting that the original premise is that it behaves “about as well or better” than more complicated measures, but the research shows mixed results. Sauro reported in Comparison of Three One-Question, Post-Task Usability Questionnaires that “SMEQ held a slight advantage over Likert” although not significantly, and that SEQ correlated better than the SMEQ with task time and errors, and worse with SUS scores and with completion rates.

Since SEQ is a single question asked at the end of the task, it has limited applicability to complex multi-step tasks. Hassenzahl and Sandweg challenge the idea that a single question can measure a complex task in From Mental Effort to Perceived Usability: Transforming Experiences into Summary Assessments, concluding that summary assessments of perceived usability do not reflect a “whole experiential episode, but rather its most recent” incidents. A complex task that was mostly easy but ended badly will score like a bad task.

SEQ measures perceived ease of use, not actual usefulness, so it may not be a good predictor of usage. Davis suggests that ease-of-use relationship with “usage all but vanishes when usefulness is controlled” in Perceived Usefulness, Perceived Ease of Use, and User Acceptance of Information Technology. Usefulness drives usage, not ease of use, and although ease of use is a component driving usefulness, it does not replace it. A task can score 6.5 on SEQ and still be something that people do not want to do, or prefer not doing, and SEQ has no way to register that.

Run a Task Ease Survey

Run a Task Ease Survey

Set up and run a task ease survey in minutes, and get a meaningful interpretation of the results. Your users answer without creating an account, and it's free to start.

Run a Task Ease Survey

Get notified about new articles & major updates
Share