A TRACES score is Menso's 0–100 summary of one AI user test on one task. It combines six weighted dimensions, each scored 0 to 5: Task Outcome 25%, Response & Recovery 15%, Affective Journey 15%, Comprehension & Mental Model 15%, Effort & Friction 20%, and Safety, Trust & Inclusion 10%. A language model scores five of them from the recorded steps, but Task Outcome stays within half a point of an anchor set by what actually happened. Effort is calculated from Menso's friction score, and the total is always recomputed from the fixed weights. The score is the index. The replay and the reasons behind it are the diagnosis.
Last updated · By Menso's founder, a former UX consultant
The six dimensions
The six TRACES dimensions, their weights and the question each answers
Dimension
Weight
The question it answers
T · Task Outcome
25%
Did this persona finish the job and reach a usable, checkable result?
R · Response & Recovery
15%
When an action reached the product, did it respond visibly, stay stable and offer a way to recover?
A · Affective Journey
15%
How did the persona's feelings change, and did something the product showed cause each change?
C · Comprehension & Mental Model
15%
Could the persona understand the product well enough to predict what would happen next?
E · Effort & Friction
20%
How much avoidable effort did the product cause: repeated attempts, backtracking, confusion, being blocked?
S · Safety, Trust & Inclusion
10%
Did the product show evidence, give control, keep safe boundaries and work for different people?
The weights follow the order of the questions: did the product deliver the job, how much avoidable effort did the journey cost, and then what response, understanding, feelings and trust explain about that outcome. They are version 0.1 weights; the limits further down say how far to trust them.
How the score is produced
The run is recorded step by step: what the AI user saw, did, thought and felt, and what the product showed back.
Evidence is attributed first, by instruction. The judge is told that only what the product caused may move a score, and that a click which never reached the product, a Menso fault or an outside service may not. That is an instruction to the model, not a guarantee: the code only guarantees that a run ending in a Menso failure is not scored.
Task Outcome is anchored to what happened. The run's outcome sets a band, and the model can only choose a score inside it (table below).
Effort is calculated, not judged: effort = 5 − friction ÷ 5, to the nearest half point. The friction score runs from 0 to 25, lower is better. It counts repeated attempts on the same target, backtracking, expressed confusion and being blocked, and it is never lower than the severity of the run's own findings.
A language model scores the other four (Response, Affect, Comprehension, Safety) and writes a reason for each. It is told to cite step numbers and never to invent something the product did not show. AI thinking time does not count toward Response; only the delay between an action reaching the product and visible feedback does. Affect is judged on the path of feelings and what caused them, not on a count of positive and negative labels.
The total is recomputed from the six scores and the fixed weights: TRACES = Σ (score ÷ 5 × weight). The model never writes the total.
Task Outcome bands
The Task Outcome band for each kind of run outcome
What happened in the run
Task Outcome (of 5)
Completed, and the finished product state confirms it
4.5 – 5
Completed, with a gap in the finished result
3.5 – 4.5
A usable partial result, task not completed
2.5 – 3.5
Real progress, task not completed
1.5 – 2.5
Stopped by something the product showed
0.5 – 1.5
No usable step or result
0 – 0.5
Menso takes the first of these that is true, in this order: the AI user reports the task done (the top two rows, split by whether the finished product state confirms it), a usable result is on screen, the run ended blocked or the AI user gave up, the run made progress. Otherwise the run gets the last row.
When a run ends in a Menso failure, it is not scored. In week 38, one of the six board runs ended that way: the cloud browser showed a blank page. It was refunded and left unscored, so it counts against no product. A Menso fault in the middle of a run that goes on to finish does not stop scoring; the judge is told to leave it out.
A worked example
The arithmetic for one real run from the Product Hunt board. The judge wrote that the homepage had no visible pricing link. A later scripted check of the live page, with no model, found Pricing in the top-right menu, which the AI user never opened. The score stays as published, and the note travels with it.
Every run here used Menso's Purchase intent template: the persona Zainab Salazar, a freelance UX consultant, reads the homepage, the pricing page and the FAQ, decides whether she would buy or sign up, and stops before paying. It is a route-specific diagnostic, not a product-wide score.
tiun.
TRACES 83.5/100
Product Hunt week 38 (Sep 14 – 20, 2026), #2 on Product Hunt's board · tested Sep 21, 2026 · one AI user, one run (n=1) · Quality test · 16 of 16 steps · tiun.io
Zainab scrolled the homepage for five steps without finding a pricing link, and at step 6 typed the pricing address herself. At step 8 she reached the Free vs Enterprise table, then read the help centre and the docs FAQ, and at step 16 decided to sign up for the free tier.
TRACES breakdown for tiun.: each dimension's score out of 5, its weight and the points it adds
Dimension
Score (of 5)
Weight
Points
T · Task Outcome
4
25
20
R · Response & Recovery
4.5
15
13.5
A · Affective Journey
4
15
12
C · Comprehension & Mental Model
4
15
12
E · Effort & Friction
4.5
20
18
S · Safety, Trust & Inclusion
4
10
8
TRACES total
100
83.5
Why each dimension got its score
Task Outcome: The person completed the task, though the finished result left a completeness gap. Evidence: Steps 1-16 covered the homepage promise, the full Free vs Enterprise pricing table, and the FAQ, and step 16 recorded a clear decision to sign up for the free tier, but the ending left a completeness gap in how the final verdict was captured.
Response & Recovery: Every scroll, click, and page change from steps 6 through 15 produced an immediate visible result — the pricing page loaded, the Help Center opened, and each FAQ entry appeared — with no failed or ambiguous responses.
Affective Journey: The journey held steady interest and anticipation with one mild annoyance at step 6 over the missing pricing link, which fully recovered into confidence and a positive sign-up decision by step 16.
Comprehension & Mental Model: The pricing table, plan limits, and FAQ structure read predictably and matched expectations, with the only mismatch being the homepage's lack of any visible pricing link forcing a guessed URL at step 6.
Effort & Friction: Calculated, not judged: Menso's friction score for this run was 3 of 25 (lower is better), and effort = 5 − 3 ÷ 5, rounded to the nearest half point.
Safety, Trust & Inclusion: The product presented plain fee figures, GDPR and hosting badges, and consistent FAQ answers with no hidden catches, and the visit stopped before any payment as intended; keyboard and assistive-technology paths were not exercised.
Review note: The actor never opened the top-right menu, where Pricing sits; this is not a claim that the site has no pricing link.
Made one of these products? Write to support@menso.io and we will correct or remove the example as soon as we can.
What a TRACES score cannot tell you
It is not a human usability score, and it does not predict adoption, retention or revenue.
It covers one persona on one route. Another persona or another task can score differently, so compare scores only under the same task and persona, as the Product Hunt board does.
The v0.1 weights are working hypotheses. Menso has not yet published a comparison with real participants on matched tasks; until it does, read a score as a relative signal, not a norm.
Untested paths are not scored as failures: the judge is told not to count accessibility, privacy or safety for or against a product unless the run tested them. Like attribution, that is an instruction the model can miss.
Attribution can fail. In week 38 the board's review found two of the five scored runs blaming the product for clicks that never reached it, and took both out of the trusted ranking.
Simulated users can miss friction that people feel and can make results look more optimistic (Liu et al., Yoon et al.). That is why every Menso result states its route, persona and limits.
TRACES draws its questions from established UX methods: ISO 9241-11, Nielsen's heuristics, the cognitive walkthrough and WCAG 2.2. Those are influences, not a validation of the score.
Questions people ask
Is a TRACES score the same as a SUS score?
No. SUS and UMUX-LITE are questionnaires that people fill in after using a product, and their norms come from people. A TRACES score comes from one AI user's recorded run, so the two numbers are not comparable. TRACES borrows the questions (can this person finish, where do they hesitate, can they recover) and changes the measurement to fit a synthetic trace.
Why doesn't the AI judge score effort?
Because effort should follow from what happened, not from how a model reads it. Effort & Friction is calculated from Menso's friction score for the run, which counts repeated attempts, backtracking, expressed confusion and being blocked. The judge is told never to score it.
Can a product score well without finishing the task?
Not on the biggest dimension. Task Outcome carries 25 points and is anchored to what happened: a run that ends blocked, with no completion reported and no usable result on screen, gets 0.5 to 1.5 of 5 there, whatever the model thinks. The reverse also shows up. A persona can finish by guessing, and then Task Outcome stays high while Comprehension falls.
Who checks the scores?
Customer runs are scored automatically, and nobody reviews them. The week-38 runs on the public Product Hunt board had one more pass: an AI agent re-checked each scored run against its replay, and scripted clicks that use no model re-tested disputed points on the live pages. Its notes are published next to the score and never change it. Pages like this one cite only runs that review kept, and the run cited here was kept with a qualification, shown as its note.
What does n=1 mean on a Menso result?
One AI user, one run. It is a diagnostic for that route, not a statistic about your users. Run another persona or task and the result can differ.