If the team asks users which design they like instead of watching the task, start by preserving one representative example. It gives the investigation a concrete reference and makes the eventual correction easier to judge.

Find the likely cause

Preference replaces evidence of usability.

Treat this as an explanation to verify against the actual work. Look at the input, the relevant decision, and the final result together. If the evidence does not support this diagnosis, investigate the mismatch before applying a convenient but unrelated fix.

Make the targeted correction

Ask participants to complete a realistic action and observe obstacles.

Small audiences require careful interpretation and often benefit from qualitative evidence alongside modest quantitative comparisons.

Check that the repair worked

Prioritize demonstrated confusion over aesthetic voting alone.

Repeat the check on the final version that the reader or customer will encounter. An approved draft, a preview, and a published result can differ; the acceptance decision should concern the version people actually use.

Prevent the next related failure

A separate issue to watch for is this: a tiny test produces an impressive percentage lift. Report actual counts and avoid strong claims from sparse evidence.

Monitor the useful outcome

Track observed task success, recurring confusion, and outcome counts with explicit uncertainty.

Keep a short record of the original symptom, the evidence behind the diagnosis, and the result of the acceptance check. That record makes the solution reusable when the same condition appears again, without assuming that every superficially similar problem has the same cause.