At a glance
Three things to know in three seconds.
The question
When researchers adapt a validated trust scale by rewording, translating, or reformatting it, does it still measure what it was validated to measure?
The method
A systematic review of how the scale is actually used across automation studies, documenting every kind of modification against the original instrument.
The stake
Cross-study comparability. Every meta-analysis, design guideline, and safety decision quietly stands on it.
The problem
Measurement drift is easy to overlook.
Trust in automation is one of the most-studied constructs in human factors, and the
strength of all that research depends on how trust is measured. In practice,
scales get adapted to fit a study's context, language, or format. Those changes
feel harmless. They are not always neutral:
validated
"I can trust the system."
reworded
"I can rely on the system." Is that trust, or reliance? Related, but not the same thing.
reformatted
Same words, but a 7-point scale becomes 5. The responses no longer line up with earlier studies.
recontextualized
"I can trust the autopilot." A general instrument has quietly become a specific one.
Each drifted study still cites the same validated source. Comparisons weaken, validity
is assumed to carry over when it may not, and the interpretation gets quietly less precise.
My approach
Read everything. Track every change.
As a graduate research assistant, my work was the systematic part of the review:
- Reviewed studies using the trust scale across automation contexts.
- Documented wording changes, translation differences, and response-format adaptations against the original validated instrument.
- Compared usage patterns to see which kinds of modification recur, and which go unacknowledged.
- Worked toward a structured synthesis of what those modifications mean for methodological strength.
It is slow, unglamorous work. It is also the kind that decides whether a whole
literature can be trusted.
What I found
Validated instruments carry design logic.
The lesson reaches past trust itself. When you change an instrument's design logic
casually, the meaning of the data changes with it. I saw this three ways:
Wording
Small wording changes shift the nuance, and with it what participants are actually rating.
Format
Response-scale changes alter response behavior and break comparability between studies.
Context
Domain-specific adaptation blurs what the instrument was originally validated to measure.
Why it matters
Use the scale. But know what you changed.
The practical takeaway: adapting a scale is sometimes necessary, but it should be
deliberate, documented, and flagged whenever results get compared across studies.
And the question underneath it, are we measuring the right thing in the right
way?, reaches well past trust scales. It is the same judgment that decides
whether a usability metric, a survey, or a safety evaluation means what its authors say.
Reflection
Why this sits at the center of my portfolio.
This project says the most about how I work: careful, methodical, and interested in how
people deal with systems that keep automating more of the job. The same thread runs
through the rest of my work. It is why my emergency-alert concept ends with a
validation study instead of a mockup, and why the water chatbot's hardest problem
turned out to be trust, not technology.
measurement rigor
critical literature review
automation trust
research synthesis