Flagship research · graduate assistantship

Trust in Automation

Hundreds of studies measure whether people trust automated systems. A lot of them lean on the same validated trust scale, then quietly change it. My research asks what those changes do to the evidence, and where our ability to compare studies quietly breaks down.

Type: research Focus: measurement validity Role: graduate research assistant
At a glance

Three things to know in three seconds.

The question

When researchers adapt a validated trust scale by rewording, translating, or reformatting it, does it still measure what it was validated to measure?

The method

A systematic review of how the scale is actually used across automation studies, documenting every kind of modification against the original instrument.

The stake

Cross-study comparability. Every meta-analysis, design guideline, and safety decision quietly stands on it.

The problem

Measurement drift is easy to overlook.

Trust in automation is one of the most-studied constructs in human factors, and the strength of all that research depends on how trust is measured. In practice, scales get adapted to fit a study's context, language, or format. Those changes feel harmless. They are not always neutral:

validated "I can trust the system."
reworded "I can rely on the system." Is that trust, or reliance? Related, but not the same thing.
reformatted Same words, but a 7-point scale becomes 5. The responses no longer line up with earlier studies.
recontextualized "I can trust the autopilot." A general instrument has quietly become a specific one.

Each drifted study still cites the same validated source. Comparisons weaken, validity is assumed to carry over when it may not, and the interpretation gets quietly less precise.

My approach

Read everything. Track every change.

As a graduate research assistant, my work was the systematic part of the review:

  • Reviewed studies using the trust scale across automation contexts.
  • Documented wording changes, translation differences, and response-format adaptations against the original validated instrument.
  • Compared usage patterns to see which kinds of modification recur, and which go unacknowledged.
  • Worked toward a structured synthesis of what those modifications mean for methodological strength.

It is slow, unglamorous work. It is also the kind that decides whether a whole literature can be trusted.

What I found

Validated instruments carry design logic.

The lesson reaches past trust itself. When you change an instrument's design logic casually, the meaning of the data changes with it. I saw this three ways:

Wording

Small wording changes shift the nuance, and with it what participants are actually rating.

Format

Response-scale changes alter response behavior and break comparability between studies.

Context

Domain-specific adaptation blurs what the instrument was originally validated to measure.

Why it matters

Use the scale. But know what you changed.

The practical takeaway: adapting a scale is sometimes necessary, but it should be deliberate, documented, and flagged whenever results get compared across studies. And the question underneath it, are we measuring the right thing in the right way?, reaches well past trust scales. It is the same judgment that decides whether a usability metric, a survey, or a safety evaluation means what its authors say.

Reflection

Why this sits at the center of my portfolio.

This project says the most about how I work: careful, methodical, and interested in how people deal with systems that keep automating more of the job. The same thread runs through the rest of my work. It is why my emergency-alert concept ends with a validation study instead of a mockup, and why the water chatbot's hardest problem turned out to be trust, not technology.

measurement rigor critical literature review automation trust research synthesis