Concept case study · self-initiated

38 Minutes

In 2018, one click on a badly designed dropdown told 1.4 million people in Hawaii that a ballistic missile was inbound. It took 38 minutes to say it wasn't. I redesigned the console behind that click. And because I am a researcher, I wrote the study that would prove the redesign works.

Type: concept study Domain: safety-critical interfaces Methods: error analysis · redesign · study design
At a glance

Three things to know in three seconds.

The failure

A live missile alert and its drill version sat one line apart in the same list of links, guarded by a generic "Are you sure?"

The fix

Visible modes, a confirmation that states the consequence, a second sign-off, and a 10-second undo window.

The difference

Most redesigns of this console stopped at mockups. Mine ends with the experiment that would validate it.

What happened

Saturday, January 13, 2018.

During a routine shift-change drill at the Hawaii Emergency Management Agency, an operator selected the live alert instead of the drill. The interface offered no meaningful resistance.

  • Every phone in the state: "BALLISTIC MISSILE THREAT INBOUND TO HAWAII. SEEK IMMEDIATE SHELTER. THIS IS NOT A DRILL."
  • Families sheltered in bathtubs, said goodbyes, put children in storm drains. There was no missile.
  • The correction finally reached phones 38 minutes later. The system had templates for launching an alert, but none for retracting one.
Why it failed

The operator made an error. The interface made it easy.

This is a textbook human-factors failure, not a careless-employee story. Three classic design violations stack on top of each other:

  • Mode invisibility. Drill and live modes looked identical: same screen, same list, same styling. Nothing in the environment said "you are about to affect the real world."
  • No forcing function. The confirmation asked "Are you sure?", the kind of question people answer on autopilot dozens of times a day. It named no consequence, so it prevented nothing.
  • No error recovery. Sending was instant and irreversible by design, with no hold window and no pre-built retraction path. That is why a one-second slip took 38 minutes to walk back.
Before: recreation of the reported UI
After: my concept console
Drill mode · simulated · nothing leaves this room
Ballistic Missile Alert · LIVE

Requires: typed confirmation → second authorization → 10-second hold. Live alerts are visually, spatially, and procedurally separate from drills.

type "LIVE ALERT 1.4M" to arm…
Awaiting second authorization · Watch Officer K. Nakamura
HOLD TO SEND · releases in 10 s · cancel anytime

"Before" is a faithful-in-spirit recreation based on published reporting and the FCC's post-incident report, not a screenshot. "After" is my concept, not an official system.

The redesign

Four principles, borrowed from places where errors kill.

  • Make the mode ambient. Drill mode changes the whole console: color, watermark, banner. Not just one label. You should be able to tell which world you're in from across the room.
  • Confirmations must state consequences. Typing "LIVE ALERT 1.4M" is not friction for friction's sake; it forces the operator's attention through the exact thing about to happen. "Are you sure?" asks nothing.
  • Two-person integrity. Nuclear weapons, bank vaults, and code deploys all require a second key for irreversible actions. Statewide alerts should too.
  • Design the undo before the send. A 10-second hold window catches slips in the moment, and a pre-drafted retraction template turns a 38-minute scramble into a 60-second correction.
How I'd prove it

A redesign is a hypothesis. Here's the experiment.

This is where my human-factors training does the work most concept studies skip: the redesign above is only as good as the evidence behind it.

  • Design: a within-subjects simulation. Operators run repeated shift-change drills on the legacy layout and on the redesigned console, with rare, unannounced "live" trials mixed in.
  • Pressure matters: trials run under time pressure and interruption, because that is the condition the real error happened in. Calm-lab results would not generalize.
  • Measures: mode-error rate (the metric that failed in 2018), time-to-send for legitimate alerts (safety can't make the real job too slow), time-to-retract, and post-trial confidence calibration.
  • The trade-off to watch: if the added friction slows a real launch by more than seconds, the design fails its own standard. That tension is the study's most important outcome.
What this shows

Blame the design, then fix it, then test it.

I chose this incident because it sits exactly where I work: the seam between human cognition and high-stakes automation. My research background is in how people trust and mistrust automated systems. This is what that thinking looks like once it is applied to pixels, procedures, and one very bad Saturday morning.

error prevention mode visibility forcing functions study design