At a glance
Three things to know in three seconds.
The problem
Arizonans worry about water quality and supply, but trustworthy answers are scattered across agencies, so people fall back on friends and social media.
The work
Statewide survey synthesis → personas → comparative heuristic evaluation of the AI engines → moderated think-aloud usability testing of the chatbot.
The outcome
A tested conversational design: guided starter questions, plain-language answers with cited agency sources, and edit / regenerate / rate controls.
The problem
Everyone talks about water. Few know who to ask.
A statewide survey run through the Arizona Water Innovation Initiative (n = 128)
showed what residents actually discuss, and how patchy their sources are.
People pieced together answers from water bills, local news, neighbors, and
Instagram Reels:
discuss water availability
64% · 82/129
discuss water quality
58% · 75/129
"Water quality. A lot of my friends scared me away from drinking the tap water."
survey respondent, AWII Arizona Water Survey, 2023
That's the gap the chatbot targets: turn anxiety and hearsay into sourced,
plain-language answers about drought, conservation, and the Colorado River.
Meet the users
Personas built from real Arizonans, not vibes.
From the survey and our own interviews we built personas of the residents the
bot had to serve. Alex is the one that drove the most design decisions, a
young, tech-savvy homeowner drowning in hard-water problems and scattered
information:
A
Alex
27 · Gilbert, AZ · Software developer · Homeowner
Five years in Arizona, two houses, recently moved from Mesa to Gilbert. Relies on Instagram, Google, and the neighborhood network for real-time information. Tech-savvy, and would absolutely use a chatbot.
Quote
"I think a chatbot would be helpful for everyone to get information. I would expect it to have location-based or neighborhood-based news."
Chatbot expectations
- Simple but accurate information
- Solutions to real-time concerns
- Sources and references
- Optional notifications
- Info by city, selectable
Motivations
- Minimize water bills
- Hard water in AZ
- New-homeowner questions
- Contamination & safety
Choosing the engine
Before designing the bot, we evaluated the brains.
We ran a comparative heuristic evaluation of the two available LLM engines
(ChatGPT and Bard) through a water-questions lens. The differences shaped the
conversation design more than any wireframe did:
- Editable answers beat re-prompting. Bard's built-in answer modification saved users a whole prompt cycle, so edit and regenerate controls became a core requirement, not a nice-to-have.
- Accessibility is a differentiator. Speech-to-text input widened who could actually use a public-service tool.
- The blank prompt box is a usability cliff. Without a guide, users need practice to phrase effective prompts, which argued for guided starter questions instead of an empty text field.
How we tested
Moderated think-alouds, with an adversarial streak.
Each of the four of us recruited at least two screened Arizona residents
(residency, age band, self-rated water knowledge, tech comfort) for moderated
30 to 45 minute think-aloud sessions, consent-formed, recorded, and structured
around goals in the five Es: efficient, effective, engaging, error-tolerant,
easy to learn. Remote results were collected through Maze.
- Scenario 1, first impressions: open exploration after hearing the bot exists.
- Scenario 2, the homeowner deep-dive: soft-water research with escalating tasks, including deliberately challenging the bot ("I've heard soft water corrodes pipes, is that true?") to probe trust and error recovery.
- Scenario 3, quality & regulation: ask, edit the question, regenerate, rate the answer, follow up.
Analysis was deliberately small-n honest: counts instead of percentages
("4 of 6 participants" says more than "67%"), every outlier reviewed
individually, and findings triangulated across the survey, sessions, and
post-test satisfaction ratings before anything earned a recommendation.
What shipped
Every persona expectation became an interface decision.
The tested design, an Arizona Water Chatbot research preview in ASU's Global
Futures Laboratory environment, answered the research directly:
- "I don't know what to ask"→Guided starter topics: Diving into Drought · Mastering Conservation · the Colorado River's role
- "Can I trust this?"→Curated source links (ADWR, AMWUA) under answers + a plain-spoken accuracy disclaimer
- "The answer isn't quite right"→Edit the question, regenerate, and rate answers inline
- "I'll need this again later"→Named conversation history (Water policy · Lake Mead & Lake Powell · Conservation practices)
What I took away
Public-service AI is a trust problem first.
The chatbot's hardest problems were never technical. They were about whether a
resident would believe, and act on, what an AI said about the water coming out
of their tap. Sources, honest disclaimers, and user control (edit, regenerate,
rate) did more for trust than any answer-quality tweak. That thread, how
people calibrate trust in automated systems, runs straight through the rest
of my research.
civic UX
conversational AI
usability testing
survey synthesis