Beyond accuracy: how to evaluate the full AI experience

Imagine asking an AI assistant to find the best dishwasher for your kitchen.
It has you enter your budget, dimensions and priorities. After asking you some follow-up questions, it returns three recommendations.
Technically, everything worked.
But do you trust its recommendations? Do you understand why it chose them? And did you enjoy your conversation with the assistant?
Those questions point to something traditional AI evaluations can easily miss.
AI isn't simply changing what software can do. It's changing our relationship with it.
AI changes our relationship with technology
Traditional software is built around two basic assumptions: the computer is a tool, and we control it through a graphical interface.
Click. Swipe. Type. Select.
AI undercuts both assumptions. Increasingly, the computer is an advisor and assistant rather than just a tool. We have back-and-forth conversations with it, and it can perform tasks for us autonomously.
That changes what a good experience looks like.
Traditional software is optimized for usability. AI software needs to be optimized for relationship.
BENCHMARK
AI Relationship Quality
The UserTesting AI Relationship Quality (ARQ™) benchmark gives you a clear measure of how well your AI experience is working for customers
Conversation changes how we communicate with computers
Research shows that when people have conversations with AI, they judge them much like they judge conversations with other humans.
Small differences in word choice, formatting, and length in AI responses can make a huge difference in how users rate the experience. They may describe the AI as polite, friendly, arrogant, or even disrespectful.
And when that happens, the central question is no longer: Was this easy to use?
It's: How do I feel about this relationship?
Agentic AI transforms software from a tool to a partner
Agentic AI takes that shift further.
As AI systems gain the ability to act autonomously, people can delegate tasks they previously performed themselves.
Think less calculator, more personal assistant, or an always available butler, that understands what you want and makes things happen on your behalf.
That further changes our relationship with software. To get comfortable using a software agent, you need to understand what it’s doing, and you need to trust its judgment and reliability. It’s the same way you’d judge a human assistant..
A spreadsheet doesn't need your trust in quite the same way an AI financial assistant does.
A travel website doesn't need to understand your intentions in quite the same way an agent booking your vacation does.
The more authority we hand over to AI, the more the quality of that relationship matters.
Guide
How to choose the right AI features
Learn how to choose and test generative AI features that solve real customer needs and build trust.
Accuracy is necessary. It isn't enough.
Traditional AI evaluation still matters.
Technical evaluations tell teams whether an AI system is accurate, appropriate and performing as expected. Product metrics such as task completion and adoption tell them whether people are using it successfully.
But neither captures the full experience.
If an AI chatbot gives the correct answer but annoys the user in the process, they won’t be back. If an agent makes the right decision but the user feels it wasn’t transparent about what it did, they will be unlikely to use it again.
AI experiences therefore have two dimensions teams need to understand: functional performance and relationship quality.
Functional performance asks whether the system works.
Relational quality asks how it feels to work with the system.
You need both.
A successful task can still be a failed experience
Imagine an AI travel agent that successfully books your flight.
Task completed.
But it chose a 6 a.m. departure without explaining why, ignored the airline you usually fly and gave a vague answer when you questioned its decision.
Traditional product metrics might record a success. The user might see something very different.
Do I understand it? Can I rely on it? Do I want to use it again?
Those reactions matter. An AI experience may function correctly while leaving people confused, uncomfortable or unwilling to rely on it. This ultimately affects adoption, continued use and perceptions of the company behind it.
Insight+
Get on-demand access to some of UserTesting's most popular events, including from Crafted Seattle and Crafted London. Register for Insight+ and start watching today.
How to test the human-AI relationship
So how do you measure the part of an AI experience that traditional evals miss?
Research shows that five elements shape how people relate to conversational and agentic AI: understanding, trust, control, outcome, and affinity. Together, they drive the user’s perception of the human-AI relationship.
- Control: Does the AI feel like a collaborative partner that stays within the boundaries set by the user? This becomes especially important with AI agents, where the system may make decisions and take actions on a user's behalf.
- Outcome: Does working with the AI produce a better result than the person could achieve alone? Completing a task isn't enough if the result doesn't reflect what the user actually wanted.
- Understanding: Do people feel they can “read” the AI’s thinking, why it behaves the way it does, and where its boundaries are? Users don't need to see every step of the machinery, but they should be able to form a useful mental model of how it works.
- Trust: Do people rely on the AI when they should—and know when to step in when they shouldn't? Good AI experiences don't demand blind trust. They help people develop the right level of trust.
- Affinity: Do people actually like interacting with it? Tone, wording and other seemingly small details can shape how people feel about AI. Just as with another person, an interaction can be competent but still leave you thinking, I don't want to work with them again.
These dimensions are reflected in UserTesting's AI Relationship Quality (ARQ™) benchmark, which assesses the human side of AI quality across the same five areas. The benchmark pairs an overall relationship-quality score with individual element scores and participant feedback videos, helping reveal not just whether an experience has a problem, but why.
That's the missing layer in many AI evaluations. Accuracy tells you whether the system functioned properly. Relationship quality helps tell you whether people will actually want to use it again.
AI evaluation needs to include the human experience
Relationship testing fills a gap that isn’t covered by technical evaluations and product metrics.
As AI products behave more like people, and work as collaborators and advisors, teams need to evaluate the relationship people have with them alongside how well they perform.
Teams need to test the relationship as rigorously as they do accuracy and efficiency. That requires bringing real people into AI evaluation and measuring the full experience over time.
Because the next generation of AI products won't succeed simply because they work.
They'll succeed because people understand them, trust them and want to keep working with them.
Additional resources
- Webinar: The AI-native product loop: build, test, learn without slowing down. Learn how to integrate real customer feedback into AI-speed product development so teams can validate continuously and ship with greater confidence.
- Guide: Human insight for the AI-driven product development process. Explore how AI changes product development and why teams need to evaluate understanding, trust, control, outcome and affinity—not just usability. This is probably the closest companion resource to the post.
- Podcast: UX research for AI: building trust in experiences. Microsoft senior UX researcher Priyanka Kuvalekar explains why evaluating AI requires looking beyond traditional usability metrics to understand trust, emotion and the human experience.
- Blog: How to test AI features: rethinking AI usability testing for conversational experiences. Learn why conversational AI requires a different approach to UX testing and how teams can evaluate AI as an interaction and relationship rather than simply another product feature.



