AWS Launches New Evaluation Method for AI Agents in Multi-Step Dialogues
AWS has launched the Agent Evaluation Metric (AEM), a new measurement method designed to identify the exact dialogue step that causes errors in multi-step AI agent conversations.

What happened?
Amazon Web Services (AWS) has unveiled the Agent Evaluation Metric (AEM), a new framework for evaluating AI agents in multi-turn conversations. The method decomposes conversation-level evaluations into individual dialogue steps to pinpoint exactly where an error occurs. Initially, AEM focuses on the dimension of accuracy to distinguish root-cause errors from downstream failure steps.
Key facts
| Teknik | Agent Evaluation Metric (AEM) |
|---|---|
| Fokusområde | Flerstegsdialoger (multi-turn conversations) |
| Första utvärderingsdimension | Korrekthet (correctness) |
”Multi-turn agents fail in ways single-turn evaluation misses: one early mistake corrupts every later turn.”
Why it matters
Multi-step agents often fail in ways that single-step evaluations fail to capture, as an early error can compromise all subsequent stages of the conversation. By isolating the exact step where a failure originated, developers can more effectively train, fine-tune, and debug their agents without the need to manually review entire conversation histories.
Who is affected?
The framework is aimed at AI developers, data scientists, and companies building autonomous agents or advanced chatbots. It facilitates the development of agents that perform complex multi-step tasks, such as those in customer service or internal business workflows.
Impact on the EU
AEM is a methodological evaluation model available globally to all developers and companies building agent-based AI systems on AWS. The method simplifies compliance with traceability and quality assurance requirements set out in the EU AI Act.
What else you should know
Traditional evaluation models often measure only the final outcome of a conversation, making it difficult to identify why an agent went off track. AEM enables the automation of troubleshooting, replacing the need for manual review of lengthy conversation logs.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vem kan använda AEM?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "AWS Launches New Evaluation Method for AI Agents in Multi-St"