Skip to content
Kodning & Utveckling· NewsAvailable

AWS Launches New Evaluation Method for AI Agents in Multi-Step Dialogues

AWS has launched the Agent Evaluation Metric (AEM), a new measurement method designed to identify the exact dialogue step that causes errors in multi-step AI agent conversations.

By the Aheadline editorial team·11 sep. 2026·2 min read·Source: AWS Machine Learning BlogVerifierad signalAI-generated
AWS Launches New Evaluation Method for AI Agents in Multi-Step Dialogues
AWS Launches New Evaluation Method for AI Agents in Multi-Step Dialogues
AWS Launches New Evaluation Method for AI Agents in Multi-Step Dialogues
By · Policy- & EU-reporter
Last updated
Vad betyder det för mig?

What happened?

Amazon Web Services (AWS) has unveiled the Agent Evaluation Metric (AEM), a new framework for evaluating AI agents in multi-turn conversations. The method decomposes conversation-level evaluations into individual dialogue steps to pinpoint exactly where an error occurs. Initially, AEM focuses on the dimension of accuracy to distinguish root-cause errors from downstream failure steps.

Key facts

TeknikAgent Evaluation Metric (AEM)
FokusområdeFlerstegsdialoger (multi-turn conversations)
Första utvärderingsdimensionKorrekthet (correctness)

”Multi-turn agents fail in ways single-turn evaluation misses: one early mistake corrupts every later turn.”

— AWS Machine Learning Blog, Utgivare · AWS Machine Learning Blog

Why it matters

Multi-step agents often fail in ways that single-step evaluations fail to capture, as an early error can compromise all subsequent stages of the conversation. By isolating the exact step where a failure originated, developers can more effectively train, fine-tune, and debug their agents without the need to manually review entire conversation histories.

Who is affected?

The framework is aimed at AI developers, data scientists, and companies building autonomous agents or advanced chatbots. It facilitates the development of agents that perform complex multi-step tasks, such as those in customer service or internal business workflows.

Impact on the EU

AEM is a methodological evaluation model available globally to all developers and companies building agent-based AI systems on AWS. The method simplifies compliance with traceability and quality assurance requirements set out in the EU AI Act.

What else you should know

Traditional evaluation models often measure only the final outcome of a conversation, making it difficult to identify why an agent went off track. AEM enables the automation of troubleshooting, replacing the need for manual review of lengthy conversation logs.

Frequently asked questions

Quick answers about this story

Vad har hänt?
AWS har offentliggjort Agent Evaluation Metric (AEM), ett nytt ramverk för att utvärdera och felsöka AI-agenter i flerstegskonversationer på dialogstegsnivå.
När hände det?
Nyheten publicerades av AWS Machine Learning Blog i maj 2024.
Varför spelar det roll?
Traditionell utvärdering missar ofta fel i flerstegsdialoger där ett tidigt misstag förstör hela samtalet. AEM gör det möjligt att isolera det exakta steget som orsakade felet.
Vem kan använda AEM?
Metoden är en generell mätmodell tillgänglig för alla utvecklare som bygger agenter på AWS, inklusive utvecklare inom EU och Sverige.
Original source
AWS Machine Learning Blog·aws.amazon.com

The link opens in a new window and leads to the publisher's own site.

Verifierad signal

Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.

AI-verktyg i artikeln

Topics

#AI-benchmarking#Agents#Machine Learning#LLM-agenter
[ STAY UP TO DATE ]

Get similar news straight to your inbox

No affiliate linksCancel anytimeGDPR-friendly
[ Frequency ]
[ What do you want to read about? ]

You'll receive updates on 2 topics.

The reader's room

Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.

Sign in to submit a comment or question.

Loading comments…
How this affects you

Read the article through your role

  • Decide whether this affects strategy over 6–12 months or is just noise.
  • Discuss with leadership: do we own the right question or does ownership need to move?
  • Ask: what risk are we taking by NOT acting on this this quarter?

Generated angle — not editorial analysis of "AWS Launches New Evaluation Method for AI Agents in Multi-St"