AI Agent Task Evaluation Metrics

Updated 

Evaluation Metrics help organisations assess the performance and quality of AI Agents across both simulated and production conversations. These metrics provide insights into how effectively an AI Agent resolves customer issues, conducts conversations independently, and delivers high-quality customer experiences.

The Evaluation Metrics module serves as a central location for viewing, enabling, disabling, and managing the metrics used to evaluate AI Agent performance.

Access Evaluation Metrics

Evaluation Metrics are available as a dedicated option within the Evaluate section.

  1. In the left navigation pane, click Evaluate under AI Agent.

  2. From the left pane of the Evaluate window, select Evaluation Metrics.

  3. The Evaluation Metrics Record Manager opens and displays all available evaluation metrics.

Evaluation Metrics Record Manager

The Record Manager provides the following details for each metric:

Column

Description

Metric Name

Name of the evaluation metric.

Description

Explanation of what the metric evaluates.

Type

Category of the metric.

You can enable or disable individual metrics directly from the Record Manager using the available toggle.

Available Evaluation Metrics

The following metrics are available to evaluate AI Agent performance:

Metric Name

Description

Type

Agent Success Rate

Percentage of simulated conversations in which the AI Agent meets all defined success criteria.

Resolution Effectiveness

Containment

Percentage of conversations resolved by the AI Agent without escalation or handoff to a human agent.

Resolution Effectiveness

CSAT

Predicted Customer Satisfaction Score.

Resolution Effectiveness

Conversation Opening

Evaluates whether the agent greeted the customer professionally and established a positive tone.

Conversation Quality

Issue Understanding & Probing

Evaluates whether the agent correctly identified and confirmed the customer's issue.

Conversation Quality

Courtesy & Professionalism

Evaluates whether the agent maintained respectful and professional communication throughout the interaction.

Conversation Quality

Empathy

Evaluates whether the agent acknowledged customer concerns and demonstrated understanding.

Conversation Quality

Active Listening

Evaluates whether the agent responded appropriately to customer inputs and conversation context.

Conversation Quality

Communication Clarity

Evaluates whether the agent communicated clearly and effectively.

Conversation Quality

Ownership

Evaluates whether the agent took responsibility for addressing the customer's issue.

Conversation Quality

Issue Resolution

Evaluates whether the customer's issue or query was successfully resolved.

Conversation Quality

Conversation Closure

Evaluates whether the agent appropriately concluded the interaction, including resolution confirmation and next steps.

Conversation Quality

Manage Metrics

Enable or Disable Metrics

Each metric includes a toggle in the Record Manager that allows administrators to:

  • Enable metrics that should be included in evaluations.
  • Disable metrics that are not relevant to their evaluation strategy.

Disabled metrics are excluded from future evaluations.

Available Actions

Delete

Removes an evaluation metric from the Record Manager. Before deleting a metric, ensure it is no longer required for AI Agent evaluation. Once deleted, the metric is not available for future evaluations unless it is recreated.

A confirmation prompt is displayed before deletion to prevent accidental removal.

Open image-20260617-202641.png

Edit Model Details

Available for standard Conversation Quality metrics. This action opens the quality model configuration screen, where administrators can modify:

  • Evaluation Question
  • Evaluation Description

These settings determine how the selected metric is assessed during evaluations.

Open image-20260617-202750.png