Microsoft Copilot Studio Introduces Automated Agent Evaluations in October Update

Microsoft Copilot Studio rolls out a public preview of automated agent evaluations, allowing makers to test and validate agents at scale using flexible test sets and robust grading frameworks.

Microsoft Copilot Studio introduces a public preview of automated agent evaluations, transforming how makers validate their AI assistants. Instead of testing scenarios one by one, developers now build and execute evaluation sets directly from the agent or the Test Pane to gain structured, repeatable insights both before and after publishing.

This new experience offers multiple ways to create evaluation sets, including uploading predefined questions and answers, reusing recent Test Pane queries, adding cases manually, or generating queries instantly with AI. This flexibility ensures comprehensive test coverage that spans organization-specific scenarios while incorporating AI-suggested questions based on agent metadata and topics.

Evaluations rely on a robust grader framework that gives makers control over how they measure accuracy, ranging from strict checks like Exact Match to semantic comparisons and AI-powered metrics such as relevance, completeness, and groundedness. Each test delivers clear pass/fail results and detailed scores, introducing a scalable framework that helps teams identify gaps early, reduce production surprises, and track quality improvements over time.

Read More at the original source →