{"id":7917,"date":"2026-06-11T09:00:00","date_gmt":"2026-06-11T16:00:00","guid":{"rendered":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/?post_type=copilot&#038;p=7917"},"modified":"2026-07-21T16:53:52","modified_gmt":"2026-07-21T23:53:52","slug":"who-evaluates-the-evaluators-the-data-science-behind-agent-evals","status":"publish","type":"copilot","link":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/","title":{"rendered":"Who evaluates the evaluators? The data science behind agent evals"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">\n  At <a href=\"https:\/\/find.codeghost.online\/en-us\/microsoft-365-copilot\/microsoft-copilot-studio\/\">Microsoft Copilot Studio<\/a>, we talk a lot about ways to make agents better. You can add knowledge sources, fine-tune models, insert prompts, and more. But how do you know whether an agent is actually getting better? How do you detect regressions before your users do? And perhaps most importantly, how do you trust the signals you&#8217;re using to make those decisions?\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  As data scientists, we don\u2019t ship a model without evaluating it. We evaluate it before the first release, and we evaluate it after every meaningful change.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  We validate offline, track metrics over time, compare variants, look for regressions, and ask a simple question: <em>Did this change actually make the model better?<\/em>\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">When we started building <a href=\"https:\/\/learn.microsoft.com\/en-us\/microsoft-copilot-studio\/analytics-agent-evaluation-intro\">evaluation features in Copilot Studio<\/a>, we treated them the same way. We asked: <em>Is the evaluation giving the right answer to \u201cis it better\u201d? <\/em><strong>This is the quality question.<\/strong><\/p>\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a606ee1108a1&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a606ee1108a1\" class=\"wp-block-image size-large wp-lightbox-container\"><img decoding=\"async\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/THE-QUALITY-QUESTION-1920px-1024x412.webp\" alt=\"Side-by-side comparison of evaluation questions: for the agent, &ldquo;Did the change make it better?&rdquo; and for the evaluation, &ldquo;Can we trust the signal?&rdquo; with note that evaluation quality is the challenge.\" class=\"wp-image-7936 webp-format\" srcset=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/THE-QUALITY-QUESTION-1920px-1024x412.webp 1024w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/THE-QUALITY-QUESTION-1920px-300x121.webp 300w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/THE-QUALITY-QUESTION-1920px-768x309.webp 768w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/THE-QUALITY-QUESTION-1920px-1536x618.webp 1536w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/THE-QUALITY-QUESTION-1920px.webp 1920w\" data-orig-src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/THE-QUALITY-QUESTION-1920px-1024x412.webp\"><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\taria-label=\"Enlarge\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.imageButtonRight\"\n\t\t\tdata-wp-style--top=\"state.imageButtonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">To answer it, we&#8217;ll explore three core areas of AI evaluation quality:<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li class=\"wp-block-list-item\">The data behind evaluation, and why generated datasets play such an important role in agent testing.<\/li>\n\n\n\n<li class=\"wp-block-list-item\">The evaluators (graders) themselves, and how we validate that graders produce reliable signals.<\/li>\n\n\n\n<li class=\"wp-block-list-item\">The metrics we use to determine whether those signals are trustworthy enough to support real-world decisions.<\/li>\n<\/ol>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Because as organizations increasingly rely on evaluation to improve AI systems, <a href=\"https:\/\/find.codeghost.online\/en-us\/ai\/responsible-ai\">confidence in the evaluation process<\/a> becomes just as important as confidence in the agent.\n<\/p>\n\n\n\n<div class=\"wp-block-buttons is-content-justification-center is-layout-flex wp-container-core-buttons-is-layout-a89b3969 wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/find.codeghost.online\/en-us\/microsoft-365-copilot\/microsoft-copilot-studio\/\" target=\"_blank\" rel=\"noreferrer noopener\">Evaluate your agents with Copilot Studio<\/a><\/div>\n<\/div>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"from-model-evaluation-to-agent-evaluation\">From model evaluation to agent evaluation<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Traditional machine learning evaluation is relatively well\u2011defined. You have:\n<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li class=\"wp-block-list-item\">Labeled data  <\/li>\n\n\n\n<li class=\"wp-block-list-item\">A clear task  <\/li>\n\n\n\n<li class=\"wp-block-list-item\">A small set of metrics  <\/li>\n\n\n\n<li class=\"wp-block-list-item\">A mostly static input\u2013output mapping  <\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/find.codeghost.online\/en-us\/microsoft-365-copilot\/agents\" target=\"_blank\" rel=\"noreferrer noopener\">AI agents<\/a> challenge many of these assumptions.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Agents operate over multi\u2011turn conversations, adapt to user behavior, use tools, and use implicit reasoning. Accordingly, they are judged across multiple quality dimensions, correctness, completeness, clarity, coherence, tone, and more.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  So Copilot Studio ships evaluation features to help makers answer questions like:\n<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li class=\"wp-block-list-item\">Did my agent regress after this change?  <\/li>\n\n\n\n<li class=\"wp-block-list-item\">Does quality degrade in longer conversations?  <\/li>\n\n\n\n<li class=\"wp-block-list-item\">Can I trust my agent to behave as expected?  <\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">But that immediately raises a second\u2011order question: <em>How do we know the evaluation features themselves are giving the right answer?<\/em><\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"building-the-right-evaluation-data-for-agents\">Building the right evaluation data for agents<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">\n  In data science, we know that evaluation quality is bounded by evaluation data.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If the data you evaluate on is narrow, unrealistic, or biased, the metrics will <strong>look confident but be wrong.<\/strong> That principle applies just as much when we evaluate evaluation features as when makers <a href=\"https:\/\/learn.microsoft.com\/en-us\/microsoft-copilot-studio\/analytics-agent-evaluation-overview\" target=\"_blank\" rel=\"noreferrer noopener\">evaluate their agents<\/a>.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"real-data-grounded-but-limited\">Real data: Grounded, but limited<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">\n  For makers, Copilot Studio supports importing real production data to evaluate agent performance. Internally, however, we do not use customer data in any form when validating evaluation features. \n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Instead, we rely on curated examples and generated datasets that allow controlled and systematic testing\u2014what we call semi\u2011real examples. These include scenarios inspired by conversations shared by design partners, as well as examples curated during feedback cycles.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  For us, and for many of our makers, these alone are not sufficient.\n<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"generated-data-scalable-targeted-and-intentional\">Generated data: Scalable, targeted, and intentional<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Real and semi\u2011real examples provide valuable grounding, but in practice, most evaluation workflows (both internally and for makers) rely primarily on generated data. And that isn\u2019t a compromise; it\u2019s an intentional design choice.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Generated data allows evaluation to start earlier, scale faster, and cover a broader range of agent behaviors.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  From our perspective as a data science team, generated datasets are essential for supporting the wide variety of agents built in Copilot Studio. They allow us to validate evaluation features across different agent types, domains, and interaction patterns, and to do so at a scale that would not be feasible with curated examples alone.\n<\/p>\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a606ee1122c1&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a606ee1122c1\" class=\"wp-block-image size-large wp-lightbox-container\"><img decoding=\"async\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-REASONS-TO-GENERATE-DATASETS-1920px-1024x412.webp\" alt=\"List of five reasons to generate datasets: evaluate agent quality before production, gain insights within compliance limits, test at larger scale with more scenarios, use evaluation-ready data without cleaning, and leverage built-in test set creation in Copilot Studio.\" class=\"wp-image-7934 webp-format\" srcset=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-REASONS-TO-GENERATE-DATASETS-1920px-1024x412.webp 1024w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-REASONS-TO-GENERATE-DATASETS-1920px-300x121.webp 300w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-REASONS-TO-GENERATE-DATASETS-1920px-768x309.webp 768w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-REASONS-TO-GENERATE-DATASETS-1920px-1536x618.webp 1536w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-REASONS-TO-GENERATE-DATASETS-1920px.webp 1920w\" data-orig-src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-REASONS-TO-GENERATE-DATASETS-1920px-1024x412.webp\"><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\taria-label=\"Enlarge\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.imageButtonRight\"\n\t\t\tdata-wp-style--top=\"state.imageButtonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">\n  For makers, the motivations are equally practical:\n<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li class=\"wp-block-list-item\"><strong>Evaluating before publishing<\/strong>: Generated datasets make it possible to <a href=\"https:\/\/learn.microsoft.com\/en-us\/microsoft-copilot-studio\/analytics-overview\" target=\"_blank\" rel=\"noreferrer noopener\">assess agent behavior and quality<\/a> before the agent is exposed to real users.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Limited or restricted access to production data<\/strong>: In many cases, makers do not have access to their agent\u2019s production conversations at all, due to compliance, governance, or organizational policies.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Working with production data selectively<\/strong>: Even when production data exists, it often needs filtering or augmentation to support systematic evaluation.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Perhaps the strongest motivation is how easy it is to generate high\u2011quality evaluation datasets. Copilot Studio enables makers to <a href=\"https:\/\/learn.microsoft.com\/en-us\/microsoft-copilot-studio\/analytics-agent-evaluation-create\" target=\"_blank\" rel=\"noreferrer noopener\">create test sets<\/a> that are targeted, repeatable, and aligned with their agent\u2019s intended behavior\u2014without requiring manual data collection.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"evaluating-data-generation-for-agent-evaluations\">Evaluating data generation for agent evaluations<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Evaluation datasets can be generated in multiple ways. We support multiple data generation strategies because they surface different aspects of agent behavior. When applied together, they give makers practical, high\u2011coverage evaluation datasets.\n<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"data-generation-strategies\">Data generation strategies<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">\n  There are four main types of dataset generation:\n<\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li class=\"wp-block-list-item\"><strong>Single\u2011turn generation<\/strong> allows makers to test specific behaviors in isolation. These datasets are easier to reason about and are well\u2011suited for validating correctness, relevance, and instruction adherence.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Multi\u2011turn generation<\/strong> adds the complexity of context tracking and conversational dependencies. This is particularly useful for makers building predefined flows or agents whose behavior depends on conversation state.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Knowledge\u2011based generation<\/strong> tends to produce very concrete, sometimes highly specific questions. These queries are effective for testing grounding and answerability against the agent\u2019s connected knowledge sources.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Topic\u2011based and instruction\u2011based generation<\/strong> often lead to more general or exploratory questions. These datasets are useful for identifying unsupported or weakly supported areas\u2014reasonable questions users may ask that fall outside the agent\u2019s main flows.<\/li>\n<\/ol>\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a606ee113669&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a606ee113669\" class=\"wp-block-image size-large wp-lightbox-container\"><img decoding=\"async\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/4-TYPES-OF-DATASET-GENERATION-1920px-1024x412.webp\" alt=\"Diagram showing four types of dataset generation: single-turn, multi-turn, knowledge-based, and topic- or instruction-based.\" class=\"wp-image-7932 webp-format\" srcset=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/4-TYPES-OF-DATASET-GENERATION-1920px-1024x412.webp 1024w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/4-TYPES-OF-DATASET-GENERATION-1920px-300x121.webp 300w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/4-TYPES-OF-DATASET-GENERATION-1920px-768x309.webp 768w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/4-TYPES-OF-DATASET-GENERATION-1920px-1536x618.webp 1536w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/4-TYPES-OF-DATASET-GENERATION-1920px.webp 1920w\" data-orig-src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/4-TYPES-OF-DATASET-GENERATION-1920px-1024x412.webp\"><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\taria-label=\"Enlarge\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.imageButtonRight\"\n\t\t\tdata-wp-style--top=\"state.imageButtonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">\n  By combining these generation types, makers can build large and diverse evaluation sets that cover both expected and unexpected usage patterns.\n<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"how-we-evaluate-generated-queries\">How we evaluate generated queries<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Because data generation itself is an evaluation feature, we explicitly assess the quality of generated queries. We use an LLM\u2011as\u2011a\u2011judge methodology to assess dataset quality along several dimensions, including:\n<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li class=\"wp-block-list-item\"><strong>Relevance.<\/strong> How well queries align with the agent\u2019s intended scope.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Interaction naturalness.<\/strong> Whether queries resemble plausible user goals, confusion, and follow\u2011ups.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Human likeness.<\/strong> The extent to which generated queries resemble questions a human would naturally ask.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Redundancy.<\/strong> Whether examples add new coverage rather than repeating similar patterns.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Intent diversity.<\/strong> The range of user intents represented in queries (for example, informational, troubleshooting, or exploratory).<\/li>\n<\/ul>\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a606ee11482e&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a606ee11482e\" class=\"wp-block-image size-large wp-lightbox-container\"><img decoding=\"async\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-DIMENSIONS-OF-DATASET-QUALITY-1920px-1024x412.webp\" alt=\"Diagram outlining five dimensions of dataset quality: relevance, interaction naturalness, human likeness, redundancy, and intent diversity.\" class=\"wp-image-7933 webp-format\" srcset=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-DIMENSIONS-OF-DATASET-QUALITY-1920px-1024x412.webp 1024w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-DIMENSIONS-OF-DATASET-QUALITY-1920px-300x121.webp 300w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-DIMENSIONS-OF-DATASET-QUALITY-1920px-768x309.webp 768w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-DIMENSIONS-OF-DATASET-QUALITY-1920px-1536x618.webp 1536w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-DIMENSIONS-OF-DATASET-QUALITY-1920px.webp 1920w\" data-orig-src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-DIMENSIONS-OF-DATASET-QUALITY-1920px-1024x412.webp\"><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\taria-label=\"Enlarge\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.imageButtonRight\"\n\t\t\tdata-wp-style--top=\"state.imageButtonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">\n  In addition, we apply <strong>generation\u2011specific measures<\/strong> where appropriate, such as topic coverage for topic\u2011based generation or grounding for knowledge\u2011based generation. These assess, for example, whether questions are answerable using the provided sources.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  These metrics allow us to reason systematically about whether a generation capability produces datasets that are broad, targeted, and useful for evaluation\u2014without relying on subjective impressions.\n<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"evaluating-graders-assessing-the-quality-of-our-evaluators\">Evaluating graders: Assessing the quality of our evaluators<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/custom-graders-in-copilot-studio-setting-high-standards-for-agent-evals\/\" target=\"_blank\" rel=\"noreferrer noopener\">Graders are the evaluators<\/a> we build to help makers assess their agents. They produce the scores and labels that makers use to understand what works well and what needs to be improved. For that reason, we assess graders explicitly and independently before they are exposed to makers.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"what-we-expect-from-a-high-quality-grader\">What we expect from a high\u2011quality grader<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">We treat graders as a system that estimates quality rather than produces one absolute answer. We assess these graders using the same principles we would apply to any automated evaluation system. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Concretely, we ask whether a grader:\n<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li class=\"wp-block-list-item\">Measures the intended dimension and only that dimension.  <\/li>\n\n\n\n<li class=\"wp-block-list-item\">Distinguishes between meaningful differences in responses.  <\/li>\n\n\n\n<li class=\"wp-block-list-item\">Behaves consistently across similar inputs.  <\/li>\n\n\n\n<li class=\"wp-block-list-item\">Produces interpretable and stable signals that can support downstream decisions.  <\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">A grader that produces <strong>reasonable explanations, but inconsistent judgments<\/strong> doesn&#8217;t meet the bar.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"purpose-built-datasets-for-grader-assessment\">Purpose\u2011built datasets for grader assessment<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">\n  To assess a grader\u2019s quality, we build purpose-built datasets, each tailored to the specific behavior or quality dimension the grader is designed to measure.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Each grader requires targeted datasets designed to measure the specific behavior being evaluated. As a result, the datasets we use for grader evaluation are intentionally designed for that purpose.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  In practice, the composition of grader\u2011specific datasets depends on the grader. For some graders, we rely primarily on human-labeled data. For others, generated data plays a central role, allowing us to construct targeted test cases with known ground truth. Most often, we use a combination of the two, balancing <strong>human judgment with scale and control.<\/strong>\n<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"a-controlled-generation-methodology\">A controlled generation methodology<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">For many graders, we use controlled synthetic datasets with known ground truth.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The process works as follows: <\/p>\n\n\n\n<ol class=\"wp-block-list\">\n<li class=\"wp-block-list-item\"><strong>Define a test agent.<\/strong> We start with a well-scoped agent configuration that represents the behavior domain the grader is intended to evaluate.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Generate high-quality queries.<\/strong> Using our data generation capabilities, we create a set of realistic, high-quality user queries aligned with the agent\u2019s scope.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Generate high-quality responses.<\/strong> For each query, we generate responses that meet the expected quality bar for the dimension under evaluation.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Introduce controlled degradations.<\/strong> We then intentionally degrade a subset of these responses in a controlled and traceable way. Each degradation targets a specific failure mode and we explicitly track whether a response was damaged and how.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>Use the dataset for evaluation.<\/strong> The resulting dataset contains both intact and intentionally degraded responses, with known ground truth about their quality.<\/li>\n<\/ol>\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a606ee115e1d&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a606ee115e1d\" class=\"wp-block-image size-large wp-lightbox-container\"><img decoding=\"async\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-STEP-DATASET-GENERATION-METHODOLOGY-1920px-1024x369.webp\" alt=\"Five-step process for dataset generation: define test agent configuration, generate high-quality queries, generate high-quality responses, introduce controlled degradations, and use the dataset for evaluation.\" class=\"wp-image-7935 webp-format\" srcset=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-STEP-DATASET-GENERATION-METHODOLOGY-1920px-1024x369.webp 1024w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-STEP-DATASET-GENERATION-METHODOLOGY-1920px-300x108.webp 300w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-STEP-DATASET-GENERATION-METHODOLOGY-1920px-768x277.webp 768w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-STEP-DATASET-GENERATION-METHODOLOGY-1920px-1536x554.webp 1536w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-STEP-DATASET-GENERATION-METHODOLOGY-1920px.webp 1920w\" data-orig-src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/5-STEP-DATASET-GENERATION-METHODOLOGY-1920px-1024x369.webp\"><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\taria-label=\"Enlarge\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.imageButtonRight\"\n\t\t\tdata-wp-style--top=\"state.imageButtonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Because we control the transformation applied to each response, we can treat this dataset as labeled. We know which responses should be flagged by the grader, and for what reason.\n<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"measuring-grader-performance\">Measuring grader performance<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Now that we know which responses were intentionally degraded and how, we can evaluate graders in a concrete and measurable way. Rather than relying on subjective inspection, we can treat grader assessment as a standard classification problem with known ground truth.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  The main metrics we track and optimize when developing graders are <strong>true positive rate (TPR)<\/strong> and <strong>true negative rate (TNR)<\/strong>.\n<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li class=\"wp-block-list-item\"><strong>TPR<\/strong> measures how often the grader correctly identifies responses that <em>should<\/em> be flagged. In our context, this reflects the grader\u2019s ability to detect intentionally damaged or low-quality responses when a problem is present.<\/li>\n\n\n\n<li class=\"wp-block-list-item\"><strong>TNR<\/strong> measures how often the grader correctly accepts responses that <em>should not<\/em> be flagged. This reflects the grader\u2019s ability to avoid false alarms and not penalize responses that meet the expected quality bar.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">\n  These metrics capture the core tradeoff every grader must manage: being sensitive enough to catch real issues, while remaining precise enough to avoid over\u2011penalizing valid responses.\n<\/p>\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a606ee117054&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a606ee117054\" class=\"wp-block-image size-large wp-lightbox-container\"><img decoding=\"async\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/HOW-DO-WE-KNOW-A-GRADER-WORKS-1920-px-1024x476.webp\" alt=\"Three-step process to validate a grader: create a baseline dataset with consistent, expected responses; introduce controlled defects such as missing information, incorrect facts, poor relevance, and safety issues; then test whether the grader flags defects and accepts correct responses, tuning as needed.\" class=\"wp-image-7939 webp-format\" srcset=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/HOW-DO-WE-KNOW-A-GRADER-WORKS-1920-px-1024x476.webp 1024w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/HOW-DO-WE-KNOW-A-GRADER-WORKS-1920-px-300x140.webp 300w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/HOW-DO-WE-KNOW-A-GRADER-WORKS-1920-px-768x357.webp 768w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/HOW-DO-WE-KNOW-A-GRADER-WORKS-1920-px-1536x714.webp 1536w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/HOW-DO-WE-KNOW-A-GRADER-WORKS-1920-px.webp 1920w\" data-orig-src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/HOW-DO-WE-KNOW-A-GRADER-WORKS-1920-px-1024x476.webp\"><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\taria-label=\"Enlarge\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.imageButtonRight\"\n\t\t\tdata-wp-style--top=\"state.imageButtonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">\n  By evaluating graders against datasets with controlled degradations, we can measure TPR and TNR directly, analyze failure modes, and iterate systematically. This allows us to tune grader behavior intentionally\u2014understanding where a grader is too permissive, where it\u2019s too strict, and how changes affect its decision boundaries.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  All together, these techniques allow us to move beyond evaluating individual grader performance and toward a broader goal: building evaluation systems whose behavior can be understood, measured, and improved over time.\n<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"bringing-rigor-to-agent-evaluation-features\">Bringing rigor to agent evaluation features<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">So, who evaluates the evaluators? At Copilot Studio, we approach evals with the same rigor we apply to models themselves. Because as teams increasingly rely on evaluation to guide real-world decisions, they need to trust the systems producing those signals. <\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  In this post, we described how we approach that challenge in Copilot Studio: constructing targeted datasets for graders, using controlled generation to create reliable ground truth, and measuring decision accuracy through metrics such as TPR and TNR. These practices help us understand not only whether an evaluation feature works, but how it behaves, where its limitations are, and how it can be improved over time.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  Feedback from design partners and customers plays an important role in this process. When real-world examples reveal gaps in a grader or generated dataset, we incorporate those learnings back into our evaluation process to continuously improve the system.\n<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">As the industry continues moving from experimental AI systems to <a href=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/new-resources-and-guidance-to-plan-build-and-operate-enterprise-ready-agents\/\" target=\"_blank\" rel=\"noreferrer noopener\">production-scale agents<\/a>, evaluation will become a foundational capability. As AI agents move into production environments, organizations need to trust not just the agents themselves, but the systems used to evaluate them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">\n  For us, rigorous evaluation is a core part of helping teams build and improve agents with confidence. Because better agent decisions start with trustworthy evaluation.\n<\/p>\n\n\n\n<div class=\"is-style-inline wp-block-bloginabox-theme-promotional\">\n\t\n<div class=\"promotional promotional--has-media promotional--media-right\">\n\t<div class=\"promotional__wrapper\">\n\t\t<div class=\"promotional__content-wrapper\">\n\t\t\t<div class=\"promotional__content\">\n\t\t\t\t\n\n<h2 class=\"wp-block-heading\" id=\"build-more-reliable-agents\">Build more reliable agents<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">See how generated datasets, grader validation, and scalable testing help improve agent evaluation quality.<\/p>\n\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button  is-full-width-on-mobile\"><a data-bi-an=\"Global CTA\" data-bi-ct=\"cta link\" data-bi-id=\"cta-block\" class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/find.codeghost.online\/en-us\/microsoft-365-copilot\/microsoft-copilot-studio\/\">Try Copilot Studio today<\/a><\/div>\n<\/div>\n\n\t\t\t<\/div>\n\t\t<\/div>\n\t\t\t\t\t<div class=\"promotional__media-wrapper\">\n\t\t\t\t<div class=\"promotional__media\">\n\t\t\t\t\t\t\t\t\t\t\t<img decoding=\"async\" width=\"600\" height=\"600\" src=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/CLO23_Collaboration_023-600px.jpg\" class=\"attachment-full size-full\" alt=\"Two IT pros in a modern office space gather around a computer that&#039;s out of frame.\" loading=\"lazy\" srcset=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/CLO23_Collaboration_023-600px.jpg 600w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/CLO23_Collaboration_023-600px-300x300.jpg 300w, https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/CLO23_Collaboration_023-600px-150x150.jpg 150w\" sizes=\"auto, (max-width: 600px) 100vw, 600px\" \/>\t\t\t\t\t\t\t\t\t<\/div>\n\t\t\t<\/div>\n\t\t\t<\/div>\n<\/div>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>An inside look at the data science and evaluation systems helping teams improve agent quality at scale in Copilot Studio.<\/p>\n","protected":false},"author":114,"featured_media":7931,"template":"","meta":{"ms_queue_id":[],"ep_exclude_from_search":false,"_classifai_error":"","_classifai_text_to_speech_error":"","_alt_title":"","ms-ems-related-posts":[],"footnotes":""},"cs-content-type":[933],"cs-topic":[939,940],"coauthors":[1036],"class_list":["post-7917","copilot","type-copilot","status-publish","has-post-thumbnail","hentry","cs-content-type-tips-and-guides","cs-topic-agent-governance","cs-topic-agentic-ai"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v27.4 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Who evaluates the evaluators? The data science behind agent evals | Microsoft Copilot Blog<\/title>\n<meta name=\"description\" content=\"Learn how Copilot Studio approaches AI agent evaluation with generated data, grader validation, and scalable testing that improve agent quality.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Who evaluates the evaluators? The data science behind agent evals | Microsoft Copilot Blog\" \/>\n<meta property=\"og:description\" content=\"Learn how Copilot Studio approaches AI agent evaluation with generated data, grader validation, and scalable testing that improve agent quality.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/\" \/>\n<meta property=\"og:site_name\" content=\"Microsoft Copilot Blog\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-21T23:53:52+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/agent-evals-2-1200.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"675\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:image\" content=\"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/agent-evals-2-1200.jpg\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"11 minutes\" \/>\n\t<meta name=\"twitter:label2\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data2\" content=\"Dikla Dotan\u2011Cohen\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/\"},\"author\":[{\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/author\\\/dikla-dotan-cohen\\\/\",\"@type\":\"Person\",\"@name\":\"Dikla Dotan\u2011Cohen\"}],\"headline\":\"Who evaluates the evaluators? The data science behind agent evals\",\"datePublished\":\"2026-06-11T16:00:00+00:00\",\"dateModified\":\"2026-07-21T23:53:52+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/\"},\"wordCount\":2012,\"publisher\":{\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/wp-content\\\/uploads\\\/2024\\\/08\\\/agent-evals-2-1260.jpg\",\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/\",\"url\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/\",\"name\":\"Who evaluates the evaluators? The data science behind agent evals | Microsoft Copilot Blog\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/wp-content\\\/uploads\\\/2024\\\/08\\\/agent-evals-2-1260.jpg\",\"datePublished\":\"2026-06-11T16:00:00+00:00\",\"dateModified\":\"2026-07-21T23:53:52+00:00\",\"description\":\"Learn how Copilot Studio approaches AI agent evaluation with generated data, grader validation, and scalable testing that improve agent quality.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/#primaryimage\",\"url\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/wp-content\\\/uploads\\\/2024\\\/08\\\/agent-evals-2-1260.jpg\",\"contentUrl\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/wp-content\\\/uploads\\\/2024\\\/08\\\/agent-evals-2-1260.jpg\",\"width\":1260,\"height\":709,\"caption\":\"A person working on a computer in an office, framed by blue graphics and stylized icons for PowerPoint, Word, and Outlook.\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Copilot Studio\",\"item\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/copilot-studio\\\/\"},{\"@type\":\"ListItem\",\"position\":3,\"name\":\"Who evaluates the evaluators? The data science behind agent evals\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/\",\"name\":\"Microsoft Copilot Blog\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/#organization\",\"name\":\"Microsoft Copilot Blog\",\"url\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/wp-content\\\/uploads\\\/2024\\\/05\\\/cropped-microsoft_logo_element.webp\",\"contentUrl\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/wp-content\\\/uploads\\\/2024\\\/05\\\/cropped-microsoft_logo_element.webp\",\"width\":512,\"height\":512,\"caption\":\"Microsoft Copilot Blog\"},\"image\":{\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/#\\\/schema\\\/person\\\/6caf63d1b9a7332f6ae74181bfe65ee4\",\"name\":\"Amy Sitaram\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/396213109a04d7ff30dddd67dbaa7aa6d6c611ec5dbbc5b13c15eeb86a7ac25a?s=96&d=microsoft&r=g5a53981b71c4f0f7cfa4cddf185617d9\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/396213109a04d7ff30dddd67dbaa7aa6d6c611ec5dbbc5b13c15eeb86a7ac25a?s=96&d=microsoft&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/396213109a04d7ff30dddd67dbaa7aa6d6c611ec5dbbc5b13c15eeb86a7ac25a?s=96&d=microsoft&r=g\",\"caption\":\"Amy Sitaram\"},\"url\":\"https:\\\/\\\/find.codeghost.online\\\/en-us\\\/microsoft-copilot\\\/blog\\\/author\\\/amysitaram\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Who evaluates the evaluators? The data science behind agent evals | Microsoft Copilot Blog","description":"Learn how Copilot Studio approaches AI agent evaluation with generated data, grader validation, and scalable testing that improve agent quality.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/","og_locale":"en_US","og_type":"article","og_title":"Who evaluates the evaluators? The data science behind agent evals | Microsoft Copilot Blog","og_description":"Learn how Copilot Studio approaches AI agent evaluation with generated data, grader validation, and scalable testing that improve agent quality.","og_url":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/","og_site_name":"Microsoft Copilot Blog","article_modified_time":"2026-07-21T23:53:52+00:00","og_image":[{"width":1200,"height":675,"url":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/agent-evals-2-1200.jpg","type":"image\/jpeg"}],"twitter_card":"summary_large_image","twitter_image":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/agent-evals-2-1200.jpg","twitter_misc":{"Est. reading time":"11 minutes","Written by":"Dikla Dotan\u2011Cohen"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/#article","isPartOf":{"@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/"},"author":[{"@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/author\/dikla-dotan-cohen\/","@type":"Person","@name":"Dikla Dotan\u2011Cohen"}],"headline":"Who evaluates the evaluators? The data science behind agent evals","datePublished":"2026-06-11T16:00:00+00:00","dateModified":"2026-07-21T23:53:52+00:00","mainEntityOfPage":{"@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/"},"wordCount":2012,"publisher":{"@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/#organization"},"image":{"@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/#primaryimage"},"thumbnailUrl":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/agent-evals-2-1260.jpg","inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/","url":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/","name":"Who evaluates the evaluators? The data science behind agent evals | Microsoft Copilot Blog","isPartOf":{"@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/#primaryimage"},"image":{"@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/#primaryimage"},"thumbnailUrl":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/agent-evals-2-1260.jpg","datePublished":"2026-06-11T16:00:00+00:00","dateModified":"2026-07-21T23:53:52+00:00","description":"Learn how Copilot Studio approaches AI agent evaluation with generated data, grader validation, and scalable testing that improve agent quality.","breadcrumb":{"@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/#primaryimage","url":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/agent-evals-2-1260.jpg","contentUrl":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/08\/agent-evals-2-1260.jpg","width":1260,"height":709,"caption":"A person working on a computer in an office, framed by blue graphics and stylized icons for PowerPoint, Word, and Outlook."},{"@type":"BreadcrumbList","@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/who-evaluates-the-evaluators-the-data-science-behind-agent-evals\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/"},{"@type":"ListItem","position":2,"name":"Copilot Studio","item":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/copilot-studio\/"},{"@type":"ListItem","position":3,"name":"Who evaluates the evaluators? The data science behind agent evals"}]},{"@type":"WebSite","@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/#website","url":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/","name":"Microsoft Copilot Blog","description":"","publisher":{"@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/#organization","name":"Microsoft Copilot Blog","url":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/05\/cropped-microsoft_logo_element.webp","contentUrl":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-content\/uploads\/2024\/05\/cropped-microsoft_logo_element.webp","width":512,"height":512,"caption":"Microsoft Copilot Blog"},"image":{"@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/#\/schema\/person\/6caf63d1b9a7332f6ae74181bfe65ee4","name":"Amy Sitaram","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/396213109a04d7ff30dddd67dbaa7aa6d6c611ec5dbbc5b13c15eeb86a7ac25a?s=96&d=microsoft&r=g5a53981b71c4f0f7cfa4cddf185617d9","url":"https:\/\/secure.gravatar.com\/avatar\/396213109a04d7ff30dddd67dbaa7aa6d6c611ec5dbbc5b13c15eeb86a7ac25a?s=96&d=microsoft&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/396213109a04d7ff30dddd67dbaa7aa6d6c611ec5dbbc5b13c15eeb86a7ac25a?s=96&d=microsoft&r=g","caption":"Amy Sitaram"},"url":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/author\/amysitaram\/"}]}},"bloginabox_animated_featured_image":null,"bloginabox_display_generated_audio":true,"_links":{"self":[{"href":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-json\/wp\/v2\/copilot\/7917","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-json\/wp\/v2\/copilot"}],"about":[{"href":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-json\/wp\/v2\/types\/copilot"}],"author":[{"embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-json\/wp\/v2\/users\/114"}],"version-history":[{"count":19,"href":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-json\/wp\/v2\/copilot\/7917\/revisions"}],"predecessor-version":[{"id":8150,"href":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-json\/wp\/v2\/copilot\/7917\/revisions\/8150"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-json\/wp\/v2\/media\/7931"}],"wp:attachment":[{"href":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-json\/wp\/v2\/media?parent=7917"}],"wp:term":[{"taxonomy":"cs-content-type","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-json\/wp\/v2\/cs-content-type?post=7917"},{"taxonomy":"cs-topic","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-json\/wp\/v2\/cs-topic?post=7917"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/find.codeghost.online\/en-us\/microsoft-copilot\/blog\/wp-json\/wp\/v2\/coauthors?post=7917"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}