Building a Golden Eval Dataset for AI Agents: From Production Traffic to Tool Calling Accuracy

Public benchmarks like MMLU measure general intelligence, but they fail to predict how your specific Agent handles real-world tool usage. This tutorial combines two critical workflows: building a version-controlled "Golden Dataset" from your production logs, and using it to rigorously test tool calling accuracy (selection, arguments, and sequencing) across different models via Celedog.io.

Introduction: Why Production Data Matters

You updated a prompt or switched a model provider, and suddenly your Agent is hallucinating tools or failing to book appointments. Public benchmarks can't catch this. You need a Golden Eval Dataset: a curated set of real user inputs with verified expected outputs, used as a regression test suite.

This guide merges two essential tutorials into one workflow: building that dataset from production traffic, and using it to specifically test Tool Calling Accuracy. We will use Celedog.io to run these evaluations across multiple models seamlessly.

Part 1: Building the Golden Dataset (5 Steps)

A Golden Dataset isn't just a list of questions; it's a version-controlled artifact in your Git repo.

Step 1: Sample Production Traffic

Extract 1-2 weeks of logs. Focus on diverse intents. If 70% of your traffic is "password reset," your dataset should reflect that distribution. Ensure you strip PII (Personally Identifiable Information) before evaluation.

Step 2: Deduplicate & Cluster

Production data is repetitive. Use embedding similarity to cluster similar queries. You want one representative sample for "How do I reset my password?", not fifty. Aim for ~100-1000 samples for a robust regression set.

Step 3: Define Expected Outputs & Rubrics

For tool calling, the "expected output" isn't just text—it's a structured function call. Define your rubric clearly:

  • Tool Selection: Did it call book_meeting instead of cancel_meeting?
  • Argument Accuracy: Did it extract the correct date and time?
  • Refusal/Safety: Did it correctly refuse invalid requests?

Step 4: The "Dry Run" & Pruning

Run your current production model against this dataset. If the model fails a case but the expectation was wrong, fix the expectation. If the case is ambiguous, delete it. A Golden Dataset must have high signal-to-noise ratio.

Step 5: Version Control

Commit your dataset.jsonl and rubric.md to Git. This allows you to trace exactly when a regression was introduced.

Part 2: Testing Tool Calling Accuracy

Once your dataset is ready, you need to evaluate how well models handle structured outputs. Tool calling errors usually fall into three categories:

  1. Hallucinated Tools: Calling a function that doesn't exist.
  2. Missing Arguments: Calling get_weather without a location.
  3. Type Errors: Passing a string where an integer is expected.

Part 3: Running Benchmarks via Celedog.io

Instead of managing separate API keys for OpenAI, Anthropic, and Google, use Celedog.io's unified API to benchmark them all against your Golden Dataset.

The Evaluation Script

Here is a minimal TypeScript example using the Celedog.io client to evaluate tool calling:

import OpenAI from "openai";

// Initialize Celedog.io Client
const client = new OpenAI({
  apiKey: process.env.CELEDG_API_KEY,
  baseURL: "https://celedog.io/v1",
});

const tools = [
  {
    type: "function",
    function: {
      name: "get_weather",
      description: "Get current weather",
      parameters: {
        type: "object",
        properties: {
          location: { type: "string", description: "City name" },
        },
        required: ["location"],
      },
    },
  },
];

const dataset = [
  { input: "What's the weather in Tokyo?", expected_tool: "get_weather" },
  // ... load from your golden dataset
];

async function runEval(modelName: string) {
  let passCount = 0;
  
  for (const item of dataset) {
    const response = await client.chat.completions.create({
      model: modelName, // e.g., "openai/gpt-4o", "anthropic/claude-3.5-sonnet"
      messages: [{ role: "user", content: item.input }],
      tools: tools,
      tool_choice: "auto",
    });

    const message = response.choices[0].message;
    const toolCall = message.tool_calls?.[0];
    
    // Simple check: Did it call the right tool?
    if (toolCall?.function.name === item.expected_tool) {
      passCount++;
    }
  }
  
  console.log(` Accuracy: ${passCount / dataset.length * 100}%`);
}

// Run across models
runEval("openai/gpt-4o");
runEval("anthropic/claude-3.5-sonnet");
runEval("google/gemini-pro");

Conclusion

By combining a production-derived Golden Dataset with rigorous tool-calling checks, you move beyond "vibes-based" development. Use Celedog.io to automate this process across the entire model landscape, ensuring your Agent works reliably before it reaches your users.


Last updated September 30, 2026

Where to go next