Building a Golden Eval Dataset for AI Agents: From Production Traffic to Tool Calling Accuracy
Public benchmarks like MMLU measure general intelligence, but they fail to predict how your specific Agent handles real-world tool usage. This tutorial combines two critical workflows: building a version-controlled "Golden Dataset" from your production logs, and using it to rigorously test tool calling accuracy (selection, arguments, and sequencing) across different models via Celedog.io.
Introduction: Why Production Data Matters
You updated a prompt or switched a model provider, and suddenly your Agent is hallucinating tools or failing to book appointments. Public benchmarks can't catch this. You need a Golden Eval Dataset: a curated set of real user inputs with verified expected outputs, used as a regression test suite.
This guide merges two essential tutorials into one workflow: building that dataset from production traffic, and using it to specifically test Tool Calling Accuracy. We will use Celedog.io to run these evaluations across multiple models seamlessly.
Part 1: Building the Golden Dataset (5 Steps)
A Golden Dataset isn't just a list of questions; it's a version-controlled artifact in your Git repo.
Step 1: Sample Production Traffic
Extract 1-2 weeks of logs. Focus on diverse intents. If 70% of your traffic is "password reset," your dataset should reflect that distribution. Ensure you strip PII (Personally Identifiable Information) before evaluation.
Step 2: Deduplicate & Cluster
Production data is repetitive. Use embedding similarity to cluster similar queries. You want one representative sample for "How do I reset my password?", not fifty. Aim for ~100-1000 samples for a robust regression set.
Step 3: Define Expected Outputs & Rubrics
For tool calling, the "expected output" isn't just text—it's a structured function call. Define your rubric clearly:
- Tool Selection: Did it call
book_meetinginstead ofcancel_meeting? - Argument Accuracy: Did it extract the correct date and time?
- Refusal/Safety: Did it correctly refuse invalid requests?
Step 4: The "Dry Run" & Pruning
Run your current production model against this dataset. If the model fails a case but the expectation was wrong, fix the expectation. If the case is ambiguous, delete it. A Golden Dataset must have high signal-to-noise ratio.
Step 5: Version Control
Commit your dataset.jsonl and rubric.md to Git. This allows you to trace exactly when a regression was introduced.
Part 2: Testing Tool Calling Accuracy
Once your dataset is ready, you need to evaluate how well models handle structured outputs. Tool calling errors usually fall into three categories:
- Hallucinated Tools: Calling a function that doesn't exist.
- Missing Arguments: Calling
get_weatherwithout a location. - Type Errors: Passing a string where an integer is expected.
Part 3: Running Benchmarks via Celedog.io
Instead of managing separate API keys for OpenAI, Anthropic, and Google, use Celedog.io's unified API to benchmark them all against your Golden Dataset.
The Evaluation Script
Here is a minimal TypeScript example using the Celedog.io client to evaluate tool calling:
import OpenAI from "openai";
// Initialize Celedog.io Client
const client = new OpenAI({
apiKey: process.env.CELEDG_API_KEY,
baseURL: "https://celedog.io/v1",
});
const tools = [
{
type: "function",
function: {
name: "get_weather",
description: "Get current weather",
parameters: {
type: "object",
properties: {
location: { type: "string", description: "City name" },
},
required: ["location"],
},
},
},
];
const dataset = [
{ input: "What's the weather in Tokyo?", expected_tool: "get_weather" },
// ... load from your golden dataset
];
async function runEval(modelName: string) {
let passCount = 0;
for (const item of dataset) {
const response = await client.chat.completions.create({
model: modelName, // e.g., "openai/gpt-4o", "anthropic/claude-3.5-sonnet"
messages: [{ role: "user", content: item.input }],
tools: tools,
tool_choice: "auto",
});
const message = response.choices[0].message;
const toolCall = message.tool_calls?.[0];
// Simple check: Did it call the right tool?
if (toolCall?.function.name === item.expected_tool) {
passCount++;
}
}
console.log(` Accuracy: ${passCount / dataset.length * 100}%`);
}
// Run across models
runEval("openai/gpt-4o");
runEval("anthropic/claude-3.5-sonnet");
runEval("google/gemini-pro");
Conclusion
By combining a production-derived Golden Dataset with rigorous tool-calling checks, you move beyond "vibes-based" development. Use Celedog.io to automate this process across the entire model landscape, ensuring your Agent works reliably before it reaches your users.
Last updated September 30, 2026
Where to go next
- Try Celedog — free credits on signup, no card required.
- API documentation
- Per-model pricing
- More Celedog Tutorials