AI Agent Tool Usage Evaluation Metric Calculator - n8n Workflow

Calculate AI evaluation metrics using this advanced n8n workflow. Determine if an n8n agent correctly utilized a required tool (like Calculator) during execution.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?


  • AI Developers and Prompt Engineers needing quantifiable metrics for agent performance.

  • Data Scientists responsible for validating AI model behavior.

  • Users building reliable n8n templates involving complex AI routing and tools.

  • Anyone looking to implement advanced AI evaluations within an n8n workflow.

Overview

The challenge in advanced AI applications is ensuring reliability and predictable behavior, especially regarding tool utilization. This specialized n8n workflow provides a robust framework for calculating a critical evaluation metric: whether the AI Agent successfully called the necessary tool for a given question. This n8n templates solution leverages the built-in n8n evaluation features to automatically pull test data (questions and expected tools) from a Google Sheet using the n8n trigger. By analyzing the agent's intermediate steps, the n8n node logic objectively records a 0 or 1 score, making it easy to generate aggregated accuracy reports. This ensures your AI Agent is performing tasks correctly and consistently.

How it Works

This n8n workflow is designed primarily for automated evaluation but also supports live testing.


  1. Triggering the Evaluation: The process begins with the "When fetching a dataset row" n8n trigger, which pulls a question and the name of the tool the AI Agent is expected to call from an external Google Sheet dataset.

  2. Input Formatting: The "Match chat format" n8n node prepares the input question for the AI Agent.

  3. AI Agent Execution: The core "AI Agent" node receives the question. It uses the "OpenAI Chat Model" (GPT-4o mini) and has access to various tools, including "Calculator" and "Fetch a webpage". Crucially, the agent is configured to return intermediate steps.

  4. Flow Control: The "Evaluating?" n8n node checks if the current execution is part of an official evaluation run. If not, the flow branches to a simple return node.

  5. Metric Calculation (Core Logic): If in evaluation mode, the "Check if tool called" n8n node executes. This powerful n8n node uses a complex expression to filter the agent's intermediate steps. It verifies if the list of executed actions contains an action where the called tool name matches the expected tool name fetched by the n8n trigger.

  6. Recording the Metric: The final "Evaluation" n8n node registers the boolean result (converted to 1 for True, 0 for False) as the tool_called metric, completing the evaluation cycle for that test row.

Installation Guide

To deploy this n8n workflow and start evaluating your AI Agent:


  1. Import the JSON: Copy the n8n workflow JSON and paste it into your n8n instance using the 'New' menu -> 'Import from JSON'.

  2. Configure Credentials: You will need two primary credentials:

OpenAI: Set up credentials for the "OpenAI Chat Model" n8n node.
Google Sheets: Configure OAuth2 credentials for the "When fetching a dataset row" n8n trigger to access the test dataset (ensure the service account has read access).

  1. Agent Setup: Verify that the "AI Agent" n8n node has the 'Return intermediate steps' option enabled, as this is essential for inspecting the tools used.

  2. Dataset Review: The "When fetching a dataset row" n8n trigger is pre-configured to point to a public example dataset. If you use your own dataset, ensure the column names match the expected inputs (especially the field containing the tooltocall value used in the conditional logic).

  3. Activation: Save the n8n workflow. Run a test evaluation to confirm the metrics are being calculated correctly.

Node Details

When fetching a dataset row (Evaluation Trigger): The starting n8n trigger for evaluation runs. It pulls rows from a specified Google Sheets URL, providing the question input and the expected tooltocall for the agent.
AI Agent (Agent Node): The central decision-making component. It uses the connected OpenAI model and available tools (Calculator, Fetch a webpage) to formulate a response. Key configuration: returnIntermediateSteps is True.
OpenAI Chat Model (LM Chat OpenAI): Provides the large language model capabilities, configured here using gpt-4o-mini via an OpenAI API credential.
Calculator / Fetch a webpage (Tool Nodes): These are the tools the AI Agent can choose to execute, depending on the input query.
Check if tool called (Set Node): This critical n8n node calculates the metric. It uses an expression ($json.intermediateSteps.filter(...)) to check if the executed steps contain the expected tool name from the dataset, resulting in a boolean toolcalled output.
Evaluation (Evaluation Node): The final step of the evaluation branch. It records the calculated metric (tool
called converted to a number) back to the n8n evaluation results dashboard.

Related n8n Workflows

Free

Nodes: 10 Nodes
Updated: December 26 2025
View all
Created by
David Roberts
David Roberts

Featured*