Use this advanced n8n workflow to automatically evaluate the tool usage accuracy of multi-agent AI systems. Leverage the n8n Evaluation node for robust testing.
Download this n8n workflow template and start using it instantly.
AI Developers and Engineers who need automated testing for multi-agent LLM systems.
Data Scientists focused on measuring and improving model reliability.
Users of n8n seeking advanced examples of evaluation capabilities.
QA specialists ensuring accuracy in complex n8n workflow configurations.
Developing reliable AI agents often requires rigorous testing, especially when the agent has access to multiple functions or tools. This sophisticated n8n workflow addresses the challenge of verifying whether an agent correctly chooses and executes the necessary tools for a given query.
By leveraging the specialized n8n Evaluation nodes and connecting them to an external dataset (Google Sheets), this n8n workflow runs predefined questions against a multi-tool agent (which uses Search Agent, Calculator, Web search, and Search_db). The key value of this n8n workflow is the automated assessment: it compares the list of tools the agent should have called (from the dataset) against the list of tools it actually called (from the agent's intermediate steps).
This makes this specific n8n workflow ideal for continuous integration and quality assurance (QA) in your AI development pipeline. If you are looking for robust n8n templates for agent evaluation, this example provides a powerful foundation.
This process starts with an n8n trigger designed specifically for testing or an interactive chat trigger for live usage.
When fetching a dataset row n8n trigger, which pulls a test case (input question and expected tools) from a Google Sheet dataset called 'Tool calling'.Match chat format n8n node to prepare the prompt for the agent.Search Agent n8n node executes the query. This powerful agent has access to multiple tools: Calculator, a Summarizer workflow, a Web search tool (using Firecrawl), and a Qdrant-backed Searchdb retrieval tool. The agent uses the OpenRouter Chat Model as its brain.Evaluating? n8n node checks if the n8n workflow is running in evaluation mode (i.e., triggered by the dataset trigger).Check if tool called n8n node executes a complex expression. This expression compares the expected tools (from the dataset row) against the actual intermediate steps recorded by the Search Agent n8n node, resulting in a boolean toolcalled metric.Set Outputs n8n node records the actual tools called back into the Google Sheet dataset for later review.Evaluation n8n node converts the boolean result into a quantifiable metric (1 for success, 0 for failure) and submits the evaluation results, allowing users to track the tool usage accuracy of the agent across many test cases. This powerful n8n node completes the automated test cycle.To deploy this evaluation n8n workflow, follow these steps:
When fetching a dataset row n8n trigger and Set Outputs n8n node to access your evaluation dataset.OpenRouter Chat Model (or substitute with an OpenAI credential). Searchdb tool.Web search n8n node.Search Agent n8n node correctly references the four tool nodes (Calculator, Summarizer, Web search, Searchdb) and the language model (OpenRouter Chat Model).When fetching a dataset row n8n trigger to point to your specific evaluation Google Sheet document and sheet name, ensuring the expected tool names are present in the dataset rows. When fetching a dataset row (Evaluation Trigger): This n8n trigger initiates the evaluation cycle by fetching a single row (test case) from a Google Sheets dataset. Key configuration includes the Document ID and Sheet Name ('Tool calling').
Search Agent (LangChain Agent n8n node): The core intelligence. It processes the input query, decides which of the four connected tools to use based on the extensive system message, and crucially, is configured to return intermediateSteps which contain the record of tools actually called.
OpenRouter Chat Model (LangChain LM Chat n8n node): Serves as the Large Language Model engine for the Search Agent. This powerful n8n node dictates the agent's reasoning.
Searchdb, Calculator, Summarizer, Web search (Tool n8n nodes): These are the functional tools provided to the agent. The Searchdb uses a Qdrant vector store and requires the Embeddings OpenAI n8n node.
Check if tool called (Set n8n node): This crucial n8n node implements the evaluation logic. It compares $json.toolstocall (from the dataset) against the agent's actual actions in $json.intermediateSteps to generate a boolean success metric.
Evaluation (n8n node): Used twice. First, the Set Outputs instance writes the agent's intermediate results back to the sheet. The final Evaluation instance sets the metric tool_called (0 or 1), completing the automated assessment for this specific n8n workflow iteration.
Use this specialized n8n workflow to rigorously evaluate the accuracy of your Retrieval-Augmented Generation (RAG) system. Measure document groundedness by comparing AI responses against retrieved context using an advanced n8n node setup and OpenAI.

Use this powerful n8n workflow to automate candidate screening. It utilizes an n8n trigger on new Google Sheets submissions, evaluates answers using Azure GPT-4o-mini, and updates final combined scores in your candidate database.

This reusable n8n workflow scores animal advocacy text for performance and preference using deployed Hugging Face Open Paws AI models. A powerful n8n template for parallel AI text analysis.

Automate the evaluation of LLM summarization quality with this n8n workflow. Measures groundedness, conciseness, and fluency using an expert AI judge, essential for production-grade AI automation.

Use this powerful n8n workflow to rigorously evaluate the performance and accuracy of an AI Agent designed for support ticket categorization. Compare generated priority and category against ground truth data using custom evaluation metrics.


Angel Menendez is a Staff Developer Advocate at n8n.io, specializing in low-code tools for cybersecurity workflows. From Puerto Rico, Angel's tech journey began by helping his father translate technical books. He later started a web development business and transitioned from a career as a flight attendant to cybersecurity engineering. His workflows have saved companies significant time. Outside work, Angel enjoys time with his two sons, riding electric bikes, reading, and exploring new places.







































