Calculate AI evaluation metrics using this advanced n8n workflow. Determine if an n8n agent correctly utilized a required tool (like Calculator) during execution.
Download this n8n workflow template and start using it instantly.
The challenge in advanced AI applications is ensuring reliability and predictable behavior, especially regarding tool utilization. This specialized n8n workflow provides a robust framework for calculating a critical evaluation metric: whether the AI Agent successfully called the necessary tool for a given question. This n8n templates solution leverages the built-in n8n evaluation features to automatically pull test data (questions and expected tools) from a Google Sheet using the n8n trigger. By analyzing the agent's intermediate steps, the n8n node logic objectively records a 0 or 1 score, making it easy to generate aggregated accuracy reports. This ensures your AI Agent is performing tasks correctly and consistently.
This n8n workflow is designed primarily for automated evaluation but also supports live testing.
tool_called metric, completing the evaluation cycle for that test row.To deploy this n8n workflow and start evaluating your AI Agent:
tooltocall value used in the conditional logic). When fetching a dataset row (Evaluation Trigger): The starting n8n trigger for evaluation runs. It pulls rows from a specified Google Sheets URL, providing the question input and the expected tooltocall for the agent.
AI Agent (Agent Node): The central decision-making component. It uses the connected OpenAI model and available tools (Calculator, Fetch a webpage) to formulate a response. Key configuration: returnIntermediateSteps is True.
OpenAI Chat Model (LM Chat OpenAI): Provides the large language model capabilities, configured here using gpt-4o-mini via an OpenAI API credential.
Calculator / Fetch a webpage (Tool Nodes): These are the tools the AI Agent can choose to execute, depending on the input query.
Check if tool called (Set Node): This critical n8n node calculates the metric. It uses an expression ($json.intermediateSteps.filter(...)) to check if the executed steps contain the expected tool name from the dataset, resulting in a boolean toolcalled output.
Evaluation (Evaluation Node): The final step of the evaluation branch. It records the calculated metric (toolcalled converted to a number) back to the n8n evaluation results dashboard.
Use this powerful n8n workflow to rigorously evaluate the performance and accuracy of an AI Agent designed for support ticket categorization. Compare generated priority and category against ground truth data using custom evaluation metrics.

Automate the evaluation of LLM summarization quality with this n8n workflow. Measures groundedness, conciseness, and fluency using an expert AI judge, essential for production-grade AI automation.

Use this advanced n8n workflow to automatically evaluate the tool usage accuracy of multi-agent AI systems. Leverage the n8n Evaluation node for robust testing.

Use this expert n8n workflow template to automatically calculate the semantic similarity and correctness score of AI-generated answers against ground truth data using the specialized n8n node for evaluations.

Use this comprehensive n8n workflow to calculate the RAG document relevance evaluation metric. Leverage the n8n Evaluation nodes and OpenAI to assess whether retrieved documents are relevant to user questions.









































