Use this expert n8n workflow template to automatically calculate the semantic similarity and correctness score of AI-generated answers against ground truth data using the specialized n8n node for evaluations.
Download this n8n workflow template and start using it instantly.
Evaluating the performance of complex AI agents is crucial. This n8n workflow provides a robust solution for calculating the 'Correctness Metric'—determining if an AI-generated response has the same semantic meaning as a known, trusted reference answer.
This specific n8n workflow is designed to run in evaluation mode, triggered by a test dataset (a Google Sheet containing questions and reference answers). The workflow first uses an AI Agent to generate a concise output based on the question. It then leverages the power of GPT-4o-mini via a specialized n8n node to act as an impartial evaluator, comparing the output against the ground truth and returning a detailed analysis and a final similarity score (1-5).
This setup ensures that your n8n workflow automations involving AI are reliable, accurate, and measurable.
This automation is primarily designed to function via the specialized n8n Evaluation Trigger, pulling test data iteratively from a Google Sheet, although it also includes a standard chat trigger path for non-evaluation runs.
question field into the chatInput variable required by the agent.gpt-4o-mini), processes the question and generates a highly concise, one-sentence answer.output and the reference_answer from the dataset. It uses a sophisticated system prompt, directing GPT-4o-mini to analyze factual accuracy, semantic similarity, and provide a score between 1 (Not similar) and 5 (Highly similar) in a strict JSON format.score from the JSON output of the previous n8n node and registers it under the metric name similarity, completing the evaluation cycle for that data row. When fetching a dataset row (n8n trigger): The primary n8n trigger for evaluation mode. It iterates over a specified Google Sheet dataset, providing inputs (question, referenceanswer) for assessment.
OpenAI Chat Model (LLM n8n node): Defines the Language Model used for generation, configured as gpt-4o-mini in this n8n workflow.
AI Agent (Langchain n8n node): The core intelligence, configured to generate a single, concise sentence answer based on the dataset question.
Evaluating? (Evaluation n8n node): A flow control utility. It checks the environment context to determine if the workflow is running in an official evaluation session. This saves cost by preventing metric calculation during simple test runs.
Calculate correctness metric (OpenAI n8n node): The crucial metric node. It employs a highly detailed system prompt using gpt-4o-mini to compare the generated output against the ground truth based on five criteria and returns the result as a JSON object containing the score and detailed extendedreasoning.
Set metrics (Evaluation n8n node): The final n8n node in the evaluation path. It captures the numerical score output by the metric calculation step and records it as the similarity metric for the current evaluation run.
Calculate AI evaluation metrics using this advanced n8n workflow. Determine if an n8n agent correctly utilized a required tool (like Calculator) during execution.

Use this powerful n8n workflow to rigorously evaluate the performance and accuracy of an AI Agent designed for support ticket categorization. Compare generated priority and category against ground truth data using custom evaluation metrics.

Automate the evaluation of LLM summarization quality with this n8n workflow. Measures groundedness, conciseness, and fluency using an expert AI judge, essential for production-grade AI automation.

Use this comprehensive n8n workflow to calculate the RAG document relevance evaluation metric. Leverage the n8n Evaluation nodes and OpenAI to assess whether retrieved documents are relevant to user questions.

Calculate AI performance using text similarity (Levenshtein distance) on extracted handwritten code. This advanced n8n workflow uses OpenAI GPT-4o and custom logic for precise metric reporting.









































