AI Output Semantic Correctness Evaluation Metric - n8n Workflow

Use this expert n8n workflow template to automatically calculate the semantic similarity and correctness score of AI-generated answers against ground truth data using the specialized n8n node for evaluations.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?


  • AI developers needing to benchmark and evaluate Language Model performance.

  • Data scientists implementing RAG systems that require rigorous quality assurance.

  • Technical writers seeking robust n8n templates for advanced metric-based testing.

  • n8n users focused on building high-quality, auditable AI agents.

Overview

Evaluating the performance of complex AI agents is crucial. This n8n workflow provides a robust solution for calculating the 'Correctness Metric'—determining if an AI-generated response has the same semantic meaning as a known, trusted reference answer.

This specific n8n workflow is designed to run in evaluation mode, triggered by a test dataset (a Google Sheet containing questions and reference answers). The workflow first uses an AI Agent to generate a concise output based on the question. It then leverages the power of GPT-4o-mini via a specialized n8n node to act as an impartial evaluator, comparing the output against the ground truth and returning a detailed analysis and a final similarity score (1-5).

This setup ensures that your n8n workflow automations involving AI are reliable, accurate, and measurable.

How it Works

This automation is primarily designed to function via the specialized n8n Evaluation Trigger, pulling test data iteratively from a Google Sheet, although it also includes a standard chat trigger path for non-evaluation runs.


  1. Dataset Input: The n8n workflow starts with the 'When fetching a dataset row' n8n trigger, which reads test data (questions and reference answers) from the configured Google Sheet URL row by row.

  2. Input Formatting: The 'Match chat format' n8n node transforms the incoming dataset question field into the chatInput variable required by the agent.

  3. AI Generation: The 'AI Agent' n8n node, powered by the 'OpenAI Chat Model' (using gpt-4o-mini), processes the question and generates a highly concise, one-sentence answer.

  4. Flow Control: The 'Evaluating?' n8n node checks if the n8n workflow is currently running in evaluation mode. If not (e.g., triggered by a live chat input), the flow is diverted to 'Return chat response' to save computation cost.

  5. Metric Calculation: If evaluating, the 'Calculate correctness metric' n8n node takes both the AI-generated output and the reference_answer from the dataset. It uses a sophisticated system prompt, directing GPT-4o-mini to analyze factual accuracy, semantic similarity, and provide a score between 1 (Not similar) and 5 (Highly similar) in a strict JSON format.

  6. Metric Storage: Finally, the 'Set metrics' n8n node extracts the numerical score from the JSON output of the previous n8n node and registers it under the metric name similarity, completing the evaluation cycle for that data row.

Installation Guide


  1. Import the n8n workflow: Copy the provided JSON and paste it into your n8n instance via the 'New' menu > 'Import from JSON'.

  2. OpenAI Credentials: Locate the 'OpenAI Chat Model' and 'Calculate correctness metric' n8n nodes. Update the 'openAiApi' credentials with your valid OpenAI API key.

  3. Google Sheets Setup: The 'When fetching a dataset row' n8n trigger is currently configured to read a specific public test sheet. If you wish to use your own test data, replace the URL and ensure your Google Sheets OAuth2 API credentials are correctly set up and selected in this n8n trigger node.

  4. Running the n8n workflow: To execute the evaluation, you must run this n8n workflow template within the n8n Evaluation environment, typically initiated from a parent evaluation view, rather than running it manually from the canvas.

Node Details

When fetching a dataset row (n8n trigger): The primary n8n trigger for evaluation mode. It iterates over a specified Google Sheet dataset, providing inputs (question, referenceanswer) for assessment.
OpenAI Chat Model (LLM n8n node): Defines the Language Model used for generation, configured as gpt-4o-mini in this n8n workflow.
AI Agent (Langchain n8n node): The core intelligence, configured to generate a single, concise sentence answer based on the dataset question.
Evaluating? (Evaluation n8n node): A flow control utility. It checks the environment context to determine if the workflow is running in an official evaluation session. This saves cost by preventing metric calculation during simple test runs.
Calculate correctness metric (OpenAI n8n node): The crucial metric node. It employs a highly detailed system prompt using gpt-4o-mini to compare the generated output against the ground truth based on five criteria and returns the result as a JSON object containing the score and detailed extendedreasoning.
Set metrics (Evaluation n8n node): The final n8n node in the evaluation path. It captures the numerical score output by the metric calculation step and records it as the similarity metric for the current evaluation run.

Related n8n Workflows

Free

Nodes: 9 Nodes
Updated: December 26 2025
View all
Created by
David Roberts
David Roberts

Featured*