Analyze AI quality with this robust n8n workflow. It uses OpenAI embeddings to calculate Answer Similarity scores against ground truth, ideal for RAG evaluations. Get started with n8n templates today.
Download this n8n workflow template and start using it instantly.
AI/ML Engineers needing automated response quality measurement.
Developers creating and testing Retrieval-Augmented Generation (RAG) systems.
Users of n8n seeking advanced evaluation tools for their AI-driven n8n workflow.
Data Scientists requiring repeatable scoring metrics for closed-ended AI questions.
Evaluating the consistency and accuracy of AI agents is crucial. This advanced n8n workflow provides a structured approach to measuring Answer Similarity, a key metric often used in RAG evaluations. The automation fetches a dataset row (input question and ground truth answer) via the specialized n8n trigger, runs the query through an AI Agent powered by an OpenAI model (gpt-4.1-mini), and then calculates the score.
Unlike simple keyword matching, this n8n template uses OpenAI’s text-embedding-3-small model to convert both the AI response and the ground truth examples into vector embeddings. The workflow then uses custom JavaScript in an n8n node to compute the cosine similarity, providing a precise numerical representation of how closely the AI’s answer matches the expected results. This is an essential n8n node combination for ensuring high quality in production AI applications.
This n8n workflow operates primarily in two modes: production (triggered by a chat message) and evaluation (triggered manually using the Evaluation Trigger n8n node).
When fetching a dataset row n8n trigger pulls test data (input query and ground truth) from a designated Google Sheet. Alternatively, the When chat message received n8n trigger handles live inputs.AI Agent, which utilizes the OpenAI Chat Model1 (gpt-4.1-mini) to generate an answer.Evaluation n8n node checks if the workflow is running in evaluation mode. If yes, it proceeds to calculation.Set Input Fields n8n node prepares the AI answer and splits the groundTruth (which may contain multiple expected answers separated by newlines) for itemization.HTTP Request n8n nodes: once for the AI answer and once for each groundTruth item, generating vector embeddings.Aggregate and Create Embeddings Result).Calculate Similarity Score n8n node (a Code node) which executes custom JavaScript to perform the cosine similarity calculation between the AI answer embedding and all ground truth embeddings, returning the average score.score, are written back to the Google Sheet using the Update Output and Update Metrics n8n nodes, completing the evaluation n8n workflow.To deploy this comprehensive n8n template, follow these steps:
OpenAI Chat Model1 n8n node and both Get Embeddings n8n nodes.When fetching a dataset row n8n trigger and the Update Output n8n node. These credentials need read/write access to your evaluation dataset.When fetching a dataset row n8n node, verify the Document ID and Sheet Name match your evaluation data sheet (as shown in the default n8n workflow configuration). When fetching a dataset row (Evaluation Trigger): This critical n8n trigger initiates the testing process by fetching rows from the 'Similarity' Google Sheet. It defines the dataset used for quality assessment.
OpenAI Chat Model1 (LangChain Chat OpenAI): Serves as the Large Language Model for the AI Agent, configured here to use gpt-4.1-mini to generate the initial response to the input query.
Set Input Fields (Set n8n node): Prepares the data structure for scoring, mapping the AI output to answer and splitting the fetched ground truth text into an array.
Get Embeddings / Get Embeddings1 (HTTP Request n8n node): Calls the https://api.openai.com/v1/embeddings endpoint using the text-embedding-3-small model. This is key for transforming both the answer and ground truths into high-dimensional vectors for similarity analysis in this n8n workflow.
GroundTruth to Items (Split Out n8n node): Ensures that if multiple ground truth answers exist, each one is processed individually to receive its own embedding vector.
Calculate Similarity Score (Code n8n node): Contains the custom JavaScript function cosineSimilarity. This n8n node performs the core mathematical calculation, comparing the answer embedding vector against all ground truth embedding vectors and calculating the average score.
score) back to the originating evaluation system or linked Google Sheet, finalizing this n8n workflow execution.Calculate AI performance using text similarity (Levenshtein distance) on extracted handwritten code. This advanced n8n workflow uses OpenAI GPT-4o and custom logic for precise metric reporting.

Use this comprehensive n8n workflow to calculate the RAG document relevance evaluation metric. Leverage the n8n Evaluation nodes and OpenAI to assess whether retrieved documents are relevant to user questions.

Use this expert n8n workflow template to automatically calculate the semantic similarity and correctness score of AI-generated answers against ground truth data using the specialized n8n node for evaluations.

Automate event registration using a powerful n8n workflow. Profile attendees with GPT-4o, route VIP alerts via Gmail, and log detailed scores to Google Sheets. Use these advanced n8n templates today.

A powerful n8n workflow for managing event ticketing end-to-end: registration, automated QR code generation via Google Sheets, Gmail delivery, and real-time check-in validation.


Freelance consultant based in the UK specialising in AI-powered automations. I work with select clients tackling their most challenging projects. For business enquiries, send me an email at [email protected] LinkedIn: https://www.linkedin.com/in/jimleuk/ X/Twitter: https://x.com/jimle_uk







































