Semantic Caching for LLMs using Redis and HuggingFace - n8n Workflow

Reduce large language model costs and improve latency using this advanced n8n workflow. It implements a semantic cache with Redis and HuggingFace embeddings.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?


  1. Users and developers highly focused on optimizing AI API costs.

  2. Organizations requiring low-latency responses for AI-powered chat interfaces.

  3. Individuals seeking advanced examples of integrating vector databases like Redis into an n8n workflow.

  4. Anyone looking for high-performance n8n templates for generative AI projects.

Overview

Calling Large Language Models repeatedly, even with slight variations in prompts, can become expensive and slow. This intelligent n8n workflow solves this by introducing a semantic caching layer. Instead of immediately routing every query to the LLM, the system first checks a Redis Vector Store for previously answered, semantically similar questions. This specific n8n workflow leverages HuggingFace for efficient text embeddings, ensuring fast and accurate similarity matching. If a cache hit occurs, the response is delivered instantly—saving both time and expense. This core logic, orchestrated perfectly within n8n, significantly boosts the efficiency of any AI application built on n8n.

How it Works

The flow begins with the When chat message received n8n trigger.


  1. Trigger and Search: Upon receiving input, the n8n workflow immediately uses the Embeddings HuggingFace Inference n8n node to vectorize the prompt and performs a semantic search via the Check for similar prompts Redis n8n node.

  2. Cache Decision: The Analyze results from store custom n8n node filters the results based on a predefined distanceThreshold (set to 0.3). This ensures only highly relevant matches qualify as a cache hit.

  3. Conditional Routing: The Is this a cache hit? n8n node acts as the decision point.

If True (Cache Hit): The cached reply is returned instantly using the Respond to Chat (from semantic cache) n8n node.
If False (Cache Miss): The flow proceeds to the LLM generation section.

  1. LLM Generation: The LLM Agent generates a fresh response using the OpenAI Chat Model (GPT-4.1-mini) while maintaining context using Redis Chat Memory.

  2. Cache Storage: Before responding, the new prompt-response pair is prepared with the Add response as metadata n8n node and then vectorized (using the second HuggingFace n8n node) and saved into the Redis index by the Store entry in cache n8n node.

  3. Final Response: The newly generated content is sent back to the user via the Respond to Chat (from LLM) n8n node. This robust n8n workflow ensures continuous learning and cost optimization.

Installation Guide

To deploy this powerful semantic caching n8n template, follow these steps:


  1. Import the n8n Workflow: Copy the provided JSON data and import it directly into your running n8n instance.

  2. Credential Setup: This n8n workflow requires credentials for three services:

OpenAI: Needed for the LLM models used in the LLM Agent n8n node.
HuggingFace Inference: Required for the embedding n8n nodes that convert text to vectors for semantic searching.
* Redis: Configure connection details for the Redis Chat Memory and the two Redis Vector Store n8n nodes. Ensure your Redis server version is compatible and the index name (chat_cache) is correctly configured.

  1. Tuning the Cache: For optimal performance, you may need to adjust the distanceThreshold value within the Analyze results from store n8n code node. A lower threshold means stricter matching (fewer cache hits, higher relevance), while a higher threshold means more cache hits but potentially less accurate replies.

  2. Activation: Activate the n8n workflow. Since the entry point is an n8n trigger, ensure your chat application is configured to send messages to the exposed webhook URL.

Node Details

When chat message received (Chat Trigger): This is the critical starting n8n trigger for the entire automation. It listens for incoming chat messages to start the cache lookup process.
Check for similar prompts (Redis Vector Store n8n node): This n8n node performs the core vector similarity search against the Redis cache to determine if a semantically equivalent query exists. It requires an ai_embedding connection.
Analyze results from store (Code n8n node): A custom n8n node that implements the scoring logic. It filters the vector search results, selecting only the best match that falls below the configurable distanceThreshold (0.3).
Is this a cache hit? (If n8n node): Controls the flow, diverting execution based on whether the Analyze results from store n8n node passed a cached document (hit) or not (miss).
LLM Agent (n8n node): Used exclusively on a cache miss. This n8n node orchestrates the LLM interaction using GPT-4.1-mini and the dedicated Redis Chat Memory to generate a fresh, contextualized response.
Embeddings HuggingFace Inference (n8n node): Two instances of this n8n node are used: one for calculating embeddings during the search phase, and one for calculating embeddings when inserting new data into the cache, maximizing the utility of the semantic search capability provided by this n8n workflow.


  • Store entry in cache (Redis Vector Store n8n node): This n8n node ensures that every newly generated LLM response is stored efficiently as an embedded vector, constantly improving the coverage of the semantic cache.

Related n8n Workflows

Free

Nodes: 12 Nodes
Updated: December 26 2025
View all
Created by

Software engineer @ Redis Inc.

Featured*