Reduce large language model costs and improve latency using this advanced n8n workflow. It implements a semantic cache with Redis and HuggingFace embeddings.
Download this n8n workflow template and start using it instantly.
Calling Large Language Models repeatedly, even with slight variations in prompts, can become expensive and slow. This intelligent n8n workflow solves this by introducing a semantic caching layer. Instead of immediately routing every query to the LLM, the system first checks a Redis Vector Store for previously answered, semantically similar questions. This specific n8n workflow leverages HuggingFace for efficient text embeddings, ensuring fast and accurate similarity matching. If a cache hit occurs, the response is delivered instantly—saving both time and expense. This core logic, orchestrated perfectly within n8n, significantly boosts the efficiency of any AI application built on n8n.
The flow begins with the When chat message received n8n trigger.
Embeddings HuggingFace Inference n8n node to vectorize the prompt and performs a semantic search via the Check for similar prompts Redis n8n node.Analyze results from store custom n8n node filters the results based on a predefined distanceThreshold (set to 0.3). This ensures only highly relevant matches qualify as a cache hit.Is this a cache hit? n8n node acts as the decision point. Respond to Chat (from semantic cache) n8n node.LLM Agent generates a fresh response using the OpenAI Chat Model (GPT-4.1-mini) while maintaining context using Redis Chat Memory.Add response as metadata n8n node and then vectorized (using the second HuggingFace n8n node) and saved into the Redis index by the Store entry in cache n8n node. Respond to Chat (from LLM) n8n node. This robust n8n workflow ensures continuous learning and cost optimization.To deploy this powerful semantic caching n8n template, follow these steps:
LLM Agent n8n node.Redis Chat Memory and the two Redis Vector Store n8n nodes. Ensure your Redis server version is compatible and the index name (chat_cache) is correctly configured.distanceThreshold value within the Analyze results from store n8n code node. A lower threshold means stricter matching (fewer cache hits, higher relevance), while a higher threshold means more cache hits but potentially less accurate replies. When chat message received (Chat Trigger): This is the critical starting n8n trigger for the entire automation. It listens for incoming chat messages to start the cache lookup process.
Check for similar prompts (Redis Vector Store n8n node): This n8n node performs the core vector similarity search against the Redis cache to determine if a semantically equivalent query exists. It requires an ai_embedding connection.
Analyze results from store (Code n8n node): A custom n8n node that implements the scoring logic. It filters the vector search results, selecting only the best match that falls below the configurable distanceThreshold (0.3).
Is this a cache hit? (If n8n node): Controls the flow, diverting execution based on whether the Analyze results from store n8n node passed a cached document (hit) or not (miss).
LLM Agent (n8n node): Used exclusively on a cache miss. This n8n node orchestrates the LLM interaction using GPT-4.1-mini and the dedicated Redis Chat Memory to generate a fresh, contextualized response.
Embeddings HuggingFace Inference (n8n node): Two instances of this n8n node are used: one for calculating embeddings during the search phase, and one for calculating embeddings when inserting new data into the cache, maximizing the utility of the semantic search capability provided by this n8n workflow.
Use this sophisticated n8n workflow template utilizing Redis to manage concurrency, prevent race conditions, and ensure reliable, idempotent execution of critical tasks triggered by webhooks.

Use this powerful n8n workflow to automatically scrape Reddit for freelance job leads and track unique opportunities directly in Google Sheets. Maximize lead generation efficiency.

Automate Reddit scraping for WWDC25 discussions, classify content, run Google Gemini sentiment analysis, and save results to Google Sheets using this powerful n8n workflow.

Use this powerful n8n workflow to automatically decrease meeting no-show rates. It extracts Calendly data using AI, calculates optimal reminder times, and sends timely email and WhatsApp reminders using a sophisticated n8n node structure.

Use this n8n workflow to automatically search, vet, and score registered companies in specific German regions using the Implisense API, generating high-quality sales leads.








































