Learn how to deploy a private Retrieval Augmented Generation (RAG) system. This n8n workflow uses local Ollama models for embeddings and LLM, integrating with Qdrant for vector storage.
Download this n8n workflow template and start using it instantly.
Users looking for self-hosted, private AI solutions, avoiding reliance on external API services like OpenAI.
Engineers and developers interested in building sophisticated Retrieval Augmented Generation (RAG) n8n templates.
Organizations needing to securely chat with internal PDF documents or knowledge bases.
Anyone seeking to leverage advanced LangChain functionalities within an n8n node environment.
This powerful n8n workflow demonstrates how to create a highly effective, local RAG pipeline. It solves the critical problem of answering questions based on proprietary, private data without sending that data to external large language models. The entire system is built upon open-source and local components: using Ollama for both vector embeddings (mxbai-embed-large) and the conversational language model, and Qdrant for scalable vector storage.
The automation is split into two main flows: data ingestion (triggered by an n8n trigger form submission) and conversational retrieval (triggered by an n8n chat trigger). This setup ensures that new knowledge can be easily indexed and subsequently queried by the AI agent, making this one of the most technical and useful n8n templates available for private AI.
This n8n workflow operates in two distinct, interconnected phases:
mxbai-embed-large:latest model to convert the text chunks into dense vector embeddings.ragcollection in the local Qdrant database, completing the ingestion portion of the n8n workflow.retrieve-as-tool mode. When the agent determines it needs external knowledge to answer a user's query, it uses this tool. The tool uses 'Embeddings Ollama1' to convert the user's query into a vector before searching the ragcollection in Qdrant.To deploy this comprehensive n8n workflow, follow these steps:
mxbai-embed-large) and Qdrant running locally or accessible via network.Embeddings Ollama and Ollama Chat Model n8n nodes, ensuring it points correctly to your local Ollama instance.rag_collection).This complex n8n workflow relies on several specialized n8n nodes:
On form submission (n8n trigger):
Function: Serves as the initial n8n trigger for data ingestion. It provides a simple web form.
Key Configuration: Requires a file input field labeled 'File', accepting .pdf format.
When chat message received (n8n trigger):
Function: The n8n trigger that initiates the RAG process when a user sends a chat message.
Key Configuration: Set up to listen for incoming chat requests.
Embeddings Ollama / Embeddings Ollama1:
Function: Generates vector embeddings for document chunks (Ingestion) and search queries (Retrieval).
Key Configuration: Uses the local model mxbai-embed-large:latest via the connected Ollama credential.
Qdrant Vector Store (Insert Mode):
Function: Takes the embedded data and inserts it into the Qdrant database, indexing the knowledge base.
Key Configuration: Mode set to 'insert', targeting the collection ID rag_collection.
Qdrant Vector Store1 (Retrieve as Tool Mode):
Function: Acts as a tool for the AI Agent, allowing it to search the vector database during a conversation.
Key Configuration: Mode set to 'retrieve-as-tool', tool name is retriever.
Recursive Character Text Splitter:
Function: Prepares large documents for embedding by splitting them into smaller, overlapping segments.
Key Configuration: Chunk Size set to 200, Chunk Overlap set to
50.
AI Agent:
Function: The central orchestrator, managing the flow between the chat model, memory, and retrieval tool.
* Key Configuration: System message directs it to use the tool to answer questions. It integrates the Ollama Chat Model, Simple Memory, and the Qdrant tool.
Use this robust n8n workflow to automate the loading of JSON documents from an FTP server, embedding them using OpenAI, and storing the resulting vectors efficiently in a Qdrant database. Essential for RAG systems.

Use this structured n8n workflow to mock complex CRM data, transform contact properties (Name, Email), and prepare the clean output for insertion into any database or Google Sheet.

Automate ETL processes by downloading CSV data, transforming fields using an n8n workflow, and loading records directly into a Snowflake table. Use this pre-built n8n template.

Use this advanced n8n workflow to build a powerful local Agentic RAG system using Ollama, PGVector, and PostgreSQL. Process PDFs, CSVs, and Excel files automatically for intelligent question answering. Ideal for self-hosted n8n templates.

Use this efficient n8n workflow template to create location-based, time-sensitive reminders sent directly to Telegram. Ideal for chores like bin collection, triggered upon arrival home.
