Local RAG Chatbot Pipeline using Ollama and Qdrant - n8n Workflow

Learn how to deploy a private Retrieval Augmented Generation (RAG) system. This n8n workflow uses local Ollama models for embeddings and LLM, integrating with Qdrant for vector storage.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?

Users looking for self-hosted, private AI solutions, avoiding reliance on external API services like OpenAI.
Engineers and developers interested in building sophisticated Retrieval Augmented Generation (RAG) n8n templates.
Organizations needing to securely chat with internal PDF documents or knowledge bases.
Anyone seeking to leverage advanced LangChain functionalities within an n8n node environment.

Overview

This powerful n8n workflow demonstrates how to create a highly effective, local RAG pipeline. It solves the critical problem of answering questions based on proprietary, private data without sending that data to external large language models. The entire system is built upon open-source and local components: using Ollama for both vector embeddings (mxbai-embed-large) and the conversational language model, and Qdrant for scalable vector storage.

The automation is split into two main flows: data ingestion (triggered by an n8n trigger form submission) and conversational retrieval (triggered by an n8n chat trigger). This setup ensures that new knowledge can be easily indexed and subsequently queried by the AI agent, making this one of the most technical and useful n8n templates available for private AI.

How it Works

This n8n workflow operates in two distinct, interconnected phases:

Phase 1: Data Ingestion


  1. Ingestion Trigger: The process starts with the 'On form submission' n8n trigger. A user uploads a file (configured specifically for PDFs) via the generated n8n form.

  2. Data Loading & Splitting: The 'Default Data Loader' handles the binary file. The data then passes through the 'Recursive Character Text Splitter' n8n node, which breaks the document into smaller, manageable chunks (200 characters with 50 character overlap).

  3. Embedding Generation: The 'Embeddings Ollama' n8n node connects to the local Ollama service, using the mxbai-embed-large:latest model to convert the text chunks into dense vector embeddings.

  4. Vector Storage: The 'Qdrant Vector Store' n8n node inserts these vectors and their corresponding text chunks into the specified ragcollection in the local Qdrant database, completing the ingestion portion of the n8n workflow.

Phase 2: Conversational Retrieval (RAG)


  1. Chat Trigger: The conversation starts when a message is received via the 'When chat message received' n8n trigger.

  2. Core Agent: The 'AI Agent' n8n node manages the conversation flow. It is configured with a system message guiding it to use the retrieval tool.

  3. LLM and Memory: The agent uses the 'Ollama Chat Model' (connected to the local Ollama service) as its brain, and the 'Simple Memory' n8n node maintains conversational context.

  4. Retrieval Tool: The agent is given access to 'Qdrant Vector Store1', which is configured in retrieve-as-tool mode. When the agent determines it needs external knowledge to answer a user's query, it uses this tool. The tool uses 'Embeddings Ollama1' to convert the user's query into a vector before searching the ragcollection in Qdrant.

  5. Response: The retrieved context is passed back to the Ollama Chat Model, allowing the agent to generate a highly accurate, context-aware answer based on the ingested documents. This sophisticated use of an n8n node setup creates a complete RAG loop.

Installation Guide

To deploy this comprehensive n8n workflow, follow these steps:


  1. Setup Prerequisites: Ensure you have local instances of Ollama (running the necessary embedding and chat models, e.g., mxbai-embed-large) and Qdrant running locally or accessible via network.

  2. Import the n8n Workflow: Copy the provided JSON data and paste it directly into your n8n instance using the 'New' -> 'Import from JSON' feature.

  3. Configure Credentials:

Ollama Credentials: Create or update your 'Local Ollama service' credential for both the Embeddings Ollama and Ollama Chat Model n8n nodes, ensuring it points correctly to your local Ollama instance.
Qdrant Credentials: Create or update your 'Local QdrantApi database' credential for both 'Qdrant Vector Store' n8n nodes, ensuring proper connectivity to your local Qdrant instance.

  1. Verify Collection Name: Confirm that both 'Qdrant Vector Store' n8n nodes are pointing to the correct collection name (rag_collection).

  2. Activate and Test: Activate the n8n workflow. Use the unique URL provided by the 'On form submission' n8n trigger to upload a PDF file first. Then, use the unique URL provided by the 'When chat message received' n8n trigger to start chatting with your data.

Node Details

This complex n8n workflow relies on several specialized n8n nodes:

On form submission (n8n trigger):
Function: Serves as the initial n8n trigger for data ingestion. It provides a simple web form.
Key Configuration: Requires a file input field labeled 'File', accepting .pdf format.

When chat message received (n8n trigger):
Function: The n8n trigger that initiates the RAG process when a user sends a chat message.
Key Configuration: Set up to listen for incoming chat requests.

Embeddings Ollama / Embeddings Ollama1:
Function: Generates vector embeddings for document chunks (Ingestion) and search queries (Retrieval).
Key Configuration: Uses the local model mxbai-embed-large:latest via the connected Ollama credential.

Qdrant Vector Store (Insert Mode):
Function: Takes the embedded data and inserts it into the Qdrant database, indexing the knowledge base.
Key Configuration: Mode set to 'insert', targeting the collection ID rag_collection.

Qdrant Vector Store1 (Retrieve as Tool Mode):
Function: Acts as a tool for the AI Agent, allowing it to search the vector database during a conversation.
Key Configuration: Mode set to 'retrieve-as-tool', tool name is retriever.

Recursive Character Text Splitter:
Function: Prepares large documents for embedding by splitting them into smaller, overlapping segments.
Key Configuration: Chunk Size set to 200, Chunk Overlap set to
50.

AI Agent:
Function: The central orchestrator, managing the flow between the chat model, memory, and retrieval tool.
* Key Configuration: System message directs it to use the tool to answer questions. It integrates the Ollama Chat Model, Simple Memory, and the Qdrant tool.

Related n8n Workflows

Free

Nodes: 10 Nodes
Updated: December 26 2025
View all
Created by