Deploy a robust RAG system using n8n to ingest PDFs via Mistral OCR, embed them with Cohere, store in Weaviate, and enable search via an n8n workflow MCP Server trigger.
Download this n8n workflow template and start using it instantly.
This advanced n8n workflow is designed to manage a complete PDF-to-search pipeline. It solves the challenge of turning complex, unstructured PDF documents into semantic data stored in a vector database (Weaviate). The ingestion side uses a dedicated Form n8n trigger for manual uploads, leveraging the Mistral AI n8n node for high-quality Optical Character Recognition (OCR). Once processed, the data is chunked, embedded using Cohere, and indexed. The search capability is exposed through an MCP Knowledge Server n8n trigger, allowing other AI agents or workflows to query the knowledge base, ensuring highly relevant results through the use of a Cohere Reranker. This is an essential n8n template for enterprise AI applications, showcasing powerful integrations within a single n8n execution.
The n8n workflow operates in two distinct modes: data ingestion and data retrieval.
Upload PDF n8n trigger (a Form Trigger). Users upload a PDF file, which is then passed to the Extract Text from PDF n8n node (Mistral AI) for OCR processing. The Prepare Document Data n8n node structures the text and adds metadata. Finally, the data is sent to the Store in Vector Database n8n node (Weaviate), utilizing the Cohere Embeddings n8n node to index the content into the vector collection.MCP Knowledge Server n8n trigger (an AI Tool endpoint). The incoming query is routed to the Search Knowledge Base n8n node (Weaviate - Retrieve as Tool mode). This powerful n8n node uses Cohere Embeddings for query vectorization and applies the Cohere Reranker to refine and select the most semantically relevant document chunks, returning the final context snippet back to the calling agent or workflow. This seamless retrieval process is managed entirely within the n8n workflow.Store in Vector Database and Search Knowledge Base n8n nodes point to the correct Weaviate collection (e.g., 'KnowledgeDocuments') and that the schema matches the input fields.Upload PDF n8n trigger URL can be used for manual document uploads, and the MCP Knowledge Server n8n trigger endpoint is ready to be used as an AI tool. Upload PDF (Form Trigger): The initial entry point and n8n trigger for the PDF ingestion pipeline. It generates a public URL form designed specifically to accept .pdf files.
Extract Text from PDF (Mistral AI n8n node): Acts as the OCR engine, converting the binary PDF file into usable text data. Essential for preparing unstructured data for vectorization.
Prepare Document Data (Set n8n node): Structures the OCR output, setting the content property and adding metadata such as filename and upload_timestamp.
Cohere Embeddings (Langchain n8n node): Used by both the ingestion and retrieval paths. It generates multilingual vector embeddings using the embed-multilingual-v3.0 model, a crucial step for semantic search.
Store in Vector Database (Weaviate n8n node): Manages the indexing and storage of the vectorized document chunks into the Weaviate database. Configured in insert mode.
MCP Knowledge Server (Langchain n8n trigger): The webhook n8n trigger that exposes this entire knowledge base as a search tool for other AI agents or master n8n workflow pipelines.
Search Knowledge Base (Weaviate n8n node): Performs the actual retrieval query. Configured in retrieve-as-tool mode, it uses the embeddings and reranking nodes to find the most accurate context.
Cohere Reranker (Langchain n8n node): An optional but highly recommended step to improve RAG quality by re-ranking the search results based on granular relevance before they are returned by the n8n workflow.
Create a fully automated n8n workflow to parse PDFs from Google Drive, normalize content, generate OpenAI embeddings, and store them in Pinecone for Retrieval-Augmented Generation (RAG) QA systems.

Create a robust Retrieval-Augmented Generation (RAG) system using an n8n workflow. Integrate Mistral OCR for PDF extraction, Qdrant vector storage, and Gemini AI for conversational Q&A.

Build a powerful multichannel customer support AI assistant using this n8n workflow. Integrates Chatwoot webhooks with OpenRouter/LLMs to automatically respond to user queries using a defined knowledge base.

Implement a powerful Retrieval Augmented Generation (RAG) system using this advanced n8n workflow. Ingest PDFs into Pinecone via an n8n trigger, utilize OpenAI embeddings, and enhance answer quality with Cohere reranking for a high-performance chat agent.

Build a powerful, multi-functional AI personal assistant using this advanced n8n workflow. It integrates Google Gemini, Gmail, Google Calendar, and Google Sheets using the Langchain Agent framework and n8n's Multi-Client Protocol (MCP).


Medical specialist in internal medicine, gastroenterology, and infectious diseases. Building innovative healthcare automation workflows with n8n, integrating AI, speech-to-text, and medical data standards for efficient clinical documentation and analysis.







































