Use this powerful n8n workflow to build a sophisticated Retrieval-Augmented Generation (RAG) system. Automatically ingest PDFs and CSVs from Google Drive, generate context-aware embeddings using OpenAI, store data in a Supabase vector database, and deploy an AI chatbot agent.
Download this n8n workflow template and start using it instantly.






















Managing vast repositories of documents, especially when scattered across platforms like Google Drive, presents a major challenge for rapid information retrieval. This specialized n8n workflow solves this by creating an automated, two-part system.
The first part is a powerful ingestion pipeline that watches a Google Drive folder. When a new document is uploaded, the n8n node system downloads it, intelligently extracts text (even from complex formats like PDF and CSV), and uses Google Gemini to generate crucial metadata. Crucially, it then uses custom JavaScript and further Gemini processing to break the document into small, context-rich chunks. OpenAI handles the vector embeddings, and finally, the data is pushed to Supabase for permanent storage.
The second part is an interactive AI Chat Agent built entirely within this n8n template. This agent uses the stored Supabase knowledge base (via the RAG technique) to answer user questions instantly, ensuring accuracy by limiting responses only to the information found in the ingested documents. This complete n8n solution eliminates manual processing and immediately transforms static files into a dynamic knowledge base, maximizing the value of your stored documents.
This complex n8n workflow operates across two main pipelines: Document Ingestion and Chat Interaction.
Google Drive Trigger File Created n8n trigger, which constantly monitors a specified Google Drive folder for new file uploads.Loop Over Items node handles sequential processing, followed by Set File ID and Download FIle. The Switch n8n node routes the workflow based on the file type (PDF or CSV) to the respective Extract from PDF or Extract from CSV nodes.Create Metadata Title & Description node, which utilizes a Google Gemini Chat Model and a Structured Output Parser to generate a defined title and description, essential for RAG filtering.Split into chunks custom code n8n node segments the document text into optimized pieces (1000 characters with 200 overlap), intelligently looking for natural breaks like paragraphs and sentences.Split Out, each chunk is processed by the Process Context node, which uses another Google Gemini Chat Model. This step adds context to fragmented chunks, corrects missing information, and prepares them for optimal retrieval performance.Embeddings OpenAI n8n node, and packaged with their metadata via the Default Data Loader. Finally, the Add Data to Supabase Vector Store n8n node inserts the vectors into the Supabase database table named documents.When chat message received n8n trigger starts the user interaction by receiving the user's query.Simple Memory n8n node maintains conversation history, enabling the AI Agent to understand context across multiple turns.Embeddings OpenAI1 n8n node. The Supabase Vector Store n8n node is configured as a search tool to retrieve the most relevant documents based on semantic similarity.AI Agent n8n node receives the user input, the Supabase RAG tool, and the conversation memory. It uses the OpenAI Chat Model (GPT-4o-mini) to synthesize an accurate, document-grounded answer, following strict instructions to prevent hallucination.To deploy this expert n8n workflow and utilize the associated n8n templates, follow these steps:
New > Import from JSON).documents must exist with the required schema (as noted in the sticky note).Google Drive Trigger File Created n8n node, update the Folder to Watch parameter with the exact URL of your target Google Drive folder.Supabase Vector Store n8n nodes point to the correct table (documents). Google Drive Trigger File Created (n8n trigger):
Function: Initiates the n8n workflow whenever a new file is uploaded to a specified Google Drive folder.
Key Configuration: Monitors the fileCreated event and is set to poll the folder every minute.
Download FIle (Google Drive n8n node):
Function: Downloads the uploaded file binary data. Critical setting for handling Google Docs conversion to PDF before extraction.
Key Configuration: Uses an expression ={{ $('Set File ID').item.json.file_id }} to dynamically fetch the file to download.
Extract from PDF / Extract from CSV (n8n node):
Function: Dedicated nodes for converting binary PDF and CSV files into plain text content, enabling the AI models to process the data.
Create Metadata Title & Description (LangChain Chain LLM n8n node):
Function: Uses the Gemini model and Structured Output Parser to generate clean, structured metadata (title and description) based on the document content.
Key Configuration: Output schema is enforced using JSON schema validation.
Split into chunks (Code n8n node):
Function: Executes custom JavaScript logic for intelligent, hierarchy-aware text splitting (paragraph, sentence, word) to create optimal chunks for RAG.
Process Context (LangChain Chain LLM n8n node):
Function: Leverages the Google Gemini model to enhance each small text chunk by providing contextual scaffolding, ensuring chunks are clear and self-contained for better vector search accuracy.
Embeddings OpenAI (LangChain Embeddings n8n node):
Function: Converts the final, enhanced text chunks into high-dimensional vector embeddings, a core component of this RAG n8n workflow.
Add Data to Supabase Vector Store (LangChain Vector Store n8n node):
Function: Inserts the vector embeddings, text content, and metadata into the designated Supabase table (documents).
When chat message received (n8n trigger):
Function: The entry point for the chat interface, receiving user questions.
AI Agent (LangChain Agent n8n node):
* Function: The central intelligence unit. It manages the RAG tool (Supabase Vector Store), memory, and uses the OpenAI Chat Model to formulate a factual response based exclusively on the retrieved context.
Instantly generate custom sales proposals using an n8n workflow triggered by a form submission. This n8n automation integrates OpenAI content generation with Google Slides text replacement for rapid document delivery.

Generate highly personalized, data-driven product videos automatically using this advanced n8n workflow. Integrates Foreplay competitor data, Google Gemini, and video generation APIs.

Use this powerful n8n workflow to automate competitor research via Google Search, generate high-quality, SEO-optimized product content using AI models, and store the output directly in Google Sheets. Explore n8n templates for digital marketing.

Implement a robust n8n workflow for advanced RAG chatbots. This n8n template uses OpenAI, Qdrant vector storage, LangChain n8n nodes, and webhooks to handle complex Q&A, calendar booking, and external API tracking.

Use this powerful n8n workflow to automate RAG-based financial analysis. Load earnings report PDFs into a Pinecone vector database and leverage a Gemini AI Agent to generate detailed reports in Google Docs.


I'm a professional software engineer and n8n expert with a passion for building scalable, no-code and low-code automation workflows. I specialize in creating seamless integrations between APIs, CRMs, and everyday tools to help businesses save time, reduce manual work, and operate smarter. Whether it's automating marketing pipelines, backend systems, or approval processes, I turn complex logic into simple, powerful workflows with n8n.







































