AI RAG Document Processing and Chatbot with Google Drive and Supabase - n8n Workflow

Use this powerful n8n workflow to build a sophisticated Retrieval-Augmented Generation (RAG) system. Automatically ingest PDFs and CSVs from Google Drive, generate context-aware embeddings using OpenAI, store data in a Supabase vector database, and deploy an AI chatbot agent.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?


  • Knowledge Managers: Who need to centralize and make internal documentation instantly searchable via AI.

  • Automation Engineers: Seeking a robust, end-to-end RAG architecture implemented using n8n templates.

  • Data Teams: Looking to automate the conversion of static files (PDFs, CSVs) into structured, queryable vector data.

  • n8n Users: Interested in advanced LangChain integration, custom code for smart text chunking, and complex data routing within an n8n node structure.

Overview

Managing vast repositories of documents, especially when scattered across platforms like Google Drive, presents a major challenge for rapid information retrieval. This specialized n8n workflow solves this by creating an automated, two-part system.

The first part is a powerful ingestion pipeline that watches a Google Drive folder. When a new document is uploaded, the n8n node system downloads it, intelligently extracts text (even from complex formats like PDF and CSV), and uses Google Gemini to generate crucial metadata. Crucially, it then uses custom JavaScript and further Gemini processing to break the document into small, context-rich chunks. OpenAI handles the vector embeddings, and finally, the data is pushed to Supabase for permanent storage.

The second part is an interactive AI Chat Agent built entirely within this n8n template. This agent uses the stored Supabase knowledge base (via the RAG technique) to answer user questions instantly, ensuring accuracy by limiting responses only to the information found in the ingested documents. This complete n8n solution eliminates manual processing and immediately transforms static files into a dynamic knowledge base, maximizing the value of your stored documents.

How it Works

This complex n8n workflow operates across two main pipelines: Document Ingestion and Chat Interaction.

1. Document Ingestion Pipeline


  1. Trigger: The process begins with the Google Drive Trigger File Created n8n trigger, which constantly monitors a specified Google Drive folder for new file uploads.

  2. File Download and Routing: The Loop Over Items node handles sequential processing, followed by Set File ID and Download FIle. The Switch n8n node routes the workflow based on the file type (PDF or CSV) to the respective Extract from PDF or Extract from CSV nodes.

  3. Metadata Generation: The extracted text is passed to the Create Metadata Title & Description node, which utilizes a Google Gemini Chat Model and a Structured Output Parser to generate a defined title and description, essential for RAG filtering.

  4. Smart Chunking: The Split into chunks custom code n8n node segments the document text into optimized pieces (1000 characters with 200 overlap), intelligently looking for natural breaks like paragraphs and sentences.

  5. Context Enhancement: After splitting the chunks using Split Out, each chunk is processed by the Process Context node, which uses another Google Gemini Chat Model. This step adds context to fragmented chunks, corrects missing information, and prepares them for optimal retrieval performance.

  6. Embedding and Storage: The enhanced chunks are summarized, embedded using the Embeddings OpenAI n8n node, and packaged with their metadata via the Default Data Loader. Finally, the Add Data to Supabase Vector Store n8n node inserts the vectors into the Supabase database table named documents.

2. Chat Agent Interaction Pipeline


  1. Chat Trigger: The When chat message received n8n trigger starts the user interaction by receiving the user's query.

  2. Conversation Memory: The Simple Memory n8n node maintains conversation history, enabling the AI Agent to understand context across multiple turns.

  3. RAG Tool Setup: The user's question is embedded by a dedicated Embeddings OpenAI1 n8n node. The Supabase Vector Store n8n node is configured as a search tool to retrieve the most relevant documents based on semantic similarity.

  4. AI Agent Response: The AI Agent n8n node receives the user input, the Supabase RAG tool, and the conversation memory. It uses the OpenAI Chat Model (GPT-4o-mini) to synthesize an accurate, document-grounded answer, following strict instructions to prevent hallucination.

Installation Guide

To deploy this expert n8n workflow and utilize the associated n8n templates, follow these steps:


  1. Import the n8n Workflow: Copy the provided JSON data and import it into your n8n instance via the Workflows section (New > Import from JSON).

  2. Configure Credentials: Ensure you have the following credentials set up in n8n:

Google Drive OAuth2 Api: Link your Google account and ensure access to the monitored folder.
Supabase API: Provide your Supabase Project URL and API Key. The table documents must exist with the required schema (as noted in the sticky note).
OpenAI API: Provide your API key for embedding generation and the final chat model.
Google Gemini API: Provide your API key for metadata generation and chunk enhancement.

  1. Google Drive Trigger Setup: In the Google Drive Trigger File Created n8n node, update the Folder to Watch parameter with the exact URL of your target Google Drive folder.

  2. Database Validation: Verify that the Supabase Vector Store n8n nodes point to the correct table (documents).

  3. Activation: Save the n8n workflow and activate it to start the automated document ingestion process and enable the AI chat agent via its dedicated webhook endpoint.

Node Details

Google Drive Trigger File Created (n8n trigger):
Function: Initiates the n8n workflow whenever a new file is uploaded to a specified Google Drive folder.
Key Configuration: Monitors the fileCreated event and is set to poll the folder every minute.
Download FIle (Google Drive n8n node):
Function: Downloads the uploaded file binary data. Critical setting for handling Google Docs conversion to PDF before extraction.
Key Configuration: Uses an expression ={{ $('Set File ID').item.json.file_id }} to dynamically fetch the file to download.
Extract from PDF / Extract from CSV (n8n node):
Function: Dedicated nodes for converting binary PDF and CSV files into plain text content, enabling the AI models to process the data.
Create Metadata Title & Description (LangChain Chain LLM n8n node):
Function: Uses the Gemini model and Structured Output Parser to generate clean, structured metadata (title and description) based on the document content.
Key Configuration: Output schema is enforced using JSON schema validation.
Split into chunks (Code n8n node):
Function: Executes custom JavaScript logic for intelligent, hierarchy-aware text splitting (paragraph, sentence, word) to create optimal chunks for RAG.
Process Context (LangChain Chain LLM n8n node):
Function: Leverages the Google Gemini model to enhance each small text chunk by providing contextual scaffolding, ensuring chunks are clear and self-contained for better vector search accuracy.
Embeddings OpenAI (LangChain Embeddings n8n node):
Function: Converts the final, enhanced text chunks into high-dimensional vector embeddings, a core component of this RAG n8n workflow.
Add Data to Supabase Vector Store (LangChain Vector Store n8n node):
Function: Inserts the vector embeddings, text content, and metadata into the designated Supabase table (documents).
When chat message received (n8n trigger):
Function: The entry point for the chat interface, receiving user questions.
AI Agent (LangChain Agent n8n node):
* Function: The central intelligence unit. It manages the RAG tool (Supabase Vector Store), memory, and uses the OpenAI Chat Model to formulate a factual response based exclusively on the retrieved context.

Related n8n Workflows

Free

Nodes: 22 Nodes
Updated: December 26 2025
View all
Created by
Billy Christi
Billy Christi

I'm a professional software engineer and n8n expert with a passion for building scalable, no-code and low-code automation workflows. I specialize in creating seamless integrations between APIs, CRMs, and everyday tools to help businesses save time, reduce manual work, and operate smarter. Whether it's automating marketing pipelines, backend systems, or approval processes, I turn complex logic into simple, powerful workflows with n8n.

Featured*