Google Drive to Supabase Contextual Hybrid RAG Knowledge Base Sync - n8n Workflow

Automate document synchronization from Google Drive to a Supabase vector database using n8n. This powerful n8n workflow features custom chunking, OpenAI contextualization, and hybrid RAG search capabilities.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?


  • AI/ML Engineers building advanced RAG solutions.

  • Data scientists needing automated, reliable data pipelines.

  • Developers looking for powerful n8n templates for vector database management.

  • Users who need to synchronize cloud files (like PDFs) with a vector store continuously.

Overview

Maintaining a synchronized and optimized knowledge base for Retrieval-Augmented Generation (RAG) is challenging. This advanced n8n workflow solves this by creating a robust pipeline between Google Drive and Supabase. The automation monitors a specified Google Drive folder for new or modified documents. It utilizes a file hashing mechanism and a record manager table in Supabase to efficiently detect changes, ensuring only necessary updates are processed (a technique known as delta processing).

Crucially, this n8n workflow goes beyond simple text splitting. It employs custom JavaScript for intelligent chunking and uses an OpenAI n8n node to generate contextual summaries for each resulting chunk. This step significantly boosts retrieval performance in the Supabase hybrid search environment. Furthermore, the overall n8n template includes a complete AI Agent setup, demonstrating how to query this newly built, optimized vector database using the RAG data it ingests, making it a complete end-to-end solution.

How it Works

This comprehensive n8n workflow operates in two main logical paths: Data Synchronization and AI Querying.

Data Synchronization Flow


  1. File Event: The synchronization starts with the Watch GD RAG Files n8n trigger, which detects new or modified files in Google Drive.

  2. Extraction and Hashing: The document is downloaded and its text is extracted (via the Extract from File n8n node). A SHA256 hash is generated using the Generate Hash n8n node to identify unique content.

  3. Delta Check: The Search Record Manager (Supabase n8n node) checks the corresponding file ID. The Switch n8n node directs the flow: if the file is new, it proceeds directly. If it is modified, the old vectors and records are deleted before proceeding. If it is unchanged, the workflow stops processing that item.

  4. Metadata Enrichment: For new or modified files, an LLM chain (Basic LLM Chain connected to an OpenAI n8n node) generates a summary and classified metadata for the document.

  5. Contextual Chunking: Custom code (Recursive Splitter2 n8n node) handles intelligent text segmentation. Another LLM chain (Add Context) leverages the entire document to write a succinct contextual sentence for each individual chunk, preparing it for superior vector search.

  6. Vector Upsert: The contextualized chunks are embedded using the Embeddings OpenAI1 n8n node and then inserted into the Supabase vector store (Supabase Vector Store1 n8n node), completing the synchronization.

Deletion Management


  • A secondary n8n trigger (Watch GD Trash) detects when a file is moved to trash, automatically deleting the corresponding vectors and tracking records from Supabase, ensuring data integrity.

AI Agent Querying


  • The workflow includes an AI Agent initialized by a When chat message received n8n trigger. This agent uses the Query Vector Store tool, which orchestrates a call to a Supabase Edge Function after converting the user's query into an embedding vector (via an HTTP Request n8n node), enabling immediate hybrid RAG search against the newly synchronized data.

Installation Guide

To deploy this n8n workflow, follow these steps:


  1. Import: Copy the provided JSON into your n8n instance using the 'New' -> 'Import from JSON' function.

  2. Credentials Setup:

Google Drive: Connect your Google Drive account credentials to the Watch GD RAG Files n8n trigger and related nodes. Specify the target folder ID for RAG files and the Trash folder ID.
OpenAI: Set up OpenAI API key credentials for the various AI n8n node operations (OpenAI Chat Model, Embeddings OpenAI1). Ensure the keys are linked.
* Supabase: Configure two sets of Supabase credentials: one for the RAG data plane (used by vector store and deletion nodes) and one for the Record Manager (used by sync tracking nodes). Note the specific table names (documentshs, recordmanagerhs).

  1. Supabase Database Preparation: As detailed in the sticky notes, you must run specific SQL commands in the Supabase SQL Editor to create the documentshs table (including vectors and tsvectors for hybrid search) and the recordmanagerHS table.

  2. Edge Function: Create and deploy the required Supabase Edge Function to handle the hybrid search queries (referencing the matchdocumentshshybrid function), and update the URL in the Edge Function HTTP Request n8n node.

Node Details


  • Watch GD RAG Files (Google Drive Trigger n8n trigger): The starting point, configured to poll a specific Google Drive folder for file additions/modifications.

  • Extract from File (n8n node): Dedicated to reliably pulling text content from documents, configured for PDF processing.

  • Generate Hash (Crypto n8n node): Essential for change detection; generates a SHA256 hash of the extracted file content.

  • Search Record Manager (Supabase n8n node): Queries the Supabase tracking table (record_managerhs) to retrieve the existing file hash for comparison.

  • Switch (n8n node): Implements the conditional logic (new, modified, or no change) based on the hash comparison, vital for an efficient n8n workflow.

  • Basic LLM Chain (Langchain n8n node): Utilizes OpenAI to analyze the document and produce a structured JSON summary, enhancing metadata quality.

  • Recursive Splitter2 (Code n8n node): A custom n8n node containing JavaScript for sophisticated text chunking, optimizing chunk boundaries (paragraph/sentence splits) and managing chunk overlap.

  • Add Context (Langchain n8n node): An LLM n8n node that reads both the chunk and the full document to prepend a contextual sentence to the chunk text, significantly improving RAG accuracy.

  • Supabase Vector Store1 (Langchain n8n node): The final destination for the data, configured to insert the contextualized, embedded documents into the documentshs table.

  • AI Agent (Langchain n8n node): The core query engine, configured with a system message to strictly answer based only on the results provided by the Query Vector Store tool.

Related n8n Workflows

Free

Nodes: 28 Nodes
Updated: December 26 2025
View all
Created by
Michael Taleb
Michael Taleb

n8n developer helping businesses save time and scale by automating complex business processes with n8n and smart integrations.

Featured*