AI Knowledge Base Chatbot for Cloud Files - n8n Workflow

Deploy a robust Retrieval-Augmented Generation (RAG) system using this advanced n8n workflow. Index files from Supabase Storage and Google Drive, vectorize data with OpenAI, and power an AI chat agent using a Supabase vector store for contextual answers.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?


  • Technical teams and developers needing to automate document indexing into a RAG system.

  • Businesses using Supabase for storage and wanting to enable AI interaction with their proprietary files.

  • Users looking for ready-to-use n8n templates for building custom AI agents.

  • Anyone who needs a comprehensive n8n workflow solution for syncing Google Drive and Supabase files for knowledge retrieval.

Overview

Manually organizing and searching information across disparate cloud storage solutions like Google Drive and Supabase Storage is inefficient. This specialized n8n workflow solves this by establishing a robust Retrieval-Augmented Generation (RAG) pipeline.

The system operates in two core modes: file ingestion and interactive querying. When new documents are uploaded to your Supabase Storage or a designated Google Drive folder, the ingestion flow is triggered. It uses the powerful LlamaParse utility to extract clean, structured markdown from complex documents (like PDFs), chunks the text, embeds it using an OpenAI n8n node, and stores the resulting vectors in a Supabase Vector Store. This ensures that your knowledge base is always up-to-date.

The second component is an AI Chat Agent, triggered by a chat message (via a webhook or chat interface). This n8n agent utilizes the indexed knowledge base as a tool to provide accurate, contextual responses, leveraging conversation history via in-built memory. This complex n8n template streamlines the creation of a high-performance, searchable knowledge system.

How it Works

This complex n8n workflow is divided into two distinct, interconnected processes: file ingestion and AI querying.

1. File Ingestion & Indexing (Scenarios 1 & 3)

A. Supabase Storage Sync (Manual Trigger):


  1. The process starts with a manual n8n trigger (When clicking ‘Test workflow’).

  2. The n8n workflow first retrieves a list of all files already indexed in the files table via a Supabase n8n node.

  3. It then fetches the full list of files from Supabase Storage via an HTTP Request n8n node.

  4. A Loop iterates over each storage item, using an If n8n node to check if the file is new (avoiding duplicates) and valid.

  5. New files are downloaded via an HTTP Request n8n node and then uploaded to the LlamaParse API.

  6. The n8n node logic polls the LlamaParse API status until successful, retrieving the extracted markdown content.

  7. The content is processed: chunked by the Recursive Character Text Splitter and embedded by an Embeddings OpenAI n8n node (text-embedding-3-small).

  8. Finally, the vectors are inserted into the Supabase Vector Store n8n node, completing the index process.

B. Google Drive Sync (Google Drive Trigger):


  1. The flow is automatically started by a Google Drive n8n trigger when a file is created or updated in a specific folder.

  2. The n8n workflow deletes any existing vectors related to that file ID in Supabase to ensure clean re-indexing.

  3. The file is downloaded (with automatic conversion to PDF for optimal parsing) using a Google Drive n8n node.

  4. The file follows the LlamaParse, chunking, embedding, and Supabase vector insertion steps identical to the Supabase Storage sync.

2. AI Agent Querying (Scenario 2)


  1. The system waits for an input via a Webhook or a When chat message received Langchain n8n trigger.

  2. Input (message and session ID) is merged and passed to the AI Agent1 Langchain n8n node.

  3. The AI Agent uses the OpenAI Chat Model (gpt-4o-mini) for reasoning.

  4. The Simple Memory n8n node maintains the conversational history using the session ID.

  5. Critically, the Supabase Vector Store is configured as a RAG tool (knowledge_base), allowing the AI Agent to retrieve relevant document chunks based on the user's query, ensuring contextual and fact-based answers using the documents indexed by the ingestion n8n workflow.

Installation Guide

To deploy this comprehensive n8n workflow, follow these steps:


  1. Import the n8n Workflow: Copy the provided JSON and import it directly into your n8n instance.

  2. Set Up Credentials: You must configure and connect several credentials within the n8n environment:

Supabase API: Required for database operations (retrieving indexed files, inserting metadata, and deleting old vectors) and vector store operations (indexing and RAG retrieval). You will need your API key and project URL. Several Supabase credential nodes are used (e.g., Supabase account info@ Demo, Supabase club).
OpenAI API: Required for embedding generation (Embeddings OpenAI n8n node) and for powering the LLM (OpenAI Chat Model).
LlamaParse API (HTTP Header Auth): Required for parsing complex documents. Configure the HTTP Header Authorization with Authorization: Bearer key>.
Google Drive OAuth2 API: Required for the Google Drive n8n trigger and file download n8n node.

  1. Configure Storage Paths: In the Supabase Storage HTTP Request n8n node (Get All files), replace id> with your actual Supabase project reference.

  2. Set Google Drive Folder: In the Google Drive n8n trigger nodes (File Created / File Updated), configure the exact folder ID you wish to monitor.

  3. Activate Triggers: Ensure the Webhook and Chat Trigger n8n nodes are active for real-time operation. Use the dedicated n8n trigger node for manually initiating the sync, if required.

Node Details

This powerful n8n workflow utilizes various specialized nodes:

When chat message received / Webhook (n8n triggers): Starts the AI querying process when a user interacts with the chatbot endpoint.
Loop Over Items (n8n node): Iterates over the list of files retrieved from Supabase Storage.
HTTP Request (n8n node): Used extensively for interacting with Supabase Storage (downloading files) and the LlamaParse API (uploading, checking status, retrieving parsed data). Note the use of form binary data for file uploads.
If / Switch (n8n nodes): Essential for flow control, determining if a file is new, and checking if the LlamaParse job status is successful (SUCCESS), PENDING, or ERROR.
Embeddings OpenAI (n8n node): Creates vector representations of the text chunks using OpenAI's text-embedding-3-small model, crucial for the RAG knowledge base.
Recursive Character Text Splitter (n8n node): Chunks the text content into smaller, manageable pieces (500 characters with 200 character overlap) suitable for vector indexing.
Supabase Vector Store (n8n node): Used for two primary functions: insert (ingestion flow) and retrieve-as-tool (RAG flow). It manages the vector database storage.
AI Agent (n8n node): The core intelligence of the chatbot, combining the LLM, memory, and the Supabase Vector Store tool to formulate context-aware responses.


  • OpenAI Chat Model (n8n node): Provides the language generation capabilities, configured here to use gpt-4o-mini.

Related n8n Workflows

Free

Nodes: 23 Nodes
Updated: December 26 2025
View all
Created by

I am a business analyst with a development background, dedicated to helping small businesses and entrepreneurs leverage cloud services for increased efficiency. My expertise lies in automating manual workflows, integrating data from multiple cloud service providers, creating insightful dashboards, and building custom CRM systems. https://www.linkedin.com/in/marklowcoding/

Featured*