AI-Powered Document Q&A System with Pinecone and Google Drive - n8n Workflow

Use this advanced n8n workflow to build a powerful Retrieval-Augmented Generation (RAG) system. Integrate Google Drive, OpenAI embeddings, and Pinecone DB for instant document Q&A via chat or a custom UI webhook.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?

Legal professionals and compliance teams needing quick answers from complex contracts.
Developers looking to deploy scalable RAG systems without custom coding.
Technical writers and automation specialists seeking advanced n8n templates.
Businesses requiring an internal knowledge base built from private documents.

Overview

This complex n8n workflow provides a complete solution for document knowledge management, transforming unstructured PDF documents from Google Drive into a searchable knowledge base using AI. It solves the challenge of manually searching long contracts by creating semantic search capabilities. The workflow operates in two phases: first, document ingestion (loading and vectorizing files into Pinecone), and second, query resolution (answering questions using an AI agent via either a native n8n trigger chat interface or a standard webhook). This n8n workflow demonstrates the full power of combining cloud storage, specialized vector databases, and cutting-edge large language models within a unified n8n automation platform. Using this n8n node combination ensures high accuracy and relevance in responses, limited only to the content of the provided documents.

How it Works

This comprehensive n8n workflow is divided into three distinct sections: Document Ingestion, Chat Querying, and Webhook Querying.

1. Document Ingestion (Data Loading)


  1. The process starts with the 'When clicking ‘Execute workflow’' n8n trigger, designed for manual execution when new documents need to be indexed.

  2. The 'Google Drive' n8n node searches a specified folder (e.g., 'contract document') for files.

  3. 'Google Drive1' downloads the files in binary format.

  4. The 'Default Data Loader' prepares the binary data for vectorization.

  5. The 'Recursive Character Text Splitter' n8n node breaks down large documents into smaller, overlapping chunks (Chunk Overlap: 100), optimizing them for semantic search.

  6. 'Embeddings OpenAI' uses an OpenAI model (like text-embedding-3-small) to convert these text chunks into numerical vector representations.

  7. Finally, the 'Pinecone Vector Store' n8n node inserts these vectors into the specified Pinecone index (package1536), completing the RAG setup.

2. Chat Querying (Internal Test/Q&A)


  1. The 'When chat message received' n8n trigger initiates the flow upon receiving a query in the n8n chat interface.

  2. The 'AI Agent' n8n node orchestrates the query process, utilizing a specific system message role (expert contracting adviser).

  3. The agent leverages the 'OpenAI Chat Model' (gpt-4.1-mini) and 'Simple Memory' for context.

  4. The core Q&A logic is handled by the 'Answer questions with a vector store' tool, which connects to the 'Pinecone Vector Store1' and 'Embeddings OpenAI1' to retrieve relevant document chunks.

3. Webhook Querying (External UI Integration)


  1. The 'Webhook' n8n trigger listens for external POST requests (e.g., from a custom UI like Lovable), expecting the user query in the request body ({{ $json.body.query }}).

  2. 'AI Agent1' processes the incoming query, similar to the chat flow, using the same expert instructions.

  3. The agent uses 'Answer questions with a vector store1' (backed by Pinecone and OpenAI embeddings) to generate the answer based only on indexed contract content.

  4. The final 'Respond to Webhook' n8n node returns the AI-generated answer back to the calling interface, effectively closing the loop for the external application. This demonstrates a production-ready application of the n8n templates.

Installation Guide

To set up this advanced n8n workflow, follow these steps:


  1. Import the n8n workflow: Copy the provided JSON data and paste it into your n8n canvas via the 'New' menu -> 'Import from JSON'.

  2. Google Drive Setup: Configure the 'Google Drive' n8n node credentials. This account must have read access to the folder containing your contract documents (Folder ID: 1NgITWoqBgLAVof9bxF0jIrVToQ9c919u).

  3. OpenAI Credentials: Set up credentials for OpenAI API keys for all 'Embeddings OpenAI' and 'OpenAI Chat Model' n8n node instances. These are crucial for both vectorization and query answering.

  4. Pinecone Credentials: Configure credentials for the Pinecone API Key and Environment. Ensure the target Pinecone index (package1536) is correctly created and specified across the 'Pinecone Vector Store' n8n nodes.

  5. Initial Ingestion: Execute the first part of the n8n workflow (Flow 1, starting with the Manual Trigger) to load and vectorize your Google Drive documents into Pinecone.

  6. Webhook Activation (Flow 3): After activating the n8n workflow, note the URL provided by the 'Webhook' n8n trigger. This URL is used to connect external interfaces like Lovable for submitting queries.

Node Details

This n8n workflow utilizes many specialized AI and data management nodes:

When clicking ‘Execute workflow’ (n8n trigger): Starts the document ingestion process manually. Ideal for testing and adding new documents to the knowledge base.
Google Drive (n8n node): Retrieves a list of files from a specific Google Drive folder ID, facilitating automated document fetching.
Embeddings OpenAI (n8n node): Uses OpenAI's embedding API (e.g., text-embedding-3-small) to convert document chunks into vectors, enabling semantic search within the Pinecone index.
Recursive Character Text Splitter (n8n node): Essential for RAG, this component breaks large documents into smaller, manageable chunks with a set overlap (100 characters), ensuring context continuity during retrieval.
Pinecone Vector Store (n8n node): Handles interaction with the Pinecone vector database. Used in Insertion Mode (Flow 1) for loading documents and in Query Mode (Flows 2 & 3) to retrieve relevant context.
When chat message received (n8n trigger): A specialized LangChain n8n trigger that initiates the query flow when a message is received in the n8n chat window.
AI Agent (n8n node): The core intelligence node. It is configured with a strict system prompt instructing it to act as an expert legal adviser and to only use the knowledge retrieved from the Pinecone vector database. This powerful n8n node manages the complex routing of LLM tools and memory.
Answer questions with a vector store (n8n node): A dedicated tool within the AI Agent that manages the RAG search. It uses the Pinecone integration to perform similarity searches based on the user's query and fetches the best document chunks.


  • Webhook (n8n trigger): Provides a public endpoint for external systems (like a UI built with Lovable) to send queries as a JSON POST request. This demonstrates how easily n8n templates can be extended to web applications.

Related n8n Workflows

Free

Nodes: 14 Nodes
Updated: December 26 2025
View all
Created by

B2B and B2C Travel App Consultant. Building AI Agent for Travel Solution.

Featured*