Brave Search Result Scraping and Gemini Structured Data Extraction - n8n Workflow

This n8n workflow extracts structured data from Brave Search using the Bright Data MCP n8n node, processes results with Google Gemini, and saves output to Google Sheets and disk.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?

Data analysts needing structured web data at scale.
Users leveraging specialized n8n templates for SERP data extraction.
Developers utilizing the Bright Data integration within the n8n node ecosystem.
Businesses automating advanced data extraction and cleansing using an n8n workflow.

Overview

This robust n8n workflow provides a powerful solution for high-volume, structured data extraction from specific Brave Search results. By integrating the Bright Data MCP Client n8n node (a community node requiring a self-hosted n8n setup), the automation can dynamically scrape specialized search types—including Images, Videos, News, or All—while handling rotation and delivery. The resulting raw HTML content is then passed to the Google Gemini model via LangChain n8n nodes, which applies a predefined structured output parser. This crucial step transforms messy, unstructured web content into clean JSON data, ready for immediate business intelligence use. This powerful n8n automation ensures reliable data sourcing and formatting, making it an essential component for any comprehensive data pipeline built on n8n.

How it Works

The n8n workflow begins with the manual n8n trigger, "When clicking ‘Test workflow’," which immediately proceeds to define the search criteria.


  1. Initial Setup: The "Set the search criteria's" n8n node defines the critical input parameters, including the searchtype (e.g., 'news') and the specific searchquery.

  2. Routing: A Switch n8n node routes the execution path dynamically based on the chosen searchtype. This ensures the flow only runs the correct path for images, videos, news, or a general 'all' search.

  3. URL Configuration: Based on the switch output, a corresponding Set n8n node sets the appropriate searchbaseurl (e.g., https://search.brave.com/video).

  4. Web Scraping: The Bright Data MCP Client n8n node executes a scrapeas_html operation against the dynamic Brave Search URL, collecting the raw SERP data. This node requires specific configuration for the Bright Data service.

  5. Data Preparation: A "Set the Search Response" n8n node standardizes the output, extracting the raw HTML content from the scraping response and preparing it for the AI model.

  6. Structured Extraction (AI): The core of the data cleansing happens using the Google Gemini Chat Model and the Structured Output Parser within the "Structured Data Extractor" n8n node. This LangChain setup uses a strict JSON schema to force the LLM to output clean, predictable data structures, including result titles, URLs, descriptions, and related queries.

  7. Output and Storage: Finally, the extracted structured data is sent to multiple destinations. The data is appended to a Google Sheets document, saved locally via the Write the structured content to disk n8n node, and simultaneously broadcasted to an external Webhook notification endpoint using the "Initiate a Webhook Notification" n8n node.

Installation Guide

To deploy this comprehensive n8n workflow, follow these steps:


  1. Import: Copy the provided n8n workflow JSON and import it into your n8n instance. Note that this specific n8n template relies on a community n8n node (Bright Data MCP Client), meaning it is primarily designed for self-hosted n8n environments.

  2. Credentials Setup:

Google Gemini (PaLM) API: Configure your API key credentials for the Google Gemini Chat Model n8n node.
Bright Data MCP Client: Set up the necessary credentials for the Bright Data MCP Client n8n node.
Google Sheets: Configure the OAuth2 credentials to allow the Google Sheets n8n node to access and write data to your specified spreadsheet.

  1. Workflow Configuration:

Search Parameters: Update the "Set the search criteria's" n8n node with your desired searchtype (must be images, videos, news, or all) and the specific searchquery.
Webhook Endpoint: Update the URL in the "Initiate a Webhook Notification for the Structured Data" n8n node to your preferred external endpoint for receiving the structured data payload.
File Path: Adjust the file path expression in the "Write the structured content to disk" n8n node if you are not running n8n on a Windows environment or prefer a different directory.

Node Details

When clicking ‘Test workflow’ (Manual Trigger n8n trigger): Serves as the starting point for manual execution of the n8n workflow.
Set the search criteria's (Set n8n node): Defines the critical input parameters like searchtype and searchquery which dictate the subsequent flow path.
Switch n8n node: Acts as a router, directing the flow to the correct Bright Data scraping configuration based on the input search_type.
Bright Data MCP Client for Brave Search (MCP Client n8n node): Connects to the Bright Data service to perform compliant, managed web scraping of the Brave Search SERP, retrieving the raw HTML content.
Structured Data Extractor (Chain LLM n8n node) with Google Gemini: This node utilizes the connected Google Gemini LLM and a detailed JSON output parser schema to transform the unstructured HTML into a predictable, structured JSON object.
Google Sheets n8n node: Used to store the final extracted structured data persistently, appending the results into a specified spreadsheet row.


  • Initiate a Webhook Notification (HTTP Request n8n node): Sends the final, cleaned JSON data payload to an external webhook for further processing or alerting.

Related n8n Workflows

Free

Nodes: 12 Nodes
Updated: December 26 2025
View all
Created by

A Professional based out of India specialized in handling AI-powered automations. Contact me at [email protected]

Featured*