Structured Data Extraction and AI Analysis with Bright Data & Google Gemini - n8n Workflow

Automate web scraping using Bright Data and process the output into structured JSON data, topics, and trends using Google Gemini in this powerful n8n workflow.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?

Data scientists and analysts needing automated, structured web data feeds.
Developers looking for advanced web scraping n8n templates integrating Bright Data.
Businesses requiring AI-driven insights (topic modeling, trend analysis) from scraped web content.
Users deploying complex AI pipelines within a robust n8n workflow environment.


  • Anyone interested in leveraging the Google Gemini n8n node for information extraction.

Overview

This expert-level n8n workflow demonstrates a cutting-edge data pipeline that combines advanced web extraction technology (Bright Data Web Unlocker) with sophisticated AI processing (Google Gemini). The core problem solved is transforming unstructured, noisy web content—received in markdown format via the Bright Data API—into clean textual data and multiple detailed structured JSON outputs. This specific n8n workflow allows for concurrent analysis, including converting markdown to clean text, performing topic modeling, and clustering emerging trends by location and category. Every step is clearly defined using an n8n node, ensuring the entire automation is transparent and highly reusable as an n8n template for data mining.

How it Works

The n8n workflow begins with a manual n8n trigger, initializing the process.


  1. Configuration: The Set URL and Bright Data Zone n8n node defines the target URL (e.g., a news site) and the specific Bright Data Web Unlocker zone to be used for scraping.

  2. Web Scraping: The Perform Bright Data Web Request n8n node executes the scrape via an HTTP request, pulling the fully unlocked page content in markdown format.

  3. Parallel Processing: The scraped markdown data is immediately sent down three parallel AI paths, leveraging various n8n nodes for LLM interaction:

Path A (Cleanup): The Markdown to Textual Data Extractor Chain LLM n8n node uses a Google Gemini model to strip the markdown of links, scripts, and formatting, outputting clean text. This result is then sent to a logging webhook.
Path B (Topic Modeling): The Topic Extractor with the structured response Information Extractor n8n node uses a separate Google Gemini connection and a predefined JSON schema to analyze the content, outputting topics, confidence scores, summaries, and keywords. This structured JSON output is converted to a binary file using an n8n Function node and saved to disk (d:\topics.json).
* Path C (Trend Analysis): The Trends by location and category with the structured response Information Extractor n8n node uses a dedicated Google Gemini model and another custom JSON schema to cluster emerging trends by geographical location and category. This result is also saved to disk (d:\trends.json) after passing through a Function n8n node.

This powerful architecture demonstrates how a single n8n trigger can initiate complex, multi-faceted data processing operations.

Installation Guide

To use this advanced n8n workflow, follow these setup steps:


  1. Import the n8n Workflow: Copy the provided JSON code and paste it into your n8n instance via the "New" menu -> "Import from JSON".

  2. Bright Data Credentials:

Locate the Perform Bright Data Web Request n8n node.
Set up a new credential of type "HTTP Header Auth."
The header name should typically be Authorization or X-BrightData-Auth. Input your Bright Data API token as the header value.

  1. Google Gemini Credentials:

You must configure the Google Gemini/PaLM API credential. This is required for all three LLM nodes (Google Gemini Chat Model for Data Extract, Google Gemini Chat Model for Sentiment Analyzer, and Google Gemini Chat Model).
Enter your valid API Key associated with the Gemini service.

  1. Configuration Check:

Update the Set URL and Bright Data Zone n8n node with the URL you wish to scrape and verify the correct Bright Data zone.
* Update the dummy https://webhook.site/... URLs in the three Initiate a Webhook Notification n8n node instances to point to your actual logging or notification endpoints.

  1. Execution: Once credentials and configurations are complete, execute the n8n workflow by clicking the Test workflow n8n trigger.

Node Details

This n8n template relies on advanced integration between external APIs and internal AI processing nodes.

When clicking ‘Test workflow’ (Manual Trigger n8n trigger):
Function: Starts the entire data pipeline upon manual execution. This is the simplest form of an n8n trigger.
Set URL and Bright Data Zone (Set n8n node):
Function: Defines crucial variables (url and zone) used downstream in the n8n workflow.
Key Configuration: URL is set to https://www.bbc.com/news/world; zone is set to webunlocker1.
Perform Bright Data Web Request (HTTP Request n8n node):
Function: Executes the web scraping task, sending parameters to the Bright Data API to fetch unlocked page content.
Key Configuration: Method POST, authentication via HTTP Header, requested data
format is markdown.
Markdown to Textual Data Extractor (Chain LLM n8n node):
Function: Converts the scraped markdown structure into clean, readable text using an LLM. This is an essential processing step in this n8n workflow.
Key Configuration: Prompt instructs the LLM to output textual data only, removing links and scripts. Connected to the Google Gemini Chat Model for Data Extract n8n node.
Topic Extractor with the structured response (Information Extractor n8n node):
Function: Analyzes the raw scraped data and forces the AI to output results structured according to a strict JSON schema, detailing topic, score, and summary.
Key Configuration: Uses a custom JSON Schema for topic modeling array output. Connected to the Google Gemini Chat Model for Sentiment Analyzer n8n node.
Trends by location and category with the structured response (Information Extractor n8n node):
Function: Performs specialized trend analysis, clustering findings by location and category as dictated by its custom JSON schema.
Key Configuration: Uses a custom JSON Schema designed for geographical and categorical trend mapping. Connected to the dedicated Google Gemini Chat Model n8n node.
Create a binary file for topics / Create a binary data for tends (Function n8n node):
Function: Converts the structured JSON output generated by the LLM information extractor n8n nodes into a base64 encoded binary data item, preparing it for local file storage.
Write the topics file to disk / Write the trends file to disk (ReadWriteFile n8n node):
* Function: Saves the processed binary data (structured JSON) locally, completing the persistence phase of this n8n workflow.

Related n8n Workflows

Free

Nodes: 9 Nodes
Updated: December 26 2025
View all
Created by
Amit Mehta
Amit Mehta

I'm a workflow automation expert with 15+ years in IT industry. I build smart, scalable n8n workflows for AI automation, marketing, CRM, and SaaS integrations. My focus is on simplifying business processes with tools like OpenAI, WhatsApp, Gmail, and Airtable. I help teams and solopreneurs automate smarter, reduce manual tasks, and grow faster—one workflow at a time.

Featured*