Glassdoor Data Extraction and Gemini AI Summarization - n8n Workflow

Automate Glassdoor company data scraping using Bright Data and summarize key findings instantly with the Google Gemini LLM. Deploy this advanced n8n workflow.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?

Recruitment and HR Professionals: Quickly summarizing sentiment and key information from employee reviews.
Market Researchers: Automating competitive analysis based on publicly available employer data.
n8n Developers: Users looking for sophisticated examples of integrating Langchain AI nodes with web scraping services.
Automation Specialists: Anyone needing a robust n8n template for data extraction and immediate AI processing.

Overview

This powerful n8n workflow demonstrates a complete data pipeline, starting from triggering a specialized web scraper to delivering an AI-generated summary. The core problem it solves is the inefficiency of manually scraping large amounts of text data (like Glassdoor reviews) and then manually analyzing that data. By utilizing Bright Data for reliable extraction and integrating the Google Gemini large language model (LLM) via an n8n node, this automation provides instant, actionable summaries.

This specific n8n template is essential for those who need high-volume, structured data, coupled with smart AI processing. The architecture includes robust flow control logic—specifically a polling mechanism (Wait and If n8n node combination) to ensure the scraping job is complete before attempting to download the results. This makes the overall n8n workflow highly resilient and production-ready.

How it Works

This n8n workflow executes through a sequential process involving web scraping, polling, and AI summarization:


  1. Trigger: The n8n workflow starts when the user clicks 'Test workflow' or manually executes the n8n trigger.

  2. Scraping Initiation: An HTTP Request n8n node sends a request to the Bright Data API to start a web scraping job for a specific Glassdoor URL, capturing the initial response containing a unique snapshotid.

  3. ID Storage: The Set Snapshot Id n8n node extracts and stores this snapshotid for subsequent steps.

  4. Status Polling Loop: The Check Snapshot Status n8n node repeatedly queries Bright Data using the stored ID. The flow enters a loop managed by the If n8n node, which checks if the status is 'ready'.

  5. Waiting Period: If the status is not 'ready', the Wait for 30 seconds n8n node pauses the n8n workflow before attempting to check the status again.

  6. Data Download: Once the job is 'ready', the Download the Snapshot Response n8n node retrieves the scraped data, formatted as JSON.

  7. AI Preparation: The raw data is processed by the Default Data Loader and segmented by the Recursive Character Text Splitter n8n node, preparing the large document for the LLM.

  8. Summarization Chain: The Summarization of Glassdoor Response Langchain n8n node, powered by the Google Gemini Chat Model (Gemini Flash Experimental model), takes the split text and generates a concise summary.

  9. Notification: Finally, the summarized output is posted via an HTTP Request n8n node to a configured webhook endpoint, concluding the execution of this complex n8n template.

Installation Guide

To install and utilize this sophisticated n8n workflow, follow these steps:


  1. Import: Copy the provided n8n workflow JSON code and paste it into your n8n instance using the 'New' -> 'Import from JSON' function.

  2. Bright Data Credentials: This n8n workflow requires an 'HTTP Header Auth' credential set up for Bright Data. This credential needs the API key included in the headers for all Bright Data HTTP Request n8n nodes.

  3. Google Gemini Credentials: Configure the 'Google Gemini (PaLM) Api account' credential. This is required for the Google Gemini Chat Model n8n node to access the LLM services.

  4. Webhook Configuration: Update the URL in the final Configure Webhook Notification n8n node to your desired endpoint where the summarized results should be sent.

  5. Activation: Once all credentials are set, activate the n8n template and run the n8n trigger to test the entire process.

Node Details

This n8n template leverages several critical n8n node types for its functionality:

When clicking ‘Test workflow’ (Manual Trigger n8n node): Serves as the starting point, manually initiating the data extraction and AI processing sequence.
HTTP Request to Glassdoor (HTTP Request n8n node): Triggers the Bright Data scraper. Key Configuration: Uses POST method, targets the Bright Data datasets/v3/trigger endpoint, and specifies the dataset ID (gdl7j0bx501ockwldaqf).
If (If n8n node): Controls the polling loop. Key Configuration: Checks if the incoming JSON property status equals 'ready'.
Wait for 30 seconds (Wait n8n node): Introduces a pause in the loop to avoid overwhelming the Bright Data API during the scraping process.
Download the Snapshot Response (HTTP Request n8n node): Retrieves the extracted Glassdoor data once the job is complete. Key Configuration: Uses expression ={{ $json.snapshotid }} to dynamically reference the job ID.
Google Gemini Chat Model (Langchain LLM n8n node): Provides the AI capability. Key Configuration: Uses models/gemini-2.0-flash-thinking-exp-01-21 (an experimental model).
Summarization of Glassdoor Response (Langchain Chain n8n node): Orchestrates the AI task. Key Configuration: Set to documentLoader mode, allowing it to process the large documents prepared by the preceding n8n nodes.
Configure Webhook Notification (HTTP Request n8n node): The final action, sending the result of the n8n workflow to an external system. Key Configuration: Posts the summary text using the expression ={{ $json.response.text }}.

Related n8n Workflows

Free

Nodes: 10 Nodes
Updated: December 26 2025
View all
Created by
Ranjan Dailata
Ranjan Dailata

Featured*