Automated Book Data Scraping, Sorting, and CSV Emailing - n8n Workflow

Use this n8n workflow to automatically scrape website data via Dumpling AI, extract book details, sort by price, and email the resulting CSV file using a Google Sheets n8n trigger.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?


  • Data Analysts needing timely and structured web data extracts.

  • E-commerce Researchers tracking competitive pricing information.

  • Automation Engineers looking for robust n8n templates for web scraping and data processing.

  • Users who need a reliable, scheduled mechanism to turn website content into actionable CSV reports.

Overview

This powerful n8n workflow solves the challenge of turning unstructured website HTML into clean, organized, and shareable data reports. By utilizing a Google Sheets n8n trigger, the automation starts immediately upon a new URL being added. The core scraping task is handled by a third-party AI scraping service (Dumpling AI), which fetches a clean version of the target page. The n8n node structure then meticulously extracts key data points (like book titles and prices), sorts them (by price, descending), converts the resulting JSON array into a standard CSV file, and finally delivers the complete report directly to an inbox using a Gmail n8n node. This makes this n8n workflow an essential tool for repeatable data gathering tasks.

How it Works

The entire process is initiated by an n8n trigger.


  1. Trigger (Google Sheets): The n8n trigger monitors a specified Google Sheet for newly added URLs. When a new row is detected, the workflow begins, passing the URL to the next step.

  2. Scraping (HTTP Request): A specialized HTTP Request n8n node forwards the URL to the Dumpling AI service for scraping, ensuring the received HTML is cleaned.

  3. Initial Extraction (HTML n8n node): The incoming HTML is processed to extract large blocks corresponding to individual books using CSS selectors (.row > li).

  4. Splitting (Split Out n8n node): The array of book HTML blocks is split so that each individual book can be processed separately.

  5. Detail Parsing (HTML n8n node): For each item, another HTML n8n node extracts specific fields—the book title (from the link's title attribute) and the price (from the price CSS class).

  6. Sorting (Sort n8n node): All extracted data items are consolidated and sorted by the 'price' field in descending order.

  7. Conversion (Convert to File n8n node): The structured, sorted data is converted into a standard CSV binary file.

  8. Delivery (Gmail n8n node): The final step sends an email containing the subject 'bookstore csv' and attaches the newly generated CSV file, completing the execution of this comprehensive n8n workflow.

Installation Guide

To deploy this efficient n8n workflow template, follow these steps:


  1. Import: Copy the provided JSON data and import it directly into your n8n instance.

  2. Google Sheets Credential Setup: Configure the Trigger- Watches For new URL in Spreadsheet n8n node with a valid Google Sheets OAuth2 API credential, pointing it to the specific spreadsheet and sheet name where URLs will be entered.

  3. Dumpling AI Credential Setup: The Scrape Website Content with Dumpling AI n8n node uses generic HTTP Header Authentication. You will need to set up the credential containing your Dumpling AI API key.

  4. Gmail Credential Setup: Configure the Send CSV via e-mail n8n node with your Gmail OAuth2 credentials and specify the recipient email address in the 'Send To' parameter.

  5. Activation: Once all credentials are set up and the target URL structure is verified against the selectors, activate the n8n workflow.

Node Details


  • Trigger- Watches For new URL in Spreadsheet (Google Sheets Trigger n8n node):

- Function: Initiates the n8n workflow whenever a new row (containing a URL) is added to the linked Google Sheet.
- Key Configuration: Event: rowAdded; Poll Times: every minute.

  • Scrape Website Content with Dumpling AI (HTTP Request n8n node):

- Function: Sends a POST request to the Dumpling AI scraping API, using the URL derived from the Google Sheets n8n trigger output.
- Key Configuration: URL: https://app.dumplingai.com/api/v1/scrape; JSON Body: Dynamic URL input, requesting format: "html" and cleaned: "True".

  • Extract all books from the page (HTML n8n node):

- Function: Extracts chunks of HTML corresponding to each book listing from the scraped content.
- Key Configuration: CSS Selector: .row > li; extracts content as an array named books.

  • Split HTML Array into Individual Books (Split Out n8n node):

- Function: Converts the single input item containing the books array into multiple output items, one for each book.
- Key Configuration: Field To Split Out: books.

  • Extract individual book price (HTML n8n node):

- Function: Parses the individual book HTML item to isolate the title and price.
- Key Configuration: Extracts title (attribute from h3 > a) and price (content from .price_color).

  • Sort by price (Sort n8n node):

- Function: Orders the extracted book records based on the price field.
- Key Configuration: Sort Field: price; Order: descending.

  • Convert to CSV File (Convert to File n8n node):

- Function: Packages the final, sorted JSON data into a binary CSV file ready for attachment.

  • Send CSV via e-mail (Gmail n8n node):

- Function: Sends the generated CSV file via email.
- Key Configuration: Subject: bookstore csv; Attaches the binary file output from the previous n8n node.

Related n8n Workflows

Free

Nodes: 8 Nodes
Updated: December 26 2025
View all
Created by
Yang
Yang

Featured*