URL Compliance & Source Verification with LLMs - n8n Workflow

Use this advanced n8n workflow to filter URLs based on robots.txt compliance, database forbidden lists, and AI-powered source desirability checks.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?

Researchers and data scientists needing to ensure ethical web scraping (respecting robots.txt).
Teams building robust web crawlers or data ingestion pipelines using n8n.


  • Automation specialists looking for advanced examples of integrating PostgreSQL caching and multi-model AI logic within a single n8n workflow.

Overview

This sophisticated n8n workflow, often deployed as one of the specialized n8n templates, acts as a dedicated "URL Officer." It solves the critical problem of ensuring data collection integrity by respecting mandatory exclusion policies (like robots.txt) and preventing ingestion from explicitly undesirable sources. The workflow first checks a local PostgreSQL database for explicitly forbidden URLs. If allowed, it then checks for a cached version of the domain's robots.txt file. If no cache exists, it fetches the file using an n8n node, analyzes the content via a custom code block to verify path allowance, and caches the result. Finally, if the URL is allowed by traditional means, it uses a powerful AI model (via the Langchain Information Extractor n8n node) to perform a final, contextual assessment of the source, allowing for flexible, semantic filtering based on quality or content type. This multilayered approach makes the resulting data ingestion highly compliant and efficient.

How it Works


  1. Start: The process begins via the Start n8n trigger, usually receiving a batch of URLs from a preceding n8n workflow.

  2. Base URL Extraction: The Get Base URL n8n node extracts the root domain from the input URL for subsequent database checks.

  3. Forbidden List Lookup: The Check forbidden urls table PostgreSQL n8n node checks if the domain or URL is globally blacklisted. If detected (If forbidden url Detected), the link is immediately routed to disallow status.

  4. Robots.txt Cache Check: If not forbidden, the n8n workflow checks the robots.txt Table (PostgreSQL) for a recent robots.txt entry. This prevents redundant HTTP requests.

  5. Fetching and Caching: If the entry is not found or expired, the Get Robots.txt HTTP Request n8n node fetches the file. Error handling ensures the workflow continues even if the file is missing.

  6. Compliance Check: The Check Robots.txt Code n8n node runs custom logic to parse the robots.txt file content and determine if the specific URL path is allowed for the crawler.

  7. AI-Powered Review: If the link is allowed by robots.txt (If Link Allowed), the n8n workflow proceeds to the AI integration section.

  8. Model Selection and Extraction: The Model Selector dynamically routes the request to a pre-configured LLM provider (Mistral, Groq, or Gemini). The Information Extractor Langchain n8n node analyzes the source characteristics for quality and suitability.

  9. AI Decision: The If Link Allowed 2 n8n node determines if the AI deems the source acceptable.

  10. Output and Cache Update: The final decision is structured using a Set n8n node, and the robots.txt status is updated in the PostgreSQL cache using the Upsert robots.txt Table n8n node before the final Output node completes the n8n workflow execution.

Installation Guide


  1. Import: Copy the provided n8n workflow JSON and paste it into your n8n instance's canvas.

  2. PostgreSQL Credentials: Set up credentials for the Postgres n8n node instances used for caching robots.txt and storing forbidden URLs. Ensure your underlying database tables (robotstxtcache, forbidden_urls) are created and accessible. You may need to run the initial Create nodes manually.

  3. AI Credentials: Configure credentials for the Langchain LLM nodes used: Mistral Cloud Chat Model, Groq Chat Model, and Google Gemini Chat Model. At least one model must be configured for the Information Extractor n8n node to function correctly.

  4. Testing: Use the Start n8n trigger to integrate this URL filtering logic into a larger data pipeline, passing the target URLs as input data.

Node Details

Start (Execute Workflow Trigger n8n node): The primary n8n trigger point, allowing this automation to be called as a sub-workflow from other processes.
Get Base URL (Code n8n node): Uses custom JavaScript to extract the domain from the full URL, standardizing the input for database checks and robots.txt fetching.
Postgres n8n node (Check/Create/Upsert): Multiple instances manage database operations. This includes Check forbidden urls table (blacklist verification) and Check robots.txt Table (cache lookup), optimizing the performance of the n8n workflow by reducing external API calls.
Get Robots.txt (HTTP Request n8n node): Responsible for fetching the /robots.txt file. The configuration allows continuation on error, essential if the file does not exist.
Check Robots.txt (Code n8n node): This crucial n8n node contains the implementation of the robots exclusion protocol logic, parsing the fetched file to determine crawling permission.
Model Selector (Langchain n8n node): Facilitates dynamic routing to various AI services (Mistral, Groq, Gemini), providing flexibility in LLM choice for the extraction task.
Information Extractor (Langchain n8n node): This AI n8n node leverages the capabilities of the chosen LLM to perform contextual analysis on the URL source, adding a layer of semantic filtering beyond structural compliance.
If n8n node (Multiple Instances): Provides core logic flow control, routing the data based on multiple checks: forbidden status, robots.txt compliance, and AI assessment, making this a highly controlled n8n workflow.

Related n8n Workflows

Paid

Nodes: 13 Nodes
Updated: December 26 2025
View all
Created by

Company dedicated to delivering tailored software solutions and data-driven experiences through effective technology. We develop workflows leveraging AI agents to maximize the productive benefits of artificial intelligence.

Featured*