Adaptive LLM Router for Cost-Optimized Chat Responses - n8n Workflow

Use this smart n8n workflow to route chat queries based on complexity, leveraging high-power models like Gemini 2.5 Pro for complex tasks and GPT-4.1 Nano for simple ones. Optimize AI costs using this flexible n8n template.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?

This n8n template is ideal for:

Overview

Managing the operational costs of large language models (LLMs) requires thoughtful resource allocation. Using a premium model like Gemini 2.5 Pro for simple, conversational questions is inefficient. This smart n8n workflow solves this problem by implementing a dynamic routing system.

This system uses a lightweight AI agent (Gemini 2.0 Flash) to classify the complexity of the user's input. Based on this classification, the query is routed through an n8n node selector to either the high-performance, higher-cost Gemini 2.5 Pro (for complex reasoning, coding, or research) or the low-cost OpenAI GPT-4.1 Nano (for simple, quick answers). This adaptive approach ensures that you only utilize premium AI compute when absolutely necessary, making your AI operations more cost-effective. Implementing this custom n8n workflow provides a significant advantage in resource management.

How it Works

This highly efficient n8n workflow operates in five distinct stages, starting with the n8n trigger:


  1. Chat Trigger: The process begins with the When chat message received n8n trigger. This n8n node listens for incoming chat messages from a user, capturing the text input for processing.

  2. Complexity Classification: The message immediately flows to the first Model Selector n8n node. This agent (powered by Gemini 2.0 Flash) analyzes the query using a structured system message, returning only '1' (Complex) or '2' (Simple).

  3. Dynamic Routing: The Model selector n8n node, acting as a router, evaluates the numerical output from the classifier. It directs the flow based on the complexity score.

  4. Model Selection: The output path leads to one of two specialized LLM nodes: If the score is '1', it selects the 2.5 pro (Gemini 2.5 Pro) model. If the score is '2', it selects the 4.1 nano (GPT-4.1 Nano) model.

  5. Final Execution: Finally, the Main Agent n8n node takes the original user input and executes the request using the specific, dynamically selected LLM from the routing step, generating an optimized response for the user.

Installation Guide

To deploy this comprehensive n8n workflow, follow these steps:


  1. Import the n8n template: Copy the provided JSON and import it directly into your n8n instance via the 'Workflows' menu.

  2. Set up OpenAI Credentials: Navigate to your n8n Credentials section. Create a new credential for 'OpenAI API' using your secret key from the OpenAI platform. This is required for the 4.1 nano n8n node.

  3. Set up Google Gemini Credentials: Create or select an existing credential for 'Google PaLM API' (used for Gemini integration). Ensure this key is linked to your Gemini service, required by the 2.5 pro, 2.0 flash, and initial Model Selector n8n nodes.

  4. Assign Credentials: Go through each LLM n8n node (4.1 nano, 2.5 pro, and the classification agent) and assign the corresponding credentials.

  5. Activate: Save and activate the n8n workflow. The chat trigger is now active and ready to handle incoming messages.

Node Details

This n8n workflow utilizes several key nodes to achieve its dynamic routing functionality:

When chat message received (n8n Trigger)
Function: Acts as the starting n8n trigger, receiving real-time user input (chatInput).
Key Configuration: Standard LangChain chat trigger setup for real-time interaction.

Model Selector (LangChain Agent)
Function: Classifies the input complexity (1 for Complex, 2 for Simple) using the Gemini 2.0 Flash model.
Key Configuration: System message is strictly defined to output only a single number (1 or 2) for programmatic consumption by the next n8n node.

Model selector (LangChain Model Selector)
Function: Routes the execution path based on the output of the classification agent.
Key Configuration: Uses conditional logic to check if the incoming value ($('Model Selector').item.json.output.toNumber()) equals 1 (routes to Gemini 2.5 Pro) or 2 (routes to GPT-4.1 Nano).

2.5 pro (LangChain Chat Google Gemini)
Function: The LLM designated for complex tasks, maximizing reasoning power.
Key Configuration: Model set to models/gemini-2.5-pro.

4.1 nano (LangChain Chat OpenAI)
Function: The LLM designated for simple, low-cost tasks.
Key Configuration: Model set to gpt-4.1-nano, requiring OpenAI credentials.

Main Agent (LangChain Agent)
Function: The final n8n node that executes the request using the model dynamically selected by the router. It pulls the original chat input for execution.
Key Configuration: Includes the current date in the system prompt for context.

Related n8n Workflows

Free

Nodes: 6 Nodes
Updated: December 26 2025
View all
Created by

AI Automation Consultant | Helping Business Owners Implement AI Systems for growth and lead gen

Featured*