Multimodal Telegram AI Bot with Claude & Gemini - n8n Workflow

Build a powerful multimodal Telegram bot using this n8n workflow. Integrates voice transcription via OpenAI, image/video analysis via Gemini, and conversational chat via Claude.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?

Automation specialists looking for advanced n8n templates.
Developers needing a customizable, multimodal chatbot infrastructure.
Anyone wanting to build and host a persistent, intelligent AI agent on Telegram.
Users who need complex logic flows within a single n8n workflow.

Overview

This advanced n8n workflow solves the complexity of creating a truly multimodal AI experience within a chat application like Telegram. Unlike simple text-only bots, this solution leverages specialized AI services—OpenAI for high-quality voice transcription, and Google Gemini for robust image and video analysis—before routing all inputs to a central language model, Anthropic's Claude. By using a sophisticated Switch n8n node, the flow intelligently detects the input type (text, voice, photo, or video) and processes it accordingly. This powerful n8n template ensures that your AI agent is versatile, context-aware, and highly responsive to diverse user inputs, providing a seamless conversational experience powered by a robust n8n backend.

How it Works

The process begins immediately upon a user sending a message, which activates the Telegram Trigger n8n node.


  1. Trigger and Routing: The Telegram Trigger initiates the n8n workflow. A subsequent Switch n8n node analyzes the incoming message object to determine if the message contains text, a voice note, a picture, or a video. It also filters out simple /start commands.

  2. Voice Processing: If a voice note is detected, the workflow uses a Telegram n8n node (Get a Audio File) to retrieve the file, which is then passed to the Transcribe a recording n8n node (OpenAI) to convert the speech into text.

  3. Visual Processing: If an image or video is detected, corresponding Telegram n8n nodes retrieve the binary data. These files are then routed to Google Gemini n8n nodes (Analyze an image or Analyze video) which extract descriptive text based on the visual content.

  4. Data Consolidation: The various paths (raw text, transcribed audio, visual analysis results) converge at the Merge n8n node. The following Edit Fields n8n node creates a unified 'User Input' variable, ensuring the core AI agent receives a single, coherent text prompt.

  5. AI Agent Execution: The data enters the AI Agent n8n node. This agent is configured with the Anthropic Chat Model (Claude Sonnet 4) for intelligence, Simple Memory to retain conversation context, and the Date & Time n8n node as an available tool.

  6. Response: The agent processes the input, generates a response, and the flow concludes with the Send Regular Message n8n node, delivering the AI's output back to the user via Telegram.

Installation Guide

To deploy this comprehensive n8n workflow, follow these steps:


  1. Import the n8n Workflow: Copy the provided JSON data and paste it into your n8n instance using the 'New' -> 'Import from JSON' function.

  2. Telegram Setup: Create a new bot using Telegram's BotFather and obtain the access token. Create a new Telegram credential in n8n and input this token. Update the credentials in both the Telegram Trigger, the Telegram file nodes (Get a Audio File, etc.), and the final Send Regular Message n8n node.

  3. AI Credentials:

OpenAI: Create an OpenAI API key for the Transcribe a recording n8n node.
Anthropic (Claude): Create an Anthropic API key and configure the Anthropic Chat Model n8n node credential.
* Google Gemini: Create a Gemini API key for the image and video analysis n8n node instances.

  1. Activate: Ensure all required n8n node credentials are set up. Activate the n8n workflow to start listening for Telegram messages.

Node Details

Telegram Trigger (n8n trigger): Starts the n8n workflow upon receiving a message update in Telegram. Crucial for real-time interaction.
Switch (n8n node): Essential flow control. Routes messages based on content type (text, voice, photo, video) using expressions like ={{ $json.message.voice }}.
Get a Audio File / Get a Photo / Get a Video File (Telegram n8n nodes): Retrieves the binary file data using the file_id extracted from the Telegram message object.
Transcribe a recording (OpenAI n8n node): Converts binary audio data into text, integrating a key OpenAI service into the n8n workflow.
Analyze an image / Analyze video (Google Gemini n8n nodes): Uses the Gemini multimodal model to analyze the visual content and generate a textual description or summary based on any user caption.
Merge (n8n node): Consolidates the outputs from the different processing branches (transcription, analysis, raw text) into a single path for the AI agent.
Edit Fields (n8n node): Standardizes the input data into a single 'User Input' field, ensuring consistency before passing to the AI Agent.
Simple Memory (n8n node): Configured with Buffer Window Memory, it uses the Telegram chat ID as the sessionKey to maintain conversational context across multiple interactions.
Anthropic Chat Model (n8n node): The chosen Large Language Model (Claude Sonnet 4) provides the core intelligence for the AI Agent. Can be swapped for other LLMs.
AI Agent (n8n node): The orchestration hub. It receives the user input, utilizes the LLM and the included tools (Date & Time), and handles the system prompt instructions.


  • Send Regular Message (n8n node): The final action, sending the AI Agent's text output back to the Telegram chat.

Related n8n Workflows

Free

Nodes: 13 Nodes
Updated: December 26 2025
View all
Created by
Keith Uy
Keith Uy

Featured*