AI Image Creation via WhatsApp Chat - n8n Workflow

Use this powerful n8n workflow to generate stunning AI images by simply sending a text message via WhatsApp. This n8n workflow leverages Google Gemini and specialized image generation APIs for instant media creation.

Workflow Preview

Ready to automate?

Download this n8n workflow template and start using it instantly.

Who is this best for?

Content Creators & Marketers: Rapidly prototype visuals based on simple text prompts from a mobile device.
Technical Users: Those looking to build advanced integrations using the WhatsApp n8n node and large language models.


  • Automation Enthusiasts: Anyone seeking practical, real-world examples of complex n8n templates combining communication and generative AI services.

Overview

This comprehensive n8n workflow transforms your WhatsApp number into an AI image generation portal. By combining the immediacy of WhatsApp automation with the intelligence of Google Gemini, this n8n template solves the problem of needing to access a desktop or specialized application just to generate an image. When a user sends a text message, the n8n trigger starts the process. Gemini first processes the potentially brief or vague request and enhances it into a highly detailed, structured prompt suitable for high-quality image generation APIs. This highly efficient n8n workflow drastically cuts down the latency between idea and visual result, demonstrating the power and flexibility available in n8n automation.

How it Works

The entire process relies on several chained n8n node operations to ensure reliability and quality.


  1. Initiation (WhatsApp Trigger): The n8n workflow starts immediately upon receiving an incoming message via the WhatsApp n8n trigger. The user's text message is passed to the next step.

  2. Prompt Refinement (Generate prompt/Gemini): The input message is fed into the Gemini 2.5 Pro language model (using the Chat LLM n8n node) via a LangChain wrapper. This step uses a Structured Prompt n8n node (Output Parser) to ensure the LLM output is a clean, optimized, and detailed image prompt, rather than just conversational text.

  3. Image Generation (Generate Image): The refined prompt is sent through an HTTP Request n8n node to a chosen image generation API (e.g., Stable Diffusion, DALL-E, or similar). This step requires configuration specific to your chosen service.

  4. File Preparation (Convert to Image): The resulting base64 or binary data from the API is processed by the Convert to File n8n node, converting it into a proper file object that the WhatsApp service can handle.

  5. Delivery (Send Image): Finally, a WhatsApp n8n node sends the newly created image back to the original contact, completing the execution of the n8n workflow.

Installation Guide

To deploy and utilize this advanced n8n workflow, follow these steps:


  1. Import the n8n Workflow: Copy the provided JSON data and paste it directly into your n8n instance's canvas using the 'Import from JSON' function.

  2. Configure WhatsApp Trigger: Set up your WhatsApp Business Account (via Meta Developers) and configure the initial WhatsApp Trigger n8n node webhook to listen for incoming messages.

  3. Configure Gemini: Establish credentials for the Google Gemini 2.5 Pro n8n node (LLM Chat Google Gemini). You will need a valid Google AI API Key.

  4. Set Up Image API: Configure the 'Generate Image' HTTP Request n8n node. You must input the specific API endpoint URL, required headers (including your API key for the chosen image service), and structure the body to correctly use the refined prompt output from the 'Generate prompt' n8n node.

  5. Configure WhatsApp Sender: Set up the credentials for the final 'Send Image' WhatsApp n8n node, ensuring it can post outgoing media.

  6. Activate: Save the n8n workflow and set the n8n trigger to 'Active' to start listening for messages.

Node Details

WhatsApp Trigger n8n node: The starting point of the n8n workflow. It initiates the sequence whenever a user sends a message to the linked WhatsApp number.
Gemini 2.5 Pro (LLM Chat Google Gemini n8n node): Provides the language model intelligence. It takes the user's raw input and works with the Chain LLM n8n node to significantly improve the prompt quality.
Structured Prompt (Output Parser Structured n8n node): Essential for prompt engineering. It ensures the output from Gemini is a machine-readable JSON structure containing the optimal text for the image generator, rather than chat dialogue.
Generate prompt (Chain LLM n8n node): Orchestrates the Gemini model and the Structured Prompt parser to create the high-fidelity prompt data.
Generate Image (HTTP Request n8n node): This core action n8n node sends the optimized prompt to an external image generation service API (configuration required for URL, headers, and body).
Convert to Image (Convert To File n8n node): Processes the raw data (often base64 encoded image data) received from the API into a standard n8n file object, ready for transmission.


  • Send Image (WhatsApp n8n node): The final communication n8n node, responsible for sending the generated image file back to the original WhatsApp sender.

Related n8n Workflows

Free

Nodes: 8 Nodes
Updated: December 26 2025
View all
Created by

Crafting Intelligent AI Solutions | AI Engineer building the next generation of intelligent workflows on n8n. Automating the complex, one node at a time.

Featured*