Multimodal telegram bot with voice, image & video analysis using Claude & Gemini
n8n template #9008Summary
Use templateQuick overview This is a starting point for building a Telegram AI agent. The base handles four input types: voice, pictures, video, and text, through the AI models of your choice. From here you connect tools to expand what the agent can do inside your n8n workflows. How it works Input: a message sent to the bot chat. A Switch node sorts the message by type: Voice message Picture message Video message Text message It currently uses OpenAI and Gemini to analyze voice, photos, and video, but you can swap in other models. The model reads the message, generates a response from the system prompt, and sends it back as a Telegram message. Setup Create the Telegram bot. In Telegram, search for "BotF
Hand off to your agent
Prompt
Help me set up the n8n workflow "Multimodal telegram bot with voice, image & video analysis using Claude & Gemini" (https://n8n.io/workflows/9008). It uses: Telegram, AI Agent, Anthropic Chat Model, Simple Memory, OpenAI, Google Gemini. Import the template JSON into my n8n instance, list every credential I need to create, and walk me through testing it.
Paste into Claude Code and it will do the rest.
Apps and nodes
Similar workflows
NameViews
- Multimodal Slack AI assistant with voice, image & video processing776
- Create & Share AI Videos with Telegram, Gemini & Post to TikTok, Instagram, FB376
- Generate AI Photos with Gemini & Auto-Post to FB, Instagram & X with Approval239
- 💅 AI Agents Generate Content & Automate Posting for Beauty Salon Social Media 📲129
- Angie, personal AI assistant with Telegram voice and text326K
- Conversational Telegram Bot with GPT-5/GPT-4o for Text and Voice Messages82K
- Telegram AI bot assistant: ready-made template for voice & text messages38K
- Personal Life Manager with Telegram, Google Services & Voice-Enabled AI34K