Keyflow

Search Keyflow

Find a shortcut, workflow, MCP server or skill

Easy Image Captioning with Gemini 1.5 Pro

n8n template #2418
This n8n workflow demonstrates how to automate image captioning tasks using Gemini 1.5 Pro - a multimodal LLM which can accept and analyse images. This is a really simple example of how easy it is to build and leverage powerful AI models in your repetitive tasks. How it works For this demo, we'll import a public image from a popular stock photography website, Pexel.com, into our workflow using the HTTP request node. With multimodal LLMs, there is little do preprocess other than ensuring the image dimensions fit within the LLMs accepted limits. Though not essential, we'll resize the image using the Edit image node to achieve fast processing. The image is used as an input to the basic LLM node

Hand off to your agent

Prompt

Help me set up the n8n workflow "Easy Image Captioning with Gemini 1.5 Pro" (https://n8n.io/workflows/2418). It uses: Edit Image, HTTP Request, Code, Basic LLM Chain, Structured Output Parser, Google Gemini Chat Model. Import the template JSON into my n8n instance, list every credential I need to create, and walk me through testing it.

Paste into Claude Code and it will do the rest.

Apps and nodes

Similar workflows

NameViews
  1. Host Your Own AI Deep Research Agent with n8n, Apify and OpenAI o348K
  2. CV Resume PDF Parsing with Multimodal Vision AI16K
  3. Transcribing Bank Statements To Markdown Using Gemini Vision AI15K
  4. 🤖💬 Conversational AI Chatbot with Google Gemini for Text & Image | Telegram11K
  5. Explore n8n Nodes in a Visual Reference Library5,197
  6. Automate AI video creation & multi-platform publishing with Gemini & Creatomate5,035
  7. Intelligent Web Query and Semantic Re-Ranking Flow using Brave and Google Gemini4,863
  8. Automated Financial Tracker: Telegram Invoices to Notion with Gemini AI Reports3,597
Buy me a coffee