Keyflow

Search Keyflow

Find a shortcut, workflow, MCP server or skill

Transcribing Bank Statements To Markdown Using Gemini Vision AI

n8n template #2421
This n8n workflow demonstrates an approach to parsing bank statement PDFs with multimodal LLMs as an alternative to traditional OCR. This allows for much more accurate data extraction from the document especially when it comes to tables and complex layouts. Multimodal Parsing is better than traditiona OCR because: It reduces complexity and overhead by avoiding the need to preprocess the document into text format such as markdown before passing to the LLM. It handles non-standard PDF formats which may produce garbled output via traditional OCR text conversion. It's orders of magnitude cheaper than premium OCR models that still require post-processing cleanup and formatting. LLMs can format to

Hand off to your agent

Prompt

Help me set up the n8n workflow "Transcribing Bank Statements To Markdown Using Gemini Vision AI" (https://n8n.io/workflows/2421). It uses: Edit Image, HTTP Request, Google Drive, Compression, Code, Basic LLM Chain, Google Gemini Chat Model, Information Extractor. Import the template JSON into my n8n instance, list every credential I need to create, and walk me through testing it.

Paste into Claude Code and it will do the rest.

Apps and nodes

Similar workflows

NameViews
  1. Scale Deal Flow with a Pitch Deck AI Vision, Chatbot and QDrant Vector Store6,822
  2. API Schema Extractor24K
  3. CV Resume PDF Parsing with Multimodal Vision AI16K
  4. Easy Image Captioning with Gemini 1.5 Pro12K
  5. Narrating over a Video using Multimodal AI8,840
  6. Explore n8n Nodes in a Visual Reference Library5,197
  7. Resume Screening & Behavioral Interviews with Gemini, Elevenlabs, & Notion ATS5,127
  8. Automatic Weekly Digital PR Stories Suggestions with Reddit and Anthropic2,911
Buy me a coffee