Reduce LLM Costs with Semantic Caching using Redis Vector Store and HuggingFace
n8n template #10887Summary
Use templateStop Paying for the Same Answer Twice Your LLM is answering the same questions over and over. "What's the weather?" "How's the weather today?" "Tell me about the weather." Same answer, three API calls, triple the cost. This workflow fixes that. What Does It Do? Semantic caching with superpowers. When someone asks a question, it checks if you've answered something similar before. Not exact matches—semantic similarity. If it finds a match, boom, instant cached response. No LLM call, no cost, no waiting. First time: "What's your refund policy?" → Calls LLM, caches answer Next time: "How do refunds work?" → Instant cached response (it knows these are the same!) Result: Faster responses + way low
Hand off to your agent
Prompt
Help me set up the n8n workflow "Reduce LLM Costs with Semantic Caching using Redis Vector Store and HuggingFace" (https://n8n.io/workflows/10887). It uses: Code, AI Agent, Embeddings Hugging Face Inference, OpenAI Chat Model, Redis Chat Memory, Recursive Character Text Splitter, Default Data Loader, Redis Vector Store. Import the template JSON into my n8n instance, list every credential I need to create, and walk me through testing it.
Paste into Claude Code and it will do the rest.
Apps and nodes
CodeAI AgentEmbeddings Hugging Face InferenceOpenAI Chat ModelRedis Chat MemoryRecursive Character Text SplitterDefault Data LoaderRedis Vector Store
Similar workflows
NameViews
- AI-Powered WhatsApp Chatbot for Text, Voice, Images, and PDF with RAG47K
- Automate sales cold calling pipeline with Apify, gpt-5.6-terra, and WhatsApp34K
- Advanced AI Demo (Presented at AI Developers #14 meetup)24K
- AI: Summarize podcast episode and enhance using Wikipedia13K
- Create a Multi-Modal Telegram Support Bot with GPT-4 and Supabase RAG10K
- Scale Deal Flow with a Pitch Deck AI Vision, Chatbot and QDrant Vector Store6,822
- Explore n8n Nodes in a Visual Reference Library5,197
- Build a Chatbot with Reinforced Learning Human Feedback (RLHF) and RAG3,306