Keyflow

Search Keyflow

Find a shortcut, workflow, MCP server or skill

Simple Eval for Legal Benchmarking

n8n template #4712
This workflow demonstrates a simple way to run evals on a set of test cases stored in a Google Sheet. The example we are using comes from an info extraction task dataset, where we tested 6 different LLMs on 18 different test cases. You can see our sample data in this spreadsheet here to get started. Once you have this working for our dataset, you can plug in your own test cases matching different LLMs to see how it works with your own data. How it works: It loads test cases from Google Sheets. For each row in our Google Sheet, it grabs the source document, converting it to text. Our "LLM judge" passes the input/output of each LLM to GPT-4.1 to evaluate each test case (Pass/Fail + Reason). It

Hand off to your agent

Prompt

Help me set up the n8n workflow "Simple Eval for Legal Benchmarking" (https://n8n.io/workflows/4712). It uses: Google Sheets, Google Drive, Basic LLM Chain, Structured Output Parser, OpenRouter Chat Model. Import the template JSON into my n8n instance, list every credential I need to create, and walk me through testing it.

Paste into Claude Code and it will do the rest.

Apps and nodes

Similar workflows

NameViews
  1. AI Email Analyzer: Process PDFs, Images & Save to Google Drive + Telegram21K
  2. Automated Viral Content Engine for LinkedIn & X with AI Generation & Publishing402
  3. Benchmark LLM Performance on Legal Documents with Google Sheets and OpenRouter210
  4. AI Automated HR Workflow for CV Analysis and Candidate Evaluation51K
  5. Automate SEO-Optimized WordPress Posts with AI & Google Sheets37K
  6. Publish WordPress Posts to Social Media X, Facebook, LinkedIn, Instagram with AI32K
  7. Invoices from Gmail to Drive and Google Sheets27K
  8. Amazon Product Search Scraper with BrightData, GPT-4, and Google Sheets11K
Buy me a coffee