Evaluate AI support agent quality and regressions with Google Sheets, Groq and Gmail

Go to Workflow
0 views
Built by Fahim Jilani Fahim Jilani
Created on September 24, 2026

Description

Quick overview
This workflow runs a customer-support AI agent against Google Sheets test cases, uses Groq-hosted LLMs to evaluate response quality, logs results back to Google Sheets, compares run metrics to an approved baseline, and sends regression or review alerts via Gmail.

How it works
Starts when you manually execute the workflow.
Reads test cases from Google Sheets, filters to only Active scenarios, and iterates through them in batches.
Builds a customer-support prompt for each test case and generates an agent response using Google Gemini.
Sends the question, expected facts/outcome, and the agent response to a Groq LLM to score relevance, reference accuracy, completeness, instruction compliance, and hallucination risk as JSON.
Parses the evaluator JSON and appends per-test results (including the agent response and overall result) to an Evaluation_Results sheet in Google Sheets.
Aggregates the current run’s results, selects the latest previously reviewed baseline from Google Sheets, and computes pass-rate regression plus metric deltas.
Writes the run-level metrics back to Google Sheets and routes the outcome to either mark the run successful or send a Gmail warning/alert to the QA team.

Setup
Create a Google Sheets file with Test_Cases, Evaluation_Results, and Baseline_Metrics tabs (including columns referenced in the workflow like Active, Test_ID, Question, Expected_Facts, Expected_Outcome, and Reviewed).
Add Google Sheets credentials in n8n and replace YOUR_GOOGLE_SHEET_ID in all Google Sheets nodes with your spreadsheet ID.
Add credentials for Google Gemini and Groq (used for the agent and evaluator chat models) and select the models you want to run.
Add Gmail credentials and update the recipient address ([email protected]) and any email content/subjects to match your QA process.
Ensure at least one baseline row in Baseline_Metrics is marked Reviewed=true so the workflow can compare the current run against an approved baseline.

Nodes Used (7)

AI Agent
@n8n/n8n-nodes-langchain.agent
Basic LLM Chain
@n8n/n8n-nodes-langchain.chainLlm
Code
n8n-nodes-base.code
Gmail
n8n-nodes-base.gmail
Google Gemini Chat Model
@n8n/n8n-nodes-langchain.lmChatGoogleGemini
Google Sheets
n8n-nodes-base.googleSheets
Groq Chat Model
@n8n/n8n-nodes-langchain.lmChatGroq