Build a website RAG support chatbot with Firecrawl, Pinecone, and Gemini

Go to Workflow
0 views
Built by Mohammad Abid Mohammad Abid
Created on September 14, 2026

Description

Quick overview
This template indexes a website into Pinecone using Firecrawl, Google Gemini embeddings, and basic HTML cleaning, then exposes a public n8n chat webhook where a Gemini-powered agent answers customer questions by retrieving relevant website chunks from Pinecone.

How it works
Starts an ingestion run when you manually execute the workflow.
Looks up the Pinecone index host and clears all vectors in the configured Pinecone namespace to avoid duplicate content.
Uses Firecrawl to map the target website and returns up to 100 discovered page URLs.
Filters, deduplicates, and normalizes the URLs, then processes them one at a time with a short delay to reduce request bursts.
Fetches each page over HTTP, strips HTML/scripts/styles into plain text, and skips pages with too little usable content.
Splits remaining text into overlapping chunks, creates Google Gemini embeddings, and stores the vectors with URL metadata in Pinecone.
When a chat message is received via webhook, a Gemini agent retrieves the most relevant chunks from Pinecone, uses short windowed memory for context, and returns the final answer.

Setup
Add credentials for Firecrawl, Pinecone, and Google Gemini (PaLM/Gemini) and ensure the same Gemini embedding model is used for both ingestion and retrieval.
Create a Pinecone index with a vector dimension that matches your chosen Gemini embedding model, then set YOUR_PINECONE_INDEX and YOUR_PINECONE_NAMESPACE in all Pinecone-related nodes.
Replace https://example.com/ with the website root URL you want to crawl, and adjust the Firecrawl URL limit, blocked keyword list, chunking, and delay settings as needed.
Review the namespace deletion step carefully and use a dedicated namespace, because each ingestion run deletes all vectors in that namespace before re-indexing.
Copy the public chat webhook URL from the chat trigger and embed/configure it in your site or chat client after ingestion completes and retrieval answers look correct.

Nodes Used (8)

AI Agent
@n8n/n8n-nodes-langchain.agent
Code
n8n-nodes-base.code
Default Data Loader
@n8n/n8n-nodes-langchain.documentDefaultDataLoader
Embeddings Google Gemini
@n8n/n8n-nodes-langchain.embeddingsGoogleGemini
Google Gemini Chat Model
@n8n/n8n-nodes-langchain.lmChatGoogleGemini
HTTP Request
n8n-nodes-base.httpRequest
Pinecone Vector Store
@n8n/n8n-nodes-langchain.vectorStorePinecone
Simple Memory
@n8n/n8n-nodes-langchain.memoryBufferWindow