How to Build an AI Voice Agent with Python, Exotel & Groq
Learn how to build a bilingual (English/Hindi) AI voice agent using Python, FastAPI, Exotel, and Groq. Includes step-by-step code and n8n WhatsApp automation.
Build an AI Voice Agent with Python, Exotel & Groq
Your school's admission office gets the same 50 calls every day. "What is the fee?" "Which documents do I need?" A staff member answers each one manually — while real work piles up.
This tutorial fixes that. You will build a bilingual AI voice agent that picks up inbound calls, answers in English or Hindi, and automatically sends a WhatsApp brochure to every interested parent. No staff involvement. No repeated answers.
By the end, you will have a live phone number powered by Python, FastAPI, Groq's Llama 3.3-70B, and Exotel — deployed and taking real calls.
Table of Contents
- [Prerequisites & Tech Stack](#prerequisites–tech-stack)
- [Architecture Overview](#architecture-overview)
- [Step 1: Collect Your API Credentials](#step-1-collect-your-api-credentials)
- [Step 2: Build the AI Brain](#step-2-build-the-ai-brain)
- [Step 3: Build the Voice Layer](#step-3-build-the-voice-layer)
- [Step 4: Detect Intent and Trigger WhatsApp](#step-4-detect-intent-and-trigger-whatsapp)
- [Step 5: Wire Everything with FastAPI](#step-5-wire-everything-with-fastapi)
- [Step 6: Set Up the n8n Workflow](#step-6-set-up-the-n8n-workflow)
- [FAQ](#faq)
- [Conclusion](#conclusion)
Prerequisites & Tech Stack
You need to know Python basics and understand what a POST request is. That is enough — everything else is explained as we go.
| Tool | Version | Why it is needed | |——|———|—————–| | Python | 3.11+ | Backend runtime | | FastAPI | 0.111 | HTTP server that Exotel POSTs call data to | | Groq SDK | 0.9.0 | API access to Llama 3.3-70B for AI responses | | OpenAI Whisper | 20231117 | Transcribes caller audio to text, detects English/Hindi | | edge-tts | 6.1.9 | Converts AI text response to natural Indian voice audio | | httpx | 0.27.0 | Async HTTP client to download Exotel recordings | | aiofiles | 23.2.1 | Required for FastAPI to serve audio files statically | | n8n | Cloud / Self-hosted | Automation layer for WhatsApp follow-up messages | | Exotel | Active account | Handles inbound phone calls across India and SE Asia | | ngrok | Latest | Tunnels your local server so Exotel can reach it |
Install everything in one command:
pip install groq==0.9.0 edge-tts==6.1.9 python-dotenv==1.0.1 \
fastapi==0.111.0 uvicorn==0.30.1 requests==2.32.3 \
openai-whisper==20231117 httpx==0.27.0 aiofiles==23.2.1
Architecture Overview
Here is the full journey of a single call — from the parent dialing to receiving a WhatsApp message:
[Parent's Phone]
|
| dials virtual number
v
[Exotel]
|
| POSTs recording URL to
v
[FastAPI Server /exotel/callback]
|
|--- Downloads audio
|--- Whisper: audio → text + detects language (en/hi)
|--- Groq Llama 3.3-70B: text → AI answer
|--- edge-tts: AI answer → MP3 audio
|
| returns MP3 URL
v
[Exotel plays MP3 to caller]
|
| call ends
v
[Background: Intent Analysis via Groq]
|
| if new guardian detected
v
[n8n Webhook → WhatsApp message in English or Hindi]
Each component's role:
- Exotel acts as the telephony layer — it receives the call, records what the parent says, and plays back whatever audio URL we return.
- Whisper does two jobs at once: it transcribes the audio and tells us whether the caller spoke in English or Hindi.
- Groq Llama 3.3-70B generates the response using only the school's knowledge base — it cannot make up facts.
- edge-tts picks the right voice: Indian English (
en-IN-NeerjaNeural) or Hindi (hi-IN-SwaraNeural). - n8n listens for a webhook and triggers a bilingual WhatsApp message when a hot lead is detected.
Step 1: Collect Your API Credentials
What we are doing: Before a single line of code, gather every API key and credential the system needs.
Create a .env file in your project root:
GROQ_API_KEY=gsk_your_key_here
N8N_WEBHOOK_URL=https://your-n8n.com/webhook/admission
EXOTEL_API_KEY=your_exotel_api_key
EXOTEL_API_TOKEN=your_exotel_api_token
EXOTEL_ACCOUNT_SID=your_account_sid
SERVER_BASE_URL=https://abc123.ngrok.io
Add .env to .gitignore immediately — never commit real credentials.
Getting each credential:
Groq: Sign up at console.groq.com → API Keys → Create API Key. The key starts with gsk_. Free tier gives 14,400 tokens/minute on Llama 3.3-70B.
Exotel: Sign up at exotel.com — approval takes 1-2 business days. After approval: Settings → API → copy your API Key, API Token, and Account SID. Then go to ExoPhones → Buy Number to get your virtual number. On your number's App Settings, set the Passthru URL to https://your-ngrok-url/exotel/callback with method POST.
n8n: Sign up at app.n8n.cloud → New Workflow → add a Webhook node → set method to POST, path to admission → click "Listen for test event" → copy the Test URL.
What just happened: SERVER_BASE_URL is the most important variable to keep updated. Every time ngrok restarts, this URL changes — update it in .env and restart uvicorn. This is the URL Exotel uses to download your audio files.
Step 2: Build the AI Brain
What we are doing: Give Groq's Llama 3.3-70B a strict role — it answers only from a school knowledge base file, never inventing facts.
First, create school_info.txt in your project root. This is the only source of truth the AI will use:
Modern Academy - Admission Information
Fees:
- Primary Section (Grade 1-5): INR 45,000 per year
- High School Section (Grade 6-10): INR 65,000 per year
Required Documents: Birth certificate, 2 passport photos,
transfer certificate, Aadhaar card (student + parent)
Timing: School 8 AM–2 PM Mon–Sat | Office 9 AM–4 PM Mon–Fri
Visit: Campus tours every Saturday 10 AM–12 PM. Call to book.
The key function in ai_brain.py:
def get_ai_response(user_question: str) -> str:
with open("school_info.txt", "r", encoding="utf-8") as f:
context = f.read()
system_prompt = f"""You are a polite AI Admission Counselor for Modern Academy.
Reply in the SAME language the parent used — English or Hindi.
Answer ONLY from the context below. Keep it under 2 sentences.
If unsure, say: "I don't have that info. Let me connect you to our team."
Context: {context}"""
response = client.chat.completions.create(
model="llama-3.3-70b-versatile",
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": user_question}
],
temperature=0.4,
max_tokens=150
)
return response.choices[0].message.content
What just happened: temperature=0.4 is deliberately low — you do not want creative variation when the AI is quoting a fee. The system prompt uses SAME language to make the AI mirror the caller's language automatically. max_tokens=150 keeps responses short enough for a phone call — nobody wants a paragraph read aloud to them.
Full code for
ai_brain.pyincluding error handling → [GitHub](#)
Step 3: Build the Voice Layer
What we are doing: Two functions handle all audio. transcribe_audio converts the caller's speech to text and tells us the language. text_to_speech_async converts the AI's answer to an MP3 in the right voice.
The key insight — Whisper already detects language during transcription. You do not need a separate detection call:
def transcribe_audio(audio_path: str) -> tuple[str, str]:
result = _whisper_model.transcribe(audio_path)
transcript = result.get("text", "").strip()
lang = result.get("language", "en") # "en", "hi", etc.
print(f"[STT] Language={lang} | Text={transcript}")
return transcript, lang
The async TTS function — uses Indian voices so callers hear a familiar accent:
async def text_to_speech_async(text: str, out_path: str, lang: str = "en") -> str | None:
voice = "hi-IN-SwaraNeural" if lang == "hi" else "en-IN-NeerjaNeural"
await edge_tts.Communicate(text, voice).save(out_path)
return out_path
What just happened: Notice text_to_speech_async is an async function. This matters — calling asyncio.run() inside a FastAPI async route raises RuntimeError: This event loop is already running. By making TTS async, we await it cleanly inside the route handler.
The Whisper model is loaded once at module import (whisper.load_model("small")) and reused across calls. Loading it per-call adds 3–5 seconds of latency — enough to make Exotel time out.
Full code for
stt_tts.py→ [GitHub](#)
Step 4: Detect Intent and Trigger WhatsApp
What we are doing: After a call ends, send the full conversation to Groq and ask if the caller is a genuine new parent. If yes, fire a webhook to n8n.
The key function in agent_server.py:
def analyze_and_trigger_n8n(caller_phone: str, conversation: str):
prompt = f"""Analyze this call. Is the caller a new parent interested in admission?
Respond ONLY as JSON: {{"is_new_guardian": true/false, "summary": "one line", "language": "en or hi"}}
Conversation: {conversation}"""
result = client.chat.completions.create(
model="llama-3.3-70b-versatile",
messages=[{"role": "user", "content": prompt}],
response_format={"type": "json_object"}
)
analysis = json.loads(result.choices[0].message.content)
if analysis.get("is_new_guardian") is True:
requests.post(os.getenv("N8N_WEBHOOK_URL"), json={
"phone": caller_phone,
"summary": analysis.get("summary", ""),
"language": analysis.get("language", "en")
}, timeout=10)
What just happened: response_format={"type": "json_object"} is non-negotiable. Without it, Groq sometimes wraps JSON in markdown code fences like `json — which crashes json.loads() silently. With it, you always get clean, parseable JSON.
The language field travels to n8n so the WhatsApp message can be sent in the caller's own language — a small detail that dramatically improves response rates.
Full code for
agent_server.py→ [GitHub](#)
Step 5: Wire Everything with FastAPI
What we are doing: Build the single endpoint Exotel calls after every recording. It runs the full pipeline — download → STT → AI → TTS — and returns an audio URL.
Two things must happen before the app starts:
os.makedirs("recordings", exist_ok=True)
app.mount("/recordings", StaticFiles(directory="recordings"), name="recordings")
Without StaticFiles, Exotel gets a 404 when it tries to download your MP3 — it plays nothing, and the caller hears silence.
The core endpoint:
@app.post("/exotel/callback")
async def exotel_callback(
bg: BackgroundTasks,
CallSid: str = Form(...),
From: str = Form(...),
RecordingUrl: str = Form(None),
CallStatus: str = Form(None),
):
# Call ended — analyze intent in background
if CallStatus == "completed":
session = call_store.pop(CallSid, {})
history = "\n".join(session.get("history", []))
bg.add_task(analyze_and_trigger_n8n, From, history)
return PlainTextResponse("Call ended.")
# First contact — play bilingual greeting
if not RecordingUrl:
greet_path = f"recordings/{CallSid}_greet.mp3"
await text_to_speech_async(
"Welcome to Modern Academy. Speak in English or Hindi. "
"Aap Hindi ya English mein baat kar sakte hain.",
greet_path, lang="en"
)
return PlainTextResponse(f"{BASE_URL}/recordings/{CallSid}_greet.mp3")
# Download recording, transcribe, respond
audio_path = f"recordings/{CallSid}_{uuid.uuid4().hex}.mp3"
async with httpx.AsyncClient() as c:
r = await c.get(RecordingUrl, timeout=60)
open(audio_path, "wb").write(r.content)
caller_text, lang = transcribe_audio(audio_path)
ai_text = get_ai_response(caller_text)
resp_path = f"recordings/{CallSid}_resp.mp3"
await text_to_speech_async(ai_text, resp_path, lang=lang)
return PlainTextResponse(f"{BASE_URL}/recordings/{CallSid}_resp.mp3")
What just happened: BackgroundTasks is what makes this production-safe. Exotel expects a response in under 3 seconds — intent analysis via Groq takes 1-2 seconds on its own. By running it as a background task, we return the response immediately and let analysis happen after. The caller is never kept waiting for something they will not hear.
Run the server:
uvicorn main:app --reload --port 8000
Then expose it with ngrok:
ngrok http 8000
Copy the HTTPS URL, update SERVER_BASE_URL in .env, restart uvicorn, and update the Passthru URL in your Exotel dashboard.
Full code for
main.py→ [GitHub](#)
Step 6: Set Up the n8n Workflow
What we are doing: Build the automation that sends a bilingual WhatsApp message every time a hot lead is detected.
In n8n, add these nodes in order:
Node 1 — Webhook (already created in Step 1) Receives this payload from agent_server.py:
{"phone": "+919999999999", "summary": "Parent interested in Grade 3", "language": "en"}
Node 2 — IF Condition: {{ $json.language }} equals hi This splits into two branches — one for Hindi callers, one for English.
Node 3a — HTTP Request (English branch) POST to your WhatsApp Business API with:
Hello! Thank you for your interest in Modern Academy.
Here is our brochure: https://modernacademy.in/visit
Our team is available Mon–Fri, 9 AM–4 PM.
Node 3b — HTTP Request (Hindi branch)
Namaste! Modern Academy mein aapki dilchaspi ke liye dhanyavaad.
Yahan hamara brochure hai: https://modernacademy.in/visit
Hamari team Mon–Fri, subah 9 baje se shaam 4 baje tak uplabdh hai.
Node 4 — Google Sheets (optional) Log each lead: phone, summary, language, timestamp.
Click the Inactive toggle in n8n's top-right to activate the workflow. Switch N8N_WEBHOOK_URL in .env from the test URL to the production URL.
FAQ
Q: Does Whisper handle Hindi accurately enough for production?
Yes — with the small model. The base model is faster but loses significant accuracy on Hindi, especially with regional accents. The small model adds about 1 second of processing time but handles Hindi, Hinglish (mixed Hindi-English), and regional Indian English well. For a call-center context where responses do not need to be instant, this trade-off is worth it.
Q: What happens if Exotel does not receive a response in time?
Exotel times out after a few seconds of no response and plays an error message to the caller. The two most common causes are: (1) Whisper taking too long on a long audio clip — keep recordings under 10 seconds by configuring Exotel's max recording duration; (2) Groq API being slow — the llama-3.3-70b-versatile model is generally under 1 second on Groq's infrastructure, but add a timeout parameter to your API call to fail fast rather than hang.
Q: Can this agent handle a conversation with follow-up questions?
Yes. The call_store dictionary stores conversation history per CallSid. Every turn appends the caller's text and the AI's response. When the AI generates its next answer, the full history is part of the system context — so it remembers what was said earlier in the same call.
Q: Is Exotel available outside India?
Exotel operates primarily in India, Southeast Asia (Singapore, Indonesia, Malaysia, Philippines), and the Middle East. If your deployment target is outside these regions, look at Twilio (global coverage) or Vonage as alternatives — the FastAPI endpoint pattern in this tutorial works identically with both; only the credential setup differs.
Conclusion
Here is what you built:
- A FastAPI server that receives Exotel call webhooks and runs a full STT → AI → TTS pipeline per turn
- A bilingual Whisper layer that auto-detects English and Hindi without any manual routing
- A Groq-powered AI counselor that answers strictly from a knowledge base — no hallucination
- A background intent analyzer that fires a bilingual WhatsApp follow-up for every hot lead
- A working n8n workflow that sends the message in the caller's own language
The modular design pays off in maintenance: when Exotel updates their webhook format, only main.py changes. When you want to switch from edge-tts to a paid voice provider, only stt_tts.py changes. Nothing cascades.
Have questions about a specific step, or did something not work as expected? Drop a comment below — I read every one and typically respond within 24 hours. If you got it working, I would love to hear what you built it for.
Coming up in Part 2: Adding sentiment detection to escalate frustrated callers to a live human agent in real time — without interrupting the AI for calm conversations.
