I Built a Real Estate Call Center Run Entirely by AI Voice Agents
What if your next phone call to a real estate company wasn't answered by a human—but by an AI agent so lifelike, you couldn't tell the difference? No hold music. No "please press 1 for sales." Just instant, intelligent conversation that books viewings, searches properties, and closes deals 24/7.
Here's the brutal truth: most AI voice demos are toys. They break the moment you ask something unexpected. They hallucinate prices. They stumble over accents. And they sure as hell can't handle a real phone line with real customers calling in.
But what if you could build something actually production-ready? Something that receives inbound calls, makes outbound calls, searches live property data in milliseconds, and speaks with a voice so natural it gives you goosebumps?
That's exactly what the neural-maze/realtime-phone-agents-course delivers. This isn't another "Hello World" tutorial. It's a 5-week deep dive into building a complete AI call center from scratch—using FastRTC for sub-second latency, Superlinked for intelligent property search, Twilio for real telephony, and RunPod for scalable GPU deployment.
By the end, you'll have deployed multiple AI avatars with distinct personalities, full conversation tracing, and the ability to handle complex queries like "Do you have apartments in Barrio de Salamanca under €900,000?"—all over a live phone call.
Ready to stop building demos and start building systems that actually work? Let's dive in.
What Is the Realtime Phone Agents Course?
The realtime-phone-agents-course is an open-source educational project created by The Neural Maze—a newsletter and YouTube channel run by senior ML engineers Miguel Otero Pedrido and Jesús Copado. Their mission? Teaching developers to build AI systems that survive contact with reality.
This course takes a radically different approach from typical tutorials. Instead of abstract examples, you build a functional real estate company where every "employee" is an AI voice agent. The curriculum spans five intensive weeks, each unlocking new capabilities through Substack articles, live coding sessions, and production-grade code pushed directly to the repository.
Why it's trending now: The voice AI space is exploding. OpenAI's Realtime API, Google's Gemini Live, and countless startups are racing to own the "voice interface." But here's what most miss: latency kills user experience. A 2-second delay between speech and response feels like talking to a machine. A 200-millisecond delay feels like magic.
This course cracks that code with FastRTC—a Python↗ Bright Coding Blog library purpose-built for real-time communication that achieves sub-second streaming. Combined with self-hosted STT/TTS models on GPU infrastructure, you get enterprise-grade voice AI without the enterprise-grade bills.
The repository has gained serious traction among ML engineers precisely because it bridges the gap between "cool demo" and "ship to production." It's not just about making an AI talk. It's about making it reliable, observable, and scalable.
Key Features That Separate This From Toy Demos
Let's dissect what makes this system production-ready rather than prototype-grade:
⚡ Sub-Second Realtime Streaming with FastRTC
FastRTC handles the brutal technical challenge of streaming audio bidirectionally with minimal latency. Unlike WebSocket hacks or polling-based approaches, it's purpose-built for real-time media. The course shows you how to integrate it with Twilio's media streams for genuine phone-call responsiveness.
🏠 Multi-Attribute Vector Search via Superlinked
Traditional RAG fails when users combine constraints: "I want a 3-bedroom, under $500K, in this specific neighborhood, with parking."* Superlinked solves this by unifying text, numerical, and categorical data into a single searchable space—with dynamic weight adjustment at query time. No more multiple searches, no more re-ranking hacks.
📞 Full Twilio Integration: Inbound AND Outbound
Most tutorials stop at web demos. This course wires your agent to actual phone numbers. Receive calls from customers. Make proactive outbound calls to leads. Handle call routing, media streaming, and webhook management—the full telephony stack.
🗣️ Production STT/TTS Pipeline with Model Choice
You're not locked into one provider. The system supports:
- STT: Moonshine (local), Groq API (fast cloud), Faster Whisper on RunPod (self-hosted quality)
- TTS: Kokoro (local), Together AI API, Orpheus 3B on RunPod (emotionally expressive speech)
Mix and match based on latency, cost, and quality requirements.
🎭 Multi-Avatar System with Personality Engineering
Each avatar has distinct YAML-defined personalities—Dan, Jess, Leah, Leo, Mia, Tara, Zac, Zoe. System prompts are versioned and tracked, enabling A/B testing of conversational styles and rollback when updates degrade performance.
📊 Full Observability with Opik Tracing
Every pipeline step is instrumented: transcription time, LLM latency, tool call duration, TTS generation speed. Complete conversation threads are stored for analysis. This isn't optional—it's essential for production debugging when your agent starts behaving strangely at 2 AM.
🚀 Scalable GPU Deployment on RunPod
Self-host your models without managing Kubernetes. The course provides Dockerfiles for Faster Whisper and Orpheus 3B with CUDA optimization, plus Makefile commands for one-click deployment.
Real-World Use Cases Where This Architecture Dominates
1. Real Estate Lead Qualification
A prospect calls your number at 11 PM. Your AI avatar answers instantly, qualifies their budget and preferences, searches live listings via Superlinked, and books a viewing—while your human agents sleep. The next morning, you have a fully qualified lead with transcribed conversation notes in Opik.
2. Healthcare Appointment Scheduling
Patients call to book, reschedule, or ask about preparation instructions. The agent handles natural conversation flow, checks real-time calendar availability, and confirms via SMS. Multi-avatar support lets you match personality to department—warm and reassuring for oncology, efficient and direct for radiology.
3. E-Commerce Order Support
"Where's my package?" "Can I return this?" "Do you have this in blue?" The agent queries live order databases, initiates returns, checks inventory across warehouses—all while maintaining conversational context across topic switches that break simpler systems.
4. Financial Services Onboarding
Walk new customers through KYC verification over the phone. The agent asks dynamic follow-up questions based on responses, validates documents in real-time, and escalates complex cases to human specialists with full context transfer.
5. Multi-Language Customer Service
With Kokoro and Orpheus supporting multiple languages, deploy avatars that natively serve global markets without hiring native speakers for every timezone. The same architecture, different YAML configuration.
Step-by-Step Installation & Setup Guide
Let's get your environment ready. This isn't a 5-minute setup—and that's the point. Production systems require proper configuration.
Prerequisites
- Python 3.10+
- Docker↗ Bright Coding Blog and Docker Compose
- Make (for build automation)
- ngrok account (for local tunneling to Twilio)
- Accounts: Twilio, RunPod, Qdrant Cloud, Opik, Groq, Together AI
Initial Setup
Follow the detailed instructions in docs/GETTINGS_STARTED.md to configure your environment. Key steps include:
# Clone the repository
git clone https://github.com/neural-maze/realtime-phone-agents-course.git
cd realtime-phone-agents-course
# Create virtual environment and install dependencies
python -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txt
# Configure environment variables
cp .env.example .env
# Edit .env with your API keys and credentials
Critical Environment Variables
Your .env file must include:
# Twilio Configuration
TWILIO_ACCOUNT_SID=your_account_sid
TWILIO_AUTH_TOKEN=your_auth_token
TWILIO_PHONE_NUMBER=+1234567890
# RunPod API Key for GPU deployment
RUNPOD_API_KEY=your_runpod_key
# Qdrant Cloud for vector storage
QDRANT_URL=https://your-cluster.qdrant.io
QDRANT_API_KEY=your_qdrant_key
# Opik for tracing
OPIK_API_KEY=your_opik_key
OPIK_WORKSPACE=your_workspace
# STT/TTS Model Endpoints (populated after RunPod deployment)
FASTER_WHISPER_URL=https://your-whisper-pod.proxy.runpod.net
ORPHEUS_URL=https://your-orpheus-pod.proxy.runpod.net
# Active Avatar
AVATAR_NAME=dan # or jess, leah, leo, mia, tara, zac, zoe
Quick Start: Gradio Demo
Test your setup locally before touching phone lines:
make start-gradio-application
Note: If you encounter
No such file or directory: 'ffprobe', install ffmpeg:brew install ffmpeg(macOS) orapt-get install ffmpeg(Linux).
Production Setup: FastAPI Call Center
# Start the call center application
make start-call-center
# Expose to internet for Twilio webhooks
make start-ngrok-tunnel
# Or manually: ngrok http 8000
Configure your Twilio TwiML App with the ngrok URL: https://your-ngrok-url/voice/telephone/incoming
REAL Code Examples: Inside the Repository
Let's examine actual implementation patterns from the course repository.
Example 1: Deploying Faster Whisper on RunPod
The course provides a purpose-built Dockerfile for self-hosted speech recognition with GPU acceleration:
# Dockerfile.faster_whisper - Production STT deployment
FROM speaches-ai/speaches:latest-cuda
# Pre-download the large-v3 model for faster cold starts
# This eliminates the 30+ second download delay on first request
RUN python -c "from faster_whisper import WhisperModel; \
model = WhisperModel('large-v3', device='cuda', compute_type='float16')"
# Expose the standard speaches port
EXPOSE 8000
# Health check for RunPod orchestration
HEALTHCHECK --interval=30s --timeout=10s --start-period=60s --retries=3 \
CMD curl -f http://localhost:8000/health || exit 1
Why this matters: Pre-downloading the model transforms user experience. Without this, your first caller after a pod restart waits 30+ seconds in silence. The speaches-ai/speaches base image provides a production-ready OpenAI-compatible API wrapper around faster-whisper, so integration is drop-in.
Deploy with a single command:
make create-faster-whisper-pod
The Makefile automates RunPod API calls, pod creation, and endpoint extraction. Once ready, copy the printed URL to your .env file.
Example 2: Multi-Avatar System with Versioned Prompts
The avatar architecture enables sophisticated personality engineering:
# src/realtime_phone_agents/avatars/base.py
from dataclasses import dataclass
from typing import Optional
import yaml
@dataclass
class Avatar:
"""Base class for all voice agent personalities.
Each avatar combines system prompt engineering with
voice characteristics for consistent persona delivery.
"""
name: str
system_prompt: str
voice_id: str
version: str = "1.0.0" # Enables A/B testing and rollback
@classmethod
def from_yaml(cls, path: str) -> "Avatar":
"""Load avatar configuration from version-controlled YAML.
This separates personality definition from implementation,
enabling non-engineers to tune conversational style.
"""
with open(path) as f:
config = yaml.safe_load(f)
return cls(**config)
def generate_prompt(self, context: dict) -> str:
"""Dynamic prompt assembly with conversation context.
Injects retrieved property data, conversation history,
and avatar-specific behavioral constraints.
"""
base = self.system_prompt
# Add retrieved context from Superlinked search
if context.get("properties"):
base += f"\n\nAvailable properties: {context['properties']}"
# Add conversation history for coherence
if context.get("history"):
base += f"\n\nConversation so far: {context['history']}"
return base
The engineering insight: Prompt versioning isn't academic—it's operational necessity. When your "Dan" avatar starts being overly aggressive after a prompt update, you need instant rollback. The YAML-based definitions let product teams iterate on personality without code changes.
Example 3: Full Pipeline Tracing with Opik
Production voice agents fail silently. Opik integration exposes every latency bottleneck:
# src/realtime_phone_agents/agent/fastrtc_agent.py
import opik
from fastrtc import ReplyOnPause
class RealtimeAgent:
"""Production voice agent with full observability.
Every pipeline stage is traced: audio→text→reasoning→
tool calls→text→speech. Identifies exactly where latency
accumulates and which components fail.
"""
def __init__(self, avatar: Avatar, stt_client, tts_client, search_tool):
self.avatar = avatar
self.stt = stt_client
self.tts = tts_client
self.search = search_tool
self.opik_trace = None
@opik.track(name="stt.transcribe")
def transcribe_audio(self, audio_chunk: bytes) -> str:
"""Convert audio to text with latency tracking.
opik.track automatically records:
- Execution duration
- Input/output sizes
- Error rates
- Model version used
"""
return self.stt.transcribe(audio_chunk)
@opik.track(name="llm.reason")
def generate_response(self, transcript: str, context: dict) -> str:
"""Generate conversational response with tool access.
If the user asks about properties, this triggers
Superlinked search and incorporates results.
"""
prompt = self.avatar.generate_prompt(context)
# Tool calling logic with search integration
if self._is_property_query(transcript):
properties = self.search.find_properties(transcript)
context["properties"] = properties
return self.llm.complete(prompt, transcript)
@opik.track(name="tts.synthesize")
def synthesize_speech(self, text: str) -> bytes:
"""Convert response text to natural speech.
Orpheus 3B adds emotional markers [sigh], [laugh]
for lifelike delivery. Kokoro provides speed.
"""
return self.tts.speak(text, voice_id=self.avatar.voice_id)
@ReplyOnPause # FastRTC: respond when user pauses
async def handle_call(self, audio_stream):
"""Main call handler orchestrating the full pipeline.
FastRTC manages the complex WebRTC/Twilio media
streaming so you focus on conversation logic.
"""
async for audio_chunk in audio_stream:
transcript = self.transcribe_audio(audio_chunk)
response_text = self.generate_response(transcript, {})
audio_response = self.synthesize_speech(response_text)
yield audio_response
Critical observation: The @opik.track decorators reveal that TTS often dominates latency budgets. With Orpheus 3B on RunPod, you might see 800ms generation for a 20-word response. This data drives infrastructure decisions—maybe Kokoro for simple confirmations, Orpheus only for emotional moments.
Example 4: Superlinked Multi-Attribute Property Search
The "missing layer" in modern RAG—unified structured and unstructured search:
# src/realtime_phone_agents/infrastructure/superlinked/index.py
from superlinked import schema, Index, TextSimilaritySpace, NumberSpace, CategoricalSimilaritySpace
class PropertySchema(schema.Schema):
"""Defines searchable attributes with appropriate vector spaces.
Unlike simple embedding approaches, Superlinked lets you
combine semantic similarity with exact numerical filtering
and categorical matching in ONE query.
"""
description: schema.String # TextSimilaritySpace for semantic search
price: schema.Float # NumberSpace for range queries
neighborhood: schema.String # CategoricalSimilaritySpace for exact/soft match
bedrooms: schema.Integer # NumberSpace
has_parking: schema.Bool # CategoricalSimilaritySpace
# Create unified index combining all spaces
property_index = Index(
PropertySchema,
spaces=[
TextSimilaritySpace(PropertySchema.description, model="sentence-transformers/all-MiniLM-L6-v2"),
NumberSpace(PropertySchema.price, min_value=0, max_value=5000000),
CategoricalSimilaritySpace(PropertySchema.neighborhood, categories=[
"Barrio de Salamanca", "Chamberí", "Malasaña", "Retiro"
]),
NumberSpace(PropertySchema.bedrooms, min_value=0, max_value=10),
CategoricalSimilaritySpace(PropertySchema.has_parking, categories=[True, False])
]
)
# Query with DYNAMIC weights adjusted per conversation
query = property_index.query(
description="spacious sunny apartment",
price=(0, 900000), # Hard constraint
neighborhood="Barrio de Salamanca", # Soft match with similarity
bedrooms=(2, 4)
).with_weights(
description=0.3, # "sunny" matters but isn't critical
price=0.4, # Budget is important
neighborhood=0.2, # Preferred but flexible
bedrooms=0.1 # Nice to match
)
Why developers miss this: Most vector databases force you to choose—semantic search OR metadata filtering. Superlinked's unified space means your agent naturally handles "something like that place in Salamanca but cheaper" without query decomposition hacks.
Advanced Usage & Best Practices
Latency Budgeting for Natural Conversation
Target <400ms total pipeline latency for perceived "instant" response. Break it down:
- STT: 50-150ms (Groq fastest, Faster Whisper best quality)
- LLM: 100-300ms (depends on model and output length)
- TTS: 100-400ms (Kokoro fast, Orpheus expressive but slower)
- Network: 50-100ms
Pro tip: Stream TTS chunks before LLM completion finishes. Start speaking the first sentence while generating the rest.
Avatar A/B Testing Framework
Version prompts in Git, deploy multiple avatars simultaneously, route 50/50 traffic, measure:
- Conversation completion rate
- User satisfaction (post-call survey)
- Average call duration (shorter often = more efficient)
- Conversion rate (booking, purchase, etc.)
Cost Optimization Strategy
| Component | Budget Option | Premium Option | When to Switch |
|---|---|---|---|
| STT | Moonshine local | Faster Whisper RunPod | High accuracy needed for names/addresses |
| TTS | Kokoro local | Orpheus RunPod | Emotional connection critical (sales, support) |
| LLM | Local 7B | GPT-4o/Claude | Complex reasoning, tool chaining |
| Compute | RunPod spot | RunPod on-demand | 24/7 production vs. development |
Failure Mode Handling
Always implement graceful degradation:
- STT fails → "I'm having trouble hearing you, could you repeat that?"
- Search returns nothing → "I don't have anything matching exactly, but what about..." (broaden query)
- LLM hallucinates prices → Ground ALL numerical claims in retrieved data, never generate
Comparison With Alternatives
| Feature | This Course (FastRTC Stack) | OpenAI Realtime API | Vapi/Retell (Managed) | DIY WebRTC |
|---|---|---|---|---|
| Latency | <400ms optimized | <300ms (theirs) | <500ms | 1-3s typical |
| Model Choice | Any STT/LLM/TTS | Locked to OpenAI | Limited selection | Any |
| Phone Integration | Native Twilio | Requires additional | Built-in | Complex |
| Vector Search | Superlinked (multi-attribute) | Basic RAG | Basic RAG | Build yourself |
| Observability | Full Opik tracing | Limited | Dashboard | Build yourself |
| Cost at Scale | Predictable (infrastructure) | $/minute adds up | $/minute + platform | Engineering time |
| Customization | Complete | Prompt-level only | Workflow-level | Complete |
| Learning Value | Maximum | Minimal | Minimal | High but scattered |
When to choose this course's approach: You need full control, want to understand every layer, plan to customize heavily, or operate at scale where per-minute SaaS pricing becomes prohibitive.
When to choose managed: Speed to first prototype matters more than long-term flexibility, or your team lacks DevOps↗ Bright Coding Blog capacity.
FAQ: What Developers Actually Ask
How much does this cost to run in production?
RunPod GPU pods start around $0.20/hour for RTX A4000 instances. For a 24/7 operation with 2-3 pods (redundancy), expect $150-300/month in compute plus Twilio call costs ($0.0085/minute inbound). Compare to managed services at $0.05-0.15/minute—at 1000 minutes/day, self-hosted saves $3,000+/month.
Can I use this for languages other than English?
Yes. Faster Whisper supports 99 languages. Kokoro covers multiple languages with quality varying. Orpheus 3B is primarily English but expanding. The architecture is language-agnostic—swap models as needed.
Do I need GPU infrastructure for development?
No. Start with Groq for STT (fast, cheap API) and Together AI for TTS. Add self-hosted GPU only when you're ready for production optimization or need specific voice characteristics.
How do I prevent the AI from hallucinating property details?
Never let the LLM generate factual claims. The course implements strict grounding: all property data comes from Superlinked search results injected into context. The LLM's job is conversational framing, not fact generation.
What's the minimum viable team to maintain this?
One senior ML engineer with DevOps comfort can handle the full stack↗ Bright Coding Blog. The course is designed to be maintainable by individuals, though production teams benefit from separating voice engineering, backend, and ML infrastructure roles.
How does this compare to fine-tuning a model for my use case?
Fine-tuning is usually unnecessary and often harmful for voice agents. The course uses prompt engineering and RAG because: (1) you need real-time data access anyway, (2) fine-tuned models lag behind frontier models, (3) iteration speed matters more than marginal accuracy gains.
Can I integrate with my existing CRM or booking system?
Absolutely. The tool-calling architecture in property_search.py is a template. Replace Superlinked queries with API calls to Salesforce, HubSpot, Calendly, or custom systems. FastRTC doesn't care what your tools do.
Conclusion: Stop Building Demos, Start Building Systems
The neural-maze/realtime-phone-agents-course isn't just another tutorial—it's a complete production blueprint for voice AI that actually works. From sub-second FastRTC streaming to multi-attribute Superlinked search, from versioned avatar personalities to full Opik observability, every component solves a real problem that breaks demos in production.
What struck me most? The intentional complexity. This course doesn't hide the hard parts. It confronts them: latency budgeting, model selection tradeoffs, prompt versioning, graceful degradation. That's what separates engineers who ship from those who post screenshots.
If you've been waiting for the moment to build voice AI that handles real phone calls with real customers—this is it. The code is open-source, the community is active, and the architecture is battle-tested.
Your next step: Clone the repository, work through Lesson 0's architecture overview, and get your first Gradio demo running locally. By Lesson 4, you'll have deployed a multi-avatar call center to RunPod with live Twilio integration.
👉 Star the repository and start building: git clone https://github.com/neural-maze/realtime-phone-agents-course.git
The future of customer interaction isn't a chatbot on a website. It's an intelligent voice that answers at 2 AM, remembers every preference, and never has a bad day. Build it now—before your competitors do.