Realtime Voice Agent Design
FortgeschrittendevelopmentMindestens 32K Kontext
Designs production voice agents for low-latency, two-way audio conversations. Covers WebRTC and WebSocket transport, voice activity detection, interruption and barge-in behavior, turn state, streaming transcripts, tool calls, specialist handoffs, consent, failure recovery, latency budgets, and evaluation of conversational quality. Complements dialogue writing with the runtime architecture required to operate a reliable realtime agent.
Anwendungsfälle
- Designing a low-latency customer support voice agent
- Adding interruption and barge-in handling to an assistant
- Coordinating voice-agent tool calls and specialist handoffs
- Defining consent and escalation behavior for recorded conversations
- Evaluating latency, task completion, and conversational recovery
Beispiel-Prompt
Design a realtime voice agent for appointment scheduling. The agent must handle natural interruptions, call tools to check and reserve time slots, confirm before committing a booking, and hand off to a human when confidence is low. Provide: - The client/server architecture and choice of WebRTC or WebSocket transport - The turn state machine, including VAD, barge-in, cancellation, and reconnect behavior - Tool-call and specialist-handoff flows - A latency budget from microphone input to first audible response - Consent, transcript retention, and escalation guardrails - An evaluation plan for task success, latency, interruptions, and recovery
Empfohlene Modelle
Kompatible Werkzeuge
claude-codecursorgemini-clikiroany
Modalitäten
Eingabe: text, audio
→Ausgabe: text, code
Ähnliche Skills
Autor
OpenModels Community