← All Categories
Category • 12 resources

Voice and Multimodal Agents

Pipecatpipecat-ai/pipecat

Production-grade voice AI framework with sub-250ms latency, WebRTC support, multimodal (voice+vision+text), real-time streaming, and 70+ language support.

Python#Streaming
GitHub ↗
qwen-audio-agentQwenAudio/qwen-audio-agent

Full-duplex voice runtime that drives coding agents (Claude Code, Codex, OpenCode, Kimi Code, and more) over ACP, keeping conversations going while background tasks run, with barge-in and a local wake word.

Desktop#Voice
GitHub ↗