Sub-second lip-synced video generated live from streaming audio — not a queued render job.
Plug into whatever STT, LLM, and TTS you already use — Deepgram, OpenAI, Cartesia, or anything else. We only handle the avatar video.
One package (koi-avatar) drops straight into a Pipecat or Pipecat + LiveKit pipeline you already run — no separate hosting to manage.
Every avatar speaks and listens in any of 39 verified-supported languages, picked per session.
Every account gets a real API key immediately — no payment required to try it.