Providers and models
Vida supports speech-to-speech through OpenAI and Google Gemini. Available providers, models, and compatible voices can vary by account as new capabilities become available. Upcoming OpenAI support includesgpt-live-1.
When to use it
Speech-to-speech is a good fit when natural turn-taking, low latency, and expressive delivery matter more than selecting separate speech-recognition and voice providers. Standard mode is a better fit when you need the broader set of speech-recognition controls, voice providers, or tuning options.Configure and test
- Open the Agent editor and select Voice.
- Choose speech-to-speech mode and an available provider.
- Select a compatible voice.
- Test realistic calls, including interruptions, names, numbers, and background noise.
- Publish only after the staged Agent behaves as expected.
Configuring speech mode through the API? See Language, voice, and speech.