Bland releases benchmark-topping new voice model
Bland released Speech v3, a text-to-speech model that ranked second in the Audio Realism Bench test, just behind human speech. The model preserves natural speech patterns like breathing and hesitations, available via API or interactive studio at a unified rate of $0.015 per 1,000 characters. The service also supports voice cloning from short audio samples and speech generation from text descriptions.
How does Bland Speech v3 differ from conventional TTS?
Bland Speech v3 is designed for natural conversational speech and preserves elements like breathing, hesitations, and natural intonation. Typical TTS systems sound artificial and overly polished, whereas Bland Speech maintains these human characteristics.
What are the voice cloning options?
The service offers two cloning types: instant cloning from 10 seconds of audio and professional cloning from 30 minutes of verified audio. Both require confirmation that you have rights to use the audio.
How does Bland Speech v3 integrate with applications?
Integration is done via a REST API endpoint POST /v1/speak that accepts text and voice ID and returns audio. Output can be streamed over HTTP chunks or WebSocket in various formats including PCM16 WAV.
- Gradium TTS: no longer choose between latency and accuracy — gradium.ai 74 % match
- Experiment with Gemma 4 as an offline translator — github.com 68 % match
- Naomi Bashkansky leaves OpenAI to develop thought-to-text AI at Conduit — naomibashkansky.com 68 % match
- Bland
- Bland Speech v3
- Audio Realism Bench
- intelligence.ai