Local speech, your service
Plug the Proto text-to-speech and speech recognition engine into your app with prepaid credits.
Two endpoints
Both authenticate with a bearer token scoped to your workspace, and both draw on the same voice-credit balance.
Text-to-speech
Synthesis built from native recordings, not an English voice bent towards a local accent.
- Native speaker recordings, not converted English models
- Voice, gender and accent selectable per language, speed 0.25–4.0
- Up to 5,000 characters a request, returned base64-encoded as MP3 or WAV
Speech recognition
Transcription built for regional accents and sentences that move between a local language and English mid-clause.
- Tuned for regional accents and mixed-language phrasing
- Transcribes up to 15MB of MP3 or WAV audio a request
- Transcript returned with an optional English translation
Pay as you go
Access the Voice API through our Starter plan. The rate decreases as usage grows.
Its own balance
Voice requests draw on prepaid voice credits, not on your platform credit volume. The two balances are billed and tracked separately.
1 credit a request
Every successful text-to-speech or speech-recognition call costs 1 credit, at any tier. A failed request is not charged.
50 credits free
Every new workspace starts with 50 voice credits, no plan or purchase required to try it.
No plan required
Voice credits can be bought on any plan, from Starter plan upward — there is no tier to reach first.
What comes with every credit
Reach the last mile
Proto works at the frontier of low-resource languages.
- Speakers
- 15M
- ASR accuracy
- 93.73%
- Word error rate
- 6.27%
- ASR fallback
- 4.93%
- ASR speed
- 1.75x
- Voice quality (MOS)
- 4.00
- Speakers
- 45M
- ASR accuracy
- 90.66%
- Word error rate
- 9.34%
- ASR fallback
- 4.21%
- ASR speed
- 1.74x
- Voice quality (MOS)
- 4.00
- Speakers
- 22M
- ASR accuracy
- 87.74%
- Word error rate
- 12.26%
- ASR fallback
- 6.36%
- ASR speed
- 1.76x
- Voice quality (MOS)
- 3.80
- Speakers
- 1.4M
- ASR accuracy
- 70.22%
- Word error rate
- 29.78%
- ASR fallback
- 8.34%
- ASR speed
- 1.77x
- Voice quality (MOS)
- 3.60
- Speakers
- 9M
- ASR accuracy
- Not yet published
- Word error rate
- Not yet published
- ASR fallback
- Not yet published
- ASR speed
- Not yet published
- Voice quality (MOS)
- Not yet published
- Speakers
- 100M
- ASR accuracy
- 51.50%
- Word error rate
- 48.50%
- ASR fallback
- Not yet published
- ASR speed
- Not yet published
- Voice quality (MOS)
- Not yet published
- Speakers
- 23M
- ASR accuracy
- Not yet published
- Word error rate
- Not yet published
- ASR fallback
- Not yet published
- ASR speed
- Not yet published
- Voice quality (MOS)
- Not yet published
- Speakers
- 100M
- ASR accuracy
- 79.40%
- Word error rate
- 20.60%
- ASR fallback
- Not yet published
- ASR speed
- Not yet published
- Voice quality (MOS)
- 4.15
Start with 50 free credits
No plan required. Get API keys and make your first request in minutes.