Voice Speech Services skills for AI agents
8 practitioner-grade voice speech services skills, each a focused Markdown document your agent loads into context on demand. Search them from Claude Desktop, Cursor or any MCP client, or pull one with the CLI.
All 8 skills
- Amazon Polly
Amazon Polly: AWS text-to-speech, neural/standard voices, SSML, lexicons, speech marks, streaming
366 lines - AssemblyAI
AssemblyAI: speech-to-text, real-time transcription, speaker diarization, content moderation, summarization, sentiment analysis
326 lines - Cartesia
Integrate Cartesia's ultra-low-latency voice API for real-time text-to-speech and voice cloning
215 lines - Deepgram
Deepgram: speech-to-text, real-time transcription, pre-recorded audio, diarization, sentiment analysis, WebSocket streaming
304 lines - ElevenLabs
ElevenLabs: AI voice synthesis, text-to-speech, voice cloning, streaming audio, voice design, multilingual, WebSocket streaming
236 lines - Google Cloud Text to Speech
Google Cloud Text-to-Speech: WaveNet/Neural2 voices, SSML, audio profiles, streaming, multilingual
312 lines - OpenAI TTS
OpenAI TTS: text-to-speech API, voice selection (alloy/echo/fable/onyx/nova/shimmer), streaming, HD voices, audio formats
277 lines - Playht
Integrate PlayHT's voice API for text-to-speech, voice cloning, and real-time audio streaming
224 lines