Learn voice AI
Understand the concepts before comparing products.
What is text-to-speech?
How modern TTS turns text into controllable speech and which quality dimensions matter.
Read guide →What is voice cloning?
How a reference voice becomes a reusable synthetic voice profile.
Read guide →TTS vs cloning vs voice changer
Three related tools that solve different production problems.
Read guide →Voice AI stack explained
STT, reasoning, tools, TTS, turn-taking, transport, and monitoring.
Read guide →What is a voice agent?
A practical explanation of conversational agents and what makes them production-ready.
Read guide →Choose an AI voice model
A buyer’s framework for balancing quality, latency, price, rights, and workflow.
Read guide →Troubleshoot voice AI
Diagnose API errors, credit usage, browser problems, request limits, and recurring pronunciation issues.
Troubleshooting guide →A map of the voice-AI system
Text-to-speech is the speaking layer. Voice cloning adds a reusable voice identity. Voice changing starts from a recorded performance rather than a script. Speech-to-text is the listening layer, while a voice agent connects listening and speaking to reasoning, tools, turn-taking, monitoring, and escalation.
Understanding those boundaries prevents category mistakes—for example, choosing a narration model for a realtime agent or expecting a voice changer to replace a cloning workflow. Learn the mechanism first, then follow the contextual links into the buying guides when you are ready to compare platforms.