AI Dictionary › AI Fundamentals

Speech-to-Text (and Text-to-Speech)

Speech-to-Text (e Text-to-Speech)

Speech-to-text (STT) converts spoken audio into written text; text-to-speech (TTS) does the reverse, generating synthetic voice from text. Together they form AI's voice interface.

Definition

Modern systems, OpenAI's Whisper for transcription, ElevenLabs and major providers' neural voices for synthesis, have reached near-human quality: they transcribe meetings with correct punctuation, recognize dozens of languages and produce natural voices with intonation and emotion.

Professional uses are immediate: automatic meeting minutes, subtitling, document dictation, voice assistants, multilingual dubbing and accessibility. Combined with an LLM, they enable entirely spoken conversations with AI, you talk, the model understands, it replies out loud.

Related terms

More in AI Fundamentals

Put it into practice

From our network

HSE Genius: AI for Safety Data Sheets

Extract SDS data, H phrases and ECHA compliance checks in seconds, powered by AI.

Visit hsegenius.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.