AI Glossary · Generative AI
Speech Recognition (STT/ASR)
Automatic Speech Recognition — converting spoken audio to text. Foundation for voice AI, transcription and analytics.
Definition
What is Speech Recognition (STT/ASR)?
Speech Recognition (STT/ASR) is automatic Speech Recognition — converting spoken audio to text. Foundation for voice AI, transcription and analytics.
- Category
- Generative AI
- Glossary set
- 10 related terms
- Audience
- Enterprise AI leaders
Why does Speech Recognition (STT/ASR) matter in enterprise AI?
Speech Recognition (STT/ASR) matters in enterprise AI programs because it helps business and technology leaders align vocabulary, scope, ownership, and measurable outcomes.
Related terms in Generative AI
- Diffusion Model
- A generative approach (used in Stable Diffusion, Imagen, Sora) that produces images or video by iteratively denoising random noise.
- Few-Shot Learning
- Guiding a model with a handful of examples in the prompt to shape output format or reasoning — no retraining required.
- Generative AI
- AI that creates new artifacts — text, code, images, audio, video — as opposed to only classifying or predicting.
- Image Generation
- Synthesizing images from text prompts (Midjourney, DALL·E, Imagen, Flux). Enterprise use: marketing, product design, synthetic training data.
- Latent Space
- The compressed internal representation where generative models operate; navigating it enables controllable generation and style transfer.
- One-Shot Learning
- Guiding a model with a single example — a common pattern for structured extraction and consistent formatting.