Gemini Audio
发布时间:2026-09-18 | 浏览:1
Our most advanced audio models push new frontiers with intuitive inputs, natural expressiveness, and the ability to take action
Best for near real-time voice interfaces with conversational capabilities optimized for both cost-effective, high-volume scale and near real-time reasoning.
Best for speech transcription. Transcribes pre-recorded audio across 85+ languages with high alphanumeric accuracy and word-level timestamps for up to three speakers.
3.5 Live Translate
Best for near real-time speech-to-speech translation. Overcomes language barriers across 70+ languages while maintaining the speaker’s natural tone and rhythm.
Best for directing intonation and inflection. Intuitive audio tags give you granular command over style, pace, and tone with unprecedented precision.
Natural and powerful audio models. Helping people communicate, developers build, and enterprises manage business.
Try Gemini Audio
Our audio models generate natural vocals at speed and scale for different developer workflows.
Engage in almost real-time conversations. Control with precision. Understand every nuance.
Fluid and natural live dialogue and translation capabilities, for powerful voice-first applications.
AI transcription
Context-aware transcription that identifies who’s talking and keeps pace with natural conversation.
Speech generation
Craft anything from short snippets to long-form narratives, with granular control over style, pace, delivery and performance.
Fluid and natural live dialogue and translation capabilities, for powerful voice-first applications.
AI transcription
Context-aware transcription that identifies who’s talking and keeps pace with natural conversation.
Speech generation
Craft anything from short snippets to long-form narratives, with granular control over style, pace, delivery and performance.
Our audio models generate natural vocals at speed and scale for different developer workflows.
3.8 Live Extended Thinking
Best for complex reasoning and high-complexity tasks. Thinks and narrates its progress in near real-time to solve difficult, multi-step challenges.
Best for near real-time voice interfaces with conversational capabilities optimized for both cost-effective, high-volume scale and near real-time reasoning.
Best for speech transcription. Transcribes pre-recorded audio across 85+ languages with high alphanumeric accuracy and timestamps for up to three speakers.
3.5 Live Translate
Best for near real-time speech-to-speech translation. Overcomes language barriers across 70+ languages while maintaining the speaker’s natural tone and rhythm.
Best for directing intonation and inflection. Intuitive audio tags give you granular command over style, pace, and tone with unprecedented precision.
Explore what you can do with Gemini Audio
Complex task management
Orchestrates multiple agents to solve background tasks, while holding a natural-sounding conversation.
Visual understanding
Understands visual input to grasp deeper context during conversations. This allows for richer, more natural, and more relevant responses.
Seamless transcription
Gemini 3.5 Transcribe handles live language switches and seamless streaming transcription.
Near real-time meeting translation
Translates multiple languages in a single session, while preserving each speaker’s original intonation, pacing and pitch.
Expressive speech generation
Best for directing intonation and inflection. Intuitive audio tags give you granular command over style, pace, and tone with unprecedented precision.
Building with responsibility at the core
We’ve proactively assessed potential risks during every stage of the development process for these native audio features, using what we’ve learned to inform our mitigation strategies. We validate these measures through rigorous internal and external safety evaluations, including comprehensive red teaming for responsible deployment.
All audio outputs from our models are marked with SynthID, our advanced watermarking technology, allowing you to detect whether speech has been created or edited using Google AI.
Try Gemini Audio
Google AI Studio
The fastest path from prompt to production
AI-powered video creation for work
Get started with cutting-edge AI models
Gemini Live API
Low-latency, real-time voice and video interactions with Gemini
Gemini Enterprise Agent Platform
Build, scale, and govern agents
Gemini Enterprise for Customer Experience
Deploy specialized agents for product discovery, shopping, and customer service
Google Translate
Understand your world and communicate across languages
Google Workspace
Collaborate, create, and communicate all in one place