Gemini Interactions API
发布时间:2026-08-25 | 浏览:1
Español – América Latina
Português – Brasil
The Gemini Interactions API allows developers to build generative AI applications using Gemini models. Gemini is our most capable model, built from the ground up to be multimodal. It can generalize and seamlessly understand, operate across, and combine different types of information including language, images, audio, video, and code. You can use the Gemini API for use cases like reasoning across text and images, content generation, dialogue agents, summarization and classification systems, and more.
Creating an interaction
Creates a new interaction.
Path / Query parameters
Path / Query Parameters
Which version of the API to use.
The request body contains data with the following structure:
The name of the `Model` used for generating the interaction. Required if `agent` is not provided.
The model that will complete your prompt.\n\nSee [models](https://ai.google.dev/gemini-api/docs/models) for additional details.
Possible values
gemini-2.5-flash Our first hybrid reasoning model which supports a 1M token context window and has thinking budgets.
Our first hybrid reasoning model which supports a 1M token context window and has thinking budgets.
gemini-2.5-pro Our state-of-the-art multipurpose model, which excels at coding and complex reasoning tasks.
Our state-of-the-art multipurpose model, which excels at coding and complex reasoning tasks.
gemma-4-26b-a4b-it Gemma 4 26B A4B IT
Gemma 4 26B A4B IT
gemma-4-31b-it Gemma 4 31B IT
gemini-flash-latest Latest release of Gemini Flash
Latest release of Gemini Flash
gemini-flash-lite-latest Latest release of Gemini Flash-Lite
Latest release of Gemini Flash-Lite
gemini-pro-latest Latest release of Gemini Pro
Latest release of Gemini Pro
gemini-2.5-flash-lite Our smallest and most cost effective model, built for at scale usage.
Our smallest and most cost effective model, built for at scale usage.
gemini-2.5-flash-image Our native image generation model, optimized for speed, flexibility, and contextual understanding. Text input and output is priced the same as 2.5 Flash.
Our native image generation model, optimized for speed, flexibility, and contextual understanding. Text input and output is priced the same as 2.5 Flash.
gemini-3-flash-preview Our most intelligent model built for speed, combining frontier intelligence with superior search and grounding.
Our most intelligent model built for speed, combining frontier intelligence with superior search and grounding.
gemini-3.1-pro-preview Our latest SOTA reasoning model with unprecedented depth and nuance, and powerful multimodal understanding and coding capabilities.
Our latest SOTA reasoning model with unprecedented depth and nuance, and powerful multimodal understanding and coding capabilities.
gemini-3.1-pro-preview-customtools Gemini 3.1 Pro Preview optimized for custom tool usage
Gemini 3.1 Pro Preview optimized for custom tool usage
gemini-3.1-flash-lite Our most cost-efficient model, optimized for high-volume agentic tasks, translation, and simple data processing.
Our most cost-efficient model, optimized for high-volume agentic tasks, translation, and simple data processing.
gemini-3-pro-image Gemini 3 Pro Image
Gemini 3 Pro Image
nano-banana-pro-preview Gemini 3 Pro Image Preview
Gemini 3 Pro Image Preview
gemini-3.1-flash-image Gemini 3.1 Flash Image.
Gemini 3.1 Flash Image.
gemini-3.5-flash Our most intelligent model for sustained frontier performance in agentic and coding tasks.
Our most intelligent model for sustained frontier performance in agentic and coding tasks.
gemini-3.6-flash Our most intelligent model for sustained frontier performance in agentic and coding tasks.
Our most intelligent model for sustained frontier performance in agentic and coding tasks.
gemini-3.7-flash Our most intelligent model for sustained frontier performance in agentic and coding tasks.
Our most intelligent model for sustained frontier performance in agentic and coding tasks.
lyria-3-clip-preview Our low-latency, music generation model optimized for high-fidelity audio clips and precise rhythmic control.
Our low-latency, music generation model optimized for high-fidelity audio clips and precise rhythmic control.
lyria-3-pro-preview Our advanced, full-song generative model with deep compositional understanding, optimized for precise structural control and complex transitions across diverse musical styles.
Our advanced, full-song generative model with deep compositional understanding, optimized for precise structural control and complex transitions across diverse musical styles.
gemini-robotics-er-1.6-preview Gemini Robotics-ER 1.6 Preview
Gemini Robotics-ER 1.6 Preview
gemini-robotics-er-2-preview Gemini Robotics Embodied Reasoning 2 Preview
Gemini Robotics Embodied Reasoning 2 Preview
The name of the `Agent` used for generating the interaction. Required if `model` is not provided.
The agent to interact with.
Possible values
deep-research-pro-preview-12-2025 Gemini Deep Research Agent
Gemini Deep Research Agent
deep-research-preview-04-2026 Gemini Deep Research Agent
Gemini Deep Research Agent
deep-research-max-preview-04-2026 Gemini Deep Research Max Agent
Gemini Deep Research Max Agent
antigravity-preview-05-2026 Use the Antigravity managed agent to perform multi-step tasks that require reasoning, file operations, and tool use.
Use the Antigravity managed agent to perform multi-step tasks that require reasoning, file operations, and tool use.
The inputs for the interaction (common to both Model and Agent).
System instruction for the interaction.
A list of tool declarations the model may call during interaction.
Enforces that the generated response is a JSON object that complies with the JSON schema specified in this field.
Input only. Whether the interaction will be streamed.
Input only. Whether to store the response and request for later retrieval.
Input only. Whether to run the model interaction in the background.
Model Configuration Configuration parameters for the model interaction. Alternative to `agent_config`. Only applicable when `model` is set.
Configuration parameters for model interactions.
The maximum number of tokens to include in the response.
Seed used in decoding for reproducibility.
Optional. Speech and multi-speaker configuration.
Configuration for multi-speaker and speech generation.
Individual speaker configurations.
The configuration for speech interaction.
The language of the speech.
The speaker's name, it should match the speaker name given in the prompt.
The voice of the speaker.
A list of character sequences that will stop output interaction.
The level of thought tokens that the model should generate.
Possible values
minimal Little to no thinking.
Little to no thinking.
low Low thinking level.
Low thinking level.
medium Medium thinking level.
Medium thinking level.
high High thinking level.
High thinking level.
Whether to include thought summaries in the response.
Possible values
auto Auto thinking summaries.
Auto thinking summaries.
none No thinking summaries.
No thinking summaries.
The tool choice configuration.
Possible values:
auto Auto tool choice.
Auto tool choice.
any Any tool choice.
Any tool choice.
none No tool choice.
No tool choice.
validated Validated tool choice.
Validated tool choice.
Optional. Configuration for speech recognition (transcription). If present, ASR is enabled.
Configuration for speech recognition (transcription).
Optional. A list of custom vocabulary phrases to bias the speech recognition model toward recognizing specific terms.
Optional. Configures speaker diarization. Supported values: "speaker".
Optional. BCP-47 language codes providing hints about the languages present in the audio. If omitted or empty, defaults to automatic language detection.
Optional. The granularity of timestamps to include in the transcription output. Supported values: "word". If empty, no timestamps are generated.
Configuration for video generation.
Configuration options for video generation.
Optional task mode for video generation. If not specified, the model automatically determines the appropriate mode based on the provided text prompt and input media.
Possible values:
text_to_video Generates video solely from a text prompt.
Generates video solely from a text prompt.
image_to_video Generates video from one or two source images. The first image defines the starting frame, and the optional second image defines the ending frame.
Generates video from one or two source images. The first image defines the starting frame, and the optional second image defines the ending frame.
reference_to_video Generates video using reference media (such as images, audio, or video).
Generates video using reference media (such as images, audio, or video).
edit Modifies an existing input video.
Modifies an existing input video.
extend Extends an existing input video.
Extends an existing input video.
Agent Configuration Configuration for the agent. Alternative to `generation_config`. Only applicable when `agent` is set.
Polymorphic discriminator: type
Configuration for the Antigravity agent runtime. Provides server-side control over the agent's execution environment and tool configuration.
Max total tokens for the agent run.
The model to use for agent reasoning.
No description provided.
Always set to "antigravity" .
Configuration for the CodeMender agent.
Parameters for finding vulnerabilities.
Request parameters specific to FIND sessions, used for discovering vulnerabilities in a codebase.
Additional context or custom instructions provided by the user to guide the vulnerability analysis.
The identifier of a specific finding to verify. This is primarily used in VERIFY mode to focus the agent's execution-based validation on a single vulnerability.
The mode of the find session.
Possible values:
scan Fast scan using only the initial classifier.
Fast scan using only the initial classifier.
verify Performs classification followed by detailed investigation.
Performs classification followed by detailed investigation.
A list of source files to provide as context for the scan.
Content of a single file in the codebase.
The UTF-8 encoded text content of the file.
The relative path of the file from the project root.
Parameters for fixing vulnerabilities.
Request parameters specific to FIX sessions, used for generating and validating security patches.
Additional context or custom instructions provided by the user to guide the patch generation process.
The identifier of the specific security finding to be remediated. This ID maps to a previously discovered vulnerability.
A list of source files providing context for the remediation. These files are typically the ones containing the identified vulnerability.
Content of a single file in the codebase.
The UTF-8 encoded text content of the file.
The relative path of the file from the project root.
The name of the model to use for the CodeMender agent. One CodeMender session will only use one model.
Optional session-specific configurations to override default agent behavior.
The configuration of CodeMender sessions.
The maximum number of interaction rounds the agent is allowed to perform before reaching a timeout.
Parameter for grouping multiple interactions that belong to the same CodeMender session.
No description provided.
Always set to "code-mender" .
Configuration for the Deep Research agent.
Enables human-in-the-loop planning for the Deep Research agent. If set to true, the Deep Research agent will provide a research plan in its response. The agent will then proceed only if the user confirms the plan in the next turn.
Enables bigquery tool for the Deep Research agent.
Whether to include thought summaries in the response.
Possible values
auto Auto thinking summaries.
Auto thinking summaries.
none No thinking summaries.
No thinking summaries.
No description provided.
Always set to "deep-research" .
Whether to include visualizations in the response.
Possible values:
off Do not include visualizations.
Do not include visualizations.
auto Automatically include visualizations.
Automatically include visualizations.
Configuration for dynamic agents.
No description provided.
Always set to "dynamic" .
The environment configuration for the interaction. Can be an object specifying remote environment sources or a string referencing an existing environment ID.
The labels with user-defined metadata for the request.
The ID of the previous interaction, if any.
Safety settings for the interaction.
A safety setting that affects the safety-blocking behavior. A SafetySetting consists of a harm category and a threshold for that category.
Optional. The method for blocking content. If not specified, the default behavior is to use the probability score.
Possible values:
severity The harm block method uses both probability and severity scores.
The harm block method uses both probability and severity scores.
probability The harm block method uses the probability score.
The harm block method uses the probability score.
Required. The threshold for blocking content. If the harm probability exceeds this threshold, the content will be blocked.
Possible values:
block_low_and_above Block content with a low harm probability or higher.
Block content with a low harm probability or higher.
block_medium_and_above Block content with a medium harm probability or higher.
Block content with a medium harm probability or higher.
block_only_high Block content with a high harm probability.
Block content with a high harm probability.
block_none Do not block any content, regardless of its harm probability.
Do not block any content, regardless of its harm probability.
off Turn off the safety filter entirely.
Turn off the safety filter entirely.
Required. The type of harm category to be blocked.
Possible values
hate_speech Content that promotes violence or incites hatred against individuals or groups based on certain attributes.
Content that promotes violence or incites hatred against individuals or groups based on certain attributes.
dangerous_content Content that promotes, facilitates, or enables dangerous activities.
Content that promotes, facilitates, or enables dangerous activities.
harassment Abusive, threatening, or content intended to bully, torment, or ridicule.
Abusive, threatening, or content intended to bully, torment, or ridicule.
sexually_explicit Content that contains sexually explicit material.
Content that contains sexually explicit material.
civic_integrity Deprecated: Election filter is not longer supported. The harm category is civic integrity.
Deprecated: Election filter is not longer supported. The harm category is civic integrity.
image_hate Images that contain hate speech.
Images that contain hate speech.
image_dangerous_content Images that contain dangerous content.
Images that contain dangerous content.
image_harassment Images that contain harassment.
Images that contain harassment.
image_sexually_explicit Images that contain sexually explicit content.
Images that contain sexually explicit content.
jailbreak Prompts designed to bypass safety filters.
Prompts designed to bypass safety filters.
The service tier for the interaction.
Possible values
flex Flex service tier.
Flex service tier.
standard Standard service tier.
Standard service tier.
priority Priority service tier.
Priority service tier.
deferred Deferred service tier.
Deferred service tier.
Optional. Webhook configuration for receiving notifications when the interaction completes.
Message for configuring webhook events for a request.
Optional. If set, these webhook URIs will be used for webhook events instead of the registered webhooks.
Optional. The user metadata that will be returned on each event emission to the webhooks.
Returns an Interaction resource.
Example Response
Example Response
Example Response
Function Calling
Example Response
Example Response
Antigravity Agent
Example Response
Reuse Environment
Example Response
Canceling an interaction
Cancels an interaction by id. This only applies to background interactions that are still running.
Path / Query parameters
Path / Query Parameters
Which version of the API to use.
The unique identifier of the interaction to cancel.
Returns an Interaction resource.
Cancel Interaction
Example Response
Retrieving an interaction
Retrieves the full details of a single interaction based on its `Interaction.id`.
Path / Query parameters
Path / Query Parameters
Which version of the API to use.
The unique identifier of the interaction to retrieve.
Optional. If set, resumes the interaction stream from the next chunk after the event marked by the event id. Can only be used if `stream` is true.
If set to true, the generated content will be streamed incrementally.
Defaults to: False
Returns an Interaction resource.
Get Interaction
Example Response
Deleting an interaction
Deletes the interaction by id.
Path / Query parameters
Path / Query Parameters
Which version of the API to use.
The unique identifier of the interaction to delete.
If successful, the response is empty.
The Interaction resource.
The name of the `Agent` used for generating the interaction.
The agent to interact with.
Possible values
deep-research-pro-preview-12-2025 Gemini Deep Research Agent
Gemini Deep Research Agent
deep-research-preview-04-2026 Gemini Deep Research Agent
Gemini Deep Research Agent
deep-research-max-preview-04-2026 Gemini Deep Research Max Agent
Gemini Deep Research Max Agent
antigravity-preview-05-2026 Use the Antigravity managed agent to perform multi-step tasks that require reasoning, file operations, and tool use.
Use the Antigravity managed agent to perform multi-step tasks that require reasoning, file operations, and tool use.
Configuration parameters for the agent interaction.
Polymorphic discriminator: type
Configuration for the Antigravity agent runtime. Provides server-side control over the agent's execution environment and tool configuration.
Max total tokens for the agent run.
The model to use for agent reasoning.
No description provided.
Always set to "antigravity" .
Configuration for the CodeMender agent.
Parameters for finding vulnerabilities.
Request parameters specific to FIND sessions, used for discovering vulnerabilities in a codebase.
Additional context or custom instructions provided by the user to guide the vulnerability analysis.
The identifier of a specific finding to verify. This is primarily used in VERIFY mode to focus the agent's execution-based validation on a single vulnerability.
The mode of the find session.
Possible values:
scan Fast scan using only the initial classifier.
Fast scan using only the initial classifier.
verify Performs classification followed by detailed investigation.
Performs classification followed by detailed investigation.
A list of source files to provide as context for the scan.
Content of a single file in the codebase.
The UTF-8 encoded text content of the file.
The relative path of the file from the project root.
Parameters for fixing vulnerabilities.
Request parameters specific to FIX sessions, used for generating and validating security patches.
Additional context or custom instructions provided by the user to guide the patch generation process.
The identifier of the specific security finding to be remediated. This ID maps to a previously discovered vulnerability.
A list of source files providing context for the remediation. These files are typically the ones containing the identified vulnerability.
Content of a single file in the codebase.
The UTF-8 encoded text content of the file.
The relative path of the file from the project root.
The name of the model to use for the CodeMender agent. One CodeMender session will only use one model.
Optional session-specific configurations to override default agent behavior.
The configuration of CodeMender sessions.
The maximum number of interaction rounds the agent is allowed to perform before reaching a timeout.
Parameter for grouping multiple interactions that belong to the same CodeMender session.
No description provided.
Always set to "code-mender" .
Configuration for the Deep Research agent.
Enables human-in-the-loop planning for the Deep Research agent. If set to true, the Deep Research agent will provide a research plan in its response. The agent will then proceed only if the user confirms the plan in the next turn.
Enables bigquery tool for the Deep Research agent.
Whether to include thought summaries in the response.
Possible values
auto Auto thinking summaries.
Auto thinking summaries.
none No thinking summaries.
No thinking summaries.
No description provided.
Always set to "deep-research" .
Whether to include visualizations in the response.
Possible values:
off Do not include visualizations.
Do not include visualizations.
auto Automatically include visualizations.
Automatically include visualizations.
Configuration for dynamic agents.
No description provided.
Always set to "dynamic" .
Output only. The time at which the response was created in ISO 8601 format (YYYY-MM-DDThh:mm:ssZ).
The environment configuration for the interaction. Can be an object specifying remote environment sources or a string referencing an existing environment ID.
Output only. The environment ID for the interaction. Only populated if environment config is set in the request.
Output only. Diagnostic faults / platform errors recorded on the interaction.
Error message from an interaction.
A URI that identifies the error type.
A human-readable error message.
Required. Output only. A unique identifier for the interaction completion.
The input for the interaction.
The labels with user-defined metadata for the request.
The name of the `Model` used for generating the interaction.
The model that will complete your prompt.\n\nSee [models](https://ai.google.dev/gemini-api/docs/models) for additional details.
Possible values
gemini-2.5-flash Our first hybrid reasoning model which supports a 1M token context window and has thinking budgets.
Our first hybrid reasoning model which supports a 1M token context window and has thinking budgets.
gemini-2.5-pro Our state-of-the-art multipurpose model, which excels at coding and complex reasoning tasks.
Our state-of-the-art multipurpose model, which excels at coding and complex reasoning tasks.
gemma-4-26b-a4b-it Gemma 4 26B A4B IT
Gemma 4 26B A4B IT
gemma-4-31b-it Gemma 4 31B IT
gemini-flash-latest Latest release of Gemini Flash
Latest release of Gemini Flash
gemini-flash-lite-latest Latest release of Gemini Flash-Lite
Latest release of Gemini Flash-Lite
gemini-pro-latest Latest release of Gemini Pro
Latest release of Gemini Pro
gemini-2.5-flash-lite Our smallest and most cost effective model, built for at scale usage.
Our smallest and most cost effective model, built for at scale usage.
gemini-2.5-flash-image Our native image generation model, optimized for speed, flexibility, and contextual understanding. Text input and output is priced the same as 2.5 Flash.
Our native image generation model, optimized for speed, flexibility, and contextual understanding. Text input and output is priced the same as 2.5 Flash.
gemini-3-flash-preview Our most intelligent model built for speed, combining frontier intelligence with superior search and grounding.
Our most intelligent model built for speed, combining frontier intelligence with superior search and grounding.
gemini-3.1-pro-preview Our latest SOTA reasoning model with unprecedented depth and nuance, and powerful multimodal understanding and coding capabilities.
Our latest SOTA reasoning model with unprecedented depth and nuance, and powerful multimodal understanding and coding capabilities.
gemini-3.1-pro-preview-customtools Gemini 3.1 Pro Preview optimized for custom tool usage
Gemini 3.1 Pro Preview optimized for custom tool usage
gemini-3.1-flash-lite Our most cost-efficient model, optimized for high-volume agentic tasks, translation, and simple data processing.
Our most cost-efficient model, optimized for high-volume agentic tasks, translation, and simple data processing.
gemini-3-pro-image Gemini 3 Pro Image
Gemini 3 Pro Image
nano-banana-pro-preview Gemini 3 Pro Image Preview
Gemini 3 Pro Image Preview
gemini-3.1-flash-image Gemini 3.1 Flash Image.
Gemini 3.1 Flash Image.
gemini-3.5-flash Our most intelligent model for sustained frontier performance in agentic and coding tasks.
Our most intelligent model for sustained frontier performance in agentic and coding tasks.
gemini-3.6-flash Our most intelligent model for sustained frontier performance in agentic and coding tasks.
Our most intelligent model for sustained frontier performance in agentic and coding tasks.
gemini-3.7-flash Our most intelligent model for sustained frontier performance in agentic and coding tasks.
Our most intelligent model for sustained frontier performance in agentic and coding tasks.
lyria-3-clip-preview Our low-latency, music generation model optimized for high-fidelity audio clips and precise rhythmic control.
Our low-latency, music generation model optimized for high-fidelity audio clips and precise rhythmic control.
lyria-3-pro-preview Our advanced, full-song generative model with deep compositional understanding, optimized for precise structural control and complex transitions across diverse musical styles.
Our advanced, full-song generative model with deep compositional understanding, optimized for precise structural control and complex transitions across diverse musical styles.
gemini-robotics-er-1.6-preview Gemini Robotics-ER 1.6 Preview
Gemini Robotics-ER 1.6 Preview
gemini-robotics-er-2-preview Gemini Robotics Embodied Reasoning 2 Preview
Gemini Robotics Embodied Reasoning 2 Preview
The last audio generated by the model in response to the current request. Note: this is added by the SDK.
An audio content block.
The number of audio channels.
The audio content.
The mime type of the audio.
Possible values:
audio/wav WAV audio format
WAV audio format
audio/mp3 MP3 audio format
MP3 audio format
audio/aiff AIFF audio format
AIFF audio format
audio/aac AAC audio format
AAC audio format
audio/ogg OGG audio format
OGG audio format
audio/flac FLAC audio format
FLAC audio format
audio/mpeg MPEG audio format
MPEG audio format
audio/m4a M4A audio format
M4A audio format
audio/l16 L16 audio format
L16 audio format
audio/opus OPUS audio format
OPUS audio format
audio/alaw ALAW audio format
ALAW audio format
audio/mulaw MULAW audio format
MULAW audio format
The sample rate of the audio.
No description provided.
Always set to "audio" .
The URI of the audio.
The last image generated by the model in response to the current request. Note: this is added by the SDK.
Concatenated text from the last model output in response to the current request. Note: this is added by the SDK.
The last video generated by the model in response to the current request. Note: this is added by the SDK.
A video content block.
The video content.
The mime type of the video.
Possible values:
video/mp4 MP4 video format
MP4 video format
video/mpeg MPEG video format
MPEG video format
video/mpg MPG video format
MPG video format
video/mov MOV video format
MOV video format
video/avi AVI video format
AVI video format
video/x-flv FLV video format
FLV video format
video/webm WebM video format
WebM video format
video/wmv WMV video format
WMV video format
video/3gpp 3GPP video format
3GPP video format
How the model processes this video for understanding.
The resolution of the media.
Possible values
low Low resolution.
Low resolution.
medium Medium resolution.
Medium resolution.
high High resolution.
High resolution.
ultra_high Ultra high resolution.
Ultra high resolution.
No description provided.
Always set to "video" .
The URI of the video.
The ID of the previous interaction, if any.
Enforces that the generated response is a JSON object that complies with the JSON schema specified in this field.
Safety settings for the interaction.
A safety setting that affects the safety-blocking behavior. A SafetySetting consists of a harm category and a threshold for that category.
Optional. The method for blocking content. If not specified, the default behavior is to use the probability score.
Possible values:
severity The harm block method uses both probability and severity scores.
The harm block method uses both probability and severity scores.
probability The harm block method uses the probability score.
The harm block method uses the probability score.
Required. The threshold for blocking content. If the harm probability exceeds this threshold, the content will be blocked.
Possible values:
block_low_and_above Block content with a low harm probability or higher.
Block content with a low harm probability or higher.
block_medium_and_above Block content with a medium harm probability or higher.
Block content with a medium harm probability or higher.
block_only_high Block content with a high harm probability.
Block content with a high harm probability.
block_none Do not block any content, regardless of its harm probability.
Do not block any content, regardless of its harm probability.
off Turn off the safety filter entirely.
Turn off the safety filter entirely.
Required. The type of harm category to be blocked.
Possible values
hate_speech Content that promotes violence or incites hatred against individuals or groups based on certain attributes.
Content that promotes violence or incites hatred against individuals or groups based on certain attributes.
dangerous_content Content that promotes, facilitates, or enables dangerous activities.
Content that promotes, facilitates, or enables dangerous activities.
harassment Abusive, threatening, or content intended to bully, torment, or ridicule.
Abusive, threatening, or content intended to bully, torment, or ridicule.
sexually_explicit Content that contains sexually explicit material.
Content that contains sexually explicit material.
civic_integrity Deprecated: Election filter is not longer supported. The harm category is civic integrity.
Deprecated: Election filter is not longer supported. The harm category is civic integrity.
image_hate Images that contain hate speech.
Images that contain hate speech.
image_dangerous_content Images that contain dangerous content.
Images that contain dangerous content.
image_harassment Images that contain harassment.
Images that contain harassment.
image_sexually_explicit Images that contain sexually explicit content.
Images that contain sexually explicit content.
jailbreak Prompts designed to bypass safety filters.
Prompts designed to bypass safety filters.
The service tier for the interaction.
Possible values
flex Flex service tier.
Flex service tier.
standard Standard service tier.
Standard service tier.
priority Priority service tier.
Priority service tier.
deferred Deferred service tier.
Deferred service tier.
Required. Output only. The status of the interaction.
Possible values:
in_progress The interaction is in progress.
The interaction is in progress.
requires_action The interaction requires action/input from the user.
The interaction requires action/input from the user.
completed The interaction is completed.
The interaction is completed.
failed The interaction failed.
The interaction failed.
cancelled The interaction was cancelled.
The interaction was cancelled.
incomplete The interaction is completed, but contains incomplete results (e.g. hitting max_tokens).
The interaction is completed, but contains incomplete results (e.g. hitting max_tokens).
budget_exceeded The interaction was halted because the token budget was exceeded.
The interaction was halted because the token budget was exceeded.
queued The interaction is queued, waiting for processing.
The interaction is queued, waiting for processing.
Output only. The steps that make up the interaction, when included in the response.
System instruction for the interaction.
A list of tool declarations the model may call during interaction.
Output only. The time at which the response was last updated in ISO 8601 format (YYYY-MM-DDThh:mm:ssZ).
Output only. Statistics on the interaction request's token usage.
Statistics on the interaction request's token usage.
A breakdown of cached token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Grounding tool count.
The number of grounding tool counts.
The number of grounding tool counts.
The grounding tool type associated with the count.
Possible values:
google_search Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
google_maps Grounding with Google Maps.
Grounding with Google Maps.
retrieval Grounding with customer's data, for example, VertexAISearch.
Grounding with customer's data, for example, VertexAISearch.
A breakdown of input token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of output token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of tool-use token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Number of tokens in the cached part of the prompt (the cached content).
Number of tokens in the prompt (context).
Total number of tokens across all the generated responses.
Number of tokens of thoughts for thinking models.
Total token count for the interaction request (prompt + responses + other internal tokens).
Number of tokens present in tool-use prompt(s).
Optional. Webhook configuration for receiving notifications when the interaction completes.
Message for configuring webhook events for a request.
Optional. If set, these webhook URIs will be used for webhook events instead of the registered webhooks.
Optional. The user metadata that will be returned on each event emission to the webhooks.
The content of the response.
An audio content block.
The number of audio channels.
The audio content.
The mime type of the audio.
Possible values:
audio/wav WAV audio format
WAV audio format
audio/mp3 MP3 audio format
MP3 audio format
audio/aiff AIFF audio format
AIFF audio format
audio/aac AAC audio format
AAC audio format
audio/ogg OGG audio format
OGG audio format
audio/flac FLAC audio format
FLAC audio format
audio/mpeg MPEG audio format
MPEG audio format
audio/m4a M4A audio format
M4A audio format
audio/l16 L16 audio format
L16 audio format
audio/opus OPUS audio format
OPUS audio format
audio/alaw ALAW audio format
ALAW audio format
audio/mulaw MULAW audio format
MULAW audio format
The sample rate of the audio.
No description provided.
Always set to "audio" .
The URI of the audio.
A document content block.
The document content.
The mime type of the document.
Possible values:
application/pdf PDF document format
PDF document format
text/csv CSV document format
CSV document format
No description provided.
Always set to "document" .
The URI of the document.
An image content block.
The image content.
The mime type of the image.
Possible values:
image/png PNG image format
PNG image format
image/jpeg JPEG image format
JPEG image format
image/webp WebP image format
WebP image format
image/heic HEIC image format
HEIC image format
image/heif HEIF image format
HEIF image format
image/gif GIF image format
GIF image format
image/bmp BMP image format
BMP image format
image/tiff TIFF image format
TIFF image format
The resolution of the media.
Possible values
low Low resolution.
Low resolution.
medium Medium resolution.
Medium resolution.
high High resolution.
High resolution.
ultra_high Ultra high resolution.
Ultra high resolution.
No description provided.
Always set to "image" .
The URI of the image.
A text content block.
Citation information for model-generated content.
Citation information for model-generated content.
A file citation annotation.
User provided metadata about the retrieved context.
The URI of the file.
End of the attributed segment, exclusive.
The name of the file.
Media ID in-case of image citations, if applicable.
Page number of the cited document, if applicable.
Source attributed for a portion of the text.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
No description provided.
Always set to "file_citation" .
A place citation annotation.
End of the attributed segment, exclusive.
Title of the place.
The ID of the place, in `places/{place_id}` format.
Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.
Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
No description provided.
Always set to "place_citation" .
URI reference of the place.
A URL citation annotation.
End of the attributed segment, exclusive.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
The title of the URL.
No description provided.
Always set to "url_citation" .
Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.
End of the attributed segment, exclusive.
End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
Optional. Speaker label for this word (e.g. "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
The transcribed word.
No description provided.
Always set to "word_info" .
Required. The text content.
No description provided.
Always set to "text" .
A video content block.
The video content.
The mime type of the video.
Possible values:
video/mp4 MP4 video format
MP4 video format
video/mpeg MPEG video format
MPEG video format
video/mpg MPG video format
MPG video format
video/mov MOV video format
MOV video format
video/avi AVI video format
AVI video format
video/x-flv FLV video format
FLV video format
video/webm WebM video format
WebM video format
video/wmv WMV video format
WMV video format
video/3gpp 3GPP video format
3GPP video format
How the model processes this video for understanding.
The resolution of the media.
Possible values
low Low resolution.
Low resolution.
medium Medium resolution.
Medium resolution.
high High resolution.
High resolution.
ultra_high Ultra high resolution.
Ultra high resolution.
No description provided.
Always set to "video" .
The URI of the video.
A tool that can be used by the model.
A tool that can be used by the model to execute code.
No description provided.
Always set to "code_execution" .
A tool that can be used by the model to interact with the computer.
Optional. Disabled safety policies for computer use.
Possible values:
financial_transactions Safety policy for financial transactions.
Safety policy for financial transactions.
sensitive_data_modification Safety policy for sensitive data modification.
Safety policy for sensitive data modification.
communication_tool Safety policy for communication tools (e.g. Gmail, Chat, Meet).
Safety policy for communication tools (e.g. Gmail, Chat, Meet).
account_creation Safety policy for account creation.
Safety policy for account creation.
data_modification Safety policy for data modification.
Safety policy for data modification.
user_consent_management Safety policy for user consent management.
Safety policy for user consent management.
legal_terms_and_agreements Safety policy for legal terms and agreements.
Safety policy for legal terms and agreements.
Whether enable the prompt injection detection check on computer-use request.
The environment being operated.
Possible values:
browser Operates in a web browser.
Operates in a web browser.
mobile Operates in a mobile environment.
Operates in a mobile environment.
desktop Operates in a desktop environment.
Operates in a desktop environment.
The list of predefined functions that are excluded from the model call.
No description provided.
Always set to "computer_use" .
A tool that can be used by the model to search files.
The file search store names to search.
Metadata filter to apply to the semantic retrieval documents and chunks.
The number of semantic retrieval chunks to retrieve.
No description provided.
Always set to "file_search" .
A tool that can be used by the model.
A description of the function.
The name of the function.
The JSON Schema for the function's parameters.
No description provided.
Always set to "function" .
A tool that can be used by the model to call Google Maps.
Whether to return a widget context token in the tool call result of the response.
The latitude of the user's location.
The longitude of the user's location.
No description provided.
Always set to "google_maps" .
A tool that can be used by the model to search Google.
The types of search grounding to enable.
Possible values:
web_search Setting this field enables web search. Only text results are returned.
Setting this field enables web search. Only text results are returned.
image_search Setting this field enables image search. Image bytes are returned.
Setting this field enables image search. Image bytes are returned.
enterprise_web_search Setting this field enables enterprise web search.
Setting this field enables enterprise web search.
No description provided.
Always set to "google_search" .
A MCPServer is a server that can be called by the model to perform actions.
The allowed tools.
The configuration for allowed tools.
The mode of the tool choice.
Possible values:
auto Auto tool choice.
Auto tool choice.
any Any tool choice.
Any tool choice.
none No tool choice.
No tool choice.
validated Validated tool choice.
Validated tool choice.
The names of the allowed tools.
Optional: Fields for authentication headers, timeouts, etc., if needed.
The name of the MCPServer.
No description provided.
Always set to "mcp_server" .
The full URL for the MCPServer endpoint. Example: "https://api.example.com/mcp"
A tool that can be used by the model to retrieve files.
Used to specify configuration for ExaAISearch.
Used to specify configuration for ExaAISearch.
Required. The API key for ExaAiSearch.
Optional. This field can be used to pass any parameter from the Exa.ai Search API.
Used to specify configuration for ParallelAISearch.
Used to specify configuration for ParallelAISearch.
Optional. The API key for ParallelAiSearch.
Optional. Custom configs for ParallelAiSearch.
Used to specify configuration for RagStore.
Use to specify configuration for RAG Store.
Optional. The representation of the rag source.
The definition of the Rag resource.
Optional. RagCorpora resource name.
Optional. rag_file_id. The files should be in the same rag_corpus set in rag_corpus field.
Optional. The retrieval config for the Rag query.
Specifies the context retrieval config.
Optional. Config for filters.
Config for filters.
Optional. String for metadata filtering.
Optional. Only returns contexts with vector distance smaller than the threshold.
Optional. Only returns contexts with vector similarity larger than the threshold.
Optional. Config for Hybrid Search.
Config for Hybrid Search.
Optional. Alpha value controls the weight between dense and sparse vector search results.
Optional. Config for ranking and reranking.
Config for ranking and reranking.
Optional. The number of contexts to retrieve.
The types of file retrieval to enable.
Possible values:
parallel_ai_search
No description provided.
Always set to "retrieval" .
A tool that can be used by the model to fetch URL context.
No description provided.
Always set to "url_context" .
No examples available for this type.
InteractionSseEvent
Polymorphic discriminator: event_type
No description provided.
Error message from an interaction.
A URI that identifies the error type.
A human-readable error message.
The event_id token to be used to resume the interaction stream, from this event.
No description provided.
Always set to "error" .
The event_id token to be used to resume the interaction stream, from this event.
No description provided.
Always set to "interaction.completed" .
Partial completed interaction resource emitted at the end of the stream.
Partial interaction resource emitted by interaction lifecycle SSE events. Streaming lifecycle payloads may omit fields that are only available on full non-streaming Interaction responses.
The agent to interact with.
Output only. The time at which the response was created in ISO 8601 format.
Required. Output only. A unique identifier for the interaction completion.
The model that will complete your prompt.
Output only. The resource type.
The service tier for the interaction.
Possible values
flex Flex service tier.
Flex service tier.
standard Standard service tier.
Standard service tier.
priority Priority service tier.
Priority service tier.
deferred Deferred service tier.
Deferred service tier.
Required. Output only. The status of the interaction.
Possible values:
in_progress The interaction is in progress.
The interaction is in progress.
requires_action The interaction requires action/input from the user.
The interaction requires action/input from the user.
completed The interaction is completed.
The interaction is completed.
failed The interaction failed.
The interaction failed.
cancelled The interaction was cancelled.
The interaction was cancelled.
incomplete The interaction is completed, but contains incomplete results (e.g. hitting max_tokens).
The interaction is completed, but contains incomplete results (e.g. hitting max_tokens).
Output only. The steps that make up the interaction, if included in this event.
Output only. The time at which the response was last updated in ISO 8601 format.
Output only. Statistics on the interaction request's token usage.
Statistics on the interaction request's token usage.
A breakdown of cached token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Grounding tool count.
The number of grounding tool counts.
The number of grounding tool counts.
The grounding tool type associated with the count.
Possible values:
google_search Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
google_maps Grounding with Google Maps.
Grounding with Google Maps.
retrieval Grounding with customer's data, for example, VertexAISearch.
Grounding with customer's data, for example, VertexAISearch.
A breakdown of input token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of output token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of tool-use token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Number of tokens in the cached part of the prompt (the cached content).
Number of tokens in the prompt (context).
Total number of tokens across all the generated responses.
Number of tokens of thoughts for thinking models.
Total token count for the interaction request (prompt + responses + other internal tokens).
Number of tokens present in tool-use prompt(s).
The event_id token to be used to resume the interaction stream, from this event.
No description provided.
Always set to "interaction.created" .
Partial interaction resource emitted when the stream is created.
Partial interaction resource emitted by interaction lifecycle SSE events. Streaming lifecycle payloads may omit fields that are only available on full non-streaming Interaction responses.
The agent to interact with.
Output only. The time at which the response was created in ISO 8601 format.
Required. Output only. A unique identifier for the interaction completion.
The model that will complete your prompt.
Output only. The resource type.
The service tier for the interaction.
Possible values
flex Flex service tier.
Flex service tier.
standard Standard service tier.
Standard service tier.
priority Priority service tier.
Priority service tier.
deferred Deferred service tier.
Deferred service tier.
Required. Output only. The status of the interaction.
Possible values:
in_progress The interaction is in progress.
The interaction is in progress.
requires_action The interaction requires action/input from the user.
The interaction requires action/input from the user.
completed The interaction is completed.
The interaction is completed.
failed The interaction failed.
The interaction failed.
cancelled The interaction was cancelled.
The interaction was cancelled.
incomplete The interaction is completed, but contains incomplete results (e.g. hitting max_tokens).
The interaction is completed, but contains incomplete results (e.g. hitting max_tokens).
Output only. The steps that make up the interaction, if included in this event.
Output only. The time at which the response was last updated in ISO 8601 format.
Output only. Statistics on the interaction request's token usage.
Statistics on the interaction request's token usage.
A breakdown of cached token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Grounding tool count.
The number of grounding tool counts.
The number of grounding tool counts.
The grounding tool type associated with the count.
Possible values:
google_search Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
google_maps Grounding with Google Maps.
Grounding with Google Maps.
retrieval Grounding with customer's data, for example, VertexAISearch.
Grounding with customer's data, for example, VertexAISearch.
A breakdown of input token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of output token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of tool-use token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Number of tokens in the cached part of the prompt (the cached content).
Number of tokens in the prompt (context).
Total number of tokens across all the generated responses.
Number of tokens of thoughts for thinking models.
Total token count for the interaction request (prompt + responses + other internal tokens).
Number of tokens present in tool-use prompt(s).
The event_id token to be used to resume the interaction stream, from this event.
No description provided.
Always set to "interaction.status_update" .
No description provided.
No description provided.
Possible values:
in_progress The interaction is in progress.
The interaction is in progress.
requires_action The interaction requires action/input from the user.
The interaction requires action/input from the user.
completed The interaction is completed.
The interaction is completed.
failed The interaction failed.
The interaction failed.
cancelled The interaction was cancelled.
The interaction was cancelled.
incomplete The interaction is completed, but contains incomplete results (e.g. hitting max_tokens).
The interaction is completed, but contains incomplete results (e.g. hitting max_tokens).
budget_exceeded The interaction was halted because the token budget was exceeded.
The interaction was halted because the token budget was exceeded.
queued The interaction is queued, waiting for processing (e.g. waiting for off-peak capacity).
The interaction is queued, waiting for processing (e.g. waiting for off-peak capacity).
No description provided.
No description provided.
No description provided.
Always set to "arguments_delta" .
The number of audio channels.
No description provided.
No description provided.
Possible values:
audio/wav WAV audio format
WAV audio format
audio/mp3 MP3 audio format
MP3 audio format
audio/aiff AIFF audio format
AIFF audio format
audio/aac AAC audio format
AAC audio format
audio/ogg OGG audio format
OGG audio format
audio/flac FLAC audio format
FLAC audio format
audio/mpeg MPEG audio format
MPEG audio format
audio/m4a M4A audio format
M4A audio format
audio/l16 L16 audio format
L16 audio format
audio/opus OPUS audio format
OPUS audio format
audio/alaw ALAW audio format
ALAW audio format
audio/mulaw MULAW audio format
MULAW audio format
The sample rate of the audio.
No description provided.
Always set to "audio" .
No description provided.
No description provided.
The arguments to pass to the code execution.
The code to be executed.
Programming language of the `code`.
Possible values:
python Python >= 3.10, with numpy and simpy available.
Python >= 3.10, with numpy and simpy available.
A signature hash for backend validation.
No description provided.
Always set to "code_execution_call" .
No description provided.
No description provided.
A signature hash for backend validation.
No description provided.
Always set to "code_execution_result" .
No description provided.
No description provided.
Possible values:
application/pdf PDF document format
PDF document format
text/csv CSV document format
CSV document format
No description provided.
Always set to "document" .
No description provided.
A signature hash for backend validation.
No description provided.
Always set to "file_search_call" .
No description provided.
The result of the File Search.
A signature hash for backend validation.
No description provided.
Always set to "file_search_result" .
Required. ID to match the ID from the function call block.
No description provided.
No description provided.
No description provided.
No description provided.
Always set to "function_result" .
The arguments to pass to the Google Maps tool.
The arguments to pass to the Google Maps tool.
The queries to be executed.
A signature hash for backend validation.
No description provided.
Always set to "google_maps_call" .
The results of the Google Maps.
The result of the Google Maps.
The places that were found.
Title of the place.
The ID of the place, in `places/{place_id}` format.
Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.
Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
URI reference of the place.
Resource name of the Google Maps widget context token.
A signature hash for backend validation.
No description provided.
Always set to "google_maps_result" .
No description provided.
The arguments to pass to Google Search.
Web search queries for the following-up web search.
A signature hash for backend validation.
No description provided.
Always set to "google_search_call" .
No description provided.
No description provided.
The result of the Google Search.
Web content snippet that can be embedded in a web page or an app webview.
A signature hash for backend validation.
No description provided.
Always set to "google_search_result" .
No description provided.
No description provided.
Possible values:
image/png PNG image format
PNG image format
image/jpeg JPEG image format
JPEG image format
image/webp WebP image format
WebP image format
image/heic HEIC image format
HEIC image format
image/heif HEIF image format
HEIF image format
image/gif GIF image format
GIF image format
image/bmp BMP image format
BMP image format
image/tiff TIFF image format
TIFF image format
The resolution of the media.
Possible values
low Low resolution.
Low resolution.
medium Medium resolution.
Medium resolution.
high High resolution.
High resolution.
ultra_high Ultra high resolution.
Ultra high resolution.
No description provided.
Always set to "image" .
No description provided.
No description provided.
No description provided.
No description provided.
No description provided.
Always set to "mcp_server_tool_call" .
No description provided.
No description provided.
No description provided.
No description provided.
Always set to "mcp_server_tool_result" .
Used by Vertex Retrieval tools such as Parallel AI, Exa AI, Vertex AI Search, etc. RetrievalType decides which tool is used.
Required. The arguments to pass to the Retrieval tool.
The arguments to pass to Retrieval tools.
Queries for Retrieval information.
The type of retrieval tools.
Possible values:
rag_store The type of retrieval tools.
The type of retrieval tools.
exa_ai_search The type of retrieval tools.
The type of retrieval tools.
parallel_ai_search The type of retrieval tools.
The type of retrieval tools.
A signature hash for backend validation.
No description provided.
Always set to "retrieval_call" .
Used by Vertex Retrieval tools such as Parallel AI, Exa AI, Vertex AI Search, etc. ToolResultDelta.type
Whether the retrieval resulted in an error.
A signature hash for backend validation.
No description provided.
Always set to "retrieval_result" .
Citation information for model-generated content.
Citation information for model-generated content.
A file citation annotation.
User provided metadata about the retrieved context.
The URI of the file.
End of the attributed segment, exclusive.
The name of the file.
Media ID in-case of image citations, if applicable.
Page number of the cited document, if applicable.
Source attributed for a portion of the text.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
No description provided.
Always set to "file_citation" .
A place citation annotation.
End of the attributed segment, exclusive.
Title of the place.
The ID of the place, in `places/{place_id}` format.
Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.
Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
No description provided.
Always set to "place_citation" .
URI reference of the place.
A URL citation annotation.
End of the attributed segment, exclusive.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
The title of the URL.
No description provided.
Always set to "url_citation" .
Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.
End of the attributed segment, exclusive.
End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
Optional. Speaker label for this word (e.g. "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
The transcribed word.
No description provided.
Always set to "word_info" .
No description provided.
Always set to "text_annotation_delta" .
No description provided.
No description provided.
Always set to "text" .
Signature to match the backend source to be part of the generation.
No description provided.
Always set to "thought_signature" .
A new summary item to be added to the thought.
No description provided.
Always set to "thought_summary" .
No description provided.
The arguments to pass to the URL context.
The URLs to fetch.
A signature hash for backend validation.
No description provided.
Always set to "url_context_call" .
No description provided.
No description provided.
The result of the URL context.
The status of the URL retrieval.
Possible values:
success Url retrieval is successful.
Url retrieval is successful.
error Url retrieval is failed due to error.
Url retrieval is failed due to error.
paywall Url retrieval is failed because the content is behind paywall.
Url retrieval is failed because the content is behind paywall.
unsafe Url retrieval is failed because the content is unsafe.
Url retrieval is failed because the content is unsafe.
The URL that was fetched.
A signature hash for backend validation.
No description provided.
Always set to "url_context_result" .
No description provided.
No description provided.
Possible values:
video/mp4 MP4 video format
MP4 video format
video/mpeg MPEG video format
MPEG video format
video/mpg MPG video format
MPG video format
video/mov MOV video format
MOV video format
video/avi AVI video format
AVI video format
video/x-flv FLV video format
FLV video format
video/webm WebM video format
WebM video format
video/wmv WMV video format
WMV video format
video/3gpp 3GPP video format
3GPP video format
The resolution of the media.
Possible values
low Low resolution.
Low resolution.
medium Medium resolution.
Medium resolution.
high High resolution.
High resolution.
ultra_high Ultra high resolution.
Ultra high resolution.
No description provided.
Always set to "video" .
No description provided.
The event_id token to be used to resume the interaction stream, from this event.
No description provided.
Always set to "step.delta" .
No description provided.
No description provided.
Optional metadata accompanying ANY streamed event.
Statistics on the interaction request's token usage.
Statistics on the interaction request's token usage.
A breakdown of cached token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Grounding tool count.
The number of grounding tool counts.
The number of grounding tool counts.
The grounding tool type associated with the count.
Possible values:
google_search Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
google_maps Grounding with Google Maps.
Grounding with Google Maps.
retrieval Grounding with customer's data, for example, VertexAISearch.
Grounding with customer's data, for example, VertexAISearch.
A breakdown of input token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of output token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of tool-use token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Number of tokens in the cached part of the prompt (the cached content).
Number of tokens in the prompt (context).
Total number of tokens across all the generated responses.
Number of tokens of thoughts for thinking models.
Total token count for the interaction request (prompt + responses + other internal tokens).
Number of tokens present in tool-use prompt(s).
The event_id token to be used to resume the interaction stream, from this event.
No description provided.
Always set to "step.start" .
No description provided.
No description provided.
The event_id token to be used to resume the interaction stream, from this event.
No description provided.
Always set to "step.stop" .
No description provided.
Model usage stats for this specific step.
Statistics on the interaction request's token usage.
A breakdown of cached token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Grounding tool count.
The number of grounding tool counts.
The number of grounding tool counts.
The grounding tool type associated with the count.
Possible values:
google_search Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
google_maps Grounding with Google Maps.
Grounding with Google Maps.
retrieval Grounding with customer's data, for example, VertexAISearch.
Grounding with customer's data, for example, VertexAISearch.
A breakdown of input token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of output token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of tool-use token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Number of tokens in the cached part of the prompt (the cached content).
Number of tokens in the prompt (context).
Total number of tokens across all the generated responses.
Number of tokens of thoughts for thinking models.
Total token count for the interaction request (prompt + responses + other internal tokens).
Number of tokens present in tool-use prompt(s).
Cumulative model usage stats from the start of the session.
Statistics on the interaction request's token usage.
A breakdown of cached token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Grounding tool count.
The number of grounding tool counts.
The number of grounding tool counts.
The grounding tool type associated with the count.
Possible values:
google_search Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
google_maps Grounding with Google Maps.
Grounding with Google Maps.
retrieval Grounding with customer's data, for example, VertexAISearch.
Grounding with customer's data, for example, VertexAISearch.
A breakdown of input token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of output token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
A breakdown of tool-use token usage by modality.
The token count for a single response modality.
The modality associated with the token count.
Possible values
text Indicates the model should return text.
Indicates the model should return text.
image Indicates the model should return images.
Indicates the model should return images.
audio Indicates the model should return audio.
Indicates the model should return audio.
video Indicates the model should return video.
Indicates the model should return video.
document Indicates the model should return documents.
Indicates the model should return documents.
Number of tokens for the modality.
Number of tokens in the cached part of the prompt (the cached content).
Number of tokens in the prompt (context).
Total number of tokens across all the generated responses.
Number of tokens of thoughts for thinking models.
Total token count for the interaction request (prompt + responses + other internal tokens).
Number of tokens present in tool-use prompt(s).
Interaction Completed
Interaction Completed
Interaction Created
Interaction Created
Interaction Status Update
Configuration for audio output format.
Bit rate in bits per second (bps). Only applicable for compressed formats (MP3, Opus).
The delivery mode for the audio output.
Possible values:
inline Audio data is returned inline in the response.
Audio data is returned inline in the response.
uri Audio data is returned as a URI.
Audio data is returned as a URI.
The MIME type of the audio output.
Possible values:
audio/mp3 MP3 audio format.
MP3 audio format.
audio/ogg_opus OGG Opus audio format.
OGG Opus audio format.
audio/l16 Raw PCM (L16) audio format.
Raw PCM (L16) audio format.
audio/wav WAV audio format.
WAV audio format.
audio/alaw A-law audio format.
A-law audio format.
audio/mulaw Mu-law audio format.
Mu-law audio format.
Sample rate in Hz.
No description provided.
Always set to "audio" .
Configuration for image output format.
The aspect ratio for the image output.
Possible values:
1:1 1:1 aspect ratio.
1:1 aspect ratio.
2:3 2:3 aspect ratio.
2:3 aspect ratio.
3:2 3:2 aspect ratio.
3:2 aspect ratio.
3:4 3:4 aspect ratio.
3:4 aspect ratio.
4:3 4:3 aspect ratio.
4:3 aspect ratio.
4:5 4:5 aspect ratio.
4:5 aspect ratio.
5:4 5:4 aspect ratio.
5:4 aspect ratio.
9:16 9:16 aspect ratio.
9:16 aspect ratio.
16:9 16:9 aspect ratio.
16:9 aspect ratio.
21:9 21:9 aspect ratio.
21:9 aspect ratio.
1:8 1:8 aspect ratio.
1:8 aspect ratio.
8:1 8:1 aspect ratio.
8:1 aspect ratio.
1:4 1:4 aspect ratio.
1:4 aspect ratio.
4:1 4:1 aspect ratio.
4:1 aspect ratio.
The delivery mode for the image output.
Possible values:
inline Image data is returned inline in the response.
Image data is returned inline in the response.
uri Image data is returned as a URI.
Image data is returned as a URI.
The size of the image output.
Possible values:
512 512px image size.
512px image size.
1K 1K image size.
2K 2K image size.
4K 4K image size.
The MIME type of the image output.
Possible values:
image/jpeg JPEG image format.
JPEG image format.
No description provided.
Always set to "image" .
Configuration for text output format.
The MIME type of the text output.
Possible values:
application/json JSON output format.
JSON output format.
text/plain Plain text output format.
Plain text output format.
The JSON schema that the output should conform to. Only applicable when mime_type is application/json.
No description provided.
Always set to "text" .
Configuration for video output format.
The aspect ratio for the video output.
Possible values:
16:9 16:9 aspect ratio.
16:9 aspect ratio.
9:16 9:16 aspect ratio.
9:16 aspect ratio.
The delivery mode for the video output.
Possible values:
inline Video data is returned inline in the response.
Video data is returned inline in the response.
uri Video data is returned as a URI.
Video data is returned as a URI.
The duration for the video output.
The Cloud Storage URI to store the video output. Required for Vertex if delivery mode is URI.
The video output resolution. Defaults to 720p.
Possible values:
360p 360p resolution.
360p resolution.
720p 720p resolution.
720p resolution.
1080p 1080p resolution.
1080p resolution.
4k 4K resolution.
No description provided.
Always set to "video" .
Text Output (JSON Schema)
A step in the interaction.
Code execution call step.
Required. The arguments to pass to the code execution.
The arguments to pass to the code execution.
The code to be executed.
Programming language of the `code`.
Possible values:
python Python >= 3.10, with numpy and simpy available.
Python >= 3.10, with numpy and simpy available.
Required. A unique ID for this specific tool call.
A signature hash for backend validation.
No description provided.
Always set to "code_execution_call" .
Code execution result step.
Required. ID to match the ID from the function call block.
Whether the code execution resulted in an error.
Required. The output of the code execution.
A signature hash for backend validation.
No description provided.
Always set to "code_execution_result" .
File Search call step.
Required. A unique ID for this specific tool call.
A signature hash for backend validation.
No description provided.
Always set to "file_search_call" .
File Search result step.
Required. ID to match the ID from the function call block.
A signature hash for backend validation.
No description provided.
Always set to "file_search_result" .
A function tool call step.
Required. The arguments to pass to the function.
Required. A unique ID for this specific tool call.
Required. The name of the tool to call.
No description provided.
Always set to "function_call" .
Result of a function tool call.
Required. ID to match the ID from the function call block.
Whether the tool call resulted in an error.
The name of the tool that was called.
Required. The result of the tool call.
No description provided.
Always set to "function_result" .
Google Maps call step.
The arguments to pass to the Google Maps tool.
The arguments to pass to the Google Maps tool.
The queries to be executed.
Required. A unique ID for this specific tool call.
A signature hash for backend validation.
No description provided.
Always set to "google_maps_call" .
Google Maps result step.
Required. ID to match the ID from the function call block.
No description provided.
The result of the Google Maps.
No description provided.
No description provided.
No description provided.
No description provided.
Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
No description provided.
No description provided.
A signature hash for backend validation.
No description provided.
Always set to "google_maps_result" .
Google Search call step.
Required. The arguments to pass to Google Search.
The arguments to pass to Google Search.
Web search queries for the following-up web search.
Required. A unique ID for this specific tool call.
The type of search grounding enabled.
Possible values:
web_search Setting this field enables web search. Only text results are returned.
Setting this field enables web search. Only text results are returned.
image_search Setting this field enables image search. Image bytes are returned.
Setting this field enables image search. Image bytes are returned.
enterprise_web_search Setting this field enables enterprise web search.
Setting this field enables enterprise web search.
A signature hash for backend validation.
No description provided.
Always set to "google_search_call" .
Google Search result step.
Required. ID to match the ID from the function call block.
Whether the Google Search resulted in an error.
Required. The results of the Google Search.
The result of the Google Search.
Web content snippet that can be embedded in a web page or an app webview.
A signature hash for backend validation.
No description provided.
Always set to "google_search_result" .
MCPServer tool call step.
Required. The JSON object of arguments for the function.
Required. A unique ID for this specific tool call.
Required. The name of the tool which was called.
Required. The name of the used MCP server.
No description provided.
Always set to "mcp_server_tool_call" .
MCPServer tool result step.
Required. ID to match the ID from the function call block.
Name of the tool which is called for this specific tool call.
Required. The output from the MCP server call. Can be simple text or rich content.
The name of the used MCP server.
No description provided.
Always set to "mcp_server_tool_result" .
Output generated by the model.
No description provided.
No description provided.
Always set to "model_output" .
A thought step.
A signature hash for backend validation.
A summary of the thought.
An image content block.
The image content.
The mime type of the image.
Possible values:
image/png PNG image format
PNG image format
image/jpeg JPEG image format
JPEG image format
image/webp WebP image format
WebP image format
image/heic HEIC image format
HEIC image format
image/heif HEIF image format
HEIF image format
image/gif GIF image format
GIF image format
image/bmp BMP image format
BMP image format
image/tiff TIFF image format
TIFF image format
The resolution of the media.
Possible values
low Low resolution.
Low resolution.
medium Medium resolution.
Medium resolution.
high High resolution.
High resolution.
ultra_high Ultra high resolution.
Ultra high resolution.
No description provided.
Always set to "image" .
The URI of the image.
A text content block.
Citation information for model-generated content.
Citation information for model-generated content.
A file citation annotation.
User provided metadata about the retrieved context.
The URI of the file.
End of the attributed segment, exclusive.
The name of the file.
Media ID in-case of image citations, if applicable.
Page number of the cited document, if applicable.
Source attributed for a portion of the text.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
No description provided.
Always set to "file_citation" .
A place citation annotation.
End of the attributed segment, exclusive.
Title of the place.
The ID of the place, in `places/{place_id}` format.
Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.
Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
No description provided.
Always set to "place_citation" .
URI reference of the place.
A URL citation annotation.
End of the attributed segment, exclusive.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
The title of the URL.
No description provided.
Always set to "url_citation" .
Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.
End of the attributed segment, exclusive.
End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
Optional. Speaker label for this word (e.g. "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
The transcribed word.
No description provided.
Always set to "word_info" .
Required. The text content.
No description provided.
Always set to "text" .
No description provided.
Always set to "thought" .
URL context call step.
Required. The arguments to pass to the URL context.
The arguments to pass to the URL context.
The URLs to fetch.
Required. A unique ID for this specific tool call.
A signature hash for backend validation.
No description provided.
Always set to "url_context_call" .
URL context result step.
Required. ID to match the ID from the function call block.
Whether the URL context resulted in an error.
Required. The results of the URL context.
The result of the URL context.
The status of the URL retrieval.
Possible values:
success Url retrieval is successful.
Url retrieval is successful.
error Url retrieval is failed due to error.
Url retrieval is failed due to error.
paywall Url retrieval is failed because the content is behind paywall.
Url retrieval is failed because the content is behind paywall.
unsafe Url retrieval is failed because the content is unsafe.
Url retrieval is failed because the content is unsafe.
The URL that was fetched.
A signature hash for backend validation.
No description provided.
Always set to "url_context_result" .
Input provided by the user.
No description provided.
No description provided.
Always set to "user_input" .
CodeExecutionCallStep
CodeExecutionResultStep
FileSearchCallStep
FileSearchResultStep
FunctionCallStep
FunctionResultStep
GoogleMapsCallStep
GoogleMapsResultStep
GoogleSearchCallStep
GoogleSearchResultStep
McpServerToolCallStep
McpServerToolResultStep
ModelOutputStep
UrlContextCallStep
UrlContextResultStep
EnvironmentConfig
Configuration for a custom environment.
Optional. The environment ID for the interaction. If specified, the request will update the existing environment instead of creating a new one.
Network configuration for the environment.
Possible values:
disabled Turns all network off.
Turns all network off.
No description provided.
A source to be mounted into the environment.
The inline content if `type` is `INLINE`.
Optional encoding for inline content (e.g. `base64`).
The source of the environment. For Cloud Storage, this is the Cloud Storage path. For GitHub, this is the GitHub path.
Where the source should appear in the environment.
No description provided.
Possible values:
gcs A Cloud Storage bucket.
A Cloud Storage bucket.
inline Inline content.
Inline content.
repository A generic repository. The protocol prefix in the source URL identifies the provider (e.g., github://, gcs://).
A generic repository. The protocol prefix in the source URL identifies the provider (e.g., github://, gcs://).
skill_registry A skill resource from the Skill Registry Service. Skill: projects/{project}/locations/{location}/skills/{skill} SkillRevision: projects/{project}/locations/{location}/skills/{skill}/revisions/{revision} Support mounting all skills under a project: projects/{project}/locations/{location}/skills.
A skill resource from the Skill Registry Service. Skill: projects/{project}/locations/{location}/skills/{skill} SkillRevision: projects/{project}/locations/{location}/skills/{skill}/revisions/{revision} Support mounting all skills under a project: projects/{project}/locations/{location}/skills.
No description provided.
Always set to "remote" .
External Sources
Network Allowlist
Proxy Credentials
EnvironmentNetworkEgressAllowlist
Outbound networking configuration for the sandbox. Accepts an object with an 'allowlist' array to restrict traffic, or the string 'disabled' to turn off all network access. Omit entirely to allow all outbound traffic with no header injection.
Outbound networking configuration for the sandbox. When specified, restricts which external domains the sandbox can reach. Omit entirely to allow all outbound traffic with no header injection.
List of allowed outbound domains. Only requests to listed domains are permitted. Use [{'domain': '*'}] to allow all domains while still injecting headers on specific ones.
A single domain allowlist rule with optional header injection.
Domain to allow outbound requests to. Supports wildcards (e.g. '*.googleapis.com'). Use '*' to allow all domains.
Headers to inject on all outbound requests matching this domain. Accepts a single dict or a list of dicts. The egress proxy injects these automatically.
Turns all network off.
Possible values
disabled Turns all network off.
Turns all network off.
ToolChoiceConfig
The tool choice configuration containing allowed tools.
The allowed tools.
The configuration for allowed tools.
The mode of the tool choice.
Possible values:
auto Auto tool choice.
Auto tool choice.
any Any tool choice.
Any tool choice.
none No tool choice.
No tool choice.
validated Validated tool choice.
Validated tool choice.
The names of the allowed tools.
An image content block.
The image content.
The mime type of the image.
Possible values:
image/png PNG image format
PNG image format
image/jpeg JPEG image format
JPEG image format
image/webp WebP image format
WebP image format
image/heic HEIC image format
HEIC image format
image/heif HEIF image format
HEIF image format
image/gif GIF image format
GIF image format
image/bmp BMP image format
BMP image format
image/tiff TIFF image format
TIFF image format
The resolution of the media.
Possible values
low Low resolution.
Low resolution.
medium Medium resolution.
Medium resolution.
high High resolution.
High resolution.
ultra_high Ultra high resolution.
Ultra high resolution.
No description provided.
Always set to "image" .
The URI of the image.
A text content block.
Citation information for model-generated content.
Citation information for model-generated content.
A file citation annotation.
User provided metadata about the retrieved context.
The URI of the file.
End of the attributed segment, exclusive.
The name of the file.
Media ID in-case of image citations, if applicable.
Page number of the cited document, if applicable.
Source attributed for a portion of the text.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
No description provided.
Always set to "file_citation" .
A place citation annotation.
End of the attributed segment, exclusive.
Title of the place.
The ID of the place, in `places/{place_id}` format.
Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.
Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
No description provided.
Always set to "place_citation" .
URI reference of the place.
A URL citation annotation.
End of the attributed segment, exclusive.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
The title of the URL.
No description provided.
Always set to "url_citation" .
Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.
End of the attributed segment, exclusive.
End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
Optional. Speaker label for this word (e.g. "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
The transcribed word.
No description provided.
Always set to "word_info" .
Required. The text content.
No description provided.
Always set to "text" .
Except as otherwise noted, the content of this page is licensed under the Creative Commons Attribution 4.0 License , and code samples are licensed under the Apache 2.0 License . For details, see the Google Developers Site Policies . Java is a registered trademark of Oracle and/or its affiliates.
Last updated 2026-08-19 UTC.