API Gemini поддерживает генерацию контента с использованием изображений, аудио, кода, инструментов и многого другого. Для получения подробной информации о каждой из этих функций читайте дальше и ознакомьтесь с примерами кода, ориентированными на решение конкретных задач, или прочитайте подробные руководства.
- Генерация текста
- Зрение
- Аудио
- Эмбеддинги
- Длинный контекст
- Выполнение кода
- Режим JSON
- Вызов функции
- Системные инструкции
Метод: models.generateContent
Генерирует ответ модели на входной запрос GenerateContentRequest . Подробную информацию об использовании см. в руководстве по генерации текста . Возможности ввода различаются для разных моделей, включая оптимизированные модели. Для получения более подробной информации см. руководство по моделям и руководство по оптимизации .
Конечная точка
posthttps: / /generativelanguage.googleapis.com /v1beta /{model=models /*}:generateContentПараметры пути
string model Обязательно. Название Model , которая будет использоваться для генерации автозавершения.
Формат: models/{model} . Он принимает вид models/{model} .
Текст запроса
Тело запроса содержит данные следующей структуры:
tools[]object ( Tool ) Необязательно. Список Tools Model может использовать для генерации следующего ответа.
Tool — это фрагмент кода, позволяющий системе взаимодействовать с внешними системами для выполнения действия или набора действий, выходящих за рамки знаний и области действия Model . Поддерживаемые Tool — это Function и codeExecution . Для получения дополнительной информации обратитесь к руководствам по вызову функций и выполнению кода .
toolConfigobject ( ToolConfig ) Необязательно. Конфигурация инструмента для любого Tool указанного в запросе. Пример использования см. в руководстве по вызову функций .
safetySettings[]object ( SafetySetting ) Необязательно. Список уникальных экземпляров SafetySetting для блокировки небезопасного контента.
Это будет применяться к GenerateContentRequest.contents и GenerateContentResponse.candidates . Для каждого типа SafetyCategory не должно быть более одной настройки. API будет блокировать любой контент и ответы, которые не соответствуют пороговым значениям, установленным этими настройками. Этот список переопределяет настройки по умолчанию для каждой SafetyCategory , указанной в safetySettings. Если для данной SafetyCategory в списке не указана SafetySetting безопасности, API будет использовать настройку безопасности по умолчанию для этой категории. Поддерживаются категории вреда HARM_CATEGORY_HATE_SPEECH, HARM_CATEGORY_SEXUALLY_EXPLICIT, HARM_CATEGORY_DANGEROUS_CONTENT, HARM_CATEGORY_HARASSMENT, HARM_CATEGORY_CIVIC_INTEGRITY, HARM_CATEGORY_JAILBREAK. Подробную информацию о доступных настройках безопасности см. в руководстве . Также ознакомьтесь с рекомендациями по безопасности , чтобы узнать, как учитывать вопросы безопасности в ваших приложениях на основе искусственного интеллекта.
systemInstructionobject ( Content )Необязательно. Разработчик задает системные инструкции . В настоящее время только текст.
generationConfigobject ( GenerationConfig )Необязательно. Параметры конфигурации для генерации модели и выходных данных.
cachedContentstring Необязательно. Название кэшированного контента, используемого в качестве контекста для выполнения прогнозирования. Формат: cachedContents/{cachedContent}
serviceTierenum ( ServiceTier )Необязательно. Уровень обслуживания запроса.
storebooleanНеобязательный параметр. Задает поведение логирования для данного запроса. Если задан, он имеет приоритет над конфигурацией логирования на уровне проекта.
Пример запроса
Текст
Python
Node.js
Идти
Оболочка
Java
Изображение
Python
Node.js
Идти
Оболочка
Java
Аудио
Python
Node.js
Идти
Оболочка
Видео
Python
Node.js
Идти
Оболочка
Python
Идти
Оболочка
Чат
Python
Node.js
Идти
Оболочка
Java
Кэш
Python
Node.js
Идти
Тюнингованная модель
Python
Режим JSON
Python
Node.js
Идти
Оболочка
Java
Выполнение кода
Python
Идти
Java
Вызов функции
Python
Идти
Node.js
Оболочка
Java
Конфигурация генерации
Python
Node.js
Идти
Оболочка
Java
Настройки безопасности
Python
Node.js
Идти
Оболочка
Java
Системная инструкция
Python
Node.js
Идти
Оболочка
Java
Ответный текст
В случае успеха тело ответа будет содержать экземпляр GenerateContentResponse .
Метод: models.streamGenerateContent
Генерирует потоковый ответ от модели на основе входных данных GenerateContentRequest .
Конечная точка
posthttps: / /generativelanguage.googleapis.com /v1beta /{model=models /*}:streamGenerateContentПараметры пути
string model Обязательно. Название Model , которая будет использоваться для генерации автозавершения.
Формат: models/{model} . Он принимает вид models/{model} .
Текст запроса
Тело запроса содержит данные следующей структуры:
tools[]object ( Tool ) Необязательно. Список Tools Model может использовать для генерации следующего ответа.
Tool — это фрагмент кода, позволяющий системе взаимодействовать с внешними системами для выполнения действия или набора действий, выходящих за рамки знаний и области действия Model . Поддерживаемые Tool — это Function и codeExecution . Для получения дополнительной информации обратитесь к руководствам по вызову функций и выполнению кода .
toolConfigobject ( ToolConfig ) Необязательно. Конфигурация инструмента для любого Tool указанного в запросе. Пример использования см. в руководстве по вызову функций .
safetySettings[]object ( SafetySetting ) Необязательно. Список уникальных экземпляров SafetySetting для блокировки небезопасного контента.
Это будет применяться к GenerateContentRequest.contents и GenerateContentResponse.candidates . Для каждого типа SafetyCategory не должно быть более одной настройки. API будет блокировать любой контент и ответы, которые не соответствуют пороговым значениям, установленным этими настройками. Этот список переопределяет настройки по умолчанию для каждой SafetyCategory , указанной в safetySettings. Если для данной SafetyCategory в списке не указана SafetySetting безопасности, API будет использовать настройку безопасности по умолчанию для этой категории. Поддерживаются категории вреда HARM_CATEGORY_HATE_SPEECH, HARM_CATEGORY_SEXUALLY_EXPLICIT, HARM_CATEGORY_DANGEROUS_CONTENT, HARM_CATEGORY_HARASSMENT, HARM_CATEGORY_CIVIC_INTEGRITY, HARM_CATEGORY_JAILBREAK. Подробную информацию о доступных настройках безопасности см. в руководстве . Также ознакомьтесь с рекомендациями по безопасности , чтобы узнать, как учитывать вопросы безопасности в ваших приложениях на основе искусственного интеллекта.
systemInstructionobject ( Content )Необязательно. Разработчик задает системные инструкции . В настоящее время только текст.
generationConfigobject ( GenerationConfig )Необязательно. Параметры конфигурации для генерации модели и выходных данных.
cachedContentstring Необязательно. Название кэшированного контента, используемого в качестве контекста для выполнения прогнозирования. Формат: cachedContents/{cachedContent}
serviceTierenum ( ServiceTier )Необязательно. Уровень обслуживания запроса.
storebooleanНеобязательный параметр. Задает поведение логирования для данного запроса. Если задан, он имеет приоритет над конфигурацией логирования на уровне проекта.
Пример запроса
Текст
Python
Node.js
Идти
Оболочка
Java
Изображение
Python
Node.js
Идти
Оболочка
Java
Аудио
Python
Идти
Оболочка
Видео
Python
Node.js
Идти
Оболочка
Python
Идти
Оболочка
Чат
Python
Node.js
Идти
Оболочка
Ответный текст
В случае успеха тело ответа содержит поток экземпляров GenerateContentResponse .
GenerateContentResponse
Ответ модели, подтверждающий наличие нескольких вариантов ответа.
Рейтинги безопасности и фильтрация контента отображаются для обоих запросов в GenerateContentResponse.prompt_feedback , а для каждого кандидата — в finishReason и safetyRatings . API: - Возвращает либо все запрошенные кандидаты, либо ни одного из них; - Не возвращает ни одного кандидата, только если с запросом что-то было не так (см. promptFeedback ); - Отображает отзывы по каждому кандидату в finishReason и safetyRatings .
candidates[]object ( Candidate )Варианты ответов, полученные от модели.
promptFeedbackobject ( PromptFeedback )Возвращает обратную связь от запроса, относящуюся к фильтрам содержимого.
usageMetadataobject ( UsageMetadata )Только вывод. Метаданные об использовании токенов в запросах на генерацию.
string modelVersionТолько выходные данные. Версия модели, использованная для генерации ответа.
responseIdstringТолько вывод. responseId используется для идентификации каждого ответа.
modelStatusobject ( ModelStatus )Только вывод. Текущее состояние данной модели.
| JSON-представление |
|---|
{ "candidates": [ { object ( |
PromptFeedback
Набор метаданных обратной связи, указанных в запросе GenerateContentRequest.content .
blockReasonenum ( BlockReason )Необязательно. Если задано, запрос был заблокирован, и варианты не возвращаются. Переформулируйте запрос.
safetyRatings[]object ( SafetyRating )Оценки безопасности подсказки. В каждой категории может быть не более одной оценки.
| JSON-представление |
|---|
{ "blockReason": enum ( |
BlockReason
Указывает причину блокировки запроса.
| Перечисления | |
|---|---|
BLOCK_REASON_UNSPECIFIED | Значение по умолчанию. Это значение не используется. |
SAFETY | Запрос был заблокирован по соображениям безопасности. Проверьте safetyRatings , чтобы узнать, какая категория безопасности привела к блокировке. |
OTHER | Запрос был заблокирован по неизвестным причинам. |
BLOCKLIST | Запрос был заблокирован из-за терминов, включенных в список заблокированных терминов. |
PROHIBITED_CONTENT | Запрос был заблокирован из-за запрещенного контента. |
IMAGE_SAFETY | Кандидаты заблокированы из-за небезопасного контента, созданного с помощью фотошопа. |
UsageMetadata
Метаданные об использовании токена в запросе на генерацию.
promptTokenCountinteger Количество токенов в запросе. Если задан параметр cachedContent , это по-прежнему общий эффективный размер запроса, то есть он включает количество токенов в кэшированном содержимом.
cachedContentTokenCountintegerКоличество токенов в кэшированной части запроса (кэшированное содержимое)
candidatesTokenCountintegerОбщее количество токенов по всем сгенерированным вариантам ответа.
toolUsePromptTokenCountintegerТолько вывод. Количество токенов, присутствующих в подсказках использования инструмента.
thoughtsTokenCountintegerТолько вывод. Количество токенов мыслей для моделей мышления.
totalTokenCountintegerОбщее количество токенов для запроса на генерацию (запрос + мысли + варианты ответа).
promptTokensDetails[]object ( ModalityTokenCount )Только выходные данные. Список модальностей, которые были обработаны во входном запросе.
cacheTokensDetails[]object ( ModalityTokenCount )Только вывод. Список вариантов содержимого, кэшированного в запросе.
candidatesTokensDetails[]object ( ModalityTokenCount )Только вывод. Список модальностей, которые были возвращены в ответе.
toolUsePromptTokensDetails[]object ( ModalityTokenCount )Только выходные данные. Список модальностей, обработанных для обработки запросов на использование инструмента.
serviceTierenum ( ServiceTier )Только вывод. Уровень обслуживания запроса.
| JSON-представление |
|---|
{ "promptTokenCount": integer, "cachedContentTokenCount": integer, "candidatesTokenCount": integer, "toolUsePromptTokenCount": integer, "thoughtsTokenCount": integer, "totalTokenCount": integer, "promptTokensDetails": [ { object ( |
ModelStatus
Статус базовой модели. Используется для обозначения стадии развития базовой модели и времени вывода из эксплуатации, если таковое имеется.
modelStageenum ( ModelStage )Этап базовой модели.
retirementTimestring ( Timestamp format)Время, когда модель будет выведена из эксплуатации.
Используется RFC 3339, согласно которому генерируемый вывод всегда будет Z-нормализован и будет содержать 0, 3, 6 или 9 дробных знаков. Допускаются также смещения, отличные от "Z". Примеры: "2014-10-02T15:01:23Z" , "2014-10-02T15:01:23.045123456Z" или "2014-10-02T15:01:23+05:30" .
string messageСообщение с пояснением статуса модели.
| JSON-представление |
|---|
{
"modelStage": enum ( |
ModelStage
Определяет стадию развития базовой модели.
| Перечисления | |
|---|---|
MODEL_STAGE_UNSPECIFIED | Неуказанная стадия модели. |
UNSTABLE_EXPERIMENTAL | Базовая модель подвергается множеству настроек. |
EXPERIMENTAL | Модели на данном этапе предназначены исключительно для экспериментальных целей. |
PREVIEW | Модели на этом этапе более зрелые, чем экспериментальные модели. |
STABLE | Модели на этом этапе считаются стабильными и готовыми к использованию в производственных целях. |
LEGACY | Если модель находится на этой стадии, это означает, что в ближайшем будущем она будет снята с поддержки. Использовать эту модель смогут только существующие клиенты. |
DEPRECATED | Модели на этом этапе устарели. Использовать эти модели нельзя. |
RETIRED | Модели на этом этапе сняты с производства. Эти модели использовать нельзя. |
Кандидат
- JSON-представление
- FinishReason
- Атрибуция заземления
- AttributionSourceId
- GroundingPassageId
- SemanticRetrieverChunk
- Метаданные заземления
- SearchEntryPoint
- GroundingChunk
- Веб
- Изображение
- Полученный контекст
- Пользовательские метаданные
- StringList
- Карты
- PlaceAnswerSources
- ReviewSnippet
- Поддержка заземления
- Сегмент
- RetrievalMetadata
- LogprobsResult
- Лучшие кандидаты
- Кандидат
- UrlContextMetadata
- UrlMetadata
- UrlRetrievalStatus
Вариант ответа, сгенерированный на основе модели.
contentobject ( Content )Только выходные данные. Сгенерированный контент, возвращаемый моделью.
finishReasonenum ( FinishReason )Необязательно. Только для вывода. Причина, по которой модель перестала генерировать токены.
Если поле пустое, модель не прекращает генерацию токенов.
safetyRatings[]object ( SafetyRating )Список оценок безопасности кандидатов на должность в оперативно-розыскной группе.
В каждой категории может быть не более одной оценки.
citationMetadataobject ( CitationMetadata )Только выходные данные. Информация об источнике информации для кандидата, сгенерированного моделью.
Это поле может быть заполнено информацией о декламации любого текста, включенного в content . Речь идет о отрывках, которые «декламируются» из материалов, защищенных авторским правом, в обучающих данных базовой магистерской программы.
tokenCountintegerТолько вывод. Количество токенов для этого кандидата.
groundingAttributions[]object ( GroundingAttribution )Только выходные данные. Информация об источниках, которые способствовали получению обоснованного ответа.
Это поле заполняется для звонков, GenerateAnswer .
groundingMetadataobject ( GroundingMetadata )Только выходные данные. Метаданные для подтверждения данных кандидата.
Это поле заполняется для вызовов GenerateContent .
avgLogprobsnumberТолько выходные данные. Средний логарифмический показатель вероятности кандидата.
logprobsResultobject ( LogprobsResult )Только вывод. Значения логарифмической функции правдоподобия для токенов ответа и лучших токенов.
object ( UrlContextMetadata ) urlContextMetadataТолько выходные данные. Метаданные, относящиеся к инструменту получения контекста URL.
indexintegerТолько вывод. Индекс кандидата в списке кандидатов на ответ.
string finishMessage Необязательно. Только для вывода. Подробно описывает причину, по которой модель перестала генерировать токены. Заполняется только при установке finishReason .
| JSON-представление |
|---|
{ "content": { object ( |
FinishReason
Указывает причину, по которой модель перестала генерировать токены.
| Перечисления | |
|---|---|
FINISH_REASON_UNSPECIFIED | Значение по умолчанию. Это значение не используется. |
STOP | Естественная точка остановки модели или заданная последовательность остановок. |
MAX_TOKENS | Достигнуто максимальное количество токенов, указанное в запросе. |
SAFETY | Содержимое предложенного варианта ответа было помечено как потенциально опасное по соображениям безопасности. |
RECITATION | Содержание предложенного варианта ответа было помечено как требующее заучивания наизусть. |
LANGUAGE | В предложенном варианте ответа было обнаружено использование неподдерживаемого языка. |
OTHER | Причина неизвестна. |
BLOCKLIST | Генерация токенов прекращена, поскольку контент содержит запрещенные термины. |
PROHIBITED_CONTENT | Выпуск токенов приостановлен из-за потенциального наличия запрещенного контента. |
SPII | Генерация токенов прекращена, поскольку содержимое потенциально содержит конфиденциальную личную информацию (SPII). |
MALFORMED_FUNCTION_CALL | Вызов функции, сгенерированный моделью, является недопустимым. |
IMAGE_SAFETY | Генерация токенов прекращена, поскольку сгенерированные изображения содержат нарушения правил безопасности. |
IMAGE_PROHIBITED_CONTENT | Генерация изображений была остановлена, поскольку сгенерированные изображения содержали запрещенный контент. |
IMAGE_OTHER | Генерация изображений прекратилась из-за других различных проблем. |
NO_IMAGE | Предполагалось, что модель сгенерирует изображение, но изображение не было сгенерировано. |
IMAGE_RECITATION | Генерация изображений прекратилась из-за чтения вслух. |
UNEXPECTED_TOOL_CALL | Модель сгенерировала вызов инструмента, но ни один инструмент не был включен в запрос. |
TOO_MANY_TOOL_CALLS | Модель вызвала слишком много инструментов подряд, в результате чего система завершила выполнение. |
MISSING_THOUGHT_SIGNATURE | В запросе отсутствует как минимум одна подпись, выражающая мысль. |
MALFORMED_RESPONSE | Завершено из-за неправильной реакции. |
ESCALATION | Запрос был отфильтрован правилом эскалации. |
Атрибуция заземления
Укажите источник, который послужил основой для ответа.
sourceIdobject ( AttributionSourceId )Только вывод. Идентификатор источника, обеспечившего данное указание.
contentobject ( Content )Исходный контент, на основе которого составлена эта атрибуция.
| JSON-представление |
|---|
{ "sourceId": { object ( |
AttributionSourceId
Идентификатор источника, предоставившего эти данные.
sourceUnion typesource может быть только один из следующих вариантов: groundingPassageobject ( GroundingPassageId )Идентификатор для встроенного фрагмента текста.
object ( SemanticRetrieverChunk ) semanticRetrieverChunk Идентификатор Chunk , полученного с помощью семантического ретривера.
| JSON-представление |
|---|
{ // source "groundingPassage": { object ( |
GroundingPassageId
Идентификатор части внутри объекта GroundingPassage .
string passageId Только вывод. Идентификатор фрагмента текста, соответствующего свойству GroundingPassage.id из GenerateAnswerRequest .
partIndexinteger Только вывод. Индекс части внутри GroundingPassage.content объекта GenerateAnswerRequest .
| JSON-представление |
|---|
{ "passageId": string, "partIndex": integer } |
SemanticRetrieverChunk
Идентификатор Chunk , полученного с помощью семантического ретривера, указанный в GenerateAnswerRequest с использованием SemanticRetrieverConfig .
sourcestring Только вывод. Имя источника, соответствующее файлу SemanticRetrieverConfig.source запроса. Пример: corpora/123 или corpora/123/documents/abc
chunkstring Только вывод. Название Chunk , содержащего атрибутированный текст. Пример: corpora/123/documents/abc/chunks/xyz
| JSON-представление |
|---|
{ "source": string, "chunk": string } |
Метаданные заземления
Метаданные возвращаются клиенту при включении заземления.
groundingChunks[]object ( GroundingChunk )Список подтверждающих ссылок, полученных из указанного источника данных. При потоковой передаче он содержит только те фрагменты данных, которые не были включены в метаданные предыдущих ответов.
groundingSupports[]object ( GroundingSupport )Список средств заземления.
webSearchQueries[]stringПоисковые запросы в интернете для последующего поиска.
imageSearchQueries[]stringПоисковые запросы по изображениям, используемые для приведения в соответствие с реальностью.
searchEntryPointobject ( SearchEntryPoint )Необязательно. Запись в поисковой выдаче Google для последующих веб-поисков.
retrievalMetadataobject ( RetrievalMetadata )Метаданные, связанные с извлечением данных в процессе заземления.
string googleMapsWidgetContextTokenНеобязательно. Имя ресурса контекстного токена виджета Google Maps, который можно использовать с виджетом PlacesContextElement для отображения контекстных данных. Заполняется только в том случае, если включена привязка к карте с помощью Google Maps.
| JSON-представление |
|---|
{ "groundingChunks": [ { object ( |
SearchEntryPoint
Точка входа в поисковую выдачу Google.
renderedContentstringНеобязательный элемент. Фрагмент веб-контента, который можно встроить в веб-страницу или веб-представление приложения.
sdkBlobstring ( bytes format)Необязательно. JSON-данные в кодировке Base64, представляющие собой массив кортежей <поисковый запрос, URL поиска>.
Строка, закодированная в формате Base64.
| JSON-представление |
|---|
{ "renderedContent": string, "sdkBlob": string } |
GroundingChunk
GroundingChunk представляет собой сегмент подтверждающих данных, который служит основой для ответа модели. Это может быть фрагмент из интернета, контекст, полученный из файла, или информация из Google Maps.
chunk_typeUnion typechunk_type может быть только одним из следующих:webobject ( Web )Фрагмент веб-страницы, предназначенный для закрепления на месте.
imageobject ( Image )Необязательно. Фрагмент для определения местоположения из результатов поиска изображений.
retrievedContextobject ( RetrievedContext )Необязательно. Фрагмент контекста, полученный с помощью инструмента поиска файлов.
mapsobject ( Maps )Необязательно. Фрагмент карты местности из Google Maps.
| JSON-представление |
|---|
{ // chunk_type "web": { object ( |
Веб
Фрагмент из интернета.
string uriТолько вывод. URI-ссылка на фрагмент.
string titleТолько вывод. Заголовок фрагмента.
| JSON-представление |
|---|
{ "uri": string, "title": string } |
Изображение
Фрагмент из поиска изображений.
string sourceUriURI веб-страницы для указания источника.
string imageUriURL-адрес графического ресурса.
string titleЗаголовок веб-страницы, с которой взято изображение.
string domainКорневой домен веб-страницы, с которой взято изображение, например, "example.com".
| JSON-представление |
|---|
{ "sourceUri": string, "imageUri": string, "title": string, "domain": string } |
Полученный контекст
Фрагмент из контекста, полученный с помощью инструмента поиска по файлу.
customMetadata[]object ( CustomMetadata )Необязательно. Предоставляемые пользователем метаданные о полученном контексте.
string uriНеобязательно. URI-ссылка на документ для семантического поиска.
string titleНеобязательно. Заголовок документа.
textstringНеобязательно. Текст фрагмента.
fileSearchStorestring Необязательно. Название FileSearchStore содержащего документ. Пример: fileSearchStores/123
pageNumberintegerНеобязательно. Номер страницы полученного контекста, если применимо.
mediaIdstringНеобязательно. Имя ресурса медиа-объекта для результатов многомодального поиска файлов. Формат: fileSearchStores/{file_search_store_id}/media/{blobId}
| JSON-представление |
|---|
{
"customMetadata": [
{
object ( |
Пользовательские метаданные
Пользователь предоставил метаданные о GroundingFact.
keystringКлюч метаданных.
valueUnion typevalue может принимать только одно из следующих значений: stringValuestringНеобязательный параметр. Строковое значение метаданных.
stringListValueobject ( StringList )Необязательно. Список строковых значений для метаданных.
numericValuenumber Необязательно. Числовое значение метаданных. Ожидаемый диапазон значений зависит от используемого key .
| JSON-представление |
|---|
{
"key": string,
// value
"stringValue": string,
"stringListValue": {
object ( |
StringList
Список строковых значений.
values[]stringСтроковые значения списка.
| JSON-представление |
|---|
{ "values": [ string ] } |
Карты
Фрагмент карты Google Maps, соответствующий одному месту.
string uriURI-ссылка на это место.
string titleНазвание места.
textstringТекстовое описание места, дающего ответ.
placeIdstring Идентификатор места в формате places/{placeId} . Пользователь может использовать этот идентификатор для поиска данного места.
placeAnswerSourcesobject ( PlaceAnswerSources )Источники, предоставляющие ответы на вопросы об особенностях того или иного места на Google Maps.
| JSON-представление |
|---|
{
"uri": string,
"title": string,
"text": string,
"placeId": string,
"placeAnswerSources": {
object ( |
PlaceAnswerSources
Коллекция источников, предоставляющих ответы на вопросы об особенностях конкретного места на Google Maps. Каждое сообщение PlaceAnswerSources соответствует определенному месту на Google Maps. Инструмент Google Maps использовал эти источники для ответа на вопросы об особенностях места (например: «Есть ли Wi-Fi в Bar Foo?» или «Доступен ли бар Foo для инвалидов-колясочников?»). В настоящее время мы поддерживаем в качестве источников только фрагменты отзывов.
reviewSnippets[]object ( ReviewSnippet )Фрагменты отзывов, используемые для генерации ответов об особенностях того или иного места в Google Maps.
| JSON-представление |
|---|
{
"reviewSnippets": [
{
object ( |
ReviewSnippet
Представляет собой фрагмент отзыва пользователя, отвечающего на вопрос об особенностях конкретного места на Google Maps.
reviewIdstringИдентификатор фрагмента отзыва.
string googleMapsUriСсылка, соответствующая отзыву пользователя на Google Maps.
string titleЗаголовок рецензии.
| JSON-представление |
|---|
{ "reviewId": string, "googleMapsUri": string, "title": string } |
Поддержка заземления
Поддержка заземления.
groundingChunkIndices[]integer Необязательно. Список индексов (в 'grounding_chunk' в response.candidate.grounding_metadata ), указывающий на цитаты, связанные с утверждением. Например, [1,3,4] означает, что grounding_chunk[1], grounding_chunk[3], grounding_chunk[4] — это полученный контент, относящийся к утверждению. Если ответ потоковый, groundingChunkIndices ссылаются на индексы во всех ответах. Клиент несет ответственность за накопление фрагментов данных из всех ответов (с сохранением того же порядка).
confidenceScores[]numberНеобязательно. Показатель достоверности ссылок на источники поддержки. Диапазон от 0 до 1. 1 означает наивысший уровень достоверности. Этот список должен иметь тот же размер, что и groundingChunkIndices.
renderedParts[]integer Только для вывода. Индексы в поле parts содержимого кандидата. Эти индексы указывают, какие отрендеренные части связаны с данным источником поддержки.
segmentobject ( Segment )Данный раздел контента относится к данной поддержке.
| JSON-представление |
|---|
{
"groundingChunkIndices": [
integer
],
"confidenceScores": [
number
],
"renderedParts": [
integer
],
"segment": {
object ( |
Сегмент
Фрагмент контента.
partIndexintegerИндекс объекта Part внутри родительского объекта Content.
startIndexintegerНачальный индекс в заданной части, измеряемый в байтах. Смещение от начала части включительно, начиная с нуля.
endIndexintegerКонечный индекс в заданной части, измеряемый в байтах. Смещение от начала части, исключая его, начиная с нуля.
textstringТекст, соответствующий фрагменту ответа.
| JSON-представление |
|---|
{ "partIndex": integer, "startIndex": integer, "endIndex": integer, "text": string } |
RetrievalMetadata
Метаданные, связанные с извлечением данных в процессе заземления.
googleSearchDynamicRetrievalScorenumberНеобязательный параметр. Оценка, указывающая на вероятность того, что информация из поиска Google поможет ответить на вопрос. Оценка находится в диапазоне [0, 1], где 0 — наименее вероятный ответ, а 1 — наиболее вероятный. Эта оценка заполняется только при включенной функции сопоставления с поиском Google и динамического поиска. Она будет сравниваться с пороговым значением для определения необходимости запуска поиска Google.
| JSON-представление |
|---|
{ "googleSearchDynamicRetrievalScore": number } |
LogprobsResult
Результат Logprobs
topCandidates[]object ( TopCandidates )Длина = общее количество шагов декодирования.
chosenCandidates[]object ( Candidate )Длина = общее количество шагов декодирования. Выбранные кандидаты могут входить или не в число лучших кандидатов.
logProbabilitySumnumberСумма логарифмических вероятностей для всех токенов.
| JSON-представление |
|---|
{ "topCandidates": [ { object ( |
Лучшие кандидаты
Кандидаты с наивысшими логарифмическими вероятностями на каждом этапе декодирования.
candidates[]object ( Candidate )Отсортировано по логарифмической вероятности в порядке убывания.
| JSON-представление |
|---|
{
"candidates": [
{
object ( |
Кандидат
Кандидат на получение токена и оценки logprobs.
string tokenСтроковое значение токена кандидата.
tokenIdintegerИдентификатор токена кандидата.
logProbabilitynumberЛогарифмическая вероятность кандидата.
| JSON-представление |
|---|
{ "token": string, "tokenId": integer, "logProbability": number } |
UrlContextMetadata
Метаданные, относящиеся к инструменту получения контекста URL-адреса.
urlMetadata[]object ( UrlMetadata )Список контекста URL-адреса.
| JSON-представление |
|---|
{
"urlMetadata": [
{
object ( |
UrlMetadata
Контекст получения одной и той же URL-ссылки.
retrievedUrlstringURL-адрес получен инструментом.
urlRetrievalStatusenum ( UrlRetrievalStatus )Статус получения URL-адреса.
| JSON-представление |
|---|
{
"retrievedUrl": string,
"urlRetrievalStatus": enum ( |
UrlRetrievalStatus
Статус получения URL-адреса.
| Перечисления | |
|---|---|
URL_RETRIEVAL_STATUS_UNSPECIFIED | Значение по умолчанию. Это значение не используется. |
URL_RETRIEVAL_STATUS_SUCCESS | Получение URL-адреса прошло успешно. |
URL_RETRIEVAL_STATUS_ERROR | Получение URL-адреса не удалось из-за ошибки. |
URL_RETRIEVAL_STATUS_PAYWALL | Не удалось получить URL-адрес, поскольку контент находится за платным доступом. |
URL_RETRIEVAL_STATUS_UNSAFE | Не удалось получить URL-адрес, поскольку содержимое небезопасно. |
Метаданные цитирования
Подборка ссылок на источники для данного контента.
citationSources[]object ( CitationSource )Ссылки на источники для конкретного ответа.
| JSON-представление |
|---|
{
"citationSources": [
{
object ( |
Источник цитаты
Ссылка на источник для части конкретного ответа.
startIndexintegerНеобязательно. Начало сегмента ответа, который относится к данному источнику.
Индекс указывает начало сегмента, измеряемое в байтах.
endIndexintegerНеобязательно. Конец атрибутированного сегмента, исключая.
string uriНеобязательно. URI, указанный в качестве источника части текста.
string licenseНеобязательно. Лицензия на проект GitHub, указанный в качестве источника для данного сегмента.
Информация о лицензии необходима для цитирования в кодексе.
| JSON-представление |
|---|
{ "startIndex": integer, "endIndex": integer, "uri": string, "license": string } |
Категория вреда
Категория рейтинга.
Эти категории охватывают различные виды вреда, которые разработчики могут пожелать исправить.
| Перечисления | |
|---|---|
HARM_CATEGORY_UNSPECIFIED | Категория не указана. |
HARM_CATEGORY_DEROGATORY | PaLM — Негативные или оскорбительные комментарии, направленные против личности и/или защищаемых атрибутов. |
HARM_CATEGORY_TOXICITY | PaLM — Контент, содержащий грубость, неуважение или нецензурную лексику. |
HARM_CATEGORY_VIOLENCE | PaLM — описывает сценарии, изображающие насилие в отношении отдельного человека или группы лиц, или содержит общие описания кровавых сцен. |
HARM_CATEGORY_SEXUAL | PaLM - Содержит упоминания о сексуальных действиях или другом непристойном контенте. |
HARM_CATEGORY_MEDICAL | PaLM — пропагандирует непроверенные медицинские рекомендации. |
HARM_CATEGORY_DANGEROUS | PaLM — Опасный контент, который пропагандирует, способствует или поощряет совершение вредоносных действий. |
HARM_CATEGORY_HARASSMENT | Близнецы - Контент, содержащий элементы домогательств. |
HARM_CATEGORY_HATE_SPEECH | Близнецы — Ненавистнические высказывания и контент. |
HARM_CATEGORY_SEXUALLY_EXPLICIT | Близнецы — Содержит материалы откровенно сексуального характера. |
HARM_CATEGORY_DANGEROUS_CONTENT | Близнецы — Опасный контент. |
HARM_CATEGORY_CIVIC_INTEGRITY | Gemini — Контент, который может быть использован для нанесения вреда гражданской целостности. УСТАРЕВШЕЕ: используйте enableEnhancedCivicAnswers вместо этого. |
HARM_CATEGORY_JAILBREAK | Близнецы — Подсказки, пытающиеся обойти или подорвать правила безопасности модели (попытки взлома). |
ModalityTokenCount
Представляет информацию о количестве токенов для одной модальности.
modalityenum ( Modality )Способ действия, связанный с этим количеством токенов.
tokenCountintegerКоличество токенов.
| JSON-представление |
|---|
{
"modality": enum ( |
Модальность
Модальность части содержимого
| Перечисления | |
|---|---|
MODALITY_UNSPECIFIED | Неуказанный способ применения. |
TEXT | Простой текст. |
IMAGE | Изображение. |
VIDEO | Видео. |
AUDIO | Аудио. |
DOCUMENT | Документ, например, PDF. |
Рейтинг безопасности
Рейтинг безопасности для данного контента.
Рейтинг безопасности содержит категорию вреда и уровень вероятности вреда в этой категории для данного контента. Контент классифицируется по безопасности по ряду категорий вреда, и здесь указывается вероятность вреда, соответствующая данной классификации.
categoryenum ( HarmCategory )Обязательно. Категория для данной оценки.
probabilityenum ( HarmProbability )Обязательно. Вероятность причинения вреда данному контенту.
blockedbooleanБыл ли этот контент заблокирован из-за этого рейтинга?
| JSON-представление |
|---|
{ "category": enum ( |
Вероятность вреда
Вероятность того, что тот или иной контент является вредным.
Система классификации показывает вероятность того, что контент небезопасен. Это не указывает на степень опасности, которую может представлять тот или иной контент.
| Перечисления | |
|---|---|
HARM_PROBABILITY_UNSPECIFIED | Вероятность не указана. |
NEGLIGIBLE | Вероятность того, что контент окажется небезопасным, ничтожно мала. |
LOW | Контент имеет низкую вероятность быть небезопасным. |
MEDIUM | Контент имеет среднюю вероятность быть небезопасным. |
HIGH | Контент с высокой вероятностью может быть небезопасным. |
Настройки безопасности
Настройки безопасности, влияющие на поведение блокировки безопасности.
Передача параметра безопасности для категории изменяет допустимую вероятность блокировки контента.
categoryenum ( HarmCategory )Обязательно. Категория для данной настройки.
thresholdenum ( HarmBlockThreshold )Обязательно. Контролирует пороговое значение вероятности, при котором предотвращается причинение вреда.
| JSON-представление |
|---|
{ "category": enum ( |
Порог блокировки вреда
Блокировать при уровне вероятности причинения вреда и выше.
| Перечисления | |
|---|---|
HARM_BLOCK_THRESHOLD_UNSPECIFIED | Пороговое значение не указано. |
BLOCK_LOW_AND_ABOVE | Контент со словом «НЕЗНАЧИТЕЛЬНО» будет разрешен. |
BLOCK_MEDIUM_AND_ABOVE | Допускается контент с оценками "НЕЗНАЧИТЕЛЬНО" и "НИЗКО". |
BLOCK_ONLY_HIGH | Допускается контент со значком «НЕЗНАЧИТЕЛЬНЫЙ», «НИЗКИЙ» и «СРЕДНИЙ». |
BLOCK_NONE | Разрешается размещать любой контент. |
OFF | Отключите защитный фильтр. |
Уровень обслуживания
Уровень обслуживания запроса.
| Перечисления | |
|---|---|
unspecified | Уровень обслуживания по умолчанию, который является стандартным. |
standard | Стандартный уровень обслуживания. |
flex | Гибкий уровень обслуживания. |
priority | Приоритетный уровень обслуживания. |
Содержание
- JSON-представление
- Часть
- Клякса
- Вызов функции
- ФункцияОтвет
- FunctionResponsePart
- FunctionResponseBlob
- Планирование
- FileData
- Исполняемый код
- Язык
- Результат выполнения кода
- Исход
- Вызов инструмента
- Тип инструмента
- ToolResponse
- Видеометаданные
- MediaResolution
- Уровень
- Обработка медиафайлов
Базовый структурированный тип данных, содержащий составное содержимое сообщения.
Объект Content включает поле role , указывающее на создателя Content , и поле « parts , содержащее многокомпонентные данные, включающие содержимое оборота сообщения.
parts[]object ( Part ) Упорядоченные Parts , составляющие единое сообщение. Части могут иметь разные MIME-типы.
string roleНеобязательно. Автор контента. Должен быть либо «пользователь», либо «модель».
Этот параметр полезен для многоэтапных диалогов, в противном случае его можно оставить пустым или не задавать.
| JSON-представление |
|---|
{
"parts": [
{
object ( |
Часть
Тип данных, содержащий медиафайлы, являющиеся частью многокомпонентного сообщения Content .
Part состоит из данных, имеющих связанный с ними тип данных. Part может содержать только один из допустимых типов в Part.data .
Если поле inlineData заполнено необработанными байтами, то для Part должен быть указан фиксированный MIME-тип IANA, определяющий тип и подтип носителя.
boolean thoughtНеобязательный параметр. Указывает, была ли деталь разработана на основе модели.
thoughtSignaturestring ( bytes format)Необязательно. Непрозрачная подпись для мысли, чтобы ее можно было использовать в последующих запросах.
Строка, закодированная в формате Base64.
partMetadataobject ( Struct format)Пользовательские метаданные, связанные с частью. Агентам, использующим genai.Part в качестве представления контента, может потребоваться отслеживать дополнительную информацию. Например, это может быть имя файла/источника, из которого происходит часть, или способ мультиплексирования нескольких потоков частей.
mediaResolutionobject ( MediaResolution )Необязательно. Разрешение входного медиафайла.
mediaProcessingenum ( MediaProcessing ) Необязательно. Способ обработки медиафайлов этой части модели для понимания. Имеет значение только для видеочастей ( inlineData или fileData с MIME-типом видео). Для частей, не содержащих видео, это поле игнорируется.
Union type datadata могут быть только одного из следующих типов:textstringВстроенный текст.
inlineDataobject ( Blob )Встроенные медиабайты.
functionCallobject ( FunctionCall ) Модель возвращает прогнозируемый FunctionCall , содержащий строку, представляющую FunctionDeclaration.name , а также аргументы и их значения.
functionResponseobject ( FunctionResponse ) Результат вызова FunctionCall , содержащий строку, представляющую FunctionDeclaration.name , и структурированный JSON-объект, содержащий любой вывод функции, используется в качестве контекста для модели.
fileDataobject ( FileData )Данные на основе URI.
executableCodeobject ( ExecutableCode )Код, сгенерированный моделью, предназначен для выполнения.
codeExecutionResultobject ( CodeExecutionResult ) Результат выполнения ExecutableCode кода.
toolCallobject ( ToolCall )Вызов инструмента на стороне сервера. Это поле заполняется, когда модель предсказывает вызов инструмента, который должен быть выполнен на сервере. Ожидается, что клиент отправит это сообщение обратно в API.
toolResponseobject ( ToolResponse ) Результат выполнения ToolCall на стороне сервера. Это поле заполняется клиентом результатами выполнения соответствующего ToolCall .
metadataUnion typemetadata могут быть только одним из следующих типов:videoMetadataobject ( VideoMetadata )Необязательно. Метаданные видео. Метаданные следует указывать только в том случае, если видеоданные представлены в формате inlineData или fileData.
| JSON-представление |
|---|
{ "thought": boolean, "thoughtSignature": string, "partMetadata": { object }, "mediaResolution": { object ( |
Клякса
Необработанные медиабайты.
Текст не следует отправлять в виде необработанных байтов, используйте поле 'text'.
mimeTypestringСтандартный MIME-тип исходных данных IANA. Примеры поддерживаемых типов: - Изображения: image/png, image/jpeg, image/jpg, image/webp, image/heic, image/heif, image/gif, image/avif - Аудио: audio/*, video/audio/s16le, video/audio/wav - Видео: video/* - Текст: text/plain, text/html, text/css, text/javascript, text/x-typescript, text/csv, text/markdown, text/x-python, text/xml, text/rtf, video/text/timestamp - Приложения: application/x-javascript, application/x-typescript, application/x-python-code, application/json, application/x-ipynb+json, application/rtf, application/pdf Для получения дополнительной информации см. раздел « Поддерживаемые форматы файлов ».
datastring ( bytes format)Необработанные байты для медиаформатов.
Строка, закодированная в формате Base64.
| JSON-представление |
|---|
{ "mimeType": string, "data": string } |
Вызов функции
Модель возвращает прогнозируемый FunctionCall , содержащий строку, представляющую FunctionDeclaration.name , а также аргументы и их значения.
string id Необязательный параметр. Уникальный идентификатор вызова функции. Если он заполнен, указывается клиент, которому следует выполнить functionCall и вернуть ответ с соответствующим id .
string nameRequired. The name of the function to call. Must be az, AZ, 0-9, or contain underscores and dashes, with a maximum length of 128.
argsobject ( Struct format)Optional. The function parameters and values in JSON object format.
| JSON-представление |
|---|
{ "id": string, "name": string, "args": { object } } |
FunctionResponse
The result output from a FunctionCall that contains a string representing the FunctionDeclaration.name and a structured JSON object containing any output from the function is used as context to the model. This should contain the result of a FunctionCall made based on model prediction.
idstring Optional. The identifier of the function call this response is for. Populated by the client to match the corresponding function call id .
string nameRequired. The name of the function to call. Must be az, AZ, 0-9, or contain underscores and dashes, with a maximum length of 128.
responseobject ( Struct format)Required. The function response in JSON object format. Callers can use any keys of their choice that fit the function's syntax to return the function output, eg "output", "result", etc. In particular, if the function call failed to execute, the response can have an "error" key to return error details to the model.
Multimedia can be included by using a subobject containing a single "$ref" key whose value is the inlineData.display_name of a FunctionResponsePart holding the multimedia. See https://ai.google.dev/gemini-api/docs/function-calling#multimodal .
parts[]object ( FunctionResponsePart ) Optional. Ordered Parts that constitute a function response. Parts may have different IANA MIME types.
willContinueboolean Optional. Signals that function call continues, and more responses will be returned, turning the function call into a generator. Is only applicable to NON_BLOCKING function calls, is ignored otherwise. If set to false, future responses will not be considered. It is allowed to return empty response with willContinue=False to signal that the function call is finished. This may still trigger the model generation. To avoid triggering the generation and finish the function call, additionally set scheduling to SILENT .
schedulingenum ( Scheduling )Optional. Specifies how the response should be scheduled in the conversation. Only applicable to NON_BLOCKING function calls, is ignored otherwise. Defaults to WHEN_IDLE.
| JSON-представление |
|---|
{ "id": string, "name": string, "response": { object }, "parts": [ { object ( |
FunctionResponsePart
A datatype containing media that is part of a FunctionResponse message.
A FunctionResponsePart consists of data which has an associated datatype. A FunctionResponsePart can only contain one of the accepted types in FunctionResponsePart.data .
A FunctionResponsePart must have a fixed IANA MIME type identifying the type and subtype of the media if the inlineData field is filled with raw bytes.
dataUnion typedata can be only one of the following: inlineDataobject ( FunctionResponseBlob )Inline media bytes.
| JSON-представление |
|---|
{
// data
"inlineData": {
object ( |
FunctionResponseBlob
Raw media bytes for function response.
Text should not be sent as raw bytes, use the 'FunctionResponse.response' field.
mimeTypestringThe IANA standard MIME type of the source data. Examples: - image/png - image/jpeg If an unsupported MIME type is provided, an error will be returned. For a complete list of supported types, see Supported file formats .
datastring ( bytes format)Raw bytes for media formats.
Строка, закодированная в формате Base64.
| JSON-представление |
|---|
{ "mimeType": string, "data": string } |
Планирование
Specifies how the response should be scheduled in the conversation.
| Перечисления | |
|---|---|
SCHEDULING_UNSPECIFIED | This value is unused. |
SILENT | Only add the result to the conversation context, do not interrupt or trigger generation. |
WHEN_IDLE | Add the result to the conversation context, and prompt to generate output without interrupting ongoing generation. |
INTERRUPT | Add the result to the conversation context, interrupt ongoing generation and prompt to generate output. |
FileData
URI based data.
mimeTypestringOptional. The IANA standard MIME type of the source data.
fileUristringRequired. URI.
| JSON-представление |
|---|
{ "mimeType": string, "fileUri": string } |
ExecutableCode
Code generated by the model that is meant to be executed, and the result returned to the model.
Only generated when using the CodeExecution tool, in which the code will be automatically executed, and a corresponding CodeExecutionResult will also be generated.
idstring Optional. Unique identifier of the ExecutableCode part. The server returns the CodeExecutionResult with the matching id .
languageenum ( Language ) Required. Programming language of the code .
codestringRequired. The code to be executed.
| JSON-представление |
|---|
{
"id": string,
"language": enum ( |
Язык
Supported programming languages for the generated code.
| Перечисления | |
|---|---|
LANGUAGE_UNSPECIFIED | Unspecified language. This value should not be used. |
PYTHON | Python >= 3.10, with numpy and simpy available. Python is the default language. |
CodeExecutionResult
Result of executing the ExecutableCode .
Generated only when the CodeExecution tool is used.
idstring Optional. The identifier of the ExecutableCode part this result is for. Only populated if the corresponding ExecutableCode has an id.
outcomeenum ( Outcome )Required. Outcome of the code execution.
outputstringOptional. Contains stdout when code execution is successful, stderr or other description otherwise.
| JSON-представление |
|---|
{
"id": string,
"outcome": enum ( |
Исход
Enumeration of possible outcomes of the code execution.
| Перечисления | |
|---|---|
OUTCOME_UNSPECIFIED | Unspecified status. This value should not be used. |
OUTCOME_OK | Code execution completed successfully. output contains the stdout, if any. |
OUTCOME_FAILED | Code execution failed. output contains the stderr and stdout, if any. |
OUTCOME_DEADLINE_EXCEEDED | Code execution ran for too long, and was cancelled. There may or may not be a partial output present. |
ToolCall
A predicted server-side ToolCall returned from the model. This message contains information about a tool that the model wants to invoke. The client is NOT expected to execute this ToolCall . Instead, the client should pass this ToolCall back to the API in a subsequent turn within a Content message, along with the corresponding ToolResponse .
idstring Optional. Unique identifier of the tool call. The server returns the tool response with the matching id .
toolNamestringOptional. The name of the tool that was called.
toolTypeenum ( ToolType )Required. The type of tool that was called.
argsobject ( Struct format)Optional. The tool call arguments. Example: {"arg1" : "value1", "arg2" : "value2" , ...}
| JSON-представление |
|---|
{
"id": string,
"toolName": string,
"toolType": enum ( |
ToolType
The type of tool in the function call.
| Перечисления | |
|---|---|
TOOL_TYPE_UNSPECIFIED | Unspecified tool type. |
GOOGLE_SEARCH_WEB | Google search tool, maps to Tool.google_search.search_types.web_search. |
GOOGLE_SEARCH_IMAGE | Image search tool, maps to Tool.google_search.search_types.image_search. |
URL_CONTEXT | URL context tool, maps to Tool.url_context. |
GOOGLE_MAPS | Google maps tool, maps to Tool.google_maps. |
FILE_SEARCH | File search tool, maps to Tool.file_search. |
ToolResponse
The output from a server-side ToolCall execution. This message contains the results of a tool invocation that was initiated by a ToolCall from the model. The client should pass this ToolResponse back to the API in a subsequent turn within a Content message, along with the corresponding ToolCall .
idstringOptional. The identifier of the tool call this response is for.
toolTypeenum ( ToolType ) Required. The type of tool that was called, matching the toolType in the corresponding ToolCall .
responseobject ( Struct format)Optional. The tool response.
| JSON-представление |
|---|
{
"id": string,
"toolType": enum ( |
VideoMetadata
Deprecated: Use GenerateContentRequest.processing_options instead. Metadata describes the input video content.
startOffsetstring ( Duration format)Optional. The start offset of the video.
Длительность в секундах, содержащая до девяти знаков после запятой, заканчивающаяся на « s ». Пример: "3.5s" .
endOffsetstring ( Duration format)Optional. The end offset of the video.
Длительность в секундах, содержащая до девяти знаков после запятой, заканчивающаяся на « s ». Пример: "3.5s" .
fpsnumberOptional. The frame rate of the video sent to the model. If not specified, the default value will be 1.0. The fps range is (0.0, 24.0].
| JSON-представление |
|---|
{ "startOffset": string, "endOffset": string, "fps": number } |
MediaResolution
Media resolution for tokenization.
valueUnion typevalue can be only one of the following:levelenum ( Level ) The tokenization quality used for given media.
| JSON-представление |
|---|
{
// value
"level": enum ( |
Уровень
The media resolution level.
| Перечисления | |
|---|---|
MEDIA_RESOLUTION_UNSPECIFIED | Media resolution has not been set. |
MEDIA_RESOLUTION_LOW | Media resolution set to low. |
MEDIA_RESOLUTION_MEDIUM | Media resolution set to medium. |
MEDIA_RESOLUTION_HIGH | Media resolution set to high. |
MEDIA_RESOLUTION_ULTRA_HIGH | Media resolution set to ultra high. |
MediaProcessing
How the model processes input media for understanding.
| Перечисления | |
|---|---|
MEDIA_PROCESSING_UNSPECIFIED | Default. Uses model-specific processing (3.5 Pro+ -> AGENTIC , older models -> STATIC ). |
STATIC | Fixed-rate frame extraction. All frames placed in context. |
AGENTIC | Model-driven dynamic navigation. Recommended for most use cases. |
Среда
An execution environment for an agent.
idstringRequired. Output only. The ID of the environment.
sources[]object ( Source )Sources to be mounted into the environment.
createdstringOutput only. The time at which the environment was created in ISO 8601 format (YYYY-MM-DDThh:mm:ssZ).
updatedstringOutput only. The time at which the environment was last updated in ISO 8601 format (YYYY-MM-DDThh:mm:ssZ).
lastAccessedstringOutput only. The time at which the environment was last accessed in ISO 8601 format (YYYY-MM-DDThh:mm:ssZ).
statusenum ( Status )Output only. The status of the environment container.
fileCountstring ( int64 format)Output only. The number of files in the environment, output only.
sizeBytesstring ( int64 format)Output only. The total size of the environment files in bytes, output only.
networkUnion typenetwork can be only one of the following:networkAllowlistobject ( EnvironmentNetworkEgressAllowlist )Allow only specific domains.
networkModeenum ( NetworkMode )Network egress mode.
| JSON-представление |
|---|
{ "id": string, "sources": [ { object ( |
Статус
Status of the environment.
| Перечисления | |
|---|---|
STATUS_UNSPECIFIED | |
ACTIVE | |
EXPIRED | |
NetworkMode
Network egress mode for non-allowlist configurations.
| Перечисления | |
|---|---|
NETWORK_MODE_UNSPECIFIED | Default value. Unused. |
DISABLED | All network egress is blocked. |
Схема
The Schema object allows the definition of input and output data types. These types can be objects, but also primitives and arrays. Represents a select subset of an OpenAPI 3.0 schema object .
typeenum ( Type )Required. Data type.
formatstringOptional. The format of the data. Any value is allowed, but most do not trigger any special functionality.
titlestringOptional. The title of the schema.
string descriptionOptional. A brief description of the parameter. This could contain examples of use. Parameter description may be formatted as Markdown.
nullablebooleanOptional. Indicates if the value may be null.
enum[]stringOptional. Possible values of the element of Type.STRING with enum format. For example we can define an Enum Direction as : {type:STRING, format:enum, enum:["EAST", NORTH", "SOUTH", "WEST"]}
maxItemsstring ( int64 format)Optional. Maximum number of the elements for Type.ARRAY.
minItemsstring ( int64 format)Optional. Minimum number of the elements for Type.ARRAY.
propertiesmap (key: string, value: object ( Schema ))Optional. Properties of Type.OBJECT.
An object containing a list of "key": value pairs. Example: { "name": "wrench", "mass": "1.3kg", "count": "3" } .
required[]stringOptional. Required properties of Type.OBJECT.
minPropertiesstring ( int64 format)Optional. Minimum number of the properties for Type.OBJECT.
maxPropertiesstring ( int64 format)Optional. Maximum number of the properties for Type.OBJECT.
minLengthstring ( int64 format)Optional. SCHEMA FIELDS FOR TYPE STRING Minimum length of the Type.STRING
maxLengthstring ( int64 format)Optional. Maximum length of the Type.STRING
patternstringOptional. Pattern of the Type.STRING to restrict a string to a regular expression.
examplevalue ( Value format)Optional. Example of the object. Will only populated when the object is the root.
anyOf[]object ( Schema )Optional. The value should be validated against any (one or more) of the subschemas in the list.
propertyOrdering[]stringOptional. The order of the properties. Not a standard field in open api spec. Used to determine the order of the properties in the response.
defaultvalue ( Value format) Optional. Default value of the field. Per JSON Schema, this field is intended for documentation generators and doesn't affect validation. Thus it's included here and ignored so that developers who send schemas with a default field don't get unknown-field errors.
itemsobject ( Schema )Optional. Schema of the elements of Type.ARRAY.
minimumnumberOptional. SCHEMA FIELDS FOR TYPE INTEGER and NUMBER Minimum value of the Type.INTEGER and Type.NUMBER
maximumnumberOptional. Maximum value of the Type.INTEGER and Type.NUMBER
| JSON-представление |
|---|
{ "type": enum ( |
Тип
Type contains the list of OpenAPI data types as defined by https://spec.openapis.org/oas/v3.0.3#data-types
| Перечисления | |
|---|---|
TYPE_UNSPECIFIED | Not specified, should not be used. |
STRING | String type. |
NUMBER | Number type. |
INTEGER | Integer type. |
BOOLEAN | Boolean type. |
ARRAY | Array type. |
OBJECT | Object type. |
NULL | Null type. |
Инструмент
- JSON-представление
- FunctionDeclaration
- Поведение
- GoogleSearchRetrieval
- DynamicRetrievalConfig
- Режим
- CodeExecution
- GoogleSearch
- Interval
- SearchTypes
- WebSearch
- ImageSearch
- ComputerUse
- Среда
- SafetyPolicy
- UrlContext
- FileSearch
- McpServer
- StreamableHttpTransport
- Google Карты
Tool details that the model may use to generate response.
A Tool is a piece of code that enables the system to interact with external systems to perform an action, or set of actions, outside of knowledge and scope of the model.
Next ID: 17
functionDeclarations[]object ( FunctionDeclaration ) Optional. A list of FunctionDeclarations available to the model that can be used for function calling.
The model or system does not execute the function. Instead the defined function may be returned as a FunctionCall with arguments to the client side for execution. The model may decide to call a subset of these functions by populating FunctionCall in the response. The next conversation turn may contain a FunctionResponse with the Content.role "function" generation context for the next model turn.
googleSearchRetrievalobject ( GoogleSearchRetrieval )Optional. Retrieval tool that is powered by Google search.
codeExecutionobject ( CodeExecution )Optional. Enables the model to execute code as part of generation.
googleSearchobject ( GoogleSearch )Optional. GoogleSearch tool type. Tool to support Google Search in Model. Powered by Google.
computerUseobject ( ComputerUse )Optional. Tool to support the model interacting directly with the computer. If enabled, it automatically populates computer-use specific Function Declarations.
urlContextobject ( UrlContext )Optional. Tool to support URL context retrieval.
fileSearchobject ( FileSearch )Optional. FileSearch tool type. Tool to retrieve knowledge from Semantic Retrieval corpora.
mcpServers[]object ( McpServer )Optional. MCP Servers to connect to.
googleMapsobject ( GoogleMaps )Optional. Tool that allows grounding the model's response with geospatial context related to the user's query.
| JSON-представление |
|---|
{ "functionDeclarations": [ { object ( |
FunctionDeclaration
Structured representation of a function declaration as defined by the OpenAPI 3.03 specification . Included in this declaration are the function name and parameters. This FunctionDeclaration is a representation of a block of code that can be used as a Tool by the model and executed by the client.
string nameRequired. The name of the function. Must be az, AZ, 0-9, or contain underscores, colons, dots, and dashes, with a maximum length of 128.
string descriptionRequired. A brief description of the function.
behaviorenum ( Behavior )Optional. Specifies the function Behavior. Currently only supported by the BidiGenerateContent method.
parametersobject ( Schema )Optional. Describes the parameters to this function. Reflects the Open API 3.03 Parameter Object string Key: the name of the parameter. Parameter names are case sensitive. Schema Value: the Schema defining the type used for the parameter.
parametersJsonSchemavalue ( Value format)Optional. Describes the parameters to the function in JSON Schema format. The schema must describe an object where the properties are the parameters to the function. For example:
{
"type": "object",
"properties": {
"name": { "type": "string" },
"age": { "type": "integer" }
},
"additionalProperties": false,
"required": ["name", "age"],
"propertyOrdering": ["name", "age"]
}
This field is mutually exclusive with parameters .
responseobject ( Schema )Optional. Describes the output from this function in JSON Schema format. Reflects the Open API 3.03 Response Object. The Schema defines the type used for the response value of the function.
responseJsonSchemavalue ( Value format)Optional. Describes the output from this function in JSON Schema format. The value specified by the schema is the response value of the function.
This field is mutually exclusive with response .
Поведение
Defines the function behavior. Defaults to BLOCKING .
| Перечисления | |
|---|---|
UNSPECIFIED | This value is unused. |
BLOCKING | If set, the system will wait to receive the function response before continuing the conversation. |
NON_BLOCKING | If set, the system will not wait to receive the function response. Instead, it will attempt to handle function responses as they become available while maintaining the conversation between the user and the model. |
GoogleSearchRetrieval
Tool to retrieve public web data for grounding, powered by Google.
dynamicRetrievalConfigobject ( DynamicRetrievalConfig )Specifies the dynamic retrieval configuration for the given source.
| JSON-представление |
|---|
{
"dynamicRetrievalConfig": {
object ( |
DynamicRetrievalConfig
Describes the options to customize dynamic retrieval.
modeenum ( Mode )The mode of the predictor to be used in dynamic retrieval.
dynamicThresholdnumberThe threshold to be used in dynamic retrieval. If not set, a system default value is used.
| JSON-представление |
|---|
{
"mode": enum ( |
Режим
The mode of the predictor to be used in dynamic retrieval.
| Перечисления | |
|---|---|
MODE_UNSPECIFIED | Always trigger retrieval. |
MODE_DYNAMIC | Run retrieval only when system decides it is necessary. |
CodeExecution
This type has no fields.
Tool that executes code generated by the model, and automatically returns the result to the model.
See also ExecutableCode and CodeExecutionResult which are only generated when using this tool.
GoogleSearch
GoogleSearch tool type. Tool to support Google Search in Model. Powered by Google.
timeRangeFilterobject ( Interval )Optional. Filter search results to a specific time range. If customers set a start time, they must set an end time (and vice versa).
searchTypesobject ( SearchTypes )Optional. The set of search types to enable. If not set, web search is enabled by default.
| JSON-представление |
|---|
{ "timeRangeFilter": { object ( |
Interval
Represents a time interval, encoded as a Timestamp start (inclusive) and a Timestamp end (exclusive).
The start must be less than or equal to the end. When the start equals the end, the interval is empty (matches no time). When both start and end are unspecified, the interval matches any time.
startTimestring ( Timestamp format)Optional. Inclusive start of the interval.
If specified, a Timestamp matching this interval will have to be the same or after the start.
Используется RFC 3339, согласно которому генерируемый вывод всегда будет Z-нормализован и будет содержать 0, 3, 6 или 9 дробных знаков. Допускаются также смещения, отличные от "Z". Примеры: "2014-10-02T15:01:23Z" , "2014-10-02T15:01:23.045123456Z" или "2014-10-02T15:01:23+05:30" .
endTimestring ( Timestamp format)Optional. Exclusive end of the interval.
If specified, a Timestamp matching this interval will have to be before the end.
Используется RFC 3339, согласно которому генерируемый вывод всегда будет Z-нормализован и будет содержать 0, 3, 6 или 9 дробных знаков. Допускаются также смещения, отличные от "Z". Примеры: "2014-10-02T15:01:23Z" , "2014-10-02T15:01:23.045123456Z" или "2014-10-02T15:01:23+05:30" .
| JSON-представление |
|---|
{ "startTime": string, "endTime": string } |
SearchTypes
Different types of search that can be enabled on the GoogleSearch tool.
webSearchobject ( WebSearch )Optional. Enables web search. Only text results are returned.
imageSearchobject ( ImageSearch )Optional. Enables image search. Image bytes are returned.
| JSON-представление |
|---|
{ "webSearch": { object ( |
WebSearch
This type has no fields.
Standard web search for grounding and related configurations.
ImageSearch
This type has no fields.
Image search for grounding and related configurations.
ComputerUse
Computer Use tool type.
environmentenum ( Environment )Required. The environment being operated.
excludedPredefinedFunctions[]stringOptional. By default, predefined functions are included in the final model call. Some of them can be explicitly excluded from being automatically included. This can serve two purposes: 1. Using a more restricted / different action space. 2. Improving the definitions / instructions of predefined functions.
enablePromptInjectionDetectionbooleanOptional. Whether enable the prompt injection detection check on computer-use request.
disabledSafetyPolicies[]enum ( SafetyPolicy )Optional. Disabled safety policies for computer use.
| JSON-представление |
|---|
{ "environment": enum ( |
Среда
Represents the environment being operated, such as a web browser.
| Перечисления | |
|---|---|
ENVIRONMENT_UNSPECIFIED | Defaults to browser. |
ENVIRONMENT_BROWSER | Operates in a web browser. |
ENVIRONMENT_MOBILE | Operates in a mobile environment. |
ENVIRONMENT_DESKTOP | Operates in a desktop environment. |
SafetyPolicy
Predefined safety policies for computer use.
| Перечисления | |
|---|---|
SAFETY_POLICY_UNSPECIFIED | Unspecified safety policy. |
FINANCIAL_TRANSACTIONS | Safety policy for financial transactions. |
SENSITIVE_DATA_MODIFICATION | Safety policy for sensitive data modification. |
COMMUNICATION_TOOL | Safety policy for communication tools (eg Gmail, Chat, Meet). |
ACCOUNT_CREATION | Safety policy for account creation. |
DATA_MODIFICATION | Safety policy for data modification. |
USER_CONSENT_MANAGEMENT | Safety policy for user consent management. |
LEGAL_TERMS_AND_AGREEMENTS | Safety policy for legal terms and agreements. |
UrlContext
This type has no fields.
Tool to support URL context retrieval.
FileSearch
The FileSearch tool that retrieves knowledge from Semantic Retrieval corpora. Files are imported to Semantic Retrieval corpora using the ImportFile API.
fileSearchStoreNames[]string Required. The names of the fileSearchStores to retrieve from. Example: fileSearchStores/my-file-search-store-123
metadataFilterstringOptional. Metadata filter to apply to the semantic retrieval documents and chunks.
topKintegerOptional. The number of semantic retrieval chunks to retrieve.
| JSON-представление |
|---|
{ "fileSearchStoreNames": [ string ], "metadataFilter": string, "topK": integer } |
McpServer
A MCPServer is a server that can be called by the model to perform actions. It is a server that implements the MCP protocol. Next ID: 6
string nameThe name of the MCPServer.
transportUnion typetransport can be only one of the following: streamableHttpTransportobject ( StreamableHttpTransport )A transport that can stream HTTP requests and responses.
| JSON-представление |
|---|
{
"name": string,
// transport
"streamableHttpTransport": {
object ( |
StreamableHttpTransport
A transport that can stream HTTP requests and responses. Next ID: 6
urlstringThe full URL for the MCPServer endpoint. Example: "https://api.example.com/mcp"
headersmap (key: string, value: string)Optional: Fields for authentication headers, timeouts, etc., if needed.
An object containing a list of "key": value pairs. Example: { "name": "wrench", "mass": "1.3kg", "count": "3" } .
timeoutstring ( Duration format)HTTP timeout for regular operations.
Длительность в секундах, содержащая до девяти знаков после запятой, заканчивающаяся на « s ». Пример: "3.5s" .
sseReadTimeoutstring ( Duration format)Timeout for SSE read operations.
Длительность в секундах, содержащая до девяти знаков после запятой, заканчивающаяся на « s ». Пример: "3.5s" .
terminateOnClosebooleanWhether to close the client session when the transport closes.
| JSON-представление |
|---|
{ "url": string, "headers": { string: string, ... }, "timeout": string, "sseReadTimeout": string, "terminateOnClose": boolean } |
Google Карты
The GoogleMaps Tool that provides geospatial context for the user's query.
enableWidgetbooleanOptional. Whether to return a widget context token in the GroundingMetadata of the response. Developers can use the widget context token to render a Google Maps widget with geospatial context related to the places that the model references in the response.
| JSON-представление |
|---|
{ "enableWidget": boolean } |
REST Resource: auth_tokens
- Resource: AuthToken
- BidiGenerateContentSetup
- GenerationConfig
- Модальность
- SpeechConfig
- VoiceConfig
- PrebuiltVoiceConfig
- MultiSpeakerVoiceConfig
- SpeakerVoiceConfig
- ThinkingConfig
- ThinkingLevel
- ImageConfig
- MediaResolution
- ResponseFormatConfig
- TextResponseFormat
- MimeType
- AudioResponseFormat
- MimeType
- Доставка
- ImageResponseFormat
- MimeType
- Доставка
- Соотношение сторон
- ImageSize
- TranslationConfig
- AudioTranscriptionConfig
- LanguageAuto
- LanguageHints
- RealtimeInputConfig
- AutomaticActivityDetection
- StartSensitivity
- EndSensitivity
- ActivityHandling
- TurnCoverage
- SessionResumptionConfig
- ContextWindowCompressionConfig
- SlidingWindow
- HistoryConfig
- Методы
Resource: AuthToken
A request to create an ephemeral authentication token.
string nameOutput only. Identifier. The token itself.
expireTimestring ( Timestamp format)Optional. Input only. Immutable. An optional time after which, when using the resulting token, messages in BidiGenerateContent sessions will be rejected. (Gemini may preemptively close the session after this time.)
If not set then this defaults to 30 minutes in the future. If set, this value must be less than 20 hours in the future.
Используется RFC 3339, согласно которому генерируемый вывод всегда будет Z-нормализован и будет содержать 0, 3, 6 или 9 дробных знаков. Допускаются также смещения, отличные от "Z". Примеры: "2014-10-02T15:01:23Z" , "2014-10-02T15:01:23.045123456Z" или "2014-10-02T15:01:23+05:30" .
newSessionExpireTimestring ( Timestamp format)Optional. Input only. Immutable. The time after which new Live API sessions using the token resulting from this request will be rejected.
If not set this defaults to 60 seconds in the future. If set, this value must be less than 20 hours in the future.
Используется RFC 3339, согласно которому генерируемый вывод всегда будет Z-нормализован и будет содержать 0, 3, 6 или 9 дробных знаков. Допускаются также смещения, отличные от "Z". Примеры: "2014-10-02T15:01:23Z" , "2014-10-02T15:01:23.045123456Z" или "2014-10-02T15:01:23+05:30" .
fieldMaskstring ( FieldMask format) Optional. Input only. Immutable. If fieldMask is empty, and bidiGenerateContentSetup is not present, then the effective BidiGenerateContentSetup message is taken from the Live API connection.
If fieldMask is empty, and bidiGenerateContentSetup is present, then the effective BidiGenerateContentSetup message is taken entirely from bidiGenerateContentSetup in this request. The setup message from the Live API connection is ignored.
If fieldMask is not empty, then the corresponding fields from bidiGenerateContentSetup will overwrite the fields from the setup message in the Live API connection.
This is a comma-separated list of fully qualified names of fields. Example: "user.displayName,photo" .
configUnion typeconfig can be only one of the following: bidiGenerateContentSetupobject ( BidiGenerateContentSetup ) Optional. Input only. Immutable. Configuration specific to BidiGenerateContent .
usesintegerOptional. Input only. Immutable. The number of times the token can be used. If this value is zero then no limit is applied. Resuming a Live API session does not count as a use. If unspecified, the default is 1.
| JSON-представление |
|---|
{
"name": string,
"expireTime": string,
"newSessionExpireTime": string,
"fieldMask": string,
// config
"bidiGenerateContentSetup": {
object ( |
BidiGenerateContentSetup
Message to be sent in the first (and only in the first) BidiGenerateContentClientMessage . Contains configuration that will apply for the duration of the streaming RPC.
Clients should wait for a BidiGenerateContentSetupComplete message before sending any additional messages.
modelstringRequired. The model's resource name. This serves as an ID for the Model to use.
Format: models/{model}
generationConfigobject ( GenerationConfig )Optional. Generation config.
The following fields are not supported:
-
responseLogprobs -
responseMimeType -
logprobs -
responseSchema -
responseJsonSchema -
stop_sequence -
skipResponseCache -
routing_config -
audio_timestamp
systemInstructionobject ( Content )Optional. The user provided system instructions for the model.
Note: Only text should be used in parts and content in each part will be in a separate paragraph.
tools[]object ( Tool ) Optional. A list of Tools the model may use to generate the next response.
A Tool is a piece of code that enables the system to interact with external systems to perform an action, or set of actions, outside of knowledge and scope of the model.
realtimeInputConfigobject ( RealtimeInputConfig )Optional. Configures the handling of realtime input.
sessionResumptionobject ( SessionResumptionConfig )Optional. Configures session resumption mechanism.
If included, the server will send SessionResumptionUpdate messages.
contextWindowCompressionobject ( ContextWindowCompressionConfig )Optional. Configures a context window compression mechanism.
If included, the server will automatically reduce the size of the context when it exceeds the configured length.
inputAudioTranscriptionobject ( AudioTranscriptionConfig )Optional. If set, enables transcription of voice input. The transcription aligns with the input audio language, if configured.
outputAudioTranscriptionobject ( AudioTranscriptionConfig )Optional. If set, enables transcription of the model's audio output. The transcription aligns with the language code specified for the output audio, if configured.
historyConfigobject ( HistoryConfig )Optional. Configures the exchange of history between the client and the server.
| JSON-представление |
|---|
{ "model": string, "generationConfig": { object ( |
GenerationConfig
Configuration options for model generation and outputs. Not all parameters are configurable for every model.
stopSequences[]string Optional. The set of character sequences (up to 5) that will stop output generation. If specified, the API will stop at the first appearance of a stop_sequence . The stop sequence will not be included as part of the response.
responseMimeTypestring Optional. MIME type of the generated candidate text. Supported MIME types are: text/plain : (default) Text output. application/json : JSON response in the response candidates. text/x.enum : ENUM as a string response in the response candidates. Refer to the docs for a list of all supported text MIME types.
responseSchema
(deprecated)object ( Schema )Optional. Output schema of the generated candidate text. Schemas must be a subset of the OpenAPI schema and can be objects, primitives or arrays.
If set, a compatible responseMimeType must also be set. Compatible MIME types: application/json : Schema for JSON response. Refer to the JSON text generation guide for more details.
_responseJsonSchema
(deprecated)value ( Value format) Optional. Output schema of the generated response. This is an alternative to responseSchema that accepts JSON Schema .
If set, responseSchema must be omitted, but responseMimeType is required.
While the full JSON Schema may be sent, not all features are supported. Specifically, only the following properties are supported:
-
$id -
$defs -
$ref -
$anchor -
type -
format -
title -
description -
enum(for strings and numbers) -
items -
prefixItems -
minItems -
maxItems -
minimum -
maximum -
anyOf -
oneOf(interpreted the same asanyOf) -
properties -
additionalProperties -
required
The non-standard propertyOrdering property may also be set.
Cyclic references are unrolled to a limited degree and, as such, may only be used within non-required properties. (Nullable properties are not sufficient.) If $ref is set on a sub-schema, no other properties, except for than those starting as a $ , may be set.
responseJsonSchemavalue ( Value format) Optional. An internal detail. Use responseJsonSchema rather than this field.
responseModalities[]enum ( Modality )Optional. The requested modalities of the response. Represents the set of modalities that the model can return, and should be expected in the response. This is an exact match to the modalities of the response.
A model may have multiple combinations of supported modalities. If the requested modalities do not match any of the supported combinations, an error will be returned.
An empty list is equivalent to requesting only text.
candidateCountintegerOptional. Number of generated responses to return. If unset, this will default to 1. Please note that this doesn't work for previous generation models (Gemini 1.0 family)
maxOutputTokensintegerOptional. The maximum number of tokens to include in a response candidate.
Note: The default value varies by model, see the Model.output_token_limit attribute of the Model returned from the getModel function.
temperaturenumberOptional. Controls the randomness of the output.
Note: The default value varies by model, see the Model.temperature attribute of the Model returned from the getModel function.
Values can range from [0.0, 2.0].
topPnumberOptional. The maximum cumulative probability of tokens to consider when sampling.
The model uses combined Top-k and Top-p (nucleus) sampling.
Tokens are sorted based on their assigned probabilities so that only the most likely tokens are considered. Top-k sampling directly limits the maximum number of tokens to consider, while Nucleus sampling limits the number of tokens based on the cumulative probability.
Note: The default value varies by Model and is specified by the Model.top_p attribute returned from the getModel function. An empty topK attribute indicates that the model doesn't apply top-k sampling and doesn't allow setting topK on requests.
topKintegerOptional. The maximum number of tokens to consider when sampling.
Gemini models use Top-p (nucleus) sampling or a combination of Top-k and nucleus sampling. Top-k sampling considers the set of topK most probable tokens. Models running with nucleus sampling don't allow topK setting.
Note: The default value varies by Model and is specified by the Model.top_p attribute returned from the getModel function. An empty topK attribute indicates that the model doesn't apply top-k sampling and doesn't allow setting topK on requests.
seedintegerOptional. Seed used in decoding. If not set, the request uses a randomly generated seed.
presencePenaltynumberOptional. Presence penalty applied to the next token's logprobs if the token has already been seen in the response.
This penalty is binary on/off and not dependant on the number of times the token is used (after the first). Use frequencyPenalty for a penalty that increases with each use.
A positive penalty will discourage the use of tokens that have already been used in the response, increasing the vocabulary.
A negative penalty will encourage the use of tokens that have already been used in the response, decreasing the vocabulary.
frequencyPenaltynumberOptional. Frequency penalty applied to the next token's logprobs, multiplied by the number of times each token has been seen in the respponse so far.
A positive penalty will discourage the use of tokens that have already been used, proportional to the number of times the token has been used: The more a token is used, the more difficult it is for the model to use that token again increasing the vocabulary of responses.
Caution: A negative penalty will encourage the model to reuse tokens proportional to the number of times the token has been used. Small negative values will reduce the vocabulary of a response. Larger negative values will cause the model to start repeating a common token until it hits the maxOutputTokens limit.
responseLogprobsbooleanOptional. If true, export the logprobs results in response.
logprobsinteger Optional. Only valid if responseLogprobs=True . This sets the number of top logprobs, including the chosen candidate, to return at each decoding step in the Candidate.logprobs_result . The number must be in the range of [0, 20].
enableEnhancedCivicAnswersbooleanOptional. Enables enhanced civic answers. It may not be available for all models.
speechConfigobject ( SpeechConfig )Optional. The speech generation config.
thinkingConfigobject ( ThinkingConfig )Optional. Config for thinking features. An error will be returned if this field is set for models that don't support thinking.
imageConfigobject ( ImageConfig )Optional. Config for image generation. An error will be returned if this field is set for models that don't support these config options.
mediaResolutionenum ( MediaResolution )Optional. If specified, the media resolution specified will be used.
enableAffectiveDialogbooleanOptional. If enabled, the model will detect emotions and adapt its responses accordingly.
responseFormatobject ( ResponseFormatConfig )Optional. Configuration for the response output format. Allows specifying output configuration per modality (text, audio, image) in a flat structure.
translationConfigobject ( TranslationConfig )Optional. Config for translation.
audioTranscriptionConfigobject ( AudioTranscriptionConfig )Optional. Config for audio transcription (speech recognition).
| JSON-представление |
|---|
{ "stopSequences": [ string ], "responseMimeType": string, "responseSchema": { object ( |
Модальность
Supported modalities of the response.
| Перечисления | |
|---|---|
MODALITY_UNSPECIFIED | Default value. |
TEXT | Indicates the model should return text. |
IMAGE | Indicates the model should return images. |
AUDIO | Indicates the model should return audio. |
SpeechConfig
Config for speech generation and transcription.
voiceConfigobject ( VoiceConfig )The configuration in case of single-voice output.
multiSpeakerVoiceConfigobject ( MultiSpeakerVoiceConfig )Optional. The configuration for the multi-speaker setup. It is mutually exclusive with the voiceConfig field.
languageCodestringOptional. The IETF BCP-47 language code that the user configured the app to use. Used for speech recognition and synthesis.
Valid values are: de-DE , en-AU , en-GB , en-IN , en-US , es-US , fr-FR , hi-IN , pt-BR , ar-XA , es-ES , fr-CA , id-ID , it-IT , ja-JP , tr-TR , vi-VN , bn-IN , gu-IN , kn-IN , ml-IN , mr-IN , ta-IN , te-IN , nl-NL , ko-KR , cmn-CN , pl-PL , ru-RU , and th-TH .
| JSON-представление |
|---|
{ "voiceConfig": { object ( |
VoiceConfig
The configuration for the voice to use.
voice_configUnion typevoice_config can be only one of the following: prebuiltVoiceConfigobject ( PrebuiltVoiceConfig )The configuration for the prebuilt voice to use.
| JSON-представление |
|---|
{
// voice_config
"prebuiltVoiceConfig": {
object ( |
PrebuiltVoiceConfig
The configuration for the prebuilt speaker to use.
voiceNamestringThe name of the preset voice to use.
| JSON-представление |
|---|
{ "voiceName": string } |
MultiSpeakerVoiceConfig
The configuration for the multi-speaker setup.
speakerVoiceConfigs[]object ( SpeakerVoiceConfig )Required. All the enabled speaker voices.
| JSON-представление |
|---|
{
"speakerVoiceConfigs": [
{
object ( |
SpeakerVoiceConfig
The configuration for a single speaker in a multi speaker setup.
speakerstringRequired. The name of the speaker to use. Should be the same as in the prompt.
voiceConfigobject ( VoiceConfig )Required. The configuration for the voice to use.
| JSON-представление |
|---|
{
"speaker": string,
"voiceConfig": {
object ( |
ThinkingConfig
Config for thinking features.
includeThoughtsbooleanIndicates whether to include thoughts in the response. If true, thoughts are returned only when available.
thinkingBudgetintegerThe number of thoughts tokens that the model should generate.
thinkingLevelenum ( ThinkingLevel )Optional. Controls the maximum depth of the model's internal reasoning process before it produces a response. The default value is model-dependent. Refer to the Thinking levels guide for more details. Recommended for Gemini 3 or later models. Use with earlier models results in an error.
| JSON-представление |
|---|
{
"includeThoughts": boolean,
"thinkingBudget": integer,
"thinkingLevel": enum ( |
ThinkingLevel
Allow user to specify how much to think using enum instead of integer budget.
| Перечисления | |
|---|---|
THINKING_LEVEL_UNSPECIFIED | Default value. |
MINIMAL | Little to no thinking. |
LOW | Low thinking level. |
MEDIUM | Medium thinking level. |
HIGH | High thinking level. |
ImageConfig
Config for image generation features.
aspectRatiostring Optional. The aspect ratio of the image to generate. Supported aspect ratios: 1:1 , 1:4 , 4:1 , 1:8 , 8:1 , 2:3 , 3:2 , 3:4 , 4:3 , 4:5 , 5:4 , 9:16 , 16:9 , or 21:9 .
If not specified, the model will choose a default aspect ratio based on any reference images provided.
imageSizestring Optional. Specifies the size of generated images. Supported values are 512 , 1K , 2K , 4K . If not specified, the model will use default value 1K .
| JSON-представление |
|---|
{ "aspectRatio": string, "imageSize": string } |
MediaResolution
Media resolution for the input media.
| Перечисления | |
|---|---|
MEDIA_RESOLUTION_UNSPECIFIED | Media resolution has not been set. |
MEDIA_RESOLUTION_LOW | Media resolution set to low (64 tokens). |
MEDIA_RESOLUTION_MEDIUM | Media resolution set to medium (256 tokens). |
MEDIA_RESOLUTION_HIGH | Media resolution set to high (zoomed reframing with 256 tokens). |
ResponseFormatConfig
Configuration for the response output format. This is a flat object where each optional sub-field configures a specific output modality.
textobject ( TextResponseFormat )Optional. Text output format configuration.
audioobject ( AudioResponseFormat )Optional. Audio output format configuration.
imageobject ( ImageResponseFormat )Optional. Image output format configuration.
| JSON-представление |
|---|
{ "text": { object ( |
TextResponseFormat
Configuration for text output format.
mimeTypeenum ( MimeType )Optional. The MIME type of the text output.
schemavalue ( Value format)Optional. The JSON schema that the output should conform to. Only applicable when mimeType is APPLICATION_JSON.
| JSON-представление |
|---|
{
"mimeType": enum ( |
MimeType
Supported MIME types for text output.
| Перечисления | |
|---|---|
MIME_TYPE_UNSPECIFIED | Default value. This value is unused. |
APPLICATION_JSON | JSON output format. |
TEXT_PLAIN | Plain text output format. |
AudioResponseFormat
Configuration for audio output format.
mimeTypeenum ( MimeType )Optional. The MIME type of the audio output.
deliveryenum ( Delivery )Optional. The delivery mode for the audio output.
sampleRateintegerOptional. Sample rate in Hz.
bitRateintegerOptional. Bit rate in bits per second (bps). Only applicable for compressed formats (MP3, Opus).
MimeType
Supported MIME types for audio output.
| Перечисления | |
|---|---|
MIME_TYPE_UNSPECIFIED | Default value. This value is unused. |
AUDIO_MP3 | MP3 audio format. |
AUDIO_OGG_OPUS | OGG Opus audio format. |
AUDIO_L16 | Raw PCM (L16) audio format. |
AUDIO_WAV | WAV audio format. |
AUDIO_ALAW | A-law audio format. |
AUDIO_MULAW | Mu-law audio format. |
Доставка
Delivery mode for audio output.
| Перечисления | |
|---|---|
DELIVERY_UNSPECIFIED | Default value. This value is unused. |
INLINE | Audio data is returned inline in the response. |
URI | Audio data is returned as a URI. |
ImageResponseFormat
Configuration for image output format.
mimeTypeenum ( MimeType )Optional. The MIME type of the image output.
deliveryenum ( Delivery )Optional. The delivery mode for the image output.
aspectRatioenum ( AspectRatio )Optional. The aspect ratio for the image output.
imageSizeenum ( ImageSize )Optional. The size of the image output.
| JSON-представление |
|---|
{ "mimeType": enum ( |
MimeType
Supported MIME types for image output.
| Перечисления | |
|---|---|
MIME_TYPE_UNSPECIFIED | Default value. This value is unused. |
IMAGE_JPEG | JPEG image format. |
Доставка
Delivery mode for image output.
| Перечисления | |
|---|---|
DELIVERY_UNSPECIFIED | Default value. This value is unused. |
INLINE | Image data is returned inline in the response. |
URI | Image data is returned as a URI. |
Соотношение сторон
Supported aspect ratios for image output.
| Перечисления | |
|---|---|
ASPECT_RATIO_UNSPECIFIED | Default value. This value is unused. |
ASPECT_RATIO_ONE_BY_ONE | 1:1 aspect ratio. |
ASPECT_RATIO_TWO_BY_THREE | 2:3 aspect ratio. |
ASPECT_RATIO_THREE_BY_TWO | 3:2 aspect ratio. |
ASPECT_RATIO_THREE_BY_FOUR | 3:4 aspect ratio. |
ASPECT_RATIO_FOUR_BY_THREE | 4:3 aspect ratio. |
ASPECT_RATIO_FOUR_BY_FIVE | 4:5 aspect ratio. |
ASPECT_RATIO_FIVE_BY_FOUR | 5:4 aspect ratio. |
ASPECT_RATIO_NINE_BY_SIXTEEN | 9:16 aspect ratio. |
ASPECT_RATIO_SIXTEEN_BY_NINE | 16:9 aspect ratio. |
ASPECT_RATIO_TWENTY_ONE_BY_NINE | 21:9 aspect ratio. |
ASPECT_RATIO_ONE_BY_EIGHT | 1:8 aspect ratio. |
ASPECT_RATIO_EIGHT_BY_ONE | 8:1 aspect ratio. |
ASPECT_RATIO_ONE_BY_FOUR | 1:4 aspect ratio. |
ASPECT_RATIO_FOUR_BY_ONE | 4:1 aspect ratio. |
ImageSize
Supported image sizes for image output.
| Перечисления | |
|---|---|
IMAGE_SIZE_UNSPECIFIED | Default value. This value is unused. |
IMAGE_SIZE_FIVE_TWELVE | 512px image size. |
IMAGE_SIZE_ONE_K | 1K image size. |
IMAGE_SIZE_TWO_K | 2K image size. |
IMAGE_SIZE_FOUR_K | 4K image size. |
TranslationConfig
Config for translation features.
targetLanguageCodestringRequired. The target language for translation. Supported values are BCP-47 language codes (eg "en", "es", "fr").
echoTargetLanguagebooleanOptional. If true, the model will generate audio when the target language is spoken, essentially it will parrot the input. If false, we will not produce audio for the target language.
| JSON-представление |
|---|
{ "targetLanguageCode": string, "echoTargetLanguage": boolean } |
AudioTranscriptionConfig
The audio transcription configuration.
languageCodes[]stringOptional. BCP-47 language codes providing hints about the languages present in the audio. If omitted or empty, defaults to automatic language detection.
adaptationPhrases[]
(deprecated)stringOptional. A list of phrases used for speech adaptation, which biases the ASR model to improve recognition of these specific terms.
customVocabulary[]stringOptional. A list of custom vocabulary phrases to bias the speech recognition model toward recognizing specific terms (product names, proper nouns, jargon).
wordTimestampbooleanOptional. Configures word-level timestamp generation.
diarizationbooleanOptional. Configures speaker diarization.
language_configUnion typelanguage_codes instead. language_config can be only one of the following: languageAuto
(deprecated)object ( LanguageAuto )Optional. The model will detect the language automatically.
languageHints
(deprecated)object ( LanguageHints )Optional. Specifies one or more languages in the audio.
| JSON-представление |
|---|
{ "languageCodes": [ string ], "adaptationPhrases": [ string ], "customVocabulary": [ string ], "wordTimestamp": boolean, "diarization": boolean, // language_config "languageAuto": { object ( |
LanguageAuto
This type has no fields.
Indicates the language of the audio should be automatically detected.
LanguageHints
Provides hints to the model about possible languages present in the audio.
languageCodes[]
(deprecated)stringRequired. BCP-47 language codes.
| JSON-представление |
|---|
{ "languageCodes": [ string ] } |
RealtimeInputConfig
Configures the realtime input behavior in BidiGenerateContent .
automaticActivityDetectionobject ( AutomaticActivityDetection )Optional. If not set, automatic activity detection is enabled by default. If automatic voice detection is disabled, the client must send activity signals.
activityHandlingenum ( ActivityHandling )Optional. Defines what effect activity has.
turnCoverageenum ( TurnCoverage )Optional. Defines which input is included in the user's turn.
| JSON-представление |
|---|
{ "automaticActivityDetection": { object ( |
AutomaticActivityDetection
Configures automatic detection of activity.
disabledbooleanOptional. If enabled (the default), detected voice and text input count as activity. If disabled, the client must send activity signals.
startOfSpeechSensitivityenum ( StartSensitivity )Optional. Determines how likely speech is to be detected.
prefixPaddingMsintegerOptional. The required duration of detected speech before start-of-speech is committed. The lower this value, the more sensitive the start-of-speech detection is and shorter speech can be recognized. However, this also increases the probability of false positives.
endOfSpeechSensitivityenum ( EndSensitivity )Optional. Determines how likely detected speech is ended.
silenceDurationMsintegerOptional. The required duration of detected non-speech (eg silence) before end-of-speech is committed. The larger this value, the longer speech gaps can be without interrupting the user's activity but this will increase the model's latency.
| JSON-представление |
|---|
{ "disabled": boolean, "startOfSpeechSensitivity": enum ( |
StartSensitivity
Determines how start of speech is detected.
| Перечисления | |
|---|---|
START_SENSITIVITY_UNSPECIFIED | The default is START_SENSITIVITY_HIGH. |
START_SENSITIVITY_HIGH | Automatic detection will detect the start of speech more often. |
START_SENSITIVITY_LOW | Automatic detection will detect the start of speech less often. |
EndSensitivity
Determines how end of speech is detected.
| Перечисления | |
|---|---|
END_SENSITIVITY_UNSPECIFIED | The default is END_SENSITIVITY_HIGH. |
END_SENSITIVITY_HIGH | Automatic detection ends speech more often. |
END_SENSITIVITY_LOW | Automatic detection ends speech less often. |
ActivityHandling
The different ways of handling user activity.
| Перечисления | |
|---|---|
ACTIVITY_HANDLING_UNSPECIFIED | If unspecified, the default behavior is START_OF_ACTIVITY_INTERRUPTS . |
START_OF_ACTIVITY_INTERRUPTS | If true, start of activity will interrupt the model's response (also called "barge in"). The model's current response will be cut-off in the moment of the interruption. This is the default behavior. |
NO_INTERRUPTION | The model's response will not be interrupted. |
TurnCoverage
Options about which input is included in the user's turn.
| Перечисления | |
|---|---|
TURN_COVERAGE_UNSPECIFIED | If unspecified, a default behavior is selected based on the model. Eg, for Gemini 2.5, the default is TURN_INCLUDES_ONLY_ACTIVITY , while for Gemini 3.1 and onwards, it's TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO . |
TURN_INCLUDES_ONLY_ACTIVITY | Includes activity since the last turn, excluding inactivity (eg silence on the audio stream). |
TURN_INCLUDES_ALL_INPUT | Includes all realtime input since the last turn, including inactivity (eg silence on the audio stream). |
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO | Includes audio activity and all video since the last turn. With automatic activity detection, audio activity means speech and excludes silence. |
SessionResumptionConfig
Session resumption configuration.
This message is included in the session configuration as BidiGenerateContentSetup.session_resumption . If configured, the server will send SessionResumptionUpdate messages.
handlestringThe handle of a previous session. If not present then a new session is created.
Session handles come from SessionResumptionUpdate.token values in previous connections.
| JSON-представление |
|---|
{ "handle": string } |
ContextWindowCompressionConfig
Enables context window compression — a mechanism for managing the model's context window so that it does not exceed a given length.
compression_mechanismUnion typecompression_mechanism can be only one of the following: slidingWindowobject ( SlidingWindow )A sliding-window mechanism.
triggerTokensstring ( int64 format)The number of tokens (before running a turn) required to trigger a context window compression.
This can be used to balance quality against latency as shorter context windows may result in faster model responses. However, any compression operation will cause a temporary latency increase, so they should not be triggered frequently.
If not set, the default is 80% of the model's context window limit. This leaves 20% for the next user request/model response.
| JSON-представление |
|---|
{
// compression_mechanism
"slidingWindow": {
object ( |
SlidingWindow
The SlidingWindow method operates by discarding content at the beginning of the context window. The resulting context will always begin at the start of a USER role turn. System instructions and any BidiGenerateContentSetup.prefix_turns will always remain at the beginning of the result.
targetTokensstring ( int64 format)The target number of tokens to keep. The default value is triggerTokens/2.
Discarding parts of the context window causes a temporary latency increase so this value should be calibrated to avoid frequent compression operations.
| JSON-представление |
|---|
{ "targetTokens": string } |
HistoryConfig
History configuration.
This message is included in the session configuration as BidiGenerateContentSetup.history_config . Configures the exchange of history messages.
initialHistoryInClientContentboolean Optional. If true, after sending setupComplete , the server will wait and at first process clientContent messages until turnComplete is true . This initial history will not trigger a model call and may end with role MODEL . After turnComplete is true , the client can start the realtime conversation via realtimeInput .
| JSON-представление |
|---|
{ "initialHistoryInClientContent": boolean } |
Method: auth_tokens.create
Creates a token that can be used to constrain the behavior of a BidiGenerateContent session.
Конечная точка
posthttps: / /generativelanguage.googleapis.com /v1beta /auth_tokensТекст запроса
The request body contains an instance of AuthToken .
expireTimestring ( Timestamp format)Optional. Input only. Immutable. An optional time after which, when using the resulting token, messages in BidiGenerateContent sessions will be rejected. (Gemini may preemptively close the session after this time.)
If not set then this defaults to 30 minutes in the future. If set, this value must be less than 20 hours in the future.
Используется RFC 3339, согласно которому генерируемый вывод всегда будет Z-нормализован и будет содержать 0, 3, 6 или 9 дробных знаков. Допускаются также смещения, отличные от "Z". Примеры: "2014-10-02T15:01:23Z" , "2014-10-02T15:01:23.045123456Z" или "2014-10-02T15:01:23+05:30" .
newSessionExpireTimestring ( Timestamp format)Optional. Input only. Immutable. The time after which new Live API sessions using the token resulting from this request will be rejected.
If not set this defaults to 60 seconds in the future. If set, this value must be less than 20 hours in the future.
Используется RFC 3339, согласно которому генерируемый вывод всегда будет Z-нормализован и будет содержать 0, 3, 6 или 9 дробных знаков. Допускаются также смещения, отличные от "Z". Примеры: "2014-10-02T15:01:23Z" , "2014-10-02T15:01:23.045123456Z" или "2014-10-02T15:01:23+05:30" .
fieldMaskstring ( FieldMask format) Optional. Input only. Immutable. If fieldMask is empty, and bidiGenerateContentSetup is not present, then the effective BidiGenerateContentSetup message is taken from the Live API connection.
If fieldMask is empty, and bidiGenerateContentSetup is present, then the effective BidiGenerateContentSetup message is taken entirely from bidiGenerateContentSetup in this request. The setup message from the Live API connection is ignored.
If fieldMask is not empty, then the corresponding fields from bidiGenerateContentSetup will overwrite the fields from the setup message in the Live API connection.
This is a comma-separated list of fully qualified names of fields. Example: "user.displayName,photo" .
configUnion typeconfig can be only one of the following: bidiGenerateContentSetupobject ( BidiGenerateContentSetup ) Optional. Input only. Immutable. Configuration specific to BidiGenerateContent .
usesintegerOptional. Input only. Immutable. The number of times the token can be used. If this value is zero then no limit is applied. Resuming a Live API session does not count as a use. If unspecified, the default is 1.
Ответный текст
If successful, the response body contains a newly created instance of AuthToken .