Live API — это API с сохранением состояния, использующий WebSockets . В этом разделе вы найдете дополнительные сведения об API WebSockets.
Сессии
Соединение WebSocket устанавливает сессию между клиентом и сервером Gemini. После того, как клиент инициирует новое соединение, сессия может обмениваться сообщениями с сервером, чтобы:
- Отправляйте текст, аудио или видео на сервер Gemini.
- Получайте аудио-, текстовые запросы или запросы на выполнение функциональных задач с сервера Gemini.
WebSocket-соединение
Для начала сеанса подключитесь к следующему адресу веб-сокета:
wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent
Настройки сессии
Первоначальное сообщение, отправляемое после установления соединения WebSocket, задает конфигурацию сессии, которая включает в себя модель, параметры генерации, системные инструкции и инструменты.
Вы не можете обновить конфигурацию, пока соединение открыто. Однако вы можете изменить параметры конфигурации, за исключением модели, при приостановке и возобновлении сессии с помощью механизма возобновления сессии .
См. следующий пример конфигурации. Обратите внимание, что регистр имен в SDK может различаться. Параметры конфигурации Python SDK можно найти здесь .
{
"model": string,
"generationConfig": {
"candidateCount": integer,
"maxOutputTokens": integer,
"temperature": number,
"topP": number,
"topK": integer,
"presencePenalty": number,
"frequencyPenalty": number,
"responseModalities": [string],
"speechConfig": object,
"mediaResolution": object,
"translationConfig": object
},
"systemInstruction": string,
"tools": [object]
}
Для получения дополнительной информации о поле API см. generationConfig .
Отправляйте сообщения
Для обмена сообщениями по WebSocket-соединению клиент должен отправить JSON-объект по открытому WebSocket-соединению. JSON-объект должен содержать ровно одно из полей из следующего набора объектов:
{
"setup": BidiGenerateContentSetup,
"clientContent": BidiGenerateContentClientContent,
"realtimeInput": BidiGenerateContentRealtimeInput,
"toolResponse": BidiGenerateContentToolResponse
}
Поддерживаемые клиентские сообщения
Список поддерживаемых клиентских сообщений приведен в следующей таблице:
| Сообщение | Описание |
|---|---|
BidiGenerateContentSetup | Настройки сессии будут отправлены в первом сообщении. |
BidiGenerateContentClientContent | Постепенное обновление содержимого текущей беседы, предоставляемое клиентом. |
BidiGenerateContentRealtimeInput | Ввод аудио, видео или текста в реальном времени. |
BidiGenerateContentToolResponse | Ответ на сообщение ToolCallMessage полученное от сервера. |
Получайте сообщения
Для получения сообщений от Gemini необходимо прослушивать событие WebSocket 'message', а затем анализировать результат в соответствии с определением поддерживаемых серверных сообщений.
См. следующее:
async with client.aio.live.connect(model='...', config=config) as session:
await session.send(input='Hello world!', end_of_turn=True)
async for message in session.receive():
print(message)
Сообщения сервера могут содержать поле usageMetadata но в остальном будут включать ровно одно из других полей сообщения BidiGenerateContentServerMessage . (Объединение messageType не выражается в формате JSON, поэтому поле будет отображаться на верхнем уровне сообщения.)
Сообщения и события
Завершение активности
Этот тип не содержит полей.
Отмечает завершение активности пользователя.
Обработка действий
Различные способы обработки активности пользователей.
| Перечисления | |
|---|---|
ACTIVITY_HANDLING_UNSPECIFIED | Если параметр не указан, по умолчанию используется поведение START_OF_ACTIVITY_INTERRUPTS . |
START_OF_ACTIVITY_INTERRUPTS | Если это так, начало активности прервёт ответ модели (также называемое «вмешательством»). Текущий ответ модели будет прерван в момент прерывания. Это поведение по умолчанию. |
NO_INTERRUPTION | Реакция модели не будет прервана. |
ActivityStart
Этот тип не содержит полей.
Отмечает начало активности пользователя.
AudioTranscriptionConfig
Настройка транскрипции аудио.
| Поля | |
|---|---|
languageCodes[] | Необязательно. Коды языков BCP-47, указывающие на языки, присутствующие в аудиозаписи. Если поле опущено или пустое, по умолчанию используется автоматическое определение языка. |
customVocabulary[] | Необязательно. Список пользовательских лексических фраз, которые помогут модели распознавания речи ориентироваться на распознавание определенных терминов (названия продуктов, имена собственные, профессиональный жаргон). |
wordTimestamp | Необязательный параметр. Настраивает генерацию временных меток на уровне слов. |
diarization | Необязательный параметр. Настраивает диаризацию речи говорящего. |
mode | Необязательный параметр. Настраивает режим транскрипции. Поддерживаемые значения: |
Режим
Режим транскрипции.
| Перечисления | |
|---|---|
MODE_UNSPECIFIED | Неуказанный режим транскрипции. |
VERBATIM | Режим дословной транскрипции. |
SMART | Интеллектуальный режим транскрипции. |
Автоматическое определение активности
Настраивает автоматическое обнаружение активности.
| Поля | |
|---|---|
disabled | Необязательно. Если включено (по умолчанию), обнаруженный голосовой и текстовый ввод засчитывается как активность. Если отключено, клиент должен отправлять сигналы активности. |
startOfSpeechSensitivity | Необязательный параметр. Определяет вероятность обнаружения речи. |
prefixPaddingMs | Необязательно. Требуемая продолжительность обнаруженной речи до момента её начала. Чем ниже это значение, тем чувствительнее определение начала речи и тем короче распознаётся речь. Однако это также увеличивает вероятность ложных срабатываний. |
endOfSpeechSensitivity | Необязательный параметр. Определяет вероятность прекращения обнаруженной речи. |
silenceDurationMs | Необязательно. Требуемая продолжительность обнаруженного отсутствия речи (например, тишины) до окончания речи. Чем больше это значение, тем дольше могут быть паузы в речи, не прерывая активность пользователя, но это увеличит задержку модели. |
BidiGenerateContentClientContent
Постепенное обновление текущей переписки, предоставляемой клиентом. Весь контент безоговорочно добавляется к истории переписки и используется в качестве запроса к модели для генерации контента.
Сообщение, размещенное здесь, прервет процесс генерации текущей модели.
| Поля | |
|---|---|
turns[] | Необязательно. Содержимое, добавляемое к текущей беседе с моделью. Для запросов с одним циклом обработки это один экземпляр. Для запросов с несколькими циклами обработки это повторяющееся поле, содержащее историю переписки и последний запрос. |
turnComplete | Необязательный параметр. Если значение равно true, указывает, что генерация контента на сервере должна начинаться с текущего накопленного запроса. В противном случае сервер ожидает дополнительных сообщений перед началом генерации. |
BidiGenerateContentRealtimeInput
Пользовательский ввод, отправляемый в режиме реального времени.
Различные форматы (аудио, видео и текст) обрабатываются как параллельные потоки. Порядок потоков не гарантируется.
Это отличается от BidiGenerateContentClientContent по нескольким параметрам:
- Может непрерывно и без перерыва отправляться на этап генерации модели.
- Если возникает необходимость смешивать данные, чередующиеся между
BidiGenerateContentClientContentиBidiGenerateContentRealtimeInput, сервер пытается оптимизировать отклик для достижения наилучшего результата, но гарантий нет. - Конец реплики явно не указан, а определяется действиями пользователя (например, окончанием речи).
- Еще до окончания хода обработки данные обрабатываются поэтапно для оптимизации и обеспечения быстрого запуска реакции модели.
| Поля | |
|---|---|
mediaChunks[] | Необязательно. Встроенные байтовые данные для ввода мультимедиа. Использование нескольких УСТАРЕВШИЙ ВАРИАНТ: Используйте вместо него |
audio | Необязательно. Они формируют поток входного аудиосигнала в реальном времени. |
video | Необязательно. Они формируют поток видеовхода в реальном времени. |
activityStart | Необязательный параметр. Отмечает начало активности пользователя. Этот параметр может быть отправлен только в том случае, если автоматическое (т.е. на стороне сервера) обнаружение активности отключено. |
activityEnd | Необязательно. Отмечает окончание активности пользователя. Это сообщение может быть отправлено только в том случае, если автоматическое (т.е. на стороне сервера) обнаружение активности отключено. |
mediaResolution | Необязательный параметр. Разрешение мультимедиа для использования. Если не указано, используется значение |
audioStreamEnd | Необязательно. Указывает на то, что аудиопоток завершился, например, из-за выключения микрофона. Это сообщение следует отправлять только при включенном автоматическом обнаружении активности (что является настройкой по умолчанию). Клиент может возобновить трансляцию, отправив аудиосообщение. |
text | Необязательно. Они формируют поток ввода текста в реальном времени. |
BidiGenerateContentServerContent
Постепенные обновления сервера, генерируемые моделью в ответ на сообщения клиента.
Контент генерируется максимально быстро, но не в режиме реального времени. Клиенты могут по своему выбору буферизовать и воспроизводить его в режиме реального времени.
| Поля | |
|---|---|
generationComplete | Только вывод. Если значение равно true, это означает, что генерация модели завершена. Если генерация модели прерывается, сообщение 'generation_complete' в прерванном цикле не появится, процесс продолжится по схеме 'interrupted > turn_complete'. Когда модель предполагает воспроизведение в реальном времени, между generation_complete и turn_complete возникает задержка, вызванная ожиданием моделью завершения воспроизведения. |
turnComplete | Только вывод. Если значение равно true, это означает, что модель завершила свой ход. Генерация начнется только в ответ на дополнительные сообщения от клиента. Обратите внимание, что если включена отчетность о состоянии воспроизведения, она отправляется только тогда, когда состояние воспроизведения указывает на его завершение. Дальнейшие отчеты о состоянии воспроизведения той же генерации будут игнорироваться. |
interrupted | Только вывод. Если значение равно true, это означает, что сообщение от клиента прервало текущую генерацию модели. Если клиент воспроизводит контент в реальном времени, это хороший сигнал для остановки и очистки текущей очереди воспроизведения. |
groundingMetadata | Только для вывода. Метаданные для сгенерированного контента. |
inputTranscription | Только вывод. Входная аудиотранскрипция. Транскрипция отправляется независимо от других сообщений сервера, и гарантированного порядка нет. |
interimInputTranscription | Только вывод. Транскрипция с низкой задержкой обновляется во время речи пользователя. Это поле часто обновляется. |
outputTranscription | Только вывод. Аудиотранскрипция. Эти транскрипции являются частью выходных данных генерации сервера. Последняя выходная транскрипция этого хода отправляется перед |
urlContextMetadata | |
waitingForInput | Только вывод. Если значение равно true, это означает, что модель не генерирует контент, поскольку ожидает дополнительной информации от пользователя, например, потому что рассчитывает на продолжение разговора. |
speechState | Только вывод. УСТАРЕВШЕЕ: используйте VoiceActivity вместо него. Указывает текущее состояние распознавания речи в |
interactionStatus | Только вывод. Текущий статус активности сессии. Всегда отправляется вместе с |
modelTurn | Только выходные данные. Контент, сгенерированный моделью в рамках текущего диалога с пользователем. |
BidiGenerateContentServerMessage
Ответное сообщение для вызова функции BidiGenerateContent.
| Поля | |
|---|---|
usageMetadata | Только вывод. Метаданные об использовании ответа(ов). |
Поле объединения messageType . Тип сообщения. messageType может принимать только одно из следующих значений: | |
setupComplete | Только для вывода. Отправляется в ответ на сообщение |
serverContent | Только выходные данные. Контент, сгенерированный моделью в ответ на сообщения клиента. |
toolCall | Только вывод. Запрос к клиенту на выполнение |
toolCallCancellation | Только вывод. Уведомление для клиента о том, что ранее отправленное сообщение |
goAway | Только вывод. Уведомление о скором отключении сервера. |
sessionResumptionUpdate | Только вывод. Обновление состояния возобновления сессии. |
BidiGenerateContentSetup
Сообщение, которое будет отправлено в первом (и только в первом) BidiGenerateContentClientMessage . Содержит конфигурацию, которая будет применяться на протяжении всего процесса потокового RPC-вызова.
Клиентам следует дождаться сообщения BidiGenerateContentSetupComplete прежде чем отправлять какие-либо дополнительные сообщения.
| Поля | |
|---|---|
model | Обязательно. Имя ресурса модели. Оно служит идентификатором для используемой модели. Формат: |
generationConfig | Необязательно. Конфигурация генерации. Следующие поля не поддерживаются:
|
systemInstruction | Необязательно. Пользователь предоставил системные инструкции для данной модели. Примечание: текст должен использоваться только по частям, содержание каждой части должно быть в отдельном абзаце. |
tools[] | Необязательно. Список |
realtimeInputConfig | Необязательный параметр. Настраивает обработку ввода в реальном времени. |
sessionResumption | Необязательный параметр. Настраивает механизм возобновления сессии. Если эта опция включена, сервер будет отправлять сообщения |
contextWindowCompression | Необязательный параметр. Настраивает механизм сжатия контекстного окна. Если эта опция включена, сервер автоматически уменьшит размер контекста, когда он превысит заданную длину. |
inputAudioTranscription | Необязательный параметр. Если задан, включает транскрипцию голосового ввода. Транскрипция соответствует языку входного аудио, если это настроено. |
outputAudioTranscription | Необязательный параметр. Если задан, включает транскрипцию аудиовыхода модели. Транскрипция соответствует языковому коду, указанному для выходного аудио, если таковой задан. |
proactivity | Необязательный параметр. Настраивает проактивность модели. Это позволяет модели заблаговременно реагировать на входные данные и игнорировать нерелевантную информацию. |
historyConfig | Необязательный параметр. Настраивает обмен историей между клиентом и сервером. |
BidiGenerateContentSetupComplete
Этот тип не содержит полей.
Отправлено в ответ на сообщение BidiGenerateContentSetup от клиента.
BidiGenerateContentToolCall
Запрос к клиенту на выполнение functionCalls и возврат ответов с соответствующими id .
| Поля | |
|---|---|
functionCalls[] | Только вывод. Вызов функции, которая будет выполнена. |
BidiGenerateContentToolCallCancellation
Уведомление для клиента о том, что ранее отправленное сообщение ToolCallMessage с указанным id не должно было быть выполнено и должно быть отменено. Если эти вызовы инструментов имели побочные эффекты, клиенты могут попытаться отменить их. Это сообщение появляется только в случаях, когда клиенты прерывают работу сервера.
| Поля | |
|---|---|
ids[] | Только вывод. Идентификаторы вызовов инструментов, которые необходимо отменить. |
BidiGenerateContentToolResponse
Ответ клиента, сгенерированный на ToolCall полученный от сервера. Отдельные объекты FunctionResponse сопоставляются с соответствующими объектами FunctionCall по полю id .
Обратите внимание, что в одностороннем и серверно-потоковом режимах вызов функции GenerateContent API происходит путем обмена частями Content , тогда как в двунаправленном режиме вызов функции GenerateContent API происходит через этот выделенный набор сообщений.
| Поля | |
|---|---|
functionResponses[] | Необязательно. Ответ на вызовы функций. |
BidiGenerateContentТранскрипция
Транскрипция аудио (входного или выходного сигнала).
| Поля | |
|---|---|
text | Текст транскрипции. |
languageCode | Языковой код транскрипции BCP-47. |
ContextWindowCompressionConfig
Включает сжатие контекстного окна — механизм управления контекстным окном модели таким образом, чтобы оно не превышало заданную длину.
| Поля | |
|---|---|
Union field compressionMechanism . Механизм сжатия контекстного окна. compressionMechanism может принимать только одно из следующих значений: | |
slidingWindow | Механизм раздвижного окна. |
triggerTokens | Количество токенов (до начала хода), необходимое для запуска сжатия контекстного окна. Это можно использовать для баланса между качеством и задержкой, поскольку более короткие контекстные окна могут привести к более быстрой реакции модели. Однако любая операция сжатия вызовет временное увеличение задержки, поэтому их не следует запускать часто. Если значение не задано, по умолчанию используется 80% от лимита контекстного окна модели. Это оставляет 20% для следующего запроса пользователя/ответа модели. |
EndSensitivity
Определяет способ определения конца речи.
| Перечисления | |
|---|---|
END_SENSITIVITY_UNSPECIFIED | Значение по умолчанию — END_SENSITIVITY_HIGH. |
END_SENSITIVITY_HIGH | Автоматическое распознавание чаще прерывает речь. |
END_SENSITIVITY_LOW | Автоматическое распознавание реже прерывает речь. |
Уходите
Уведомление о том, что сервер скоро отключится.
| Поля | |
|---|---|
timeLeft | Оставшееся время до завершения соединения будет помечено как «ПРЕРЫВНО». Эта продолжительность никогда не будет меньше минимального значения, установленного для конкретной модели, которое будет указано вместе с ограничениями скорости для данной модели. |
HistoryConfig
Настройки истории.
Это сообщение включается в конфигурацию сессии как BidiGenerateContentSetup.historyConfig . Оно настраивает обмен историческими сообщениями.
| Поля | |
|---|---|
initialHistoryInClientContent | Необязательно. Если true, после отправки |
ProactivityConfig
Настройка функций проактивного управления.
| Поля | |
|---|---|
proactiveAudio | Необязательно. Если включено, модель может отказаться отвечать на последний запрос. Например, это позволяет модели игнорировать речь вне контекста или молчать, если пользователь еще не отправлял запрос. |
RealtimeInputConfig
Настраивает поведение ввода в реальном времени в BidiGenerateContent .
| Поля | |
|---|---|
automaticActivityDetection | Необязательно. Если не задано, автоматическое определение активности включено по умолчанию. Если автоматическое определение голоса отключено, клиент должен отправлять сигналы активности. |
activityHandling | Необязательный параметр. Определяет, какое воздействие оказывает данное действие. |
turnCoverage | Необязательный параметр. Определяет, какой ввод будет включен в ход пользователя. |
SessionResumptionConfig
Настройка возобновления сессии.
Это сообщение включается в конфигурацию сессии как BidiGenerateContentSetup.sessionResumption . Если это настроено, сервер будет отправлять сообщения SessionResumptionUpdate .
| Поля | |
|---|---|
handle | Идентификатор предыдущей сессии. Если он отсутствует, создается новая сессия. Дескрипторы сессий берутся из значений |
Обновление возобновления сессии
Обновление состояния возобновления сессии.
Отправляется только в том случае, если задан параметр BidiGenerateContentSetup.sessionResumption .
| Поля | |
|---|---|
newHandle | Новый дескриптор, представляющий состояние, которое можно возобновить. Пусто, если |
resumable | Истина, если текущую сессию можно возобновить на данном этапе. Возобновление сессии невозможно на некоторых этапах. Например, когда модель выполняет вызовы функций или генерирует данные. Возобновление сессии (с использованием токена предыдущей сессии) в таком состоянии приведет к потере данных. В этих случаях |
СлайдингВанно
Метод SlidingWindow работает путем удаления содержимого в начале контекстного окна. Полученный контекст всегда будет начинаться с начала хода роли USER. Системные инструкции и любые BidiGenerateContentSetup.prefixTurns всегда будут оставаться в начале результата.
| Поля | |
|---|---|
targetTokens | Целевое количество токенов для сохранения. Значение по умолчанию: trigger_tokens/2. Удаление части контекстного окна приводит к временному увеличению задержки, поэтому это значение следует калибровать, чтобы избежать частых операций сжатия. |
Начальная чувствительность
Определяет способ определения начала речи.
| Перечисления | |
|---|---|
START_SENSITIVITY_UNSPECIFIED | Значение по умолчанию — START_SENSITIVITY_HIGH. |
START_SENSITIVITY_HIGH | Автоматическое распознавание будет чаще определять начало речи. |
START_SENSITIVITY_LOW | Автоматическое распознавание будет реже обнаруживать начало речи. |
TurnCoverage
Варианты выбора того, какой ввод будет включен в ход пользователя.
| Перечисления | |
|---|---|
TURN_COVERAGE_UNSPECIFIED | Если параметр не указан, выбирается поведение по умолчанию в зависимости от модели. Например, для Gemini 2.5 по умолчанию используется TURN_INCLUDES_ONLY_ACTIVITY , а для Gemini 3.1 и более поздних версий — TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO . |
TURN_INCLUDES_ONLY_ACTIVITY | Включает активность с момента последнего хода, исключая бездействие (например, тишину в аудиопотоке). |
TURN_INCLUDES_ALL_INPUT | Включает весь ввод данных в реальном времени с момента последнего хода, в том числе и в периоды бездействия (например, тишина в аудиопотоке). |
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO | Включает в себя аудиоактивность и все видео с момента последнего поворота. При автоматическом определении активности аудиоактивность означает речь и исключает тишину. |
TranslationConfig
Настройки функций перевода.
| Поля | |
|---|---|
targetLanguageCode | Обязательно. Целевой язык для перевода. Поддерживаемые значения — коды языков BCP-47 (например, "en", "es", "fr"). |
echoTargetLanguage | Необязательно. Если значение равно true, модель будет генерировать аудио при произнесении целевого языка, по сути, она будет повторять входные данные. Если значение равно false, мы не будем генерировать аудио для целевого языка. |
UrlContextMetadata
Метаданные, относящиеся к инструменту получения контекста URL-адреса.
| Поля | |
|---|---|
urlMetadata[] | Список контекста URL-адреса. |
UsageMetadata
Метаданные об использовании ответа(ов).
| Поля | |
|---|---|
promptTokenCount | Только вывод. Количество токенов в приглашении. Если задан параметр |
cachedContentTokenCount | Количество токенов в кэшированной части запроса (кэшированное содержимое) |
responseTokenCount | Только вывод. Общее количество токенов по всем сгенерированным вариантам ответа. |
toolUsePromptTokenCount | Только вывод. Количество токенов, присутствующих в подсказках использования инструмента. |
thoughtsTokenCount | Только вывод. Количество токенов мыслей для моделей мышления. |
totalTokenCount | Только вывод. Общее количество токенов для запроса на генерацию (кандидаты запроса + ответа). |
promptTokensDetails[] | Только выходные данные. Список модальностей, которые были обработаны во входном запросе. |
cacheTokensDetails[] | Только вывод. Список вариантов содержимого, кэшированного в запросе. |
responseTokensDetails[] | Только вывод. Список модальностей, которые были возвращены в ответе. |
toolUsePromptTokensDetails[] | Только выходные данные. Список модальностей, обработанных для обработки запросов на использование инструмента. |
Временные токены аутентификации
Временные токены аутентификации можно получить, вызвав AuthTokenService.CreateToken , а затем использовать с GenerativeService.BidiGenerateContentConstrained , либо передав токен в параметре запроса access_token , либо в заголовке HTTP Authorization с префиксом " Token ".
CreateAuthTokenRequest
Создайте временный токен аутентификации.
| Поля | |
|---|---|
authToken | Обязательно. Токен для создания. |
AuthToken
Запрос на создание временного токена аутентификации.
| Поля | |
|---|---|
name | Только вывод. Идентификатор. Сам токен. |
expireTime | Необязательно. Только для ввода. Неизменяемо. Необязательное время, по истечении которого при использовании полученного токена сообщения в сессиях BidiGenerateContent будут отклонены. (Gemini может превентивно закрыть сессию по истечении этого времени.) Если значение не задано, то по умолчанию устанавливается интервал в 30 минут. Если значение задано, оно должно быть меньше 20 часов. |
newSessionExpireTime | Необязательный параметр. Только для ввода. Неизменяемый. Время, по истечении которого новые сессии Live API, использующие токен, полученный в результате этого запроса, будут отклонены. Если значение не задано, по умолчанию устанавливается интервал в 60 секунд. Если задано, это значение должно быть меньше 20 часов. |
fieldMask | Необязательно. Только для ввода. Неизменяемо. Если field_mask пуст и Если field_mask пуст, а Если field_mask не пуст, то соответствующие поля из |
config поля объединения. Конфигурация, специфичная для метода, для результирующего токена. config может принимать только одно из следующих значений: | |
bidiGenerateContentSetup | Необязательный параметр. Только для ввода. Неизменяемый. Конфигурация, специфичная для |
uses | Необязательный параметр. Только для ввода. Неизменяемый. Количество раз, которое можно использовать токен. Если это значение равно нулю, ограничение не применяется. Возобновление сессии Live API не считается использованием. Если значение не указано, по умолчанию используется 1. |
Более подробная информация о распространенных типах.
Для получения дополнительной информации о часто используемых типах ресурсов API: Blob , Content , FunctionCall , FunctionResponse , GenerationConfig , GroundingMetadata , ModalityTokenCount и Tool , см. раздел «Генерация контента» .
Live API — это API с сохранением состояния, использующий WebSockets . В этом разделе вы найдете дополнительные сведения об API WebSockets.
Сессии
Соединение WebSocket устанавливает сессию между клиентом и сервером Gemini. После того, как клиент инициирует новое соединение, сессия может обмениваться сообщениями с сервером, чтобы:
- Отправляйте текст, аудио или видео на сервер Gemini.
- Получайте аудио-, текстовые запросы или запросы на выполнение функциональных задач с сервера Gemini.
WebSocket-соединение
Для начала сеанса подключитесь к следующему адресу веб-сокета:
wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent
Настройки сессии
Первоначальное сообщение, отправляемое после установления соединения WebSocket, задает конфигурацию сессии, которая включает в себя модель, параметры генерации, системные инструкции и инструменты.
Вы не можете обновить конфигурацию, пока соединение открыто. Однако вы можете изменить параметры конфигурации, за исключением модели, при приостановке и возобновлении сессии с помощью механизма возобновления сессии .
См. следующий пример конфигурации. Обратите внимание, что регистр имен в SDK может различаться. Параметры конфигурации Python SDK можно найти здесь .
{
"model": string,
"generationConfig": {
"candidateCount": integer,
"maxOutputTokens": integer,
"temperature": number,
"topP": number,
"topK": integer,
"presencePenalty": number,
"frequencyPenalty": number,
"responseModalities": [string],
"speechConfig": object,
"mediaResolution": object,
"translationConfig": object
},
"systemInstruction": string,
"tools": [object]
}
Для получения дополнительной информации о поле API см. generationConfig .
Отправляйте сообщения
Для обмена сообщениями по WebSocket-соединению клиент должен отправить JSON-объект по открытому WebSocket-соединению. JSON-объект должен содержать ровно одно из полей из следующего набора объектов:
{
"setup": BidiGenerateContentSetup,
"clientContent": BidiGenerateContentClientContent,
"realtimeInput": BidiGenerateContentRealtimeInput,
"toolResponse": BidiGenerateContentToolResponse
}
Поддерживаемые клиентские сообщения
Список поддерживаемых клиентских сообщений приведен в следующей таблице:
| Сообщение | Описание |
|---|---|
BidiGenerateContentSetup | Настройки сессии будут отправлены в первом сообщении. |
BidiGenerateContentClientContent | Постепенное обновление содержимого текущей беседы, предоставляемое клиентом. |
BidiGenerateContentRealtimeInput | Ввод аудио, видео или текста в реальном времени. |
BidiGenerateContentToolResponse | Ответ на сообщение ToolCallMessage полученное от сервера. |
Получайте сообщения
Для получения сообщений от Gemini необходимо прослушивать событие WebSocket 'message', а затем анализировать результат в соответствии с определением поддерживаемых серверных сообщений.
См. следующее:
async with client.aio.live.connect(model='...', config=config) as session:
await session.send(input='Hello world!', end_of_turn=True)
async for message in session.receive():
print(message)
Сообщения сервера могут содержать поле usageMetadata но в остальном будут включать ровно одно из других полей сообщения BidiGenerateContentServerMessage . (Объединение messageType не выражается в формате JSON, поэтому поле будет отображаться на верхнем уровне сообщения.)
Сообщения и события
Завершение активности
Этот тип не содержит полей.
Отмечает завершение активности пользователя.
Обработка действий
Различные способы обработки активности пользователей.
| Перечисления | |
|---|---|
ACTIVITY_HANDLING_UNSPECIFIED | Если параметр не указан, по умолчанию используется поведение START_OF_ACTIVITY_INTERRUPTS . |
START_OF_ACTIVITY_INTERRUPTS | Если это так, начало активности прервёт ответ модели (также называемое «вмешательством»). Текущий ответ модели будет прерван в момент прерывания. Это поведение по умолчанию. |
NO_INTERRUPTION | Реакция модели не будет прервана. |
ActivityStart
Этот тип не содержит полей.
Отмечает начало активности пользователя.
AudioTranscriptionConfig
Настройка транскрипции аудио.
| Поля | |
|---|---|
languageCodes[] | Необязательно. Коды языков BCP-47, указывающие на языки, присутствующие в аудиозаписи. Если поле опущено или пустое, по умолчанию используется автоматическое определение языка. |
customVocabulary[] | Необязательно. Список пользовательских лексических фраз, которые помогут модели распознавания речи ориентироваться на распознавание определенных терминов (названия продуктов, имена собственные, профессиональный жаргон). |
wordTimestamp | Необязательный параметр. Настраивает генерацию временных меток на уровне слов. |
diarization | Необязательный параметр. Настраивает диаризацию речи говорящего. |
mode | Необязательный параметр. Настраивает режим транскрипции. Поддерживаемые значения: |
Режим
Режим транскрипции.
| Перечисления | |
|---|---|
MODE_UNSPECIFIED | Неуказанный режим транскрипции. |
VERBATIM | Режим дословной транскрипции. |
SMART | Интеллектуальный режим транскрипции. |
Автоматическое определение активности
Настраивает автоматическое обнаружение активности.
| Поля | |
|---|---|
disabled | Необязательно. Если включено (по умолчанию), обнаруженный голосовой и текстовый ввод засчитывается как активность. Если отключено, клиент должен отправлять сигналы активности. |
startOfSpeechSensitivity | Необязательный параметр. Определяет вероятность обнаружения речи. |
prefixPaddingMs | Необязательно. Требуемая продолжительность обнаруженной речи до момента её начала. Чем ниже это значение, тем чувствительнее определение начала речи и тем короче распознаётся речь. Однако это также увеличивает вероятность ложных срабатываний. |
endOfSpeechSensitivity | Необязательный параметр. Определяет вероятность прекращения обнаруженной речи. |
silenceDurationMs | Необязательно. Требуемая продолжительность обнаруженного отсутствия речи (например, тишины) до окончания речи. Чем больше это значение, тем дольше могут быть паузы в речи, не прерывая активность пользователя, но это увеличит задержку модели. |
BidiGenerateContentClientContent
Постепенное обновление текущей переписки, предоставляемой клиентом. Весь контент безоговорочно добавляется к истории переписки и используется в качестве запроса к модели для генерации контента.
Сообщение, размещенное здесь, прервет процесс генерации текущей модели.
| Поля | |
|---|---|
turns[] | Необязательно. Содержимое, добавляемое к текущей беседе с моделью. Для запросов с одним циклом обработки это один экземпляр. Для запросов с несколькими циклами обработки это повторяющееся поле, содержащее историю переписки и последний запрос. |
turnComplete | Необязательный параметр. Если значение равно true, указывает, что генерация контента на сервере должна начинаться с текущего накопленного запроса. В противном случае сервер ожидает дополнительных сообщений перед началом генерации. |
BidiGenerateContentRealtimeInput
Пользовательский ввод, отправляемый в режиме реального времени.
Различные форматы (аудио, видео и текст) обрабатываются как параллельные потоки. Порядок потоков не гарантируется.
Это отличается от BidiGenerateContentClientContent по нескольким параметрам:
- Может непрерывно и без перерыва отправляться на этап генерации модели.
- Если возникает необходимость смешивать данные, чередующиеся между
BidiGenerateContentClientContentиBidiGenerateContentRealtimeInput, сервер пытается оптимизировать отклик для достижения наилучшего результата, но гарантий нет. - Конец реплики явно не указан, а определяется действиями пользователя (например, окончанием речи).
- Еще до окончания хода обработки данные обрабатываются поэтапно для оптимизации и обеспечения быстрого запуска реакции модели.
| Поля | |
|---|---|
mediaChunks[] | Необязательно. Встроенные байтовые данные для ввода мультимедиа. Использование нескольких УСТАРЕВШИЙ ВАРИАНТ: Используйте вместо него |
audio | Необязательно. Они формируют поток входного аудиосигнала в реальном времени. |
video | Необязательно. Они формируют поток видеовхода в реальном времени. |
activityStart | Необязательный параметр. Отмечает начало активности пользователя. Этот параметр может быть отправлен только в том случае, если автоматическое (т.е. на стороне сервера) обнаружение активности отключено. |
activityEnd | Необязательно. Отмечает окончание активности пользователя. Это сообщение может быть отправлено только в том случае, если автоматическое (т.е. на стороне сервера) обнаружение активности отключено. |
mediaResolution | Необязательный параметр. Разрешение мультимедиа для использования. Если не указано, используется значение |
audioStreamEnd | Необязательно. Указывает на то, что аудиопоток завершился, например, из-за выключения микрофона. Это сообщение следует отправлять только при включенном автоматическом обнаружении активности (что является настройкой по умолчанию). Клиент может возобновить трансляцию, отправив аудиосообщение. |
text | Необязательно. Они формируют поток ввода текста в реальном времени. |
BidiGenerateContentServerContent
Постепенные обновления сервера, генерируемые моделью в ответ на сообщения клиента.
Контент генерируется максимально быстро, но не в режиме реального времени. Клиенты могут по своему выбору буферизовать и воспроизводить его в режиме реального времени.
| Поля | |
|---|---|
generationComplete | Только вывод. Если значение равно true, это означает, что генерация модели завершена. Если генерация модели прерывается, сообщение 'generation_complete' в прерванном цикле не появится, процесс продолжится по схеме 'interrupted > turn_complete'. Когда модель предполагает воспроизведение в реальном времени, между generation_complete и turn_complete возникает задержка, вызванная ожиданием моделью завершения воспроизведения. |
turnComplete | Только вывод. Если значение равно true, это означает, что модель завершила свой ход. Генерация начнется только в ответ на дополнительные сообщения от клиента. Обратите внимание, что если включена отчетность о состоянии воспроизведения, она отправляется только тогда, когда состояние воспроизведения указывает на его завершение. Дальнейшие отчеты о состоянии воспроизведения той же генерации будут игнорироваться. |
interrupted | Только вывод. Если значение равно true, это означает, что сообщение от клиента прервало текущую генерацию модели. Если клиент воспроизводит контент в реальном времени, это хороший сигнал для остановки и очистки текущей очереди воспроизведения. |
groundingMetadata | Только для вывода. Метаданные для сгенерированного контента. |
inputTranscription | Только вывод. Входная аудиотранскрипция. Транскрипция отправляется независимо от других сообщений сервера, и гарантированного порядка нет. |
interimInputTranscription | Только вывод. Транскрипция с низкой задержкой обновляется во время речи пользователя. Это поле часто обновляется. |
outputTranscription | Только вывод. Аудиотранскрипция. Эти транскрипции являются частью выходных данных генерации сервера. Последняя выходная транскрипция этого хода отправляется перед |
urlContextMetadata | |
waitingForInput | Только вывод. Если значение равно true, это означает, что модель не генерирует контент, поскольку ожидает дополнительной информации от пользователя, например, потому что рассчитывает на продолжение разговора. |
speechState | Только вывод. УСТАРЕВШЕЕ: используйте VoiceActivity вместо него. Указывает текущее состояние распознавания речи в |
interactionStatus | Только вывод. Текущий статус активности сессии. Всегда отправляется вместе с |
modelTurn | Только выходные данные. Контент, сгенерированный моделью в рамках текущего диалога с пользователем. |
BidiGenerateContentServerMessage
Ответное сообщение для вызова функции BidiGenerateContent.
| Поля | |
|---|---|
usageMetadata | Только вывод. Метаданные об использовании ответа(ов). |
Поле объединения messageType . Тип сообщения. messageType может принимать только одно из следующих значений: | |
setupComplete | Только для вывода. Отправляется в ответ на сообщение |
serverContent | Только выходные данные. Контент, сгенерированный моделью в ответ на сообщения клиента. |
toolCall | Только вывод. Запрос к клиенту на выполнение |
toolCallCancellation | Только вывод. Уведомление для клиента о том, что ранее отправленное сообщение |
goAway | Только вывод. Уведомление о скором отключении сервера. |
sessionResumptionUpdate | Только вывод. Обновление состояния возобновления сессии. |
BidiGenerateContentSetup
Сообщение, которое будет отправлено в первом (и только в первом) BidiGenerateContentClientMessage . Содержит конфигурацию, которая будет применяться на протяжении всего процесса потокового RPC-вызова.
Клиентам следует дождаться сообщения BidiGenerateContentSetupComplete прежде чем отправлять какие-либо дополнительные сообщения.
| Поля | |
|---|---|
model | Обязательно. Имя ресурса модели. Оно служит идентификатором для используемой модели. Формат: |
generationConfig | Необязательно. Конфигурация генерации. Следующие поля не поддерживаются:
|
systemInstruction | Необязательно. Пользователь предоставил системные инструкции для данной модели. Примечание: текст должен использоваться только по частям, содержание каждой части должно быть в отдельном абзаце. |
tools[] | Необязательно. Список |
realtimeInputConfig | Необязательный параметр. Настраивает обработку ввода в реальном времени. |
sessionResumption | Необязательный параметр. Настраивает механизм возобновления сессии. Если эта опция включена, сервер будет отправлять сообщения |
contextWindowCompression | Необязательный параметр. Настраивает механизм сжатия контекстного окна. Если эта опция включена, сервер автоматически уменьшит размер контекста, когда он превысит заданную длину. |
inputAudioTranscription | Необязательный параметр. Если задан, включает транскрипцию голосового ввода. Транскрипция соответствует языку входного аудио, если это настроено. |
outputAudioTranscription | Необязательный параметр. Если задан, включает транскрипцию аудиовыхода модели. Транскрипция соответствует языковому коду, указанному для выходного аудио, если таковой задан. |
proactivity | Optional. Configures the proactivity of the model. This allows the model to respond proactively to the input and to ignore irrelevant input. |
historyConfig | Optional. Configures the exchange of history between the client and the server. |
BidiGenerateContentSetupComplete
Этот тип не содержит полей.
Sent in response to a BidiGenerateContentSetup message from the client.
BidiGenerateContentToolCall
Request for the client to execute the functionCalls and return the responses with the matching id s.
| Поля | |
|---|---|
functionCalls[] | Output only. The function call to be executed. |
BidiGenerateContentToolCallCancellation
Notification for the client that a previously issued ToolCallMessage with the specified id s should not have been executed and should be cancelled. If there were side-effects to those tool calls, clients may attempt to undo the tool calls. This message occurs only in cases where the clients interrupt server turns.
| Поля | |
|---|---|
ids[] | Output only. The ids of the tool calls to be cancelled. |
BidiGenerateContentToolResponse
Client generated response to a ToolCall received from the server. Individual FunctionResponse objects are matched to the respective FunctionCall objects by the id field.
Note that in the unary and server-streaming GenerateContent APIs function calling happens by exchanging the Content parts, while in the bidi GenerateContent APIs function calling happens over these dedicated set of messages.
| Поля | |
|---|---|
functionResponses[] | Optional. The response to the function calls. |
BidiGenerateContentTranscription
Transcription of audio (input or output).
| Поля | |
|---|---|
text | Transcription text. |
languageCode | The BCP-47 language code of the transcription. |
ContextWindowCompressionConfig
Enables context window compression — a mechanism for managing the model's context window so that it does not exceed a given length.
| Поля | |
|---|---|
Union field compressionMechanism . The context window compression mechanism used. compressionMechanism can be only one of the following: | |
slidingWindow | A sliding-window mechanism. |
triggerTokens | The number of tokens (before running a turn) required to trigger a context window compression. This can be used to balance quality against latency as shorter context windows may result in faster model responses. However, any compression operation will cause a temporary latency increase, so they should not be triggered frequently. If not set, the default is 80% of the model's context window limit. This leaves 20% for the next user request/model response. |
EndSensitivity
Determines how end of speech is detected.
| Перечисления | |
|---|---|
END_SENSITIVITY_UNSPECIFIED | The default is END_SENSITIVITY_HIGH. |
END_SENSITIVITY_HIGH | Automatic detection ends speech more often. |
END_SENSITIVITY_LOW | Automatic detection ends speech less often. |
Уходите
A notice that the server will soon disconnect.
| Поля | |
|---|---|
timeLeft | The remaining time before the connection will be terminated as ABORTED. This duration will never be less than a model-specific minimum, which will be specified together with the rate limits for the model. |
HistoryConfig
History configuration.
This message is included in the session configuration as BidiGenerateContentSetup.historyConfig . Configures the exchange of history messages.
| Поля | |
|---|---|
initialHistoryInClientContent | Optional. If true, after sending |
ProactivityConfig
Config for proactivity features.
| Поля | |
|---|---|
proactiveAudio | Optional. If enabled, the model can reject responding to the last prompt. For example, this allows the model to ignore out of context speech or to stay silent if the user did not make a request, yet. |
RealtimeInputConfig
Configures the realtime input behavior in BidiGenerateContent .
| Поля | |
|---|---|
automaticActivityDetection | Optional. If not set, automatic activity detection is enabled by default. If automatic voice detection is disabled, the client must send activity signals. |
activityHandling | Optional. Defines what effect activity has. |
turnCoverage | Optional. Defines which input is included in the user's turn. |
SessionResumptionConfig
Session resumption configuration.
This message is included in the session configuration as BidiGenerateContentSetup.sessionResumption . If configured, the server will send SessionResumptionUpdate messages.
| Поля | |
|---|---|
handle | The handle of a previous session. If not present then a new session is created. Session handles come from |
SessionResumptionUpdate
Update of the session resumption state.
Only sent if BidiGenerateContentSetup.sessionResumption was set.
| Поля | |
|---|---|
newHandle | New handle that represents a state that can be resumed. Empty if |
resumable | True if the current session can be resumed at this point. Resumption is not possible at some points in the session. For example, when the model is executing function calls or generating. Resuming the session (using a previous session token) in such a state will result in some data loss. In these cases, |
SlidingWindow
The SlidingWindow method operates by discarding content at the beginning of the context window. The resulting context will always begin at the start of a USER role turn. System instructions and any BidiGenerateContentSetup.prefixTurns will always remain at the beginning of the result.
| Поля | |
|---|---|
targetTokens | The target number of tokens to keep. The default value is trigger_tokens/2. Discarding parts of the context window causes a temporary latency increase so this value should be calibrated to avoid frequent compression operations. |
StartSensitivity
Determines how start of speech is detected.
| Перечисления | |
|---|---|
START_SENSITIVITY_UNSPECIFIED | The default is START_SENSITIVITY_HIGH. |
START_SENSITIVITY_HIGH | Automatic detection will detect the start of speech more often. |
START_SENSITIVITY_LOW | Automatic detection will detect the start of speech less often. |
TurnCoverage
Options about which input is included in the user's turn.
| Перечисления | |
|---|---|
TURN_COVERAGE_UNSPECIFIED | If unspecified, a default behavior is selected based on the model. Eg, for Gemini 2.5, the default is TURN_INCLUDES_ONLY_ACTIVITY , while for Gemini 3.1 and onwards, it's TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO . |
TURN_INCLUDES_ONLY_ACTIVITY | Includes activity since the last turn, excluding inactivity (eg silence on the audio stream). |
TURN_INCLUDES_ALL_INPUT | Includes all realtime input since the last turn, including inactivity (eg silence on the audio stream). |
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO | Includes audio activity and all video since the last turn. With automatic activity detection, audio activity means speech and excludes silence. |
TranslationConfig
Config for translation features.
| Поля | |
|---|---|
targetLanguageCode | Required. The target language for translation. Supported values are BCP-47 language codes (eg "en", "es", "fr"). |
echoTargetLanguage | Optional. If true, the model will generate audio when the target language is spoken, essentially it will parrot the input. If false, we will not produce audio for the target language. |
UrlContextMetadata
Metadata related to url context retrieval tool.
| Поля | |
|---|---|
urlMetadata[] | List of url context. |
UsageMetadata
Usage metadata about response(s).
| Поля | |
|---|---|
promptTokenCount | Output only. Number of tokens in the prompt. When |
cachedContentTokenCount | Количество токенов в кэшированной части запроса (кэшированное содержимое) |
responseTokenCount | Output only. Total number of tokens across all the generated response candidates. |
toolUsePromptTokenCount | Только вывод. Количество токенов, присутствующих в подсказках использования инструмента. |
thoughtsTokenCount | Только вывод. Количество токенов мыслей для моделей мышления. |
totalTokenCount | Output only. Total token count for the generation request (prompt + response candidates). |
promptTokensDetails[] | Только выходные данные. Список модальностей, которые были обработаны во входном запросе. |
cacheTokensDetails[] | Только вывод. Список вариантов содержимого, кэшированного в запросе. |
responseTokensDetails[] | Только вывод. Список модальностей, которые были возвращены в ответе. |
toolUsePromptTokensDetails[] | Только выходные данные. Список модальностей, обработанных для обработки запросов на использование инструмента. |
Ephemeral authentication tokens
Ephemeral authentication tokens can be obtained by calling AuthTokenService.CreateToken and then used with GenerativeService.BidiGenerateContentConstrained , either by passing the token in an access_token query parameter, or in an HTTP Authorization header with " Token " prefixed to it.
CreateAuthTokenRequest
Create an ephemeral authentication token.
| Поля | |
|---|---|
authToken | Required. The token to create. |
AuthToken
A request to create an ephemeral authentication token.
| Поля | |
|---|---|
name | Output only. Identifier. The token itself. |
expireTime | Optional. Input only. Immutable. An optional time after which, when using the resulting token, messages in BidiGenerateContent sessions will be rejected. (Gemini may preemptively close the session after this time.) If not set then this defaults to 30 minutes in the future. If set, this value must be less than 20 hours in the future. |
newSessionExpireTime | Optional. Input only. Immutable. The time after which new Live API sessions using the token resulting from this request will be rejected. If not set this defaults to 60 seconds in the future. If set, this value must be less than 20 hours in the future. |
fieldMask | Optional. Input only. Immutable. If field_mask is empty, and If field_mask is empty, and If field_mask is not empty, then the corresponding fields from |
Union field config . The method-specific configuration for the resulting token. config can be only one of the following: | |
bidiGenerateContentSetup | Optional. Input only. Immutable. Configuration specific to |
uses | Optional. Input only. Immutable. The number of times the token can be used. If this value is zero then no limit is applied. Resuming a Live API session does not count as a use. If unspecified, the default is 1. |
More information on common types
For more information on the commonly-used API resource types Blob , Content , FunctionCall , FunctionResponse , GenerationConfig , GroundingMetadata , ModalityTokenCount , and Tool , see Generating content .
The Live API is a stateful API that uses WebSockets . In this section, you'll find additional details regarding the WebSockets API.
Сессии
A WebSocket connection establishes a session between the client and the Gemini server. After a client initiates a new connection the session can exchange messages with the server to:
- Send text, audio, or video to the Gemini server.
- Receive audio, text, or function call requests from the Gemini server.
WebSocket connection
Для начала сеанса подключитесь к следующему адресу веб-сокета:
wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent
Session configuration
The initial message sent after establishing the WebSocket connection sets the session configuration, which includes the model, generation parameters, system instructions, and tools.
You cannot update the configuration while the connection is open. However, you can change the configuration parameters, except the model, when pausing and resuming via the session resumption mechanism .
См. следующий пример конфигурации. Обратите внимание, что регистр имен в SDK может различаться. Параметры конфигурации Python SDK можно найти здесь .
{
"model": string,
"generationConfig": {
"candidateCount": integer,
"maxOutputTokens": integer,
"temperature": number,
"topP": number,
"topK": integer,
"presencePenalty": number,
"frequencyPenalty": number,
"responseModalities": [string],
"speechConfig": object,
"mediaResolution": object,
"translationConfig": object
},
"systemInstruction": string,
"tools": [object]
}
For more information on the API field, see generationConfig .
Отправляйте сообщения
Для обмена сообщениями по WebSocket-соединению клиент должен отправить JSON-объект по открытому WebSocket-соединению. JSON-объект должен содержать ровно одно из полей из следующего набора объектов:
{
"setup": BidiGenerateContentSetup,
"clientContent": BidiGenerateContentClientContent,
"realtimeInput": BidiGenerateContentRealtimeInput,
"toolResponse": BidiGenerateContentToolResponse
}
Поддерживаемые клиентские сообщения
Список поддерживаемых клиентских сообщений приведен в следующей таблице:
| Сообщение | Описание |
|---|---|
BidiGenerateContentSetup | Session configuration to be sent in the first message |
BidiGenerateContentClientContent | Incremental content update of the current conversation delivered from the client |
BidiGenerateContentRealtimeInput | Real time audio, video, or text input |
BidiGenerateContentToolResponse | Response to a ToolCallMessage received from the server |
Получайте сообщения
To receive messages from Gemini, listen for the WebSocket 'message' event, and then parse the result according to the definition of the supported server messages.
См. следующее:
async with client.aio.live.connect(model='...', config=config) as session:
await session.send(input='Hello world!', end_of_turn=True)
async for message in session.receive():
print(message)
Server messages may have a usageMetadata field but will otherwise include exactly one of the other fields from the BidiGenerateContentServerMessage message. (The messageType union is not expressed in JSON so the field will appear at the top-level of the message.)
Сообщения и события
ActivityEnd
Этот тип не содержит полей.
Marks the end of user activity.
ActivityHandling
The different ways of handling user activity.
| Перечисления | |
|---|---|
ACTIVITY_HANDLING_UNSPECIFIED | If unspecified, the default behavior is START_OF_ACTIVITY_INTERRUPTS . |
START_OF_ACTIVITY_INTERRUPTS | If true, start of activity will interrupt the model's response (also called "barge in"). The model's current response will be cut-off in the moment of the interruption. This is the default behavior. |
NO_INTERRUPTION | The model's response will not be interrupted. |
ActivityStart
Этот тип не содержит полей.
Marks the start of user activity.
AudioTranscriptionConfig
The audio transcription configuration.
| Поля | |
|---|---|
languageCodes[] | Optional. BCP-47 language codes providing hints about the languages present in the audio. If omitted or empty, defaults to automatic language detection. |
customVocabulary[] | Optional. A list of custom vocabulary phrases to bias the speech recognition model toward recognizing specific terms (product names, proper nouns, jargon). |
wordTimestamp | Optional. Configures word-level timestamp generation. |
diarization | Optional. Configures speaker diarization. |
mode | Optional. Configures transcription mode. Supported values: |
Режим
Transcription mode.
| Перечисления | |
|---|---|
MODE_UNSPECIFIED | Unspecified transcription mode. |
VERBATIM | Verbatim transcription mode. |
SMART | Smart transcription mode. |
AutomaticActivityDetection
Configures automatic detection of activity.
| Поля | |
|---|---|
disabled | Optional. If enabled (the default), detected voice and text input count as activity. If disabled, the client must send activity signals. |
startOfSpeechSensitivity | Optional. Determines how likely speech is to be detected. |
prefixPaddingMs | Optional. The required duration of detected speech before start-of-speech is committed. The lower this value, the more sensitive the start-of-speech detection is and shorter speech can be recognized. However, this also increases the probability of false positives. |
endOfSpeechSensitivity | Optional. Determines how likely detected speech is ended. |
silenceDurationMs | Optional. The required duration of detected non-speech (eg silence) before end-of-speech is committed. The larger this value, the longer speech gaps can be without interrupting the user's activity but this will increase the model's latency. |
BidiGenerateContentClientContent
Incremental update of the current conversation delivered from the client. All of the content here is unconditionally appended to the conversation history and used as part of the prompt to the model to generate content.
A message here will interrupt any current model generation.
| Поля | |
|---|---|
turns[] | Optional. The content appended to the current conversation with the model. For single-turn queries, this is a single instance. For multi-turn queries, this is a repeated field that contains conversation history and the latest request. |
turnComplete | Optional. If true, indicates that the server content generation should start with the currently accumulated prompt. Otherwise, the server awaits additional messages before starting generation. |
BidiGenerateContentRealtimeInput
User input that is sent in real time.
The different modalities (audio, video and text) are handled as concurrent streams. The ordering across these streams is not guaranteed.
This is different from BidiGenerateContentClientContent in a few ways:
- Can be sent continuously without interruption to model generation.
- If there is a need to mix data interleaved across the
BidiGenerateContentClientContentand theBidiGenerateContentRealtimeInput, the server attempts to optimize for best response, but there are no guarantees. - End of turn is not explicitly specified, but is rather derived from user activity (for example, end of speech).
- Even before the end of turn, the data is processed incrementally to optimize for a fast start of the response from the model.
| Поля | |
|---|---|
mediaChunks[] | Optional. Inlined bytes data for media input. Multiple DEPRECATED: Use one of |
audio | Optional. These form the realtime audio input stream. |
video | Optional. These form the realtime video input stream. |
activityStart | Optional. Marks the start of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled. |
activityEnd | Optional. Marks the end of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled. |
mediaResolution | Optional. The media resolution to use. If not specified, |
audioStreamEnd | Optional. Indicates that the audio stream has ended, eg because the microphone was turned off. This should only be sent when automatic activity detection is enabled (which is the default). The client can reopen the stream by sending an audio message. |
text | Optional. These form the realtime text input stream. |
BidiGenerateContentServerContent
Постепенные обновления сервера, генерируемые моделью в ответ на сообщения клиента.
Контент генерируется максимально быстро, но не в режиме реального времени. Клиенты могут по своему выбору буферизовать и воспроизводить его в режиме реального времени.
| Поля | |
|---|---|
generationComplete | Output only. If true, indicates that the model is done generating. When model is interrupted while generating there will be no 'generation_complete' message in interrupted turn, it will go through 'interrupted > turn_complete'. When model assumes realtime playback there will be delay between generation_complete and turn_complete that is caused by model waiting for playback to finish. |
turnComplete | Output only. If true, indicates that the model has completed its turn. Generation will only start in response to additional client messages. Note when playback status reporting is enabled, this is emitted only when the playback status indicates that the playback is done. Future playback status of the same generation will be ignored. |
interrupted | Output only. If true, indicates that a client message has interrupted current model generation. If the client is playing out the content in real time, this is a good signal to stop and empty the current playback queue. |
groundingMetadata | Output only. Grounding metadata for the generated content. |
inputTranscription | Output only. Input audio transcription. The transcription is sent independently of the other server messages and there is no guaranteed ordering. |
interimInputTranscription | Output only. Low latency transcription updated while the user is speaking. This field is subject to frequent updates. |
outputTranscription | Output only. Output audio transcription. These transcriptions are part of the Generation output of the server. The last output transcription of this turn is sent before either |
urlContextMetadata | |
waitingForInput | Output only. If true, indicates that the model is not generating content because it is waiting for more input from the user, eg because it expects the user to continue talking. |
speechState | Output only. DEPRECATED: Use VoiceActivity instead. Indicates the current state of speech detection on |
interactionStatus | Output only. The current activity status of the live session. Always sent alongside |
modelTurn | Output only. The content that the model has generated as part of the current conversation with the user. |
BidiGenerateContentServerMessage
Response message for the BidiGenerateContent call.
| Поля | |
|---|---|
usageMetadata | Output only. Usage metadata about the response(s). |
Поле объединения messageType . Тип сообщения. messageType может принимать только одно из следующих значений: | |
setupComplete | Output only. Sent in response to a |
serverContent | Только выходные данные. Контент, сгенерированный моделью в ответ на сообщения клиента. |
toolCall | Output only. Request for the client to execute the |
toolCallCancellation | Output only. Notification for the client that a previously issued |
goAway | Output only. A notice that the server will soon disconnect. |
sessionResumptionUpdate | Output only. Update of the session resumption state. |
BidiGenerateContentSetup
Message to be sent in the first (and only in the first) BidiGenerateContentClientMessage . Contains configuration that will apply for the duration of the streaming RPC.
Clients should wait for a BidiGenerateContentSetupComplete message before sending any additional messages.
| Поля | |
|---|---|
model | Required. The model's resource name. This serves as an ID for the Model to use. Формат: |
generationConfig | Необязательно. Конфигурация генерации. Следующие поля не поддерживаются:
|
systemInstruction | Optional. The user provided system instructions for the model. Note: Only text should be used in parts and content in each part will be in a separate paragraph. |
tools[] | Optional. A list of A |
realtimeInputConfig | Optional. Configures the handling of realtime input. |
sessionResumption | Optional. Configures session resumption mechanism. If included, the server will send |
contextWindowCompression | Optional. Configures a context window compression mechanism. If included, the server will automatically reduce the size of the context when it exceeds the configured length. |
inputAudioTranscription | Optional. If set, enables transcription of voice input. The transcription aligns with the input audio language, if configured. |
outputAudioTranscription | Optional. If set, enables transcription of the model's audio output. The transcription aligns with the language code specified for the output audio, if configured. |
proactivity | Optional. Configures the proactivity of the model. This allows the model to respond proactively to the input and to ignore irrelevant input. |
historyConfig | Optional. Configures the exchange of history between the client and the server. |
BidiGenerateContentSetupComplete
Этот тип не содержит полей.
Sent in response to a BidiGenerateContentSetup message from the client.
BidiGenerateContentToolCall
Request for the client to execute the functionCalls and return the responses with the matching id s.
| Поля | |
|---|---|
functionCalls[] | Output only. The function call to be executed. |
BidiGenerateContentToolCallCancellation
Notification for the client that a previously issued ToolCallMessage with the specified id s should not have been executed and should be cancelled. If there were side-effects to those tool calls, clients may attempt to undo the tool calls. This message occurs only in cases where the clients interrupt server turns.
| Поля | |
|---|---|
ids[] | Output only. The ids of the tool calls to be cancelled. |
BidiGenerateContentToolResponse
Client generated response to a ToolCall received from the server. Individual FunctionResponse objects are matched to the respective FunctionCall objects by the id field.
Note that in the unary and server-streaming GenerateContent APIs function calling happens by exchanging the Content parts, while in the bidi GenerateContent APIs function calling happens over these dedicated set of messages.
| Поля | |
|---|---|
functionResponses[] | Optional. The response to the function calls. |
BidiGenerateContentTranscription
Transcription of audio (input or output).
| Поля | |
|---|---|
text | Transcription text. |
languageCode | The BCP-47 language code of the transcription. |
ContextWindowCompressionConfig
Enables context window compression — a mechanism for managing the model's context window so that it does not exceed a given length.
| Поля | |
|---|---|
Union field compressionMechanism . The context window compression mechanism used. compressionMechanism can be only one of the following: | |
slidingWindow | A sliding-window mechanism. |
triggerTokens | The number of tokens (before running a turn) required to trigger a context window compression. This can be used to balance quality against latency as shorter context windows may result in faster model responses. However, any compression operation will cause a temporary latency increase, so they should not be triggered frequently. If not set, the default is 80% of the model's context window limit. This leaves 20% for the next user request/model response. |
EndSensitivity
Determines how end of speech is detected.
| Перечисления | |
|---|---|
END_SENSITIVITY_UNSPECIFIED | The default is END_SENSITIVITY_HIGH. |
END_SENSITIVITY_HIGH | Automatic detection ends speech more often. |
END_SENSITIVITY_LOW | Automatic detection ends speech less often. |
Уходите
A notice that the server will soon disconnect.
| Поля | |
|---|---|
timeLeft | The remaining time before the connection will be terminated as ABORTED. This duration will never be less than a model-specific minimum, which will be specified together with the rate limits for the model. |
HistoryConfig
History configuration.
This message is included in the session configuration as BidiGenerateContentSetup.historyConfig . Configures the exchange of history messages.
| Поля | |
|---|---|
initialHistoryInClientContent | Optional. If true, after sending |
ProactivityConfig
Config for proactivity features.
| Поля | |
|---|---|
proactiveAudio | Optional. If enabled, the model can reject responding to the last prompt. For example, this allows the model to ignore out of context speech or to stay silent if the user did not make a request, yet. |
RealtimeInputConfig
Configures the realtime input behavior in BidiGenerateContent .
| Поля | |
|---|---|
automaticActivityDetection | Optional. If not set, automatic activity detection is enabled by default. If automatic voice detection is disabled, the client must send activity signals. |
activityHandling | Optional. Defines what effect activity has. |
turnCoverage | Optional. Defines which input is included in the user's turn. |
SessionResumptionConfig
Session resumption configuration.
This message is included in the session configuration as BidiGenerateContentSetup.sessionResumption . If configured, the server will send SessionResumptionUpdate messages.
| Поля | |
|---|---|
handle | The handle of a previous session. If not present then a new session is created. Session handles come from |
SessionResumptionUpdate
Update of the session resumption state.
Only sent if BidiGenerateContentSetup.sessionResumption was set.
| Поля | |
|---|---|
newHandle | New handle that represents a state that can be resumed. Empty if |
resumable | True if the current session can be resumed at this point. Resumption is not possible at some points in the session. For example, when the model is executing function calls or generating. Resuming the session (using a previous session token) in such a state will result in some data loss. In these cases, |
SlidingWindow
The SlidingWindow method operates by discarding content at the beginning of the context window. The resulting context will always begin at the start of a USER role turn. System instructions and any BidiGenerateContentSetup.prefixTurns will always remain at the beginning of the result.
| Поля | |
|---|---|
targetTokens | The target number of tokens to keep. The default value is trigger_tokens/2. Discarding parts of the context window causes a temporary latency increase so this value should be calibrated to avoid frequent compression operations. |
StartSensitivity
Determines how start of speech is detected.
| Перечисления | |
|---|---|
START_SENSITIVITY_UNSPECIFIED | The default is START_SENSITIVITY_HIGH. |
START_SENSITIVITY_HIGH | Automatic detection will detect the start of speech more often. |
START_SENSITIVITY_LOW | Automatic detection will detect the start of speech less often. |
TurnCoverage
Options about which input is included in the user's turn.
| Перечисления | |
|---|---|
TURN_COVERAGE_UNSPECIFIED | If unspecified, a default behavior is selected based on the model. Eg, for Gemini 2.5, the default is TURN_INCLUDES_ONLY_ACTIVITY , while for Gemini 3.1 and onwards, it's TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO . |
TURN_INCLUDES_ONLY_ACTIVITY | Includes activity since the last turn, excluding inactivity (eg silence on the audio stream). |
TURN_INCLUDES_ALL_INPUT | Includes all realtime input since the last turn, including inactivity (eg silence on the audio stream). |
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO | Includes audio activity and all video since the last turn. With automatic activity detection, audio activity means speech and excludes silence. |
TranslationConfig
Config for translation features.
| Поля | |
|---|---|
targetLanguageCode | Required. The target language for translation. Supported values are BCP-47 language codes (eg "en", "es", "fr"). |
echoTargetLanguage | Optional. If true, the model will generate audio when the target language is spoken, essentially it will parrot the input. If false, we will not produce audio for the target language. |
UrlContextMetadata
Metadata related to url context retrieval tool.
| Поля | |
|---|---|
urlMetadata[] | List of url context. |
UsageMetadata
Usage metadata about response(s).
| Поля | |
|---|---|
promptTokenCount | Output only. Number of tokens in the prompt. When |
cachedContentTokenCount | Количество токенов в кэшированной части запроса (кэшированное содержимое) |
responseTokenCount | Output only. Total number of tokens across all the generated response candidates. |
toolUsePromptTokenCount | Только вывод. Количество токенов, присутствующих в подсказках использования инструмента. |
thoughtsTokenCount | Только вывод. Количество токенов мыслей для моделей мышления. |
totalTokenCount | Output only. Total token count for the generation request (prompt + response candidates). |
promptTokensDetails[] | Только выходные данные. Список модальностей, которые были обработаны во входном запросе. |
cacheTokensDetails[] | Только вывод. Список вариантов содержимого, кэшированного в запросе. |
responseTokensDetails[] | Только вывод. Список модальностей, которые были возвращены в ответе. |
toolUsePromptTokensDetails[] | Только выходные данные. Список модальностей, обработанных для обработки запросов на использование инструмента. |
Ephemeral authentication tokens
Ephemeral authentication tokens can be obtained by calling AuthTokenService.CreateToken and then used with GenerativeService.BidiGenerateContentConstrained , either by passing the token in an access_token query parameter, or in an HTTP Authorization header with " Token " prefixed to it.
CreateAuthTokenRequest
Create an ephemeral authentication token.
| Поля | |
|---|---|
authToken | Required. The token to create. |
AuthToken
A request to create an ephemeral authentication token.
| Поля | |
|---|---|
name | Output only. Identifier. The token itself. |
expireTime | Optional. Input only. Immutable. An optional time after which, when using the resulting token, messages in BidiGenerateContent sessions will be rejected. (Gemini may preemptively close the session after this time.) If not set then this defaults to 30 minutes in the future. If set, this value must be less than 20 hours in the future. |
newSessionExpireTime | Optional. Input only. Immutable. The time after which new Live API sessions using the token resulting from this request will be rejected. If not set this defaults to 60 seconds in the future. If set, this value must be less than 20 hours in the future. |
fieldMask | Optional. Input only. Immutable. If field_mask is empty, and If field_mask is empty, and If field_mask is not empty, then the corresponding fields from |
Union field config . The method-specific configuration for the resulting token. config can be only one of the following: | |
bidiGenerateContentSetup | Optional. Input only. Immutable. Configuration specific to |
uses | Optional. Input only. Immutable. The number of times the token can be used. If this value is zero then no limit is applied. Resuming a Live API session does not count as a use. If unspecified, the default is 1. |
More information on common types
For more information on the commonly-used API resource types Blob , Content , FunctionCall , FunctionResponse , GenerationConfig , GroundingMetadata , ModalityTokenCount , and Tool , see Generating content .
The Live API is a stateful API that uses WebSockets . In this section, you'll find additional details regarding the WebSockets API.
Сессии
A WebSocket connection establishes a session between the client and the Gemini server. After a client initiates a new connection the session can exchange messages with the server to:
- Send text, audio, or video to the Gemini server.
- Receive audio, text, or function call requests from the Gemini server.
WebSocket connection
Для начала сеанса подключитесь к следующему адресу веб-сокета:
wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent
Session configuration
The initial message sent after establishing the WebSocket connection sets the session configuration, which includes the model, generation parameters, system instructions, and tools.
You cannot update the configuration while the connection is open. However, you can change the configuration parameters, except the model, when pausing and resuming via the session resumption mechanism .
См. следующий пример конфигурации. Обратите внимание, что регистр имен в SDK может различаться. Параметры конфигурации Python SDK можно найти здесь .
{
"model": string,
"generationConfig": {
"candidateCount": integer,
"maxOutputTokens": integer,
"temperature": number,
"topP": number,
"topK": integer,
"presencePenalty": number,
"frequencyPenalty": number,
"responseModalities": [string],
"speechConfig": object,
"mediaResolution": object,
"translationConfig": object
},
"systemInstruction": string,
"tools": [object]
}
For more information on the API field, see generationConfig .
Отправляйте сообщения
Для обмена сообщениями по WebSocket-соединению клиент должен отправить JSON-объект по открытому WebSocket-соединению. JSON-объект должен содержать ровно одно из полей из следующего набора объектов:
{
"setup": BidiGenerateContentSetup,
"clientContent": BidiGenerateContentClientContent,
"realtimeInput": BidiGenerateContentRealtimeInput,
"toolResponse": BidiGenerateContentToolResponse
}
Поддерживаемые клиентские сообщения
Список поддерживаемых клиентских сообщений приведен в следующей таблице:
| Сообщение | Описание |
|---|---|
BidiGenerateContentSetup | Session configuration to be sent in the first message |
BidiGenerateContentClientContent | Incremental content update of the current conversation delivered from the client |
BidiGenerateContentRealtimeInput | Real time audio, video, or text input |
BidiGenerateContentToolResponse | Response to a ToolCallMessage received from the server |
Получайте сообщения
To receive messages from Gemini, listen for the WebSocket 'message' event, and then parse the result according to the definition of the supported server messages.
См. следующее:
async with client.aio.live.connect(model='...', config=config) as session:
await session.send(input='Hello world!', end_of_turn=True)
async for message in session.receive():
print(message)
Server messages may have a usageMetadata field but will otherwise include exactly one of the other fields from the BidiGenerateContentServerMessage message. (The messageType union is not expressed in JSON so the field will appear at the top-level of the message.)
Сообщения и события
ActivityEnd
Этот тип не содержит полей.
Marks the end of user activity.
ActivityHandling
The different ways of handling user activity.
| Перечисления | |
|---|---|
ACTIVITY_HANDLING_UNSPECIFIED | If unspecified, the default behavior is START_OF_ACTIVITY_INTERRUPTS . |
START_OF_ACTIVITY_INTERRUPTS | If true, start of activity will interrupt the model's response (also called "barge in"). The model's current response will be cut-off in the moment of the interruption. This is the default behavior. |
NO_INTERRUPTION | The model's response will not be interrupted. |
ActivityStart
Этот тип не содержит полей.
Marks the start of user activity.
AudioTranscriptionConfig
The audio transcription configuration.
| Поля | |
|---|---|
languageCodes[] | Optional. BCP-47 language codes providing hints about the languages present in the audio. If omitted or empty, defaults to automatic language detection. |
customVocabulary[] | Optional. A list of custom vocabulary phrases to bias the speech recognition model toward recognizing specific terms (product names, proper nouns, jargon). |
wordTimestamp | Optional. Configures word-level timestamp generation. |
diarization | Optional. Configures speaker diarization. |
mode | Optional. Configures transcription mode. Supported values: |
Режим
Transcription mode.
| Перечисления | |
|---|---|
MODE_UNSPECIFIED | Unspecified transcription mode. |
VERBATIM | Verbatim transcription mode. |
SMART | Smart transcription mode. |
AutomaticActivityDetection
Configures automatic detection of activity.
| Поля | |
|---|---|
disabled | Optional. If enabled (the default), detected voice and text input count as activity. If disabled, the client must send activity signals. |
startOfSpeechSensitivity | Optional. Determines how likely speech is to be detected. |
prefixPaddingMs | Optional. The required duration of detected speech before start-of-speech is committed. The lower this value, the more sensitive the start-of-speech detection is and shorter speech can be recognized. However, this also increases the probability of false positives. |
endOfSpeechSensitivity | Optional. Determines how likely detected speech is ended. |
silenceDurationMs | Optional. The required duration of detected non-speech (eg silence) before end-of-speech is committed. The larger this value, the longer speech gaps can be without interrupting the user's activity but this will increase the model's latency. |
BidiGenerateContentClientContent
Incremental update of the current conversation delivered from the client. All of the content here is unconditionally appended to the conversation history and used as part of the prompt to the model to generate content.
A message here will interrupt any current model generation.
| Поля | |
|---|---|
turns[] | Optional. The content appended to the current conversation with the model. For single-turn queries, this is a single instance. For multi-turn queries, this is a repeated field that contains conversation history and the latest request. |
turnComplete | Optional. If true, indicates that the server content generation should start with the currently accumulated prompt. Otherwise, the server awaits additional messages before starting generation. |
BidiGenerateContentRealtimeInput
User input that is sent in real time.
The different modalities (audio, video and text) are handled as concurrent streams. The ordering across these streams is not guaranteed.
This is different from BidiGenerateContentClientContent in a few ways:
- Can be sent continuously without interruption to model generation.
- If there is a need to mix data interleaved across the
BidiGenerateContentClientContentand theBidiGenerateContentRealtimeInput, the server attempts to optimize for best response, but there are no guarantees. - End of turn is not explicitly specified, but is rather derived from user activity (for example, end of speech).
- Even before the end of turn, the data is processed incrementally to optimize for a fast start of the response from the model.
| Поля | |
|---|---|
mediaChunks[] | Optional. Inlined bytes data for media input. Multiple DEPRECATED: Use one of |
audio | Optional. These form the realtime audio input stream. |
video | Optional. These form the realtime video input stream. |
activityStart | Optional. Marks the start of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled. |
activityEnd | Optional. Marks the end of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled. |
mediaResolution | Optional. The media resolution to use. If not specified, |
audioStreamEnd | Optional. Indicates that the audio stream has ended, eg because the microphone was turned off. This should only be sent when automatic activity detection is enabled (which is the default). The client can reopen the stream by sending an audio message. |
text | Optional. These form the realtime text input stream. |
BidiGenerateContentServerContent
Постепенные обновления сервера, генерируемые моделью в ответ на сообщения клиента.
Контент генерируется максимально быстро, но не в режиме реального времени. Клиенты могут по своему выбору буферизовать и воспроизводить его в режиме реального времени.
| Поля | |
|---|---|
generationComplete | Output only. If true, indicates that the model is done generating. When model is interrupted while generating there will be no 'generation_complete' message in interrupted turn, it will go through 'interrupted > turn_complete'. When model assumes realtime playback there will be delay between generation_complete and turn_complete that is caused by model waiting for playback to finish. |
turnComplete | Output only. If true, indicates that the model has completed its turn. Generation will only start in response to additional client messages. Note when playback status reporting is enabled, this is emitted only when the playback status indicates that the playback is done. Future playback status of the same generation will be ignored. |
interrupted | Output only. If true, indicates that a client message has interrupted current model generation. If the client is playing out the content in real time, this is a good signal to stop and empty the current playback queue. |
groundingMetadata | Output only. Grounding metadata for the generated content. |
inputTranscription | Output only. Input audio transcription. The transcription is sent independently of the other server messages and there is no guaranteed ordering. |
interimInputTranscription | Output only. Low latency transcription updated while the user is speaking. This field is subject to frequent updates. |
outputTranscription | Output only. Output audio transcription. These transcriptions are part of the Generation output of the server. The last output transcription of this turn is sent before either |
urlContextMetadata | |
waitingForInput | Output only. If true, indicates that the model is not generating content because it is waiting for more input from the user, eg because it expects the user to continue talking. |
speechState | Output only. DEPRECATED: Use VoiceActivity instead. Indicates the current state of speech detection on |
interactionStatus | Output only. The current activity status of the live session. Always sent alongside |
modelTurn | Output only. The content that the model has generated as part of the current conversation with the user. |
BidiGenerateContentServerMessage
Response message for the BidiGenerateContent call.
| Поля | |
|---|---|
usageMetadata | Output only. Usage metadata about the response(s). |
Поле объединения messageType . Тип сообщения. messageType может принимать только одно из следующих значений: | |
setupComplete | Output only. Sent in response to a |
serverContent | Только выходные данные. Контент, сгенерированный моделью в ответ на сообщения клиента. |
toolCall | Output only. Request for the client to execute the |
toolCallCancellation | Output only. Notification for the client that a previously issued |
goAway | Output only. A notice that the server will soon disconnect. |
sessionResumptionUpdate | Output only. Update of the session resumption state. |
BidiGenerateContentSetup
Message to be sent in the first (and only in the first) BidiGenerateContentClientMessage . Contains configuration that will apply for the duration of the streaming RPC.
Clients should wait for a BidiGenerateContentSetupComplete message before sending any additional messages.
| Поля | |
|---|---|
model | Required. The model's resource name. This serves as an ID for the Model to use. Формат: |
generationConfig | Необязательно. Конфигурация генерации. Следующие поля не поддерживаются:
|
systemInstruction | Optional. The user provided system instructions for the model. Note: Only text should be used in parts and content in each part will be in a separate paragraph. |
tools[] | Optional. A list of A |
realtimeInputConfig | Optional. Configures the handling of realtime input. |
sessionResumption | Optional. Configures session resumption mechanism. If included, the server will send |
contextWindowCompression | Optional. Configures a context window compression mechanism. If included, the server will automatically reduce the size of the context when it exceeds the configured length. |
inputAudioTranscription | Optional. If set, enables transcription of voice input. The transcription aligns with the input audio language, if configured. |
outputAudioTranscription | Optional. If set, enables transcription of the model's audio output. The transcription aligns with the language code specified for the output audio, if configured. |
proactivity | Optional. Configures the proactivity of the model. This allows the model to respond proactively to the input and to ignore irrelevant input. |
historyConfig | Optional. Configures the exchange of history between the client and the server. |
BidiGenerateContentSetupComplete
Этот тип не содержит полей.
Sent in response to a BidiGenerateContentSetup message from the client.
BidiGenerateContentToolCall
Request for the client to execute the functionCalls and return the responses with the matching id s.
| Поля | |
|---|---|
functionCalls[] | Output only. The function call to be executed. |
BidiGenerateContentToolCallCancellation
Notification for the client that a previously issued ToolCallMessage with the specified id s should not have been executed and should be cancelled. If there were side-effects to those tool calls, clients may attempt to undo the tool calls. This message occurs only in cases where the clients interrupt server turns.
| Поля | |
|---|---|
ids[] | Output only. The ids of the tool calls to be cancelled. |
BidiGenerateContentToolResponse
Client generated response to a ToolCall received from the server. Individual FunctionResponse objects are matched to the respective FunctionCall objects by the id field.
Note that in the unary and server-streaming GenerateContent APIs function calling happens by exchanging the Content parts, while in the bidi GenerateContent APIs function calling happens over these dedicated set of messages.
| Поля | |
|---|---|
functionResponses[] | Optional. The response to the function calls. |
BidiGenerateContentTranscription
Transcription of audio (input or output).
| Поля | |
|---|---|
text | Transcription text. |
languageCode | The BCP-47 language code of the transcription. |
ContextWindowCompressionConfig
Enables context window compression — a mechanism for managing the model's context window so that it does not exceed a given length.
| Поля | |
|---|---|
Union field compressionMechanism . The context window compression mechanism used. compressionMechanism can be only one of the following: | |
slidingWindow | A sliding-window mechanism. |
triggerTokens | The number of tokens (before running a turn) required to trigger a context window compression. This can be used to balance quality against latency as shorter context windows may result in faster model responses. However, any compression operation will cause a temporary latency increase, so they should not be triggered frequently. If not set, the default is 80% of the model's context window limit. This leaves 20% for the next user request/model response. |
EndSensitivity
Determines how end of speech is detected.
| Перечисления | |
|---|---|
END_SENSITIVITY_UNSPECIFIED | The default is END_SENSITIVITY_HIGH. |
END_SENSITIVITY_HIGH | Automatic detection ends speech more often. |
END_SENSITIVITY_LOW | Automatic detection ends speech less often. |
Уходите
A notice that the server will soon disconnect.
| Поля | |
|---|---|
timeLeft | The remaining time before the connection will be terminated as ABORTED. This duration will never be less than a model-specific minimum, which will be specified together with the rate limits for the model. |
HistoryConfig
History configuration.
This message is included in the session configuration as BidiGenerateContentSetup.historyConfig . Configures the exchange of history messages.
| Поля | |
|---|---|
initialHistoryInClientContent | Optional. If true, after sending |
ProactivityConfig
Config for proactivity features.
| Поля | |
|---|---|
proactiveAudio | Optional. If enabled, the model can reject responding to the last prompt. For example, this allows the model to ignore out of context speech or to stay silent if the user did not make a request, yet. |
RealtimeInputConfig
Configures the realtime input behavior in BidiGenerateContent .
| Поля | |
|---|---|
automaticActivityDetection | Optional. If not set, automatic activity detection is enabled by default. If automatic voice detection is disabled, the client must send activity signals. |
activityHandling | Optional. Defines what effect activity has. |
turnCoverage | Optional. Defines which input is included in the user's turn. |
SessionResumptionConfig
Session resumption configuration.
This message is included in the session configuration as BidiGenerateContentSetup.sessionResumption . If configured, the server will send SessionResumptionUpdate messages.
| Поля | |
|---|---|
handle | The handle of a previous session. If not present then a new session is created. Session handles come from |
SessionResumptionUpdate
Update of the session resumption state.
Only sent if BidiGenerateContentSetup.sessionResumption was set.
| Поля | |
|---|---|
newHandle | New handle that represents a state that can be resumed. Empty if |
resumable | True if the current session can be resumed at this point. Resumption is not possible at some points in the session. For example, when the model is executing function calls or generating. Resuming the session (using a previous session token) in such a state will result in some data loss. In these cases, |
SlidingWindow
The SlidingWindow method operates by discarding content at the beginning of the context window. The resulting context will always begin at the start of a USER role turn. System instructions and any BidiGenerateContentSetup.prefixTurns will always remain at the beginning of the result.
| Поля | |
|---|---|
targetTokens | The target number of tokens to keep. The default value is trigger_tokens/2. Discarding parts of the context window causes a temporary latency increase so this value should be calibrated to avoid frequent compression operations. |
StartSensitivity
Determines how start of speech is detected.
| Перечисления | |
|---|---|
START_SENSITIVITY_UNSPECIFIED | The default is START_SENSITIVITY_HIGH. |
START_SENSITIVITY_HIGH | Automatic detection will detect the start of speech more often. |
START_SENSITIVITY_LOW | Automatic detection will detect the start of speech less often. |
TurnCoverage
Options about which input is included in the user's turn.
| Перечисления | |
|---|---|
TURN_COVERAGE_UNSPECIFIED | If unspecified, a default behavior is selected based on the model. Eg, for Gemini 2.5, the default is TURN_INCLUDES_ONLY_ACTIVITY , while for Gemini 3.1 and onwards, it's TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO . |
TURN_INCLUDES_ONLY_ACTIVITY | Includes activity since the last turn, excluding inactivity (eg silence on the audio stream). |
TURN_INCLUDES_ALL_INPUT | Includes all realtime input since the last turn, including inactivity (eg silence on the audio stream). |
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO | Includes audio activity and all video since the last turn. With automatic activity detection, audio activity means speech and excludes silence. |
TranslationConfig
Config for translation features.
| Поля | |
|---|---|
targetLanguageCode | Required. The target language for translation. Supported values are BCP-47 language codes (eg "en", "es", "fr"). |
echoTargetLanguage | Optional. If true, the model will generate audio when the target language is spoken, essentially it will parrot the input. If false, we will not produce audio for the target language. |
UrlContextMetadata
Metadata related to url context retrieval tool.
| Поля | |
|---|---|
urlMetadata[] | List of url context. |
UsageMetadata
Usage metadata about response(s).
| Поля | |
|---|---|
promptTokenCount | Output only. Number of tokens in the prompt. When |
cachedContentTokenCount | Количество токенов в кэшированной части запроса (кэшированное содержимое) |
responseTokenCount | Output only. Total number of tokens across all the generated response candidates. |
toolUsePromptTokenCount | Только вывод. Количество токенов, присутствующих в подсказках использования инструмента. |
thoughtsTokenCount | Только вывод. Количество токенов мыслей для моделей мышления. |
totalTokenCount | Output only. Total token count for the generation request (prompt + response candidates). |
promptTokensDetails[] | Только выходные данные. Список модальностей, которые были обработаны во входном запросе. |
cacheTokensDetails[] | Только вывод. Список вариантов содержимого, кэшированного в запросе. |
responseTokensDetails[] | Только вывод. Список модальностей, которые были возвращены в ответе. |
toolUsePromptTokensDetails[] | Только выходные данные. Список модальностей, обработанных для обработки запросов на использование инструмента. |
Ephemeral authentication tokens
Ephemeral authentication tokens can be obtained by calling AuthTokenService.CreateToken and then used with GenerativeService.BidiGenerateContentConstrained , either by passing the token in an access_token query parameter, or in an HTTP Authorization header with " Token " prefixed to it.
CreateAuthTokenRequest
Create an ephemeral authentication token.
| Поля | |
|---|---|
authToken | Required. The token to create. |
AuthToken
A request to create an ephemeral authentication token.
| Поля | |
|---|---|
name | Output only. Identifier. The token itself. |
expireTime | Optional. Input only. Immutable. An optional time after which, when using the resulting token, messages in BidiGenerateContent sessions will be rejected. (Gemini may preemptively close the session after this time.) If not set then this defaults to 30 minutes in the future. If set, this value must be less than 20 hours in the future. |
newSessionExpireTime | Optional. Input only. Immutable. The time after which new Live API sessions using the token resulting from this request will be rejected. If not set this defaults to 60 seconds in the future. If set, this value must be less than 20 hours in the future. |
fieldMask | Optional. Input only. Immutable. If field_mask is empty, and If field_mask is empty, and If field_mask is not empty, then the corresponding fields from |
Union field config . The method-specific configuration for the resulting token. config can be only one of the following: | |
bidiGenerateContentSetup | Optional. Input only. Immutable. Configuration specific to |
uses | Optional. Input only. Immutable. The number of times the token can be used. If this value is zero then no limit is applied. Resuming a Live API session does not count as a use. If unspecified, the default is 1. |
More information on common types
For more information on the commonly-used API resource types Blob , Content , FunctionCall , FunctionResponse , GenerationConfig , GroundingMetadata , ModalityTokenCount , and Tool , see Generating content .