Live API - WebSockets API reference

Live API یک API با قابلیت Stateful است که از WebSockets استفاده می‌کند. در این بخش، جزئیات بیشتری در مورد WebSockets API خواهید یافت.

جلسات

یک اتصال WebSocket یک جلسه بین کلاینت و سرور Gemini برقرار می‌کند. پس از اینکه کلاینت یک اتصال جدید را آغاز می‌کند، جلسه می‌تواند پیام‌هایی را با سرور رد و بدل کند تا:

  • متن، صدا یا ویدیو را به سرور جمینی ارسال کنید.
  • درخواست‌های صوتی، متنی یا تماس عملکردی را از سرور Gemini دریافت کنید.

اتصال وب‌سوکت

برای شروع یک جلسه، به این نقطه پایانی وب سوکت متصل شوید:

wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent

پیکربندی جلسه

پیام اولیه‌ای که پس از برقراری اتصال WebSocket ارسال می‌شود، پیکربندی جلسه را تنظیم می‌کند که شامل مدل، پارامترهای تولید، دستورالعمل‌های سیستم و ابزارها می‌شود.

شما نمی‌توانید پیکربندی را در حالی که اتصال باز است، به‌روزرسانی کنید. با این حال، می‌توانید پارامترهای پیکربندی، به جز مدل، را هنگام مکث و از سرگیری از طریق مکانیسم از سرگیری جلسه تغییر دهید.

به مثال پیکربندی زیر توجه کنید. توجه داشته باشید که نحوه‌ی نام‌گذاری در SDKها ممکن است متفاوت باشد. می‌توانید گزینه‌های پیکربندی SDK پایتون را اینجا جستجو کنید.


{
  "model": string,
  "generationConfig": {
    "candidateCount": integer,
    "maxOutputTokens": integer,
    "temperature": number,
    "topP": number,
    "topK": integer,
    "presencePenalty": number,
    "frequencyPenalty": number,
    "responseModalities": [string],
    "speechConfig": object,
    "mediaResolution": object,
    "translationConfig": object
  },
  "systemInstruction": string,
  "tools": [object]
}

برای اطلاعات بیشتر در مورد فیلد API، به generationConfig مراجعه کنید.

ارسال پیام

برای تبادل پیام از طریق اتصال WebSocket، کلاینت باید یک شیء JSON را از طریق یک اتصال WebSocket باز ارسال کند. شیء JSON باید دقیقاً یکی از فیلدهای مجموعه شیء زیر را داشته باشد:


{
  "setup": BidiGenerateContentSetup,
  "clientContent": BidiGenerateContentClientContent,
  "realtimeInput": BidiGenerateContentRealtimeInput,
  "toolResponse": BidiGenerateContentToolResponse
}

پیام‌های کلاینت پشتیبانی‌شده

پیام‌های کلاینت پشتیبانی‌شده را در جدول زیر مشاهده کنید:

پیام توضیحات
BidiGenerateContentSetup پیکربندی جلسه که باید در اولین پیام ارسال شود
BidiGenerateContentClientContent به‌روزرسانی تدریجی محتوای مکالمه فعلی ارائه شده از طرف کلاینت
BidiGenerateContentRealtimeInput ورودی صوتی، تصویری یا متنی بلادرنگ
BidiGenerateContentToolResponse پاسخ به یک ToolCallMessage دریافت شده از سرور

دریافت پیام‌ها

برای دریافت پیام‌ها از Gemini، به رویداد 'message' در WebSocket گوش دهید و سپس نتیجه را طبق تعریف پیام‌های سرور پشتیبانی‌شده تجزیه کنید.

موارد زیر را ببینید:

async with client.aio.live.connect(model='...', config=config) as session:
    await session.send(input='Hello world!', end_of_turn=True)
    async for message in session.receive():
        print(message)

پیام‌های سرور ممکن است دارای فیلد usageMetadata باشند، اما در غیر این صورت دقیقاً یکی از فیلدهای دیگر پیام BidiGenerateContentServerMessage را شامل می‌شوند. (اتحادیه messageType در JSON بیان نشده است، بنابراین این فیلد در سطح بالای پیام ظاهر می‌شود.)

پیام‌ها و رویدادها

پایان فعالیت

این نوع هیچ فیلدی ندارد.

پایان فعالیت کاربر را نشان می‌دهد.

مدیریت فعالیت

روش‌های مختلف مدیریت فعالیت کاربران

انوم‌ها
ACTIVITY_HANDLING_UNSPECIFIED اگر مشخص نشده باشد، رفتار پیش‌فرض START_OF_ACTIVITY_INTERRUPTS است.
START_OF_ACTIVITY_INTERRUPTS اگر درست باشد، شروع فعالیت، پاسخ مدل را قطع می‌کند (که به آن "barge in" نیز می‌گویند). پاسخ فعلی مدل در لحظه وقفه قطع می‌شود. این رفتار پیش‌فرض است.
NO_INTERRUPTION پاسخ مدل قطع نخواهد شد.

شروع فعالیت

این نوع هیچ فیلدی ندارد.

شروع فعالیت کاربر را نشان می‌دهد.

پیکربندی رونویسی صوتی

پیکربندی رونویسی صوتی.

فیلدها
languageCodes[]

string

اختیاری. کدهای زبان BCP-47 نکاتی در مورد زبان‌های موجود در صدا ارائه می‌دهند. در صورت حذف یا خالی بودن، به طور پیش‌فرض روی تشخیص خودکار زبان تنظیم می‌شود.

customVocabulary[]

string

اختیاری. فهرستی از عبارات واژگانی سفارشی برای جهت‌دهی مدل تشخیص گفتار به سمت تشخیص اصطلاحات خاص (نام محصولات، اسم‌های خاص، اصطلاحات تخصصی).

wordTimestamp

bool

اختیاری. تولید مهر زمانی در سطح کلمه را پیکربندی می‌کند.

diarization

bool

اختیاری. تنظیم فاصله بین بلندگوها را پیکربندی می‌کند.

mode

Mode

اختیاری. حالت رونویسی را پیکربندی می‌کند. مقادیر پشتیبانی‌شده: VERBATIM ، SMART . اگر مشخص نشود، پیش‌فرض روی رونویسی VERBATIM است. در حالت SMART ، مدل حذف ناروانی (حذف کلمات پرکننده، تکرارها و شروع‌های نادرست)، پاکسازی سبک دستوری، قالب‌بندی خودکار (پاراگراف‌ها، نقاط گلوله‌ای، فهرست‌های شماره‌گذاری‌شده) و ویرایش‌های جزئی کاربر (اصلاحات درون‌خطی) را انجام می‌دهد. مهرهای زمانی و تنظیم خودکار تاریخ با حالت SMART سازگار نیستند.

حالت

حالت رونویسی.

انوم‌ها
MODE_UNSPECIFIED حالت رونویسی نامشخص.
VERBATIM حالت رونویسی کلمه به کلمه.
SMART حالت رونویسی هوشمند.

تشخیص خودکار فعالیت

تشخیص خودکار فعالیت را پیکربندی می‌کند.

فیلدها
disabled

bool

اختیاری. در صورت فعال بودن (پیش‌فرض)، ورودی‌های صوتی و متنی شناسایی‌شده به عنوان فعالیت شمارش می‌شوند. در صورت غیرفعال بودن، کلاینت باید سیگنال‌های فعالیت ارسال کند.

startOfSpeechSensitivity

StartSensitivity

اختیاری. تعیین می‌کند که احتمال تشخیص گفتار چقدر است.

prefixPaddingMs

int32

اختیاری. مدت زمان مورد نیاز برای تشخیص گفتار قبل از شروع گفتار. هرچه این مقدار کمتر باشد، تشخیص شروع گفتار حساس‌تر است و گفتار کوتاه‌تر قابل تشخیص است. با این حال، این امر احتمال تشخیص‌های مثبت کاذب را نیز افزایش می‌دهد.

endOfSpeechSensitivity

EndSensitivity

اختیاری. تعیین می‌کند که احتمال پایان یافتن گفتار شناسایی‌شده چقدر است.

silenceDurationMs

int32

اختیاری. مدت زمان مورد نیاز برای تشخیص عدم گفتار (مثلاً سکوت) قبل از پایان گفتار. هرچه این مقدار بزرگتر باشد، می‌توان فواصل گفتاری را بدون ایجاد وقفه در فعالیت کاربر طولانی‌تر کرد، اما این امر تأخیر مدل را افزایش می‌دهد.

BidiGenerateContentClientContent

به‌روزرسانی افزایشی مکالمه فعلی ارائه شده از کلاینت. تمام محتوای اینجا بدون قید و شرط به تاریخچه مکالمه اضافه می‌شود و به عنوان بخشی از اعلان مدل برای تولید محتوا استفاده می‌شود.

یک پیام در اینجا هرگونه تولید مدل فعلی را قطع می‌کند.

فیلدها
turns[]

Content

اختیاری. محتوایی که به مکالمه فعلی با مدل اضافه شده است.

برای پرس‌وجوهای تک نوبتی، این یک نمونه واحد است. برای پرس‌وجوهای چند نوبتی، این یک فیلد تکراری است که شامل سابقه مکالمه و آخرین درخواست است.

turnComplete

bool

اختیاری. اگر درست باشد، نشان می‌دهد که تولید محتوای سرور باید با اعلان انباشته‌شده‌ی فعلی شروع شود. در غیر این صورت، سرور قبل از شروع تولید منتظر پیام‌های اضافی می‌ماند.

BidiGenerateContentRealtimeInput

ورودی کاربر که به صورت بلادرنگ ارسال می‌شود.

روش‌های مختلف (صوت، تصویر و متن) به صورت جریان‌های همزمان مدیریت می‌شوند. ترتیب قرارگیری در بین این جریان‌ها تضمین شده نیست.

این از چند جهت با BidiGenerateContentClientContent متفاوت است:

  • می‌تواند به طور مداوم و بدون وقفه برای تولید مدل ارسال شود.
  • اگر نیاز به ترکیب داده‌های موجود در بین BidiGenerateContentClientContent و BidiGenerateContentRealtimeInput باشد، سرور تلاش می‌کند تا بهترین پاسخ را بهینه‌سازی کند، اما هیچ تضمینی وجود ندارد.
  • پایان نوبت به صراحت مشخص نشده است، بلکه از فعالیت کاربر (مثلاً پایان سخنرانی) استنباط می‌شود.
  • حتی قبل از پایان نوبت، داده‌ها به صورت تدریجی پردازش می‌شوند تا برای شروع سریع پاسخ از مدل بهینه شوند.
فیلدها
mediaChunks[]

Blob

اختیاری. داده‌های بایت درون‌خطی برای ورودی رسانه. چندین mediaChunks پشتیبانی نمی‌شوند، همه به جز اولی نادیده گرفته می‌شوند.

منسوخ شده: به جای آن از یکی از audio ، video یا text استفاده کنید.

audio

Blob

اختیاری. اینها جریان ورودی صوتی بلادرنگ را تشکیل می‌دهند.

video

Blob

اختیاری. اینها جریان ورودی ویدیوی بلادرنگ را تشکیل می‌دهند.

activityStart

ActivityStart

اختیاری. شروع فعالیت کاربر را نشان می‌دهد. این فقط در صورتی قابل ارسال است که تشخیص خودکار فعالیت (یعنی سمت سرور) غیرفعال باشد.

activityEnd

ActivityEnd

اختیاری. پایان فعالیت کاربر را نشان می‌دهد. این فقط در صورتی قابل ارسال است که تشخیص خودکار فعالیت (یعنی سمت سرور) غیرفعال باشد.

mediaResolution

MediaResolution

اختیاری. وضوح رسانه‌ای که باید استفاده شود. اگر مشخص نشده باشد، setup.generationConfig.mediaResolution استفاده می‌شود یا اگر تنظیمات ارائه نشده باشد، پیش‌فرض است.

audioStreamEnd

bool

اختیاری. نشان می‌دهد که جریان صوتی پایان یافته است، مثلاً به دلیل خاموش شدن میکروفون.

این فقط باید زمانی ارسال شود که تشخیص خودکار فعالیت فعال باشد (که پیش‌فرض است).

کلاینت می‌تواند با ارسال یک پیام صوتی، استریم را دوباره باز کند.

text

string

اختیاری. اینها جریان ورودی متن بلادرنگ را تشکیل می‌دهند.

BidiGenerateContentServerContent

به‌روزرسانی افزایشی سرور که توسط مدل در پاسخ به پیام‌های کلاینت ایجاد می‌شود.

محتوا در سریع‌ترین زمان ممکن تولید می‌شود و نه به صورت بلادرنگ. مشتریان می‌توانند آن را ذخیره کرده و به صورت بلادرنگ پخش کنند.

فیلدها
generationComplete

bool

فقط خروجی. اگر درست باشد، نشان می‌دهد که مدل تولید را تمام کرده است.

وقتی مدل هنگام تولید دچار وقفه شود، هیچ پیام «generation_complete» در نوبت وقفه‌دار وجود نخواهد داشت، و از طریق «interrupted > turn_complete» اجرا می‌شود.

وقتی مدل فرض می‌کند که پخش در زمان واقعی انجام می‌شود، بین generation_complete و turn_complete تأخیری وجود خواهد داشت که ناشی از انتظار مدل برای پایان پخش است.

turnComplete

bool

فقط خروجی. اگر درست باشد، نشان می‌دهد که مدل نوبت خود را تکمیل کرده است. تولید فقط در پاسخ به پیام‌های اضافی کلاینت شروع می‌شود. توجه داشته باشید که وقتی گزارش وضعیت پخش فعال است، این فقط زمانی منتشر می‌شود که وضعیت پخش نشان دهد که پخش انجام شده است. وضعیت پخش آینده در همان نسل نادیده گرفته می‌شود.

interrupted

bool

فقط خروجی. اگر درست باشد، نشان می‌دهد که یک پیام کلاینت، تولید مدل فعلی را متوقف کرده است. اگر کلاینت در حال پخش محتوا به صورت بلادرنگ است، این سیگنال خوبی برای توقف و خالی کردن صف پخش فعلی است.

groundingMetadata

GroundingMetadata

فقط خروجی. فراداده‌های زمینه‌ای برای محتوای تولید شده.

inputTranscription

BidiGenerateContentTranscription

فقط خروجی. رونویسی صوتی ورودی. رونویسی مستقل از سایر پیام‌های سرور ارسال می‌شود و هیچ ترتیب تضمین‌شده‌ای وجود ندارد.

interimInputTranscription

BidiGenerateContentTranscription

فقط خروجی. رونویسی با تأخیر کم هنگام صحبت کاربر به‌روزرسانی می‌شود. این فیلد مرتباً به‌روزرسانی می‌شود.

outputTranscription

BidiGenerateContentTranscription

فقط خروجی. رونویسی صوتی خروجی. این رونویسی‌ها بخشی از خروجی Generation سرور هستند. آخرین رونویسی خروجی این نوبت قبل از generationComplete یا interrupted ارسال می‌شود که به نوبه خود با turnComplete دنبال می‌شوند. هیچ ترتیب دقیقی بین رونویسی‌ها و سایر خروجی‌های modelTurn تضمین نشده است، اما سرور سعی می‌کند رونویسی‌ها را نزدیک به خروجی صوتی مربوطه ارسال کند.

urlContextMetadata

UrlContextMetadata

waitingForInput

bool

فقط خروجی. اگر درست باشد، نشان می‌دهد که مدل محتوا تولید نمی‌کند زیرا منتظر ورودی بیشتر از کاربر است، مثلاً چون انتظار دارد کاربر به صحبت کردن ادامه دهد.

speechState
(deprecated)

SpeechState

فقط خروجی. منسوخ شده: به جای آن از VoiceActivity استفاده کنید.

وضعیت فعلی تشخیص گفتار را در realtimeInput.audio نشان می‌دهد. اگر وضعیت بدون تغییر باشد، تنظیم نشده یا صفر می‌شود.

interactionStatus

InteractionStatus

فقط خروجی. وضعیت فعالیت فعلی جلسه زنده. همیشه در کنار turnComplete ارسال می‌شود.

modelTurn

Content

فقط خروجی. محتوایی که مدل به عنوان بخشی از مکالمه فعلی با کاربر تولید کرده است.

BidiGenerateContentServerMessage

پیام پاسخ برای فراخوانی BidiGenerateContent.

فیلدها
usageMetadata

UsageMetadata

فقط خروجی. استفاده از فراداده در مورد پاسخ(ها).

فیلد union messageType . نوع پیام. messageType فقط می‌تواند یکی از موارد زیر باشد:
setupComplete

BidiGenerateContentSetupComplete

فقط خروجی. در پاسخ به پیام BidiGenerateContentSetup از کلاینت، پس از اتمام راه‌اندازی، ارسال می‌شود.

serverContent

BidiGenerateContentServerContent

فقط خروجی. محتوایی که توسط مدل در پاسخ به پیام‌های کلاینت تولید می‌شود.

toolCall

BidiGenerateContentToolCall

فقط خروجی. از کلاینت درخواست کنید تا functionCalls را اجرا کند و پاسخ‌ها را با id منطبق برگرداند.

toolCallCancellation

BidiGenerateContentToolCallCancellation

فقط خروجی. اعلانی برای کلاینت مبنی بر اینکه ToolCallMessage قبلاً صادر شده با id مشخص شده باید لغو شود.

goAway

GoAway

فقط خروجی. اخطاری مبنی بر اینکه سرور به زودی قطع خواهد شد.

sessionResumptionUpdate

SessionResumptionUpdate

فقط خروجی. به‌روزرسانی وضعیت از سرگیری جلسه.

تنظیمات محتوای BidiGenerate

پیامی که قرار است در اولین (و فقط در اولین) BidiGenerateContentClientMessage ارسال شود. حاوی پیکربندی است که در طول مدت RPC استریمینگ اعمال خواهد شد.

کلاینت‌ها باید قبل از ارسال هرگونه پیام اضافی، منتظر پیام BidiGenerateContentSetupComplete باشند.

فیلدها
model

string

الزامی. نام منبع مدل. این به عنوان شناسه‌ای برای استفاده مدل عمل می‌کند.

قالب: models/{model}

generationConfig

GenerationConfig

اختیاری. پیکربندی نسل.

فیلدهای زیر پشتیبانی نمی‌شوند:

  • responseLogprobs
  • responseMimeType
  • logprobs
  • responseSchema
  • responseJsonSchema
  • stopSequence
  • skipResponseCache
  • routingConfig
  • audioTimestamp
systemInstruction

Content

اختیاری. کاربر دستورالعمل‌های سیستمی را برای مدل ارائه داده است.

توجه: فقط متن باید در بخش‌ها استفاده شود و محتوای هر بخش در یک پاراگراف جداگانه قرار گیرد.

tools[]

Tool

اختیاری. فهرستی از Tools مدل ممکن است برای تولید پاسخ بعدی استفاده کند.

Tool ، قطعه کدی است که سیستم را قادر می‌سازد تا با سیستم‌های خارجی تعامل داشته باشد تا یک یا مجموعه‌ای از اقدامات را خارج از دانش و محدوده مدل انجام دهد.

realtimeInputConfig

RealtimeInputConfig

اختیاری. نحوه‌ی مدیریت ورودی‌های بی‌درنگ را پیکربندی می‌کند.

sessionResumption

SessionResumptionConfig

اختیاری. مکانیزم از سرگیری جلسه را پیکربندی می‌کند.

در صورت وجود، سرور پیام‌های SessionResumptionUpdate را ارسال خواهد کرد.

contextWindowCompression

ContextWindowCompressionConfig

اختیاری. مکانیزم فشرده‌سازی پنجره زمینه را پیکربندی می‌کند.

در صورت وجود، سرور به طور خودکار اندازه context را هنگامی که از طول پیکربندی شده تجاوز کند، کاهش می‌دهد.

inputAudioTranscription

AudioTranscriptionConfig

اختیاری. در صورت تنظیم، رونویسی ورودی صوتی را فعال می‌کند. در صورت پیکربندی، رونویسی با زبان صوتی ورودی هم‌تراز می‌شود.

outputAudioTranscription

AudioTranscriptionConfig

اختیاری. در صورت تنظیم، رونویسی خروجی صدای مدل را فعال می‌کند. در صورت پیکربندی، رونویسی با کد زبان مشخص شده برای صدای خروجی هم‌تراز می‌شود.

proactivity

ProactivityConfig

اختیاری. میزان فعالیت مدل را پیکربندی می‌کند.

این به مدل اجازه می‌دهد تا به ورودی‌ها به صورت پیشگیرانه پاسخ دهد و ورودی‌های نامربوط را نادیده بگیرد.

historyConfig

HistoryConfig

اختیاری. تبادل تاریخچه بین کلاینت و سرور را پیکربندی می‌کند.

BidiGenerateContentSetupComplete

این نوع هیچ فیلدی ندارد.

در پاسخ به پیام BidiGenerateContentSetup از کلاینت ارسال شده است.

تماس با ابزار تولید محتوا (BidiGenerateContentToolCall)

از کلاینت درخواست کنید تا functionCalls را اجرا کند و پاسخ‌ها را با id منطبق s برگرداند.

فیلدها
functionCalls[]

FunctionCall

فقط خروجی. فراخوانی تابعی که قرار است اجرا شود.

BidiGenerateContentToolCallCancellation

اعلانی برای کلاینت مبنی بر اینکه ToolCallMessage قبلاً صادر شده با id مشخص شده نباید اجرا می‌شد و باید لغو شود. اگر عوارض جانبی در آن فراخوانی‌های ابزار وجود داشته باشد، کلاینت‌ها می‌توانند سعی کنند فراخوانی‌های ابزار را لغو کنند. این پیام فقط در مواردی رخ می‌دهد که کلاینت‌ها چرخش‌های سرور را قطع می‌کنند.

فیلدها
ids[]

string

فقط خروجی. شناسه‌های فراخوانی‌های ابزار باید لغو شوند.

BidiGenerateContentToolResponse

پاسخ تولید شده توسط کلاینت به یک ToolCall دریافتی از سرور. اشیاء FunctionResponse به صورت جداگانه توسط فیلد id با اشیاء FunctionCall مربوطه تطبیق داده می‌شوند.

توجه داشته باشید که در APIهای GenerateContent تکی و جریانی سرور، فراخوانی تابع با تبادل بخش‌های Content انجام می‌شود، در حالی که در APIهای GenerateContent بیدی، فراخوانی تابع روی این مجموعه اختصاصی از پیام‌ها انجام می‌شود.

فیلدها
functionResponses[]

FunctionResponse

اختیاری. پاسخ به فراخوانی‌های تابع.

BidiGenerateContentTranscription

رونویسی صدا (ورودی یا خروجی).

فیلدها
text

string

متن رونویسی.

languageCode

string

کد زبانی BCP-47 مربوط به رونویسی.

پیکربندی ContextWindowCompression

فشرده‌سازی پنجره زمینه را فعال می‌کند - مکانیزمی برای مدیریت پنجره زمینه مدل به طوری که از طول مشخصی تجاوز نکند.

فیلدها
compressionMechanism فیلد یونیون. مکانیسم فشرده‌سازی پنجره زمینه مورد استفاده. compressionMechanism می‌تواند فقط یکی از موارد زیر باشد:
slidingWindow

SlidingWindow

مکانیزم پنجره کشویی.

triggerTokens

int64

تعداد توکن‌هایی (قبل از اجرای یک نوبت) که برای فشرده‌سازی پنجره‌ی زمینه لازم است.

این می‌تواند برای ایجاد تعادل بین کیفیت و تأخیر استفاده شود، زیرا پنجره‌های متنی کوتاه‌تر ممکن است منجر به پاسخ‌های سریع‌تر مدل شوند. با این حال، هرگونه عملیات فشرده‌سازی باعث افزایش موقت تأخیر می‌شود، بنابراین نباید مرتباً فعال شوند.

اگر تنظیم نشود، پیش‌فرض ۸۰٪ از محدودیت پنجره‌ی زمینه‌ی مدل است. این مقدار ۲۰٪ را برای درخواست/پاسخ مدل بعدی کاربر باقی می‌گذارد.

پایان حساسیت

نحوه تشخیص پایان گفتار را تعیین می‌کند.

انوم‌ها
END_SENSITIVITY_UNSPECIFIED مقدار پیش‌فرض END_SENSITIVITY_HIGH است.
END_SENSITIVITY_HIGH تشخیص خودکار، گفتار را بیشتر قطع می‌کند.
END_SENSITIVITY_LOW تشخیص خودکار، گفتار را کمتر قطع می‌کند.

برو کنار

اخطاری مبنی بر اینکه سرور به زودی قطع خواهد شد.

فیلدها
timeLeft

Duration

زمان باقی مانده قبل از اتصال به عنوان لغو شده (ABORTED) خاتمه خواهد یافت.

این مدت هرگز کمتر از حداقل مشخص شده برای مدل نخواهد بود، که همراه با محدودیت‌های نرخ برای مدل مشخص می‌شود.

پیکربندی تاریخچه

پیکربندی تاریخچه.

این پیام در پیکربندی جلسه با عنوان BidiGenerateContentSetup.historyConfig گنجانده شده است. تبادل پیام‌های تاریخچه را پیکربندی می‌کند.

فیلدها
initialHistoryInClientContent

bool

اختیاری. اگر مقدار setupComplete درست باشد، پس از ارسال setupComplete ، سرور منتظر می‌ماند و ابتدا پیام‌های clientContent را پردازش می‌کند تا turnComplete true شود. این تاریخچه اولیه فراخوانی مدل را آغاز نمی‌کند و ممکن است با role MODEL پایان یابد. پس از اینکه turnComplete true باشد، کلاینت می‌تواند مکالمه بلادرنگ را از طریق realtimeInput آغاز کند.

پرواکتیویته کانفیگ

پیکربندی برای ویژگی‌های پیشگیرانه.

فیلدها
proactiveAudio

bool

اختیاری. در صورت فعال بودن، مدل می‌تواند از پاسخ دادن به آخرین درخواست خودداری کند. برای مثال، این به مدل اجازه می‌دهد تا سخنان خارج از متن را نادیده بگیرد یا اگر کاربر هنوز درخواستی نداده است، سکوت کند.

پیکربندی ورودی بلادرنگ

رفتار ورودی بی‌درنگ را در BidiGenerateContent پیکربندی می‌کند.

فیلدها
automaticActivityDetection

AutomaticActivityDetection

اختیاری. اگر تنظیم نشود، تشخیص خودکار فعالیت به طور پیش‌فرض فعال است. اگر تشخیص خودکار صدا غیرفعال باشد، کلاینت باید سیگنال‌های فعالیت ارسال کند.

activityHandling

ActivityHandling

اختیاری. تعریف می‌کند که فعالیت چه تأثیری دارد.

turnCoverage

TurnCoverage

اختیاری. مشخص می‌کند که کدام ورودی در نوبت کاربر لحاظ شود.

پیکربندی SessionResum

پیکربندی از سرگیری جلسه.

این پیام در پیکربندی جلسه با عنوان BidiGenerateContentSetup.sessionResumption گنجانده شده است. در صورت پیکربندی، سرور پیام‌های SessionResumptionUpdate ارسال خواهد کرد.

فیلدها
handle

string

شناسه‌ی جلسه‌ی قبلی. اگر موجود نباشد، یک جلسه‌ی جدید ایجاد می‌شود.

شناسه‌های نشست (Session Handles) از مقادیر SessionResumptionUpdate.token در اتصالات قبلی می‌آیند.

به‌روزرسانی از سرگیری جلسه

به‌روزرسانی وضعیت از سرگیری جلسه.

فقط در صورتی ارسال می‌شود که BidiGenerateContentSetup.sessionResumption تنظیم شده باشد.

فیلدها
newHandle

string

یک شناسه جدید که نشان‌دهنده وضعیتی است که می‌تواند از سر گرفته شود. در صورت resumable =false بودن، خالی است.

resumable

bool

اگر بتوان نشست فعلی را در این مرحله از سر گرفت، صحیح است.

در برخی از نقاط session، از سرگیری امکان‌پذیر نیست. برای مثال، هنگامی که مدل در حال اجرای فراخوانی‌های تابع یا تولید است. از سرگیری session (با استفاده از توکن session قبلی) در چنین حالتی منجر به از دست رفتن مقداری از داده‌ها خواهد شد. در این موارد، newHandle خالی و resumable نادرست خواهد بود.

پنجره کشویی

متد SlidingWindow با حذف محتوا در ابتدای پنجره context عمل می‌کند. context حاصل همیشه از ابتدای چرخش نقش USER شروع می‌شود. دستورالعمل‌های سیستم و هرگونه BidiGenerateContentSetup.prefixTurns همیشه در ابتدای نتیجه باقی می‌مانند.

فیلدها
targetTokens

int64

تعداد هدف توکن‌هایی که باید نگه داشته شوند. مقدار پیش‌فرض trigger_tokens/2 است.

حذف بخش‌هایی از پنجره زمینه باعث افزایش موقت تأخیر می‌شود، بنابراین این مقدار باید کالیبره شود تا از عملیات فشرده‌سازی مکرر جلوگیری شود.

حساسیت شروع

نحوه تشخیص شروع گفتار را تعیین می‌کند.

انوم‌ها
START_SENSITIVITY_UNSPECIFIED مقدار پیش‌فرض START_SENSITIVITY_HIGH است.
START_SENSITIVITY_HIGH تشخیص خودکار، شروع گفتار را بیشتر تشخیص می‌دهد.
START_SENSITIVITY_LOW تشخیص خودکار، شروع گفتار را کمتر تشخیص می‌دهد.

پوشش نوبتی

گزینه‌هایی در مورد اینکه کدام ورودی‌ها در نوبت کاربر لحاظ شوند.

انوم‌ها
TURN_COVERAGE_UNSPECIFIED اگر مشخص نشده باشد، یک رفتار پیش‌فرض بر اساس مدل انتخاب می‌شود. به عنوان مثال، برای Gemini 2.5، پیش‌فرض TURN_INCLUDES_ONLY_ACTIVITY است، در حالی که برای Gemini 3.1 و بالاتر، TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO است.
TURN_INCLUDES_ONLY_ACTIVITY شامل فعالیت از آخرین نوبت به بعد، به استثنای عدم فعالیت (مثلاً سکوت در جریان صوتی) می‌شود.
TURN_INCLUDES_ALL_INPUT شامل تمام ورودی‌های بلادرنگ از آخرین نوبت، از جمله عدم فعالیت (مثلاً سکوت در جریان صوتی) می‌شود.
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO شامل فعالیت‌های صوتی و تمام ویدیوهای ضبط شده از آخرین نوبت. با تشخیص خودکار فعالیت، فعالیت صوتی شامل گفتار می‌شود و سکوت را شامل نمی‌شود.

پیکربندی ترجمه

پیکربندی برای ویژگی‌های ترجمه.

فیلدها
targetLanguageCode

string

الزامی. زبان مقصد برای ترجمه. مقادیر پشتیبانی‌شده کدهای زبان BCP-47 هستند (مثلاً "en"، "es"، "fr").

echoTargetLanguage

bool

اختیاری. اگر مقدار آن درست باشد، مدل هنگام صحبت به زبان مقصد، صدا تولید می‌کند، اساساً ورودی را طوطی‌وار تکرار می‌کند. اگر مقدار آن نادرست باشد، ما برای زبان مقصد صدا تولید نمی‌کنیم.

فراداده‌ی UrlContext

فراداده مربوط به ابزار بازیابی متن url.

فیلدها
urlMetadata[]

UrlMetadata

فهرست زمینه آدرس اینترنتی.

کاربردفراداده

فراداده‌های مربوط به پاسخ(ها).

فیلدها
promptTokenCount

int32

فقط خروجی. تعداد توکن‌های موجود در اعلان. وقتی cachedContent تنظیم شده باشد، این همچنان اندازه کل مؤثر اعلان است، به این معنی که شامل تعداد توکن‌های موجود در محتوای ذخیره شده نیز می‌شود.

cachedContentTokenCount

int32

تعداد توکن‌ها در بخش ذخیره‌شده‌ی اعلان (محتوای ذخیره‌شده)

responseTokenCount

int32

فقط خروجی. تعداد کل توکن‌ها در بین تمام کاندیدهای پاسخ تولید شده.

toolUsePromptTokenCount

int32

فقط خروجی. تعداد توکن‌های موجود در اعلان(های) استفاده از ابزار.

thoughtsTokenCount

int32

فقط خروجی. تعداد توکن‌های افکار برای مدل‌های تفکر.

totalTokenCount

int32

فقط خروجی. تعداد کل توکن‌ها برای درخواست تولید (نامزدهای اعلان + پاسخ).

promptTokensDetails[]

ModalityTokenCount

فقط خروجی. فهرست روش‌هایی که در ورودی درخواست پردازش شده‌اند.

cacheTokensDetails[]

ModalityTokenCount

فقط خروجی. فهرستی از روش‌های محتوای ذخیره‌شده در ورودی درخواست.

responseTokensDetails[]

ModalityTokenCount

فقط خروجی. فهرست روش‌هایی که در پاسخ برگردانده شده‌اند.

toolUsePromptTokensDetails[]

ModalityTokenCount

فقط خروجی. فهرست روش‌هایی که برای ورودی‌های درخواست استفاده از ابزار پردازش شده‌اند.

توکن‌های احراز هویت موقت

توکن‌های احراز هویت موقت را می‌توان با فراخوانی AuthTokenService.CreateToken به دست آورد و سپس با GenerativeService.BidiGenerateContentConstrained استفاده کرد، یا با ارسال توکن در یک پارامتر پرس‌وجوی access_token ، یا در یک هدر HTTP Authorization با پیشوند " Token " به آن.

درخواست ایجاد توکن احراز هویت

یک توکن احراز هویت موقت ایجاد کنید.

فیلدها
authToken

AuthToken

الزامی. توکنی که باید ایجاد شود.

توکن احراز هویت

درخواستی برای ایجاد یک توکن احراز هویت موقت.

فیلدها
name

string

فقط خروجی. شناسه. خود توکن.

expireTime

Timestamp

اختیاری. فقط ورودی. تغییرناپذیر. یک زمان اختیاری که پس از آن، هنگام استفاده از توکن حاصل، پیام‌های موجود در جلسات BidiGenerateContent رد می‌شوند. (Gemini ممکن است پس از این زمان، جلسه را به صورت پیشگیرانه ببندد.)

اگر تنظیم نشده باشد، این مقدار به صورت پیش‌فرض روی ۳۰ دقیقه در آینده تنظیم می‌شود. در صورت تنظیم، این مقدار باید کمتر از ۲۰ ساعت در آینده باشد.

newSessionExpireTime

Timestamp

اختیاری. فقط ورودی. تغییرناپذیر. زمانی که پس از آن، جلسات جدید Live API با استفاده از توکن حاصل از این درخواست رد می‌شوند.

اگر تنظیم نشود، پیش‌فرض‌ها روی ۶۰ ثانیه در آینده تنظیم می‌شوند. اگر تنظیم شود، این مقدار باید کمتر از ۲۰ ساعت در آینده باشد.

fieldMask

FieldMask

اختیاری. فقط ورودی. تغییرناپذیر. اگر field_mask خالی باشد و bidiGenerateContentSetup وجود نداشته باشد، پیام مؤثر BidiGenerateContentSetup از اتصال Live API گرفته می‌شود.

اگر field_mask خالی باشد و bidiGenerateContentSetup موجود باشد ، پیام مؤثر BidiGenerateContentSetup به طور کامل از bidiGenerateContentSetup در این درخواست گرفته می‌شود. پیام راه‌اندازی از اتصال Live API نادیده گرفته می‌شود.

اگر field_mask خالی نباشد، فیلدهای مربوطه از bidiGenerateContentSetup فیلدهای پیام راه‌اندازی در اتصال Live API را بازنویسی می‌کنند.

config فیلد Union. پیکربندی مختص متد برای token.config config می‌تواند فقط یکی از موارد زیر باشد:
bidiGenerateContentSetup

BidiGenerateContentSetup

اختیاری. فقط ورودی. تغییرناپذیر. پیکربندی مختص BidiGenerateContent .

uses

int32

اختیاری. فقط ورودی. تغییرناپذیر. تعداد دفعاتی که می‌توان از توکن استفاده کرد. اگر این مقدار صفر باشد، هیچ محدودیتی اعمال نمی‌شود. از سرگیری یک جلسه Live API به عنوان یک استفاده حساب نمی‌شود. اگر مشخص نشود، پیش‌فرض ۱ است.

اطلاعات بیشتر در مورد انواع رایج

برای اطلاعات بیشتر در مورد انواع منابع API رایج Blob ، Content ، FunctionCall ، FunctionResponse ، GenerationConfig ، GroundingMetadata ، ModalityTokenCount و Tool ، به بخش تولید محتوا مراجعه کنید.

،

Live API یک API با قابلیت Stateful است که از WebSockets استفاده می‌کند. در این بخش، جزئیات بیشتری در مورد WebSockets API خواهید یافت.

جلسات

یک اتصال WebSocket یک جلسه بین کلاینت و سرور Gemini برقرار می‌کند. پس از اینکه کلاینت یک اتصال جدید را آغاز می‌کند، جلسه می‌تواند پیام‌هایی را با سرور رد و بدل کند تا:

  • متن، صدا یا ویدیو را به سرور جمینی ارسال کنید.
  • درخواست‌های صوتی، متنی یا تماس عملکردی را از سرور Gemini دریافت کنید.

اتصال وب‌سوکت

برای شروع یک جلسه، به این نقطه پایانی وب سوکت متصل شوید:

wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent

پیکربندی جلسه

پیام اولیه‌ای که پس از برقراری اتصال WebSocket ارسال می‌شود، پیکربندی جلسه را تنظیم می‌کند که شامل مدل، پارامترهای تولید، دستورالعمل‌های سیستم و ابزارها می‌شود.

شما نمی‌توانید پیکربندی را در حالی که اتصال باز است، به‌روزرسانی کنید. با این حال، می‌توانید پارامترهای پیکربندی، به جز مدل، را هنگام مکث و از سرگیری از طریق مکانیسم از سرگیری جلسه تغییر دهید.

به مثال پیکربندی زیر توجه کنید. توجه داشته باشید که نحوه‌ی نام‌گذاری در SDKها ممکن است متفاوت باشد. می‌توانید گزینه‌های پیکربندی SDK پایتون را اینجا جستجو کنید.


{
  "model": string,
  "generationConfig": {
    "candidateCount": integer,
    "maxOutputTokens": integer,
    "temperature": number,
    "topP": number,
    "topK": integer,
    "presencePenalty": number,
    "frequencyPenalty": number,
    "responseModalities": [string],
    "speechConfig": object,
    "mediaResolution": object,
    "translationConfig": object
  },
  "systemInstruction": string,
  "tools": [object]
}

برای اطلاعات بیشتر در مورد فیلد API، به generationConfig مراجعه کنید.

ارسال پیام

برای تبادل پیام از طریق اتصال WebSocket، کلاینت باید یک شیء JSON را از طریق یک اتصال WebSocket باز ارسال کند. شیء JSON باید دقیقاً یکی از فیلدهای مجموعه شیء زیر را داشته باشد:


{
  "setup": BidiGenerateContentSetup,
  "clientContent": BidiGenerateContentClientContent,
  "realtimeInput": BidiGenerateContentRealtimeInput,
  "toolResponse": BidiGenerateContentToolResponse
}

پیام‌های کلاینت پشتیبانی‌شده

پیام‌های کلاینت پشتیبانی‌شده را در جدول زیر مشاهده کنید:

پیام توضیحات
BidiGenerateContentSetup پیکربندی جلسه که باید در اولین پیام ارسال شود
BidiGenerateContentClientContent به‌روزرسانی تدریجی محتوای مکالمه فعلی ارائه شده از طرف کلاینت
BidiGenerateContentRealtimeInput ورودی صوتی، تصویری یا متنی بلادرنگ
BidiGenerateContentToolResponse پاسخ به یک ToolCallMessage دریافت شده از سرور

دریافت پیام‌ها

برای دریافت پیام‌ها از Gemini، به رویداد 'message' در WebSocket گوش دهید و سپس نتیجه را طبق تعریف پیام‌های سرور پشتیبانی‌شده تجزیه کنید.

موارد زیر را ببینید:

async with client.aio.live.connect(model='...', config=config) as session:
    await session.send(input='Hello world!', end_of_turn=True)
    async for message in session.receive():
        print(message)

پیام‌های سرور ممکن است دارای فیلد usageMetadata باشند، اما در غیر این صورت دقیقاً یکی از فیلدهای دیگر پیام BidiGenerateContentServerMessage را شامل می‌شوند. (اتحادیه messageType در JSON بیان نشده است، بنابراین این فیلد در سطح بالای پیام ظاهر می‌شود.)

پیام‌ها و رویدادها

پایان فعالیت

این نوع هیچ فیلدی ندارد.

پایان فعالیت کاربر را نشان می‌دهد.

مدیریت فعالیت

روش‌های مختلف مدیریت فعالیت کاربران

انوم‌ها
ACTIVITY_HANDLING_UNSPECIFIED اگر مشخص نشده باشد، رفتار پیش‌فرض START_OF_ACTIVITY_INTERRUPTS است.
START_OF_ACTIVITY_INTERRUPTS اگر درست باشد، شروع فعالیت، پاسخ مدل را قطع می‌کند (که به آن "barge in" نیز می‌گویند). پاسخ فعلی مدل در لحظه وقفه قطع می‌شود. این رفتار پیش‌فرض است.
NO_INTERRUPTION پاسخ مدل قطع نخواهد شد.

شروع فعالیت

این نوع هیچ فیلدی ندارد.

شروع فعالیت کاربر را نشان می‌دهد.

پیکربندی رونویسی صوتی

پیکربندی رونویسی صوتی.

فیلدها
languageCodes[]

string

اختیاری. کدهای زبان BCP-47 نکاتی در مورد زبان‌های موجود در صدا ارائه می‌دهند. در صورت حذف یا خالی بودن، به طور پیش‌فرض روی تشخیص خودکار زبان تنظیم می‌شود.

customVocabulary[]

string

اختیاری. فهرستی از عبارات واژگانی سفارشی برای جهت‌دهی مدل تشخیص گفتار به سمت تشخیص اصطلاحات خاص (نام محصولات، اسم‌های خاص، اصطلاحات تخصصی).

wordTimestamp

bool

اختیاری. تولید مهر زمانی در سطح کلمه را پیکربندی می‌کند.

diarization

bool

اختیاری. تنظیم فاصله بین بلندگوها را پیکربندی می‌کند.

mode

Mode

اختیاری. حالت رونویسی را پیکربندی می‌کند. مقادیر پشتیبانی‌شده: VERBATIM ، SMART . اگر مشخص نشود، پیش‌فرض روی رونویسی VERBATIM است. در حالت SMART ، مدل حذف ناروانی (حذف کلمات پرکننده، تکرارها و شروع‌های نادرست)، پاکسازی سبک دستوری، قالب‌بندی خودکار (پاراگراف‌ها، نقاط گلوله‌ای، فهرست‌های شماره‌گذاری‌شده) و ویرایش‌های جزئی کاربر (اصلاحات درون‌خطی) را انجام می‌دهد. مهرهای زمانی و تنظیم خودکار تاریخ با حالت SMART سازگار نیستند.

حالت

حالت رونویسی.

انوم‌ها
MODE_UNSPECIFIED حالت رونویسی نامشخص.
VERBATIM حالت رونویسی کلمه به کلمه.
SMART حالت رونویسی هوشمند.

تشخیص خودکار فعالیت

تشخیص خودکار فعالیت را پیکربندی می‌کند.

فیلدها
disabled

bool

اختیاری. در صورت فعال بودن (پیش‌فرض)، ورودی‌های صوتی و متنی شناسایی‌شده به عنوان فعالیت شمارش می‌شوند. در صورت غیرفعال بودن، کلاینت باید سیگنال‌های فعالیت ارسال کند.

startOfSpeechSensitivity

StartSensitivity

اختیاری. تعیین می‌کند که احتمال تشخیص گفتار چقدر است.

prefixPaddingMs

int32

اختیاری. مدت زمان مورد نیاز برای تشخیص گفتار قبل از شروع گفتار. هرچه این مقدار کمتر باشد، تشخیص شروع گفتار حساس‌تر است و گفتار کوتاه‌تر قابل تشخیص است. با این حال، این امر احتمال تشخیص‌های مثبت کاذب را نیز افزایش می‌دهد.

endOfSpeechSensitivity

EndSensitivity

اختیاری. تعیین می‌کند که احتمال پایان یافتن گفتار شناسایی‌شده چقدر است.

silenceDurationMs

int32

اختیاری. مدت زمان مورد نیاز برای تشخیص عدم گفتار (مثلاً سکوت) قبل از پایان گفتار. هرچه این مقدار بزرگتر باشد، می‌توان فواصل گفتاری را بدون ایجاد وقفه در فعالیت کاربر طولانی‌تر کرد، اما این امر تأخیر مدل را افزایش می‌دهد.

BidiGenerateContentClientContent

به‌روزرسانی افزایشی مکالمه فعلی ارائه شده از کلاینت. تمام محتوای اینجا بدون قید و شرط به تاریخچه مکالمه اضافه می‌شود و به عنوان بخشی از اعلان مدل برای تولید محتوا استفاده می‌شود.

یک پیام در اینجا هرگونه تولید مدل فعلی را قطع می‌کند.

فیلدها
turns[]

Content

اختیاری. محتوایی که به مکالمه فعلی با مدل اضافه شده است.

برای پرس‌وجوهای تک نوبتی، این یک نمونه واحد است. برای پرس‌وجوهای چند نوبتی، این یک فیلد تکراری است که شامل سابقه مکالمه و آخرین درخواست است.

turnComplete

bool

اختیاری. اگر درست باشد، نشان می‌دهد که تولید محتوای سرور باید با اعلان انباشته‌شده‌ی فعلی شروع شود. در غیر این صورت، سرور قبل از شروع تولید منتظر پیام‌های اضافی می‌ماند.

BidiGenerateContentRealtimeInput

ورودی کاربر که به صورت بلادرنگ ارسال می‌شود.

روش‌های مختلف (صوت، تصویر و متن) به صورت جریان‌های همزمان مدیریت می‌شوند. ترتیب قرارگیری در بین این جریان‌ها تضمین شده نیست.

این از چند جهت با BidiGenerateContentClientContent متفاوت است:

  • می‌تواند به طور مداوم و بدون وقفه برای تولید مدل ارسال شود.
  • اگر نیاز به ترکیب داده‌های موجود در بین BidiGenerateContentClientContent و BidiGenerateContentRealtimeInput باشد، سرور تلاش می‌کند تا بهترین پاسخ را بهینه‌سازی کند، اما هیچ تضمینی وجود ندارد.
  • پایان نوبت به صراحت مشخص نشده است، بلکه از فعالیت کاربر (مثلاً پایان سخنرانی) استنباط می‌شود.
  • حتی قبل از پایان نوبت، داده‌ها به صورت تدریجی پردازش می‌شوند تا برای شروع سریع پاسخ از مدل بهینه شوند.
فیلدها
mediaChunks[]

Blob

اختیاری. داده‌های بایت درون‌خطی برای ورودی رسانه. چندین mediaChunks پشتیبانی نمی‌شوند، همه به جز اولی نادیده گرفته می‌شوند.

منسوخ شده: به جای آن از یکی از audio ، video یا text استفاده کنید.

audio

Blob

اختیاری. اینها جریان ورودی صوتی بلادرنگ را تشکیل می‌دهند.

video

Blob

اختیاری. اینها جریان ورودی ویدیوی بلادرنگ را تشکیل می‌دهند.

activityStart

ActivityStart

اختیاری. شروع فعالیت کاربر را نشان می‌دهد. این فقط در صورتی قابل ارسال است که تشخیص خودکار فعالیت (یعنی سمت سرور) غیرفعال باشد.

activityEnd

ActivityEnd

اختیاری. پایان فعالیت کاربر را نشان می‌دهد. این فقط در صورتی قابل ارسال است که تشخیص خودکار فعالیت (یعنی سمت سرور) غیرفعال باشد.

mediaResolution

MediaResolution

اختیاری. وضوح رسانه‌ای که باید استفاده شود. اگر مشخص نشده باشد، setup.generationConfig.mediaResolution استفاده می‌شود یا اگر تنظیمات ارائه نشده باشد، پیش‌فرض است.

audioStreamEnd

bool

اختیاری. نشان می‌دهد که جریان صوتی پایان یافته است، مثلاً به دلیل خاموش شدن میکروفون.

این فقط باید زمانی ارسال شود که تشخیص خودکار فعالیت فعال باشد (که پیش‌فرض است).

کلاینت می‌تواند با ارسال یک پیام صوتی، استریم را دوباره باز کند.

text

string

اختیاری. اینها جریان ورودی متن بلادرنگ را تشکیل می‌دهند.

BidiGenerateContentServerContent

به‌روزرسانی افزایشی سرور که توسط مدل در پاسخ به پیام‌های کلاینت ایجاد می‌شود.

محتوا در سریع‌ترین زمان ممکن تولید می‌شود و نه به صورت بلادرنگ. مشتریان می‌توانند آن را ذخیره کرده و به صورت بلادرنگ پخش کنند.

فیلدها
generationComplete

bool

فقط خروجی. اگر درست باشد، نشان می‌دهد که مدل تولید را تمام کرده است.

وقتی مدل هنگام تولید دچار وقفه شود، هیچ پیام «generation_complete» در نوبت وقفه‌دار وجود نخواهد داشت، و از طریق «interrupted > turn_complete» اجرا می‌شود.

وقتی مدل فرض می‌کند که پخش در زمان واقعی انجام می‌شود، بین generation_complete و turn_complete تأخیری وجود خواهد داشت که ناشی از انتظار مدل برای پایان پخش است.

turnComplete

bool

فقط خروجی. اگر درست باشد، نشان می‌دهد که مدل نوبت خود را تکمیل کرده است. تولید فقط در پاسخ به پیام‌های اضافی کلاینت شروع می‌شود. توجه داشته باشید که وقتی گزارش وضعیت پخش فعال است، این فقط زمانی منتشر می‌شود که وضعیت پخش نشان دهد که پخش انجام شده است. وضعیت پخش آینده در همان نسل نادیده گرفته می‌شود.

interrupted

bool

فقط خروجی. اگر درست باشد، نشان می‌دهد که یک پیام کلاینت، تولید مدل فعلی را متوقف کرده است. اگر کلاینت در حال پخش محتوا به صورت بلادرنگ است، این سیگنال خوبی برای توقف و خالی کردن صف پخش فعلی است.

groundingMetadata

GroundingMetadata

فقط خروجی. فراداده‌های زمینه‌ای برای محتوای تولید شده.

inputTranscription

BidiGenerateContentTranscription

فقط خروجی. رونویسی صوتی ورودی. رونویسی مستقل از سایر پیام‌های سرور ارسال می‌شود و هیچ ترتیب تضمین‌شده‌ای وجود ندارد.

interimInputTranscription

BidiGenerateContentTranscription

فقط خروجی. رونویسی با تأخیر کم هنگام صحبت کاربر به‌روزرسانی می‌شود. این فیلد مرتباً به‌روزرسانی می‌شود.

outputTranscription

BidiGenerateContentTranscription

Output only. Output audio transcription. These transcriptions are part of the Generation output of the server. The last output transcription of this turn is sent before either generationComplete or interrupted , which in turn are followed by turnComplete . There is no guaranteed exact ordering between transcriptions and other modelTurn output but the server tries to send the transcripts close to the corresponding audio output.

urlContextMetadata

UrlContextMetadata

waitingForInput

bool

Output only. If true, indicates that the model is not generating content because it is waiting for more input from the user, eg because it expects the user to continue talking.

speechState
(deprecated)

SpeechState

Output only. DEPRECATED: Use VoiceActivity instead.

Indicates the current state of speech detection on realtimeInput.audio . Not set or zero if the state is unchanged.

interactionStatus

InteractionStatus

Output only. The current activity status of the live session. Always sent alongside turnComplete .

modelTurn

Content

Output only. The content that the model has generated as part of the current conversation with the user.

BidiGenerateContentServerMessage

Response message for the BidiGenerateContent call.

Fields
usageMetadata

UsageMetadata

Output only. Usage metadata about the response(s).

Union field messageType . The type of the message. messageType can be only one of the following:
setupComplete

BidiGenerateContentSetupComplete

Output only. Sent in response to a BidiGenerateContentSetup message from the client when setup is complete.

serverContent

BidiGenerateContentServerContent

Output only. Content generated by the model in response to client messages.

toolCall

BidiGenerateContentToolCall

Output only. Request for the client to execute the functionCalls and return the responses with the matching id s.

toolCallCancellation

BidiGenerateContentToolCallCancellation

Output only. Notification for the client that a previously issued ToolCallMessage with the specified id s should be cancelled.

goAway

GoAway

Output only. A notice that the server will soon disconnect.

sessionResumptionUpdate

SessionResumptionUpdate

Output only. Update of the session resumption state.

BidiGenerateContentSetup

Message to be sent in the first (and only in the first) BidiGenerateContentClientMessage . Contains configuration that will apply for the duration of the streaming RPC.

Clients should wait for a BidiGenerateContentSetupComplete message before sending any additional messages.

Fields
model

string

Required. The model's resource name. This serves as an ID for the Model to use.

Format: models/{model}

generationConfig

GenerationConfig

Optional. Generation config.

The following fields are not supported:

  • responseLogprobs
  • responseMimeType
  • logprobs
  • responseSchema
  • responseJsonSchema
  • stopSequence
  • skipResponseCache
  • routingConfig
  • audioTimestamp
systemInstruction

Content

Optional. The user provided system instructions for the model.

Note: Only text should be used in parts and content in each part will be in a separate paragraph.

tools[]

Tool

Optional. A list of Tools the model may use to generate the next response.

A Tool is a piece of code that enables the system to interact with external systems to perform an action, or set of actions, outside of knowledge and scope of the model.

realtimeInputConfig

RealtimeInputConfig

Optional. Configures the handling of realtime input.

sessionResumption

SessionResumptionConfig

Optional. Configures session resumption mechanism.

If included, the server will send SessionResumptionUpdate messages.

contextWindowCompression

ContextWindowCompressionConfig

Optional. Configures a context window compression mechanism.

If included, the server will automatically reduce the size of the context when it exceeds the configured length.

inputAudioTranscription

AudioTranscriptionConfig

Optional. If set, enables transcription of voice input. The transcription aligns with the input audio language, if configured.

outputAudioTranscription

AudioTranscriptionConfig

Optional. If set, enables transcription of the model's audio output. The transcription aligns with the language code specified for the output audio, if configured.

proactivity

ProactivityConfig

Optional. Configures the proactivity of the model.

This allows the model to respond proactively to the input and to ignore irrelevant input.

historyConfig

HistoryConfig

Optional. Configures the exchange of history between the client and the server.

BidiGenerateContentSetupComplete

This type has no fields.

Sent in response to a BidiGenerateContentSetup message from the client.

BidiGenerateContentToolCall

Request for the client to execute the functionCalls and return the responses with the matching id s.

Fields
functionCalls[]

FunctionCall

Output only. The function call to be executed.

BidiGenerateContentToolCallCancellation

Notification for the client that a previously issued ToolCallMessage with the specified id s should not have been executed and should be cancelled. If there were side-effects to those tool calls, clients may attempt to undo the tool calls. This message occurs only in cases where the clients interrupt server turns.

Fields
ids[]

string

Output only. The ids of the tool calls to be cancelled.

BidiGenerateContentToolResponse

Client generated response to a ToolCall received from the server. Individual FunctionResponse objects are matched to the respective FunctionCall objects by the id field.

Note that in the unary and server-streaming GenerateContent APIs function calling happens by exchanging the Content parts, while in the bidi GenerateContent APIs function calling happens over these dedicated set of messages.

Fields
functionResponses[]

FunctionResponse

Optional. The response to the function calls.

BidiGenerateContentTranscription

Transcription of audio (input or output).

Fields
text

string

Transcription text.

languageCode

string

The BCP-47 language code of the transcription.

ContextWindowCompressionConfig

Enables context window compression — a mechanism for managing the model's context window so that it does not exceed a given length.

Fields
Union field compressionMechanism . The context window compression mechanism used. compressionMechanism can be only one of the following:
slidingWindow

SlidingWindow

A sliding-window mechanism.

triggerTokens

int64

The number of tokens (before running a turn) required to trigger a context window compression.

This can be used to balance quality against latency as shorter context windows may result in faster model responses. However, any compression operation will cause a temporary latency increase, so they should not be triggered frequently.

If not set, the default is 80% of the model's context window limit. This leaves 20% for the next user request/model response.

EndSensitivity

Determines how end of speech is detected.

Enums
END_SENSITIVITY_UNSPECIFIED The default is END_SENSITIVITY_HIGH.
END_SENSITIVITY_HIGH Automatic detection ends speech more often.
END_SENSITIVITY_LOW Automatic detection ends speech less often.

GoAway

A notice that the server will soon disconnect.

Fields
timeLeft

Duration

The remaining time before the connection will be terminated as ABORTED.

This duration will never be less than a model-specific minimum, which will be specified together with the rate limits for the model.

HistoryConfig

History configuration.

This message is included in the session configuration as BidiGenerateContentSetup.historyConfig . Configures the exchange of history messages.

Fields
initialHistoryInClientContent

bool

Optional. If true, after sending setupComplete , the server will wait and at first process clientContent messages until turnComplete is true . This initial history will not trigger a model call and may end with role MODEL . After turnComplete is true , the client can start the realtime conversation via realtimeInput .

ProactivityConfig

Config for proactivity features.

Fields
proactiveAudio

bool

Optional. If enabled, the model can reject responding to the last prompt. For example, this allows the model to ignore out of context speech or to stay silent if the user did not make a request, yet.

RealtimeInputConfig

Configures the realtime input behavior in BidiGenerateContent .

Fields
automaticActivityDetection

AutomaticActivityDetection

Optional. If not set, automatic activity detection is enabled by default. If automatic voice detection is disabled, the client must send activity signals.

activityHandling

ActivityHandling

Optional. Defines what effect activity has.

turnCoverage

TurnCoverage

Optional. Defines which input is included in the user's turn.

SessionResumptionConfig

Session resumption configuration.

This message is included in the session configuration as BidiGenerateContentSetup.sessionResumption . If configured, the server will send SessionResumptionUpdate messages.

Fields
handle

string

The handle of a previous session. If not present then a new session is created.

Session handles come from SessionResumptionUpdate.token values in previous connections.

SessionResumptionUpdate

Update of the session resumption state.

Only sent if BidiGenerateContentSetup.sessionResumption was set.

Fields
newHandle

string

New handle that represents a state that can be resumed. Empty if resumable =false.

resumable

bool

True if the current session can be resumed at this point.

Resumption is not possible at some points in the session. For example, when the model is executing function calls or generating. Resuming the session (using a previous session token) in such a state will result in some data loss. In these cases, newHandle will be empty and resumable will be false.

SlidingWindow

The SlidingWindow method operates by discarding content at the beginning of the context window. The resulting context will always begin at the start of a USER role turn. System instructions and any BidiGenerateContentSetup.prefixTurns will always remain at the beginning of the result.

Fields
targetTokens

int64

The target number of tokens to keep. The default value is trigger_tokens/2.

Discarding parts of the context window causes a temporary latency increase so this value should be calibrated to avoid frequent compression operations.

StartSensitivity

Determines how start of speech is detected.

Enums
START_SENSITIVITY_UNSPECIFIED The default is START_SENSITIVITY_HIGH.
START_SENSITIVITY_HIGH Automatic detection will detect the start of speech more often.
START_SENSITIVITY_LOW Automatic detection will detect the start of speech less often.

TurnCoverage

Options about which input is included in the user's turn.

Enums
TURN_COVERAGE_UNSPECIFIED If unspecified, a default behavior is selected based on the model. Eg, for Gemini 2.5, the default is TURN_INCLUDES_ONLY_ACTIVITY , while for Gemini 3.1 and onwards, it's TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO .
TURN_INCLUDES_ONLY_ACTIVITY Includes activity since the last turn, excluding inactivity (eg silence on the audio stream).
TURN_INCLUDES_ALL_INPUT Includes all realtime input since the last turn, including inactivity (eg silence on the audio stream).
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO Includes audio activity and all video since the last turn. With automatic activity detection, audio activity means speech and excludes silence.

TranslationConfig

Config for translation features.

Fields
targetLanguageCode

string

Required. The target language for translation. Supported values are BCP-47 language codes (eg "en", "es", "fr").

echoTargetLanguage

bool

Optional. If true, the model will generate audio when the target language is spoken, essentially it will parrot the input. If false, we will not produce audio for the target language.

UrlContextMetadata

Metadata related to url context retrieval tool.

Fields
urlMetadata[]

UrlMetadata

List of url context.

UsageMetadata

Usage metadata about response(s).

Fields
promptTokenCount

int32

Output only. Number of tokens in the prompt. When cachedContent is set, this is still the total effective prompt size meaning this includes the number of tokens in the cached content.

cachedContentTokenCount

int32

Number of tokens in the cached part of the prompt (the cached content)

responseTokenCount

int32

Output only. Total number of tokens across all the generated response candidates.

toolUsePromptTokenCount

int32

Output only. Number of tokens present in tool-use prompt(s).

thoughtsTokenCount

int32

Output only. Number of tokens of thoughts for thinking models.

totalTokenCount

int32

Output only. Total token count for the generation request (prompt + response candidates).

promptTokensDetails[]

ModalityTokenCount

Output only. List of modalities that were processed in the request input.

cacheTokensDetails[]

ModalityTokenCount

Output only. List of modalities of the cached content in the request input.

responseTokensDetails[]

ModalityTokenCount

Output only. List of modalities that were returned in the response.

toolUsePromptTokensDetails[]

ModalityTokenCount

Output only. List of modalities that were processed for tool-use request inputs.

Ephemeral authentication tokens

Ephemeral authentication tokens can be obtained by calling AuthTokenService.CreateToken and then used with GenerativeService.BidiGenerateContentConstrained , either by passing the token in an access_token query parameter, or in an HTTP Authorization header with " Token " prefixed to it.

CreateAuthTokenRequest

Create an ephemeral authentication token.

Fields
authToken

AuthToken

Required. The token to create.

AuthToken

A request to create an ephemeral authentication token.

Fields
name

string

Output only. Identifier. The token itself.

expireTime

Timestamp

Optional. Input only. Immutable. An optional time after which, when using the resulting token, messages in BidiGenerateContent sessions will be rejected. (Gemini may preemptively close the session after this time.)

If not set then this defaults to 30 minutes in the future. If set, this value must be less than 20 hours in the future.

newSessionExpireTime

Timestamp

Optional. Input only. Immutable. The time after which new Live API sessions using the token resulting from this request will be rejected.

If not set this defaults to 60 seconds in the future. If set, this value must be less than 20 hours in the future.

fieldMask

FieldMask

Optional. Input only. Immutable. If field_mask is empty, and bidiGenerateContentSetup is not present, then the effective BidiGenerateContentSetup message is taken from the Live API connection.

If field_mask is empty, and bidiGenerateContentSetup is present, then the effective BidiGenerateContentSetup message is taken entirely from bidiGenerateContentSetup in this request. The setup message from the Live API connection is ignored.

If field_mask is not empty, then the corresponding fields from bidiGenerateContentSetup will overwrite the fields from the setup message in the Live API connection.

Union field config . The method-specific configuration for the resulting token. config can be only one of the following:
bidiGenerateContentSetup

BidiGenerateContentSetup

Optional. Input only. Immutable. Configuration specific to BidiGenerateContent .

uses

int32

Optional. Input only. Immutable. The number of times the token can be used. If this value is zero then no limit is applied. Resuming a Live API session does not count as a use. If unspecified, the default is 1.

More information on common types

For more information on the commonly-used API resource types Blob , Content , FunctionCall , FunctionResponse , GenerationConfig , GroundingMetadata , ModalityTokenCount , and Tool , see Generating content .

،

The Live API is a stateful API that uses WebSockets . In this section, you'll find additional details regarding the WebSockets API.

جلسات

A WebSocket connection establishes a session between the client and the Gemini server. After a client initiates a new connection the session can exchange messages with the server to:

  • Send text, audio, or video to the Gemini server.
  • Receive audio, text, or function call requests from the Gemini server.

WebSocket connection

To start a session, connect to this websocket endpoint:

wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent

Session configuration

The initial message sent after establishing the WebSocket connection sets the session configuration, which includes the model, generation parameters, system instructions, and tools.

You cannot update the configuration while the connection is open. However, you can change the configuration parameters, except the model, when pausing and resuming via the session resumption mechanism .

See the following example configuration. Note that the name casing in SDKs may vary. You can look up the Python SDK configuration options here .


{
  "model": string,
  "generationConfig": {
    "candidateCount": integer,
    "maxOutputTokens": integer,
    "temperature": number,
    "topP": number,
    "topK": integer,
    "presencePenalty": number,
    "frequencyPenalty": number,
    "responseModalities": [string],
    "speechConfig": object,
    "mediaResolution": object,
    "translationConfig": object
  },
  "systemInstruction": string,
  "tools": [object]
}

For more information on the API field, see generationConfig .

Send messages

To exchange messages over the WebSocket connection, the client must send a JSON object over an open WebSocket connection. The JSON object must have exactly one of the fields from the following object set:


{
  "setup": BidiGenerateContentSetup,
  "clientContent": BidiGenerateContentClientContent,
  "realtimeInput": BidiGenerateContentRealtimeInput,
  "toolResponse": BidiGenerateContentToolResponse
}

Supported client messages

See the supported client messages in the following table:

پیام توضیحات
BidiGenerateContentSetup Session configuration to be sent in the first message
BidiGenerateContentClientContent Incremental content update of the current conversation delivered from the client
BidiGenerateContentRealtimeInput Real time audio, video, or text input
BidiGenerateContentToolResponse Response to a ToolCallMessage received from the server

Receive messages

To receive messages from Gemini, listen for the WebSocket 'message' event, and then parse the result according to the definition of the supported server messages.

See the following:

async with client.aio.live.connect(model='...', config=config) as session:
    await session.send(input='Hello world!', end_of_turn=True)
    async for message in session.receive():
        print(message)

Server messages may have a usageMetadata field but will otherwise include exactly one of the other fields from the BidiGenerateContentServerMessage message. (The messageType union is not expressed in JSON so the field will appear at the top-level of the message.)

Messages and events

ActivityEnd

This type has no fields.

Marks the end of user activity.

ActivityHandling

The different ways of handling user activity.

Enums
ACTIVITY_HANDLING_UNSPECIFIED If unspecified, the default behavior is START_OF_ACTIVITY_INTERRUPTS .
START_OF_ACTIVITY_INTERRUPTS If true, start of activity will interrupt the model's response (also called "barge in"). The model's current response will be cut-off in the moment of the interruption. This is the default behavior.
NO_INTERRUPTION The model's response will not be interrupted.

ActivityStart

This type has no fields.

Marks the start of user activity.

AudioTranscriptionConfig

The audio transcription configuration.

Fields
languageCodes[]

string

Optional. BCP-47 language codes providing hints about the languages present in the audio. If omitted or empty, defaults to automatic language detection.

customVocabulary[]

string

Optional. A list of custom vocabulary phrases to bias the speech recognition model toward recognizing specific terms (product names, proper nouns, jargon).

wordTimestamp

bool

Optional. Configures word-level timestamp generation.

diarization

bool

Optional. Configures speaker diarization.

mode

Mode

Optional. Configures transcription mode. Supported values: VERBATIM , SMART . If unspecified, defaults to VERBATIM transcription. In SMART mode, the model performs disfluency removal (eliminating filler words, repetitions, and false starts), light grammatical cleanup, automatic formatting (paragraphs, bullet points, numbered lists), and minor user edits (inline self-corrections). Timestamps and diarization are incompatible with mode SMART .

حالت

Transcription mode.

Enums
MODE_UNSPECIFIED Unspecified transcription mode.
VERBATIM Verbatim transcription mode.
SMART Smart transcription mode.

AutomaticActivityDetection

Configures automatic detection of activity.

Fields
disabled

bool

Optional. If enabled (the default), detected voice and text input count as activity. If disabled, the client must send activity signals.

startOfSpeechSensitivity

StartSensitivity

Optional. Determines how likely speech is to be detected.

prefixPaddingMs

int32

Optional. The required duration of detected speech before start-of-speech is committed. The lower this value, the more sensitive the start-of-speech detection is and shorter speech can be recognized. However, this also increases the probability of false positives.

endOfSpeechSensitivity

EndSensitivity

Optional. Determines how likely detected speech is ended.

silenceDurationMs

int32

Optional. The required duration of detected non-speech (eg silence) before end-of-speech is committed. The larger this value, the longer speech gaps can be without interrupting the user's activity but this will increase the model's latency.

BidiGenerateContentClientContent

Incremental update of the current conversation delivered from the client. All of the content here is unconditionally appended to the conversation history and used as part of the prompt to the model to generate content.

A message here will interrupt any current model generation.

Fields
turns[]

Content

Optional. The content appended to the current conversation with the model.

For single-turn queries, this is a single instance. For multi-turn queries, this is a repeated field that contains conversation history and the latest request.

turnComplete

bool

Optional. If true, indicates that the server content generation should start with the currently accumulated prompt. Otherwise, the server awaits additional messages before starting generation.

BidiGenerateContentRealtimeInput

User input that is sent in real time.

The different modalities (audio, video and text) are handled as concurrent streams. The ordering across these streams is not guaranteed.

This is different from BidiGenerateContentClientContent in a few ways:

  • Can be sent continuously without interruption to model generation.
  • If there is a need to mix data interleaved across the BidiGenerateContentClientContent and the BidiGenerateContentRealtimeInput , the server attempts to optimize for best response, but there are no guarantees.
  • End of turn is not explicitly specified, but is rather derived from user activity (for example, end of speech).
  • Even before the end of turn, the data is processed incrementally to optimize for a fast start of the response from the model.
Fields
mediaChunks[]

Blob

Optional. Inlined bytes data for media input. Multiple mediaChunks are not supported, all but the first will be ignored.

DEPRECATED: Use one of audio , video , or text instead.

audio

Blob

Optional. These form the realtime audio input stream.

video

Blob

Optional. These form the realtime video input stream.

activityStart

ActivityStart

Optional. Marks the start of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled.

activityEnd

ActivityEnd

Optional. Marks the end of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled.

mediaResolution

MediaResolution

Optional. The media resolution to use. If not specified, setup.generationConfig.mediaResolution is used or a default if the setup is not provided.

audioStreamEnd

bool

Optional. Indicates that the audio stream has ended, eg because the microphone was turned off.

This should only be sent when automatic activity detection is enabled (which is the default).

The client can reopen the stream by sending an audio message.

text

string

Optional. These form the realtime text input stream.

BidiGenerateContentServerContent

Incremental server update generated by the model in response to client messages.

Content is generated as quickly as possible, and not in real time. Clients may choose to buffer and play it out in real time.

Fields
generationComplete

bool

Output only. If true, indicates that the model is done generating.

When model is interrupted while generating there will be no 'generation_complete' message in interrupted turn, it will go through 'interrupted > turn_complete'.

When model assumes realtime playback there will be delay between generation_complete and turn_complete that is caused by model waiting for playback to finish.

turnComplete

bool

Output only. If true, indicates that the model has completed its turn. Generation will only start in response to additional client messages. Note when playback status reporting is enabled, this is emitted only when the playback status indicates that the playback is done. Future playback status of the same generation will be ignored.

interrupted

bool

Output only. If true, indicates that a client message has interrupted current model generation. If the client is playing out the content in real time, this is a good signal to stop and empty the current playback queue.

groundingMetadata

GroundingMetadata

Output only. Grounding metadata for the generated content.

inputTranscription

BidiGenerateContentTranscription

Output only. Input audio transcription. The transcription is sent independently of the other server messages and there is no guaranteed ordering.

interimInputTranscription

BidiGenerateContentTranscription

Output only. Low latency transcription updated while the user is speaking. This field is subject to frequent updates.

outputTranscription

BidiGenerateContentTranscription

Output only. Output audio transcription. These transcriptions are part of the Generation output of the server. The last output transcription of this turn is sent before either generationComplete or interrupted , which in turn are followed by turnComplete . There is no guaranteed exact ordering between transcriptions and other modelTurn output but the server tries to send the transcripts close to the corresponding audio output.

urlContextMetadata

UrlContextMetadata

waitingForInput

bool

Output only. If true, indicates that the model is not generating content because it is waiting for more input from the user, eg because it expects the user to continue talking.

speechState
(deprecated)

SpeechState

Output only. DEPRECATED: Use VoiceActivity instead.

Indicates the current state of speech detection on realtimeInput.audio . Not set or zero if the state is unchanged.

interactionStatus

InteractionStatus

Output only. The current activity status of the live session. Always sent alongside turnComplete .

modelTurn

Content

Output only. The content that the model has generated as part of the current conversation with the user.

BidiGenerateContentServerMessage

Response message for the BidiGenerateContent call.

Fields
usageMetadata

UsageMetadata

Output only. Usage metadata about the response(s).

Union field messageType . The type of the message. messageType can be only one of the following:
setupComplete

BidiGenerateContentSetupComplete

Output only. Sent in response to a BidiGenerateContentSetup message from the client when setup is complete.

serverContent

BidiGenerateContentServerContent

Output only. Content generated by the model in response to client messages.

toolCall

BidiGenerateContentToolCall

Output only. Request for the client to execute the functionCalls and return the responses with the matching id s.

toolCallCancellation

BidiGenerateContentToolCallCancellation

Output only. Notification for the client that a previously issued ToolCallMessage with the specified id s should be cancelled.

goAway

GoAway

Output only. A notice that the server will soon disconnect.

sessionResumptionUpdate

SessionResumptionUpdate

Output only. Update of the session resumption state.

BidiGenerateContentSetup

Message to be sent in the first (and only in the first) BidiGenerateContentClientMessage . Contains configuration that will apply for the duration of the streaming RPC.

Clients should wait for a BidiGenerateContentSetupComplete message before sending any additional messages.

Fields
model

string

Required. The model's resource name. This serves as an ID for the Model to use.

Format: models/{model}

generationConfig

GenerationConfig

Optional. Generation config.

The following fields are not supported:

  • responseLogprobs
  • responseMimeType
  • logprobs
  • responseSchema
  • responseJsonSchema
  • stopSequence
  • skipResponseCache
  • routingConfig
  • audioTimestamp
systemInstruction

Content

Optional. The user provided system instructions for the model.

Note: Only text should be used in parts and content in each part will be in a separate paragraph.

tools[]

Tool

Optional. A list of Tools the model may use to generate the next response.

A Tool is a piece of code that enables the system to interact with external systems to perform an action, or set of actions, outside of knowledge and scope of the model.

realtimeInputConfig

RealtimeInputConfig

Optional. Configures the handling of realtime input.

sessionResumption

SessionResumptionConfig

Optional. Configures session resumption mechanism.

If included, the server will send SessionResumptionUpdate messages.

contextWindowCompression

ContextWindowCompressionConfig

Optional. Configures a context window compression mechanism.

If included, the server will automatically reduce the size of the context when it exceeds the configured length.

inputAudioTranscription

AudioTranscriptionConfig

Optional. If set, enables transcription of voice input. The transcription aligns with the input audio language, if configured.

outputAudioTranscription

AudioTranscriptionConfig

Optional. If set, enables transcription of the model's audio output. The transcription aligns with the language code specified for the output audio, if configured.

proactivity

ProactivityConfig

Optional. Configures the proactivity of the model.

This allows the model to respond proactively to the input and to ignore irrelevant input.

historyConfig

HistoryConfig

Optional. Configures the exchange of history between the client and the server.

BidiGenerateContentSetupComplete

This type has no fields.

Sent in response to a BidiGenerateContentSetup message from the client.

BidiGenerateContentToolCall

Request for the client to execute the functionCalls and return the responses with the matching id s.

Fields
functionCalls[]

FunctionCall

Output only. The function call to be executed.

BidiGenerateContentToolCallCancellation

Notification for the client that a previously issued ToolCallMessage with the specified id s should not have been executed and should be cancelled. If there were side-effects to those tool calls, clients may attempt to undo the tool calls. This message occurs only in cases where the clients interrupt server turns.

Fields
ids[]

string

Output only. The ids of the tool calls to be cancelled.

BidiGenerateContentToolResponse

Client generated response to a ToolCall received from the server. Individual FunctionResponse objects are matched to the respective FunctionCall objects by the id field.

Note that in the unary and server-streaming GenerateContent APIs function calling happens by exchanging the Content parts, while in the bidi GenerateContent APIs function calling happens over these dedicated set of messages.

Fields
functionResponses[]

FunctionResponse

Optional. The response to the function calls.

BidiGenerateContentTranscription

Transcription of audio (input or output).

Fields
text

string

Transcription text.

languageCode

string

The BCP-47 language code of the transcription.

ContextWindowCompressionConfig

Enables context window compression — a mechanism for managing the model's context window so that it does not exceed a given length.

Fields
Union field compressionMechanism . The context window compression mechanism used. compressionMechanism can be only one of the following:
slidingWindow

SlidingWindow

A sliding-window mechanism.

triggerTokens

int64

The number of tokens (before running a turn) required to trigger a context window compression.

This can be used to balance quality against latency as shorter context windows may result in faster model responses. However, any compression operation will cause a temporary latency increase, so they should not be triggered frequently.

If not set, the default is 80% of the model's context window limit. This leaves 20% for the next user request/model response.

EndSensitivity

Determines how end of speech is detected.

Enums
END_SENSITIVITY_UNSPECIFIED The default is END_SENSITIVITY_HIGH.
END_SENSITIVITY_HIGH Automatic detection ends speech more often.
END_SENSITIVITY_LOW Automatic detection ends speech less often.

GoAway

A notice that the server will soon disconnect.

Fields
timeLeft

Duration

The remaining time before the connection will be terminated as ABORTED.

This duration will never be less than a model-specific minimum, which will be specified together with the rate limits for the model.

HistoryConfig

History configuration.

This message is included in the session configuration as BidiGenerateContentSetup.historyConfig . Configures the exchange of history messages.

Fields
initialHistoryInClientContent

bool

Optional. If true, after sending setupComplete , the server will wait and at first process clientContent messages until turnComplete is true . This initial history will not trigger a model call and may end with role MODEL . After turnComplete is true , the client can start the realtime conversation via realtimeInput .

ProactivityConfig

Config for proactivity features.

Fields
proactiveAudio

bool

Optional. If enabled, the model can reject responding to the last prompt. For example, this allows the model to ignore out of context speech or to stay silent if the user did not make a request, yet.

RealtimeInputConfig

Configures the realtime input behavior in BidiGenerateContent .

Fields
automaticActivityDetection

AutomaticActivityDetection

Optional. If not set, automatic activity detection is enabled by default. If automatic voice detection is disabled, the client must send activity signals.

activityHandling

ActivityHandling

Optional. Defines what effect activity has.

turnCoverage

TurnCoverage

Optional. Defines which input is included in the user's turn.

SessionResumptionConfig

Session resumption configuration.

This message is included in the session configuration as BidiGenerateContentSetup.sessionResumption . If configured, the server will send SessionResumptionUpdate messages.

Fields
handle

string

The handle of a previous session. If not present then a new session is created.

Session handles come from SessionResumptionUpdate.token values in previous connections.

SessionResumptionUpdate

Update of the session resumption state.

Only sent if BidiGenerateContentSetup.sessionResumption was set.

Fields
newHandle

string

New handle that represents a state that can be resumed. Empty if resumable =false.

resumable

bool

True if the current session can be resumed at this point.

Resumption is not possible at some points in the session. For example, when the model is executing function calls or generating. Resuming the session (using a previous session token) in such a state will result in some data loss. In these cases, newHandle will be empty and resumable will be false.

SlidingWindow

The SlidingWindow method operates by discarding content at the beginning of the context window. The resulting context will always begin at the start of a USER role turn. System instructions and any BidiGenerateContentSetup.prefixTurns will always remain at the beginning of the result.

Fields
targetTokens

int64

The target number of tokens to keep. The default value is trigger_tokens/2.

Discarding parts of the context window causes a temporary latency increase so this value should be calibrated to avoid frequent compression operations.

StartSensitivity

Determines how start of speech is detected.

Enums
START_SENSITIVITY_UNSPECIFIED The default is START_SENSITIVITY_HIGH.
START_SENSITIVITY_HIGH Automatic detection will detect the start of speech more often.
START_SENSITIVITY_LOW Automatic detection will detect the start of speech less often.

TurnCoverage

Options about which input is included in the user's turn.

Enums
TURN_COVERAGE_UNSPECIFIED If unspecified, a default behavior is selected based on the model. Eg, for Gemini 2.5, the default is TURN_INCLUDES_ONLY_ACTIVITY , while for Gemini 3.1 and onwards, it's TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO .
TURN_INCLUDES_ONLY_ACTIVITY Includes activity since the last turn, excluding inactivity (eg silence on the audio stream).
TURN_INCLUDES_ALL_INPUT Includes all realtime input since the last turn, including inactivity (eg silence on the audio stream).
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO Includes audio activity and all video since the last turn. With automatic activity detection, audio activity means speech and excludes silence.

TranslationConfig

Config for translation features.

Fields
targetLanguageCode

string

Required. The target language for translation. Supported values are BCP-47 language codes (eg "en", "es", "fr").

echoTargetLanguage

bool

Optional. If true, the model will generate audio when the target language is spoken, essentially it will parrot the input. If false, we will not produce audio for the target language.

UrlContextMetadata

Metadata related to url context retrieval tool.

Fields
urlMetadata[]

UrlMetadata

List of url context.

UsageMetadata

Usage metadata about response(s).

Fields
promptTokenCount

int32

Output only. Number of tokens in the prompt. When cachedContent is set, this is still the total effective prompt size meaning this includes the number of tokens in the cached content.

cachedContentTokenCount

int32

Number of tokens in the cached part of the prompt (the cached content)

responseTokenCount

int32

Output only. Total number of tokens across all the generated response candidates.

toolUsePromptTokenCount

int32

Output only. Number of tokens present in tool-use prompt(s).

thoughtsTokenCount

int32

Output only. Number of tokens of thoughts for thinking models.

totalTokenCount

int32

Output only. Total token count for the generation request (prompt + response candidates).

promptTokensDetails[]

ModalityTokenCount

Output only. List of modalities that were processed in the request input.

cacheTokensDetails[]

ModalityTokenCount

Output only. List of modalities of the cached content in the request input.

responseTokensDetails[]

ModalityTokenCount

Output only. List of modalities that were returned in the response.

toolUsePromptTokensDetails[]

ModalityTokenCount

Output only. List of modalities that were processed for tool-use request inputs.

Ephemeral authentication tokens

Ephemeral authentication tokens can be obtained by calling AuthTokenService.CreateToken and then used with GenerativeService.BidiGenerateContentConstrained , either by passing the token in an access_token query parameter, or in an HTTP Authorization header with " Token " prefixed to it.

CreateAuthTokenRequest

Create an ephemeral authentication token.

Fields
authToken

AuthToken

Required. The token to create.

AuthToken

A request to create an ephemeral authentication token.

Fields
name

string

Output only. Identifier. The token itself.

expireTime

Timestamp

Optional. Input only. Immutable. An optional time after which, when using the resulting token, messages in BidiGenerateContent sessions will be rejected. (Gemini may preemptively close the session after this time.)

If not set then this defaults to 30 minutes in the future. If set, this value must be less than 20 hours in the future.

newSessionExpireTime

Timestamp

Optional. Input only. Immutable. The time after which new Live API sessions using the token resulting from this request will be rejected.

If not set this defaults to 60 seconds in the future. If set, this value must be less than 20 hours in the future.

fieldMask

FieldMask

Optional. Input only. Immutable. If field_mask is empty, and bidiGenerateContentSetup is not present, then the effective BidiGenerateContentSetup message is taken from the Live API connection.

If field_mask is empty, and bidiGenerateContentSetup is present, then the effective BidiGenerateContentSetup message is taken entirely from bidiGenerateContentSetup in this request. The setup message from the Live API connection is ignored.

If field_mask is not empty, then the corresponding fields from bidiGenerateContentSetup will overwrite the fields from the setup message in the Live API connection.

Union field config . The method-specific configuration for the resulting token. config can be only one of the following:
bidiGenerateContentSetup

BidiGenerateContentSetup

Optional. Input only. Immutable. Configuration specific to BidiGenerateContent .

uses

int32

Optional. Input only. Immutable. The number of times the token can be used. If this value is zero then no limit is applied. Resuming a Live API session does not count as a use. If unspecified, the default is 1.

More information on common types

For more information on the commonly-used API resource types Blob , Content , FunctionCall , FunctionResponse , GenerationConfig , GroundingMetadata , ModalityTokenCount , and Tool , see Generating content .

،

The Live API is a stateful API that uses WebSockets . In this section, you'll find additional details regarding the WebSockets API.

جلسات

A WebSocket connection establishes a session between the client and the Gemini server. After a client initiates a new connection the session can exchange messages with the server to:

  • Send text, audio, or video to the Gemini server.
  • Receive audio, text, or function call requests from the Gemini server.

WebSocket connection

To start a session, connect to this websocket endpoint:

wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent

Session configuration

The initial message sent after establishing the WebSocket connection sets the session configuration, which includes the model, generation parameters, system instructions, and tools.

You cannot update the configuration while the connection is open. However, you can change the configuration parameters, except the model, when pausing and resuming via the session resumption mechanism .

See the following example configuration. Note that the name casing in SDKs may vary. You can look up the Python SDK configuration options here .


{
  "model": string,
  "generationConfig": {
    "candidateCount": integer,
    "maxOutputTokens": integer,
    "temperature": number,
    "topP": number,
    "topK": integer,
    "presencePenalty": number,
    "frequencyPenalty": number,
    "responseModalities": [string],
    "speechConfig": object,
    "mediaResolution": object,
    "translationConfig": object
  },
  "systemInstruction": string,
  "tools": [object]
}

For more information on the API field, see generationConfig .

Send messages

To exchange messages over the WebSocket connection, the client must send a JSON object over an open WebSocket connection. The JSON object must have exactly one of the fields from the following object set:


{
  "setup": BidiGenerateContentSetup,
  "clientContent": BidiGenerateContentClientContent,
  "realtimeInput": BidiGenerateContentRealtimeInput,
  "toolResponse": BidiGenerateContentToolResponse
}

Supported client messages

See the supported client messages in the following table:

پیام توضیحات
BidiGenerateContentSetup Session configuration to be sent in the first message
BidiGenerateContentClientContent Incremental content update of the current conversation delivered from the client
BidiGenerateContentRealtimeInput Real time audio, video, or text input
BidiGenerateContentToolResponse Response to a ToolCallMessage received from the server

Receive messages

To receive messages from Gemini, listen for the WebSocket 'message' event, and then parse the result according to the definition of the supported server messages.

See the following:

async with client.aio.live.connect(model='...', config=config) as session:
    await session.send(input='Hello world!', end_of_turn=True)
    async for message in session.receive():
        print(message)

Server messages may have a usageMetadata field but will otherwise include exactly one of the other fields from the BidiGenerateContentServerMessage message. (The messageType union is not expressed in JSON so the field will appear at the top-level of the message.)

Messages and events

ActivityEnd

This type has no fields.

Marks the end of user activity.

ActivityHandling

The different ways of handling user activity.

Enums
ACTIVITY_HANDLING_UNSPECIFIED If unspecified, the default behavior is START_OF_ACTIVITY_INTERRUPTS .
START_OF_ACTIVITY_INTERRUPTS If true, start of activity will interrupt the model's response (also called "barge in"). The model's current response will be cut-off in the moment of the interruption. This is the default behavior.
NO_INTERRUPTION The model's response will not be interrupted.

ActivityStart

This type has no fields.

Marks the start of user activity.

AudioTranscriptionConfig

The audio transcription configuration.

Fields
languageCodes[]

string

Optional. BCP-47 language codes providing hints about the languages present in the audio. If omitted or empty, defaults to automatic language detection.

customVocabulary[]

string

Optional. A list of custom vocabulary phrases to bias the speech recognition model toward recognizing specific terms (product names, proper nouns, jargon).

wordTimestamp

bool

Optional. Configures word-level timestamp generation.

diarization

bool

Optional. Configures speaker diarization.

mode

Mode

Optional. Configures transcription mode. Supported values: VERBATIM , SMART . If unspecified, defaults to VERBATIM transcription. In SMART mode, the model performs disfluency removal (eliminating filler words, repetitions, and false starts), light grammatical cleanup, automatic formatting (paragraphs, bullet points, numbered lists), and minor user edits (inline self-corrections). Timestamps and diarization are incompatible with mode SMART .

حالت

Transcription mode.

Enums
MODE_UNSPECIFIED Unspecified transcription mode.
VERBATIM Verbatim transcription mode.
SMART Smart transcription mode.

AutomaticActivityDetection

Configures automatic detection of activity.

Fields
disabled

bool

Optional. If enabled (the default), detected voice and text input count as activity. If disabled, the client must send activity signals.

startOfSpeechSensitivity

StartSensitivity

Optional. Determines how likely speech is to be detected.

prefixPaddingMs

int32

Optional. The required duration of detected speech before start-of-speech is committed. The lower this value, the more sensitive the start-of-speech detection is and shorter speech can be recognized. However, this also increases the probability of false positives.

endOfSpeechSensitivity

EndSensitivity

Optional. Determines how likely detected speech is ended.

silenceDurationMs

int32

Optional. The required duration of detected non-speech (eg silence) before end-of-speech is committed. The larger this value, the longer speech gaps can be without interrupting the user's activity but this will increase the model's latency.

BidiGenerateContentClientContent

Incremental update of the current conversation delivered from the client. All of the content here is unconditionally appended to the conversation history and used as part of the prompt to the model to generate content.

A message here will interrupt any current model generation.

Fields
turns[]

Content

Optional. The content appended to the current conversation with the model.

For single-turn queries, this is a single instance. For multi-turn queries, this is a repeated field that contains conversation history and the latest request.

turnComplete

bool

Optional. If true, indicates that the server content generation should start with the currently accumulated prompt. Otherwise, the server awaits additional messages before starting generation.

BidiGenerateContentRealtimeInput

User input that is sent in real time.

The different modalities (audio, video and text) are handled as concurrent streams. The ordering across these streams is not guaranteed.

This is different from BidiGenerateContentClientContent in a few ways:

  • Can be sent continuously without interruption to model generation.
  • If there is a need to mix data interleaved across the BidiGenerateContentClientContent and the BidiGenerateContentRealtimeInput , the server attempts to optimize for best response, but there are no guarantees.
  • End of turn is not explicitly specified, but is rather derived from user activity (for example, end of speech).
  • Even before the end of turn, the data is processed incrementally to optimize for a fast start of the response from the model.
Fields
mediaChunks[]

Blob

Optional. Inlined bytes data for media input. Multiple mediaChunks are not supported, all but the first will be ignored.

DEPRECATED: Use one of audio , video , or text instead.

audio

Blob

Optional. These form the realtime audio input stream.

video

Blob

Optional. These form the realtime video input stream.

activityStart

ActivityStart

Optional. Marks the start of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled.

activityEnd

ActivityEnd

Optional. Marks the end of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled.

mediaResolution

MediaResolution

Optional. The media resolution to use. If not specified, setup.generationConfig.mediaResolution is used or a default if the setup is not provided.

audioStreamEnd

bool

Optional. Indicates that the audio stream has ended, eg because the microphone was turned off.

This should only be sent when automatic activity detection is enabled (which is the default).

The client can reopen the stream by sending an audio message.

text

string

Optional. These form the realtime text input stream.

BidiGenerateContentServerContent

Incremental server update generated by the model in response to client messages.

Content is generated as quickly as possible, and not in real time. Clients may choose to buffer and play it out in real time.

Fields
generationComplete

bool

Output only. If true, indicates that the model is done generating.

When model is interrupted while generating there will be no 'generation_complete' message in interrupted turn, it will go through 'interrupted > turn_complete'.

When model assumes realtime playback there will be delay between generation_complete and turn_complete that is caused by model waiting for playback to finish.

turnComplete

bool

Output only. If true, indicates that the model has completed its turn. Generation will only start in response to additional client messages. Note when playback status reporting is enabled, this is emitted only when the playback status indicates that the playback is done. Future playback status of the same generation will be ignored.

interrupted

bool

Output only. If true, indicates that a client message has interrupted current model generation. If the client is playing out the content in real time, this is a good signal to stop and empty the current playback queue.

groundingMetadata

GroundingMetadata

Output only. Grounding metadata for the generated content.

inputTranscription

BidiGenerateContentTranscription

Output only. Input audio transcription. The transcription is sent independently of the other server messages and there is no guaranteed ordering.

interimInputTranscription

BidiGenerateContentTranscription

Output only. Low latency transcription updated while the user is speaking. This field is subject to frequent updates.

outputTranscription

BidiGenerateContentTranscription

Output only. Output audio transcription. These transcriptions are part of the Generation output of the server. The last output transcription of this turn is sent before either generationComplete or interrupted , which in turn are followed by turnComplete . There is no guaranteed exact ordering between transcriptions and other modelTurn output but the server tries to send the transcripts close to the corresponding audio output.

urlContextMetadata

UrlContextMetadata

waitingForInput

bool

Output only. If true, indicates that the model is not generating content because it is waiting for more input from the user, eg because it expects the user to continue talking.

speechState
(deprecated)

SpeechState

Output only. DEPRECATED: Use VoiceActivity instead.

Indicates the current state of speech detection on realtimeInput.audio . Not set or zero if the state is unchanged.

interactionStatus

InteractionStatus

Output only. The current activity status of the live session. Always sent alongside turnComplete .

modelTurn

Content

Output only. The content that the model has generated as part of the current conversation with the user.

BidiGenerateContentServerMessage

Response message for the BidiGenerateContent call.

Fields
usageMetadata

UsageMetadata

Output only. Usage metadata about the response(s).

Union field messageType . The type of the message. messageType can be only one of the following:
setupComplete

BidiGenerateContentSetupComplete

Output only. Sent in response to a BidiGenerateContentSetup message from the client when setup is complete.

serverContent

BidiGenerateContentServerContent

Output only. Content generated by the model in response to client messages.

toolCall

BidiGenerateContentToolCall

Output only. Request for the client to execute the functionCalls and return the responses with the matching id s.

toolCallCancellation

BidiGenerateContentToolCallCancellation

Output only. Notification for the client that a previously issued ToolCallMessage with the specified id s should be cancelled.

goAway

GoAway

Output only. A notice that the server will soon disconnect.

sessionResumptionUpdate

SessionResumptionUpdate

Output only. Update of the session resumption state.

BidiGenerateContentSetup

Message to be sent in the first (and only in the first) BidiGenerateContentClientMessage . Contains configuration that will apply for the duration of the streaming RPC.

Clients should wait for a BidiGenerateContentSetupComplete message before sending any additional messages.

Fields
model

string

Required. The model's resource name. This serves as an ID for the Model to use.

Format: models/{model}

generationConfig

GenerationConfig

Optional. Generation config.

The following fields are not supported:

  • responseLogprobs
  • responseMimeType
  • logprobs
  • responseSchema
  • responseJsonSchema
  • stopSequence
  • skipResponseCache
  • routingConfig
  • audioTimestamp
systemInstruction

Content

Optional. The user provided system instructions for the model.

Note: Only text should be used in parts and content in each part will be in a separate paragraph.

tools[]

Tool

Optional. A list of Tools the model may use to generate the next response.

A Tool is a piece of code that enables the system to interact with external systems to perform an action, or set of actions, outside of knowledge and scope of the model.

realtimeInputConfig

RealtimeInputConfig

Optional. Configures the handling of realtime input.

sessionResumption

SessionResumptionConfig

Optional. Configures session resumption mechanism.

If included, the server will send SessionResumptionUpdate messages.

contextWindowCompression

ContextWindowCompressionConfig

Optional. Configures a context window compression mechanism.

If included, the server will automatically reduce the size of the context when it exceeds the configured length.

inputAudioTranscription

AudioTranscriptionConfig

Optional. If set, enables transcription of voice input. The transcription aligns with the input audio language, if configured.

outputAudioTranscription

AudioTranscriptionConfig

Optional. If set, enables transcription of the model's audio output. The transcription aligns with the language code specified for the output audio, if configured.

proactivity

ProactivityConfig

Optional. Configures the proactivity of the model.

This allows the model to respond proactively to the input and to ignore irrelevant input.

historyConfig

HistoryConfig

Optional. Configures the exchange of history between the client and the server.

BidiGenerateContentSetupComplete

This type has no fields.

Sent in response to a BidiGenerateContentSetup message from the client.

BidiGenerateContentToolCall

Request for the client to execute the functionCalls and return the responses with the matching id s.

Fields
functionCalls[]

FunctionCall

Output only. The function call to be executed.

BidiGenerateContentToolCallCancellation

Notification for the client that a previously issued ToolCallMessage with the specified id s should not have been executed and should be cancelled. If there were side-effects to those tool calls, clients may attempt to undo the tool calls. This message occurs only in cases where the clients interrupt server turns.

Fields
ids[]

string

Output only. The ids of the tool calls to be cancelled.

BidiGenerateContentToolResponse

Client generated response to a ToolCall received from the server. Individual FunctionResponse objects are matched to the respective FunctionCall objects by the id field.

Note that in the unary and server-streaming GenerateContent APIs function calling happens by exchanging the Content parts, while in the bidi GenerateContent APIs function calling happens over these dedicated set of messages.

Fields
functionResponses[]

FunctionResponse

Optional. The response to the function calls.

BidiGenerateContentTranscription

Transcription of audio (input or output).

Fields
text

string

Transcription text.

languageCode

string

The BCP-47 language code of the transcription.

ContextWindowCompressionConfig

Enables context window compression — a mechanism for managing the model's context window so that it does not exceed a given length.

Fields
Union field compressionMechanism . The context window compression mechanism used. compressionMechanism can be only one of the following:
slidingWindow

SlidingWindow

A sliding-window mechanism.

triggerTokens

int64

The number of tokens (before running a turn) required to trigger a context window compression.

This can be used to balance quality against latency as shorter context windows may result in faster model responses. However, any compression operation will cause a temporary latency increase, so they should not be triggered frequently.

If not set, the default is 80% of the model's context window limit. This leaves 20% for the next user request/model response.

EndSensitivity

Determines how end of speech is detected.

Enums
END_SENSITIVITY_UNSPECIFIED The default is END_SENSITIVITY_HIGH.
END_SENSITIVITY_HIGH Automatic detection ends speech more often.
END_SENSITIVITY_LOW Automatic detection ends speech less often.

GoAway

A notice that the server will soon disconnect.

Fields
timeLeft

Duration

The remaining time before the connection will be terminated as ABORTED.

This duration will never be less than a model-specific minimum, which will be specified together with the rate limits for the model.

HistoryConfig

History configuration.

This message is included in the session configuration as BidiGenerateContentSetup.historyConfig . Configures the exchange of history messages.

Fields
initialHistoryInClientContent

bool

Optional. If true, after sending setupComplete , the server will wait and at first process clientContent messages until turnComplete is true . This initial history will not trigger a model call and may end with role MODEL . After turnComplete is true , the client can start the realtime conversation via realtimeInput .

ProactivityConfig

Config for proactivity features.

Fields
proactiveAudio

bool

Optional. If enabled, the model can reject responding to the last prompt. For example, this allows the model to ignore out of context speech or to stay silent if the user did not make a request, yet.

RealtimeInputConfig

Configures the realtime input behavior in BidiGenerateContent .

Fields
automaticActivityDetection

AutomaticActivityDetection

Optional. If not set, automatic activity detection is enabled by default. If automatic voice detection is disabled, the client must send activity signals.

activityHandling

ActivityHandling

Optional. Defines what effect activity has.

turnCoverage

TurnCoverage

Optional. Defines which input is included in the user's turn.

SessionResumptionConfig

Session resumption configuration.

This message is included in the session configuration as BidiGenerateContentSetup.sessionResumption . If configured, the server will send SessionResumptionUpdate messages.

Fields
handle

string

The handle of a previous session. If not present then a new session is created.

Session handles come from SessionResumptionUpdate.token values in previous connections.

SessionResumptionUpdate

Update of the session resumption state.

Only sent if BidiGenerateContentSetup.sessionResumption was set.

Fields
newHandle

string

New handle that represents a state that can be resumed. Empty if resumable =false.

resumable

bool

True if the current session can be resumed at this point.

Resumption is not possible at some points in the session. For example, when the model is executing function calls or generating. Resuming the session (using a previous session token) in such a state will result in some data loss. In these cases, newHandle will be empty and resumable will be false.

SlidingWindow

The SlidingWindow method operates by discarding content at the beginning of the context window. The resulting context will always begin at the start of a USER role turn. System instructions and any BidiGenerateContentSetup.prefixTurns will always remain at the beginning of the result.

Fields
targetTokens

int64

The target number of tokens to keep. The default value is trigger_tokens/2.

Discarding parts of the context window causes a temporary latency increase so this value should be calibrated to avoid frequent compression operations.

StartSensitivity

Determines how start of speech is detected.

Enums
START_SENSITIVITY_UNSPECIFIED The default is START_SENSITIVITY_HIGH.
START_SENSITIVITY_HIGH Automatic detection will detect the start of speech more often.
START_SENSITIVITY_LOW Automatic detection will detect the start of speech less often.

TurnCoverage

Options about which input is included in the user's turn.

Enums
TURN_COVERAGE_UNSPECIFIED If unspecified, a default behavior is selected based on the model. Eg, for Gemini 2.5, the default is TURN_INCLUDES_ONLY_ACTIVITY , while for Gemini 3.1 and onwards, it's TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO .
TURN_INCLUDES_ONLY_ACTIVITY Includes activity since the last turn, excluding inactivity (eg silence on the audio stream).
TURN_INCLUDES_ALL_INPUT Includes all realtime input since the last turn, including inactivity (eg silence on the audio stream).
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO Includes audio activity and all video since the last turn. With automatic activity detection, audio activity means speech and excludes silence.

TranslationConfig

Config for translation features.

Fields
targetLanguageCode

string

Required. The target language for translation. Supported values are BCP-47 language codes (eg "en", "es", "fr").

echoTargetLanguage

bool

Optional. If true, the model will generate audio when the target language is spoken, essentially it will parrot the input. If false, we will not produce audio for the target language.

UrlContextMetadata

Metadata related to url context retrieval tool.

Fields
urlMetadata[]

UrlMetadata

List of url context.

UsageMetadata

Usage metadata about response(s).

Fields
promptTokenCount

int32

Output only. Number of tokens in the prompt. When cachedContent is set, this is still the total effective prompt size meaning this includes the number of tokens in the cached content.

cachedContentTokenCount

int32

Number of tokens in the cached part of the prompt (the cached content)

responseTokenCount

int32

Output only. Total number of tokens across all the generated response candidates.

toolUsePromptTokenCount

int32

Output only. Number of tokens present in tool-use prompt(s).

thoughtsTokenCount

int32

Output only. Number of tokens of thoughts for thinking models.

totalTokenCount

int32

Output only. Total token count for the generation request (prompt + response candidates).

promptTokensDetails[]

ModalityTokenCount

Output only. List of modalities that were processed in the request input.

cacheTokensDetails[]

ModalityTokenCount

Output only. List of modalities of the cached content in the request input.

responseTokensDetails[]

ModalityTokenCount

Output only. List of modalities that were returned in the response.

toolUsePromptTokensDetails[]

ModalityTokenCount

Output only. List of modalities that were processed for tool-use request inputs.

Ephemeral authentication tokens

Ephemeral authentication tokens can be obtained by calling AuthTokenService.CreateToken and then used with GenerativeService.BidiGenerateContentConstrained , either by passing the token in an access_token query parameter, or in an HTTP Authorization header with " Token " prefixed to it.

CreateAuthTokenRequest

Create an ephemeral authentication token.

Fields
authToken

AuthToken

Required. The token to create.

AuthToken

A request to create an ephemeral authentication token.

Fields
name

string

Output only. Identifier. The token itself.

expireTime

Timestamp

Optional. Input only. Immutable. An optional time after which, when using the resulting token, messages in BidiGenerateContent sessions will be rejected. (Gemini may preemptively close the session after this time.)

If not set then this defaults to 30 minutes in the future. If set, this value must be less than 20 hours in the future.

newSessionExpireTime

Timestamp

Optional. Input only. Immutable. The time after which new Live API sessions using the token resulting from this request will be rejected.

If not set this defaults to 60 seconds in the future. If set, this value must be less than 20 hours in the future.

fieldMask

FieldMask

Optional. Input only. Immutable. If field_mask is empty, and bidiGenerateContentSetup is not present, then the effective BidiGenerateContentSetup message is taken from the Live API connection.

If field_mask is empty, and bidiGenerateContentSetup is present, then the effective BidiGenerateContentSetup message is taken entirely from bidiGenerateContentSetup in this request. The setup message from the Live API connection is ignored.

If field_mask is not empty, then the corresponding fields from bidiGenerateContentSetup will overwrite the fields from the setup message in the Live API connection.

Union field config . The method-specific configuration for the resulting token. config can be only one of the following:
bidiGenerateContentSetup

BidiGenerateContentSetup

Optional. Input only. Immutable. Configuration specific to BidiGenerateContent .

uses

int32

Optional. Input only. Immutable. The number of times the token can be used. If this value is zero then no limit is applied. Resuming a Live API session does not count as a use. If unspecified, the default is 1.

More information on common types

For more information on the commonly-used API resource types Blob , Content , FunctionCall , FunctionResponse , GenerationConfig , GroundingMetadata , ModalityTokenCount , and Tool , see Generating content .