Live API یک API با قابلیت Stateful است که از WebSockets استفاده میکند. در این بخش، جزئیات بیشتری در مورد WebSockets API خواهید یافت.
جلسات
یک اتصال WebSocket یک جلسه بین کلاینت و سرور Gemini برقرار میکند. پس از اینکه کلاینت یک اتصال جدید را آغاز میکند، جلسه میتواند پیامهایی را با سرور رد و بدل کند تا:
- متن، صدا یا ویدیو را به سرور جمینی ارسال کنید.
- درخواستهای صوتی، متنی یا تماس عملکردی را از سرور Gemini دریافت کنید.
اتصال وبسوکت
برای شروع یک جلسه، به این نقطه پایانی وب سوکت متصل شوید:
wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent
پیکربندی جلسه
پیام اولیهای که پس از برقراری اتصال WebSocket ارسال میشود، پیکربندی جلسه را تنظیم میکند که شامل مدل، پارامترهای تولید، دستورالعملهای سیستم و ابزارها میشود.
شما نمیتوانید پیکربندی را در حالی که اتصال باز است، بهروزرسانی کنید. با این حال، میتوانید پارامترهای پیکربندی، به جز مدل، را هنگام مکث و از سرگیری از طریق مکانیسم از سرگیری جلسه تغییر دهید.
به مثال پیکربندی زیر توجه کنید. توجه داشته باشید که نحوهی نامگذاری در SDKها ممکن است متفاوت باشد. میتوانید گزینههای پیکربندی SDK پایتون را اینجا جستجو کنید.
{
"model": string,
"generationConfig": {
"candidateCount": integer,
"maxOutputTokens": integer,
"temperature": number,
"topP": number,
"topK": integer,
"presencePenalty": number,
"frequencyPenalty": number,
"responseModalities": [string],
"speechConfig": object,
"mediaResolution": object,
"translationConfig": object
},
"systemInstruction": string,
"tools": [object]
}
برای اطلاعات بیشتر در مورد فیلد API، به generationConfig مراجعه کنید.
ارسال پیام
برای تبادل پیام از طریق اتصال WebSocket، کلاینت باید یک شیء JSON را از طریق یک اتصال WebSocket باز ارسال کند. شیء JSON باید دقیقاً یکی از فیلدهای مجموعه شیء زیر را داشته باشد:
{
"setup": BidiGenerateContentSetup,
"clientContent": BidiGenerateContentClientContent,
"realtimeInput": BidiGenerateContentRealtimeInput,
"toolResponse": BidiGenerateContentToolResponse
}
پیامهای کلاینت پشتیبانیشده
پیامهای کلاینت پشتیبانیشده را در جدول زیر مشاهده کنید:
| پیام | توضیحات |
|---|---|
BidiGenerateContentSetup | پیکربندی جلسه که باید در اولین پیام ارسال شود |
BidiGenerateContentClientContent | بهروزرسانی تدریجی محتوای مکالمه فعلی ارائه شده از طرف کلاینت |
BidiGenerateContentRealtimeInput | ورودی صوتی، تصویری یا متنی بلادرنگ |
BidiGenerateContentToolResponse | پاسخ به یک ToolCallMessage دریافت شده از سرور |
دریافت پیامها
برای دریافت پیامها از Gemini، به رویداد 'message' در WebSocket گوش دهید و سپس نتیجه را طبق تعریف پیامهای سرور پشتیبانیشده تجزیه کنید.
موارد زیر را ببینید:
async with client.aio.live.connect(model='...', config=config) as session:
await session.send(input='Hello world!', end_of_turn=True)
async for message in session.receive():
print(message)
پیامهای سرور ممکن است دارای فیلد usageMetadata باشند، اما در غیر این صورت دقیقاً یکی از فیلدهای دیگر پیام BidiGenerateContentServerMessage را شامل میشوند. (اتحادیه messageType در JSON بیان نشده است، بنابراین این فیلد در سطح بالای پیام ظاهر میشود.)
پیامها و رویدادها
پایان فعالیت
این نوع هیچ فیلدی ندارد.
پایان فعالیت کاربر را نشان میدهد.
مدیریت فعالیت
روشهای مختلف مدیریت فعالیت کاربران
| انومها | |
|---|---|
ACTIVITY_HANDLING_UNSPECIFIED | اگر مشخص نشده باشد، رفتار پیشفرض START_OF_ACTIVITY_INTERRUPTS است. |
START_OF_ACTIVITY_INTERRUPTS | اگر درست باشد، شروع فعالیت، پاسخ مدل را قطع میکند (که به آن "barge in" نیز میگویند). پاسخ فعلی مدل در لحظه وقفه قطع میشود. این رفتار پیشفرض است. |
NO_INTERRUPTION | پاسخ مدل قطع نخواهد شد. |
شروع فعالیت
این نوع هیچ فیلدی ندارد.
شروع فعالیت کاربر را نشان میدهد.
پیکربندی رونویسی صوتی
پیکربندی رونویسی صوتی.
| فیلدها | |
|---|---|
languageCodes[] | اختیاری. کدهای زبان BCP-47 نکاتی در مورد زبانهای موجود در صدا ارائه میدهند. در صورت حذف یا خالی بودن، به طور پیشفرض روی تشخیص خودکار زبان تنظیم میشود. |
customVocabulary[] | اختیاری. فهرستی از عبارات واژگانی سفارشی برای جهتدهی مدل تشخیص گفتار به سمت تشخیص اصطلاحات خاص (نام محصولات، اسمهای خاص، اصطلاحات تخصصی). |
wordTimestamp | اختیاری. تولید مهر زمانی در سطح کلمه را پیکربندی میکند. |
diarization | اختیاری. تنظیم فاصله بین بلندگوها را پیکربندی میکند. |
mode | اختیاری. حالت رونویسی را پیکربندی میکند. مقادیر پشتیبانیشده: |
حالت
حالت رونویسی.
| انومها | |
|---|---|
MODE_UNSPECIFIED | حالت رونویسی نامشخص. |
VERBATIM | حالت رونویسی کلمه به کلمه. |
SMART | حالت رونویسی هوشمند. |
تشخیص خودکار فعالیت
تشخیص خودکار فعالیت را پیکربندی میکند.
| فیلدها | |
|---|---|
disabled | اختیاری. در صورت فعال بودن (پیشفرض)، ورودیهای صوتی و متنی شناساییشده به عنوان فعالیت شمارش میشوند. در صورت غیرفعال بودن، کلاینت باید سیگنالهای فعالیت ارسال کند. |
startOfSpeechSensitivity | اختیاری. تعیین میکند که احتمال تشخیص گفتار چقدر است. |
prefixPaddingMs | اختیاری. مدت زمان مورد نیاز برای تشخیص گفتار قبل از شروع گفتار. هرچه این مقدار کمتر باشد، تشخیص شروع گفتار حساستر است و گفتار کوتاهتر قابل تشخیص است. با این حال، این امر احتمال تشخیصهای مثبت کاذب را نیز افزایش میدهد. |
endOfSpeechSensitivity | اختیاری. تعیین میکند که احتمال پایان یافتن گفتار شناساییشده چقدر است. |
silenceDurationMs | اختیاری. مدت زمان مورد نیاز برای تشخیص عدم گفتار (مثلاً سکوت) قبل از پایان گفتار. هرچه این مقدار بزرگتر باشد، میتوان فواصل گفتاری را بدون ایجاد وقفه در فعالیت کاربر طولانیتر کرد، اما این امر تأخیر مدل را افزایش میدهد. |
BidiGenerateContentClientContent
بهروزرسانی افزایشی مکالمه فعلی ارائه شده از کلاینت. تمام محتوای اینجا بدون قید و شرط به تاریخچه مکالمه اضافه میشود و به عنوان بخشی از اعلان مدل برای تولید محتوا استفاده میشود.
یک پیام در اینجا هرگونه تولید مدل فعلی را قطع میکند.
| فیلدها | |
|---|---|
turns[] | اختیاری. محتوایی که به مکالمه فعلی با مدل اضافه شده است. برای پرسوجوهای تک نوبتی، این یک نمونه واحد است. برای پرسوجوهای چند نوبتی، این یک فیلد تکراری است که شامل سابقه مکالمه و آخرین درخواست است. |
turnComplete | اختیاری. اگر درست باشد، نشان میدهد که تولید محتوای سرور باید با اعلان انباشتهشدهی فعلی شروع شود. در غیر این صورت، سرور قبل از شروع تولید منتظر پیامهای اضافی میماند. |
BidiGenerateContentRealtimeInput
ورودی کاربر که به صورت بلادرنگ ارسال میشود.
روشهای مختلف (صوت، تصویر و متن) به صورت جریانهای همزمان مدیریت میشوند. ترتیب قرارگیری در بین این جریانها تضمین شده نیست.
این از چند جهت با BidiGenerateContentClientContent متفاوت است:
- میتواند به طور مداوم و بدون وقفه برای تولید مدل ارسال شود.
- اگر نیاز به ترکیب دادههای موجود در بین
BidiGenerateContentClientContentوBidiGenerateContentRealtimeInputباشد، سرور تلاش میکند تا بهترین پاسخ را بهینهسازی کند، اما هیچ تضمینی وجود ندارد. - پایان نوبت به صراحت مشخص نشده است، بلکه از فعالیت کاربر (مثلاً پایان سخنرانی) استنباط میشود.
- حتی قبل از پایان نوبت، دادهها به صورت تدریجی پردازش میشوند تا برای شروع سریع پاسخ از مدل بهینه شوند.
| فیلدها | |
|---|---|
mediaChunks[] | اختیاری. دادههای بایت درونخطی برای ورودی رسانه. چندین منسوخ شده: به جای آن از یکی از |
audio | اختیاری. اینها جریان ورودی صوتی بلادرنگ را تشکیل میدهند. |
video | اختیاری. اینها جریان ورودی ویدیوی بلادرنگ را تشکیل میدهند. |
activityStart | اختیاری. شروع فعالیت کاربر را نشان میدهد. این فقط در صورتی قابل ارسال است که تشخیص خودکار فعالیت (یعنی سمت سرور) غیرفعال باشد. |
activityEnd | اختیاری. پایان فعالیت کاربر را نشان میدهد. این فقط در صورتی قابل ارسال است که تشخیص خودکار فعالیت (یعنی سمت سرور) غیرفعال باشد. |
mediaResolution | اختیاری. وضوح رسانهای که باید استفاده شود. اگر مشخص نشده باشد، |
audioStreamEnd | اختیاری. نشان میدهد که جریان صوتی پایان یافته است، مثلاً به دلیل خاموش شدن میکروفون. این فقط باید زمانی ارسال شود که تشخیص خودکار فعالیت فعال باشد (که پیشفرض است). کلاینت میتواند با ارسال یک پیام صوتی، استریم را دوباره باز کند. |
text | اختیاری. اینها جریان ورودی متن بلادرنگ را تشکیل میدهند. |
BidiGenerateContentServerContent
بهروزرسانی افزایشی سرور که توسط مدل در پاسخ به پیامهای کلاینت ایجاد میشود.
محتوا در سریعترین زمان ممکن تولید میشود و نه به صورت بلادرنگ. مشتریان میتوانند آن را ذخیره کرده و به صورت بلادرنگ پخش کنند.
| فیلدها | |
|---|---|
generationComplete | فقط خروجی. اگر درست باشد، نشان میدهد که مدل تولید را تمام کرده است. وقتی مدل هنگام تولید دچار وقفه شود، هیچ پیام «generation_complete» در نوبت وقفهدار وجود نخواهد داشت، و از طریق «interrupted > turn_complete» اجرا میشود. وقتی مدل فرض میکند که پخش در زمان واقعی انجام میشود، بین generation_complete و turn_complete تأخیری وجود خواهد داشت که ناشی از انتظار مدل برای پایان پخش است. |
turnComplete | فقط خروجی. اگر درست باشد، نشان میدهد که مدل نوبت خود را تکمیل کرده است. تولید فقط در پاسخ به پیامهای اضافی کلاینت شروع میشود. توجه داشته باشید که وقتی گزارش وضعیت پخش فعال است، این فقط زمانی منتشر میشود که وضعیت پخش نشان دهد که پخش انجام شده است. وضعیت پخش آینده در همان نسل نادیده گرفته میشود. |
interrupted | فقط خروجی. اگر درست باشد، نشان میدهد که یک پیام کلاینت، تولید مدل فعلی را متوقف کرده است. اگر کلاینت در حال پخش محتوا به صورت بلادرنگ است، این سیگنال خوبی برای توقف و خالی کردن صف پخش فعلی است. |
groundingMetadata | فقط خروجی. فرادادههای زمینهای برای محتوای تولید شده. |
inputTranscription | فقط خروجی. رونویسی صوتی ورودی. رونویسی مستقل از سایر پیامهای سرور ارسال میشود و هیچ ترتیب تضمینشدهای وجود ندارد. |
interimInputTranscription | فقط خروجی. رونویسی با تأخیر کم هنگام صحبت کاربر بهروزرسانی میشود. این فیلد مرتباً بهروزرسانی میشود. |
outputTranscription | فقط خروجی. رونویسی صوتی خروجی. این رونویسیها بخشی از خروجی Generation سرور هستند. آخرین رونویسی خروجی این نوبت قبل از |
urlContextMetadata | |
waitingForInput | فقط خروجی. اگر درست باشد، نشان میدهد که مدل محتوا تولید نمیکند زیرا منتظر ورودی بیشتر از کاربر است، مثلاً چون انتظار دارد کاربر به صحبت کردن ادامه دهد. |
speechState | فقط خروجی. منسوخ شده: به جای آن از VoiceActivity استفاده کنید. وضعیت فعلی تشخیص گفتار را در |
interactionStatus | فقط خروجی. وضعیت فعالیت فعلی جلسه زنده. همیشه در کنار |
modelTurn | فقط خروجی. محتوایی که مدل به عنوان بخشی از مکالمه فعلی با کاربر تولید کرده است. |
BidiGenerateContentServerMessage
پیام پاسخ برای فراخوانی BidiGenerateContent.
| فیلدها | |
|---|---|
usageMetadata | فقط خروجی. استفاده از فراداده در مورد پاسخ(ها). |
فیلد union messageType . نوع پیام. messageType فقط میتواند یکی از موارد زیر باشد: | |
setupComplete | فقط خروجی. در پاسخ به پیام |
serverContent | فقط خروجی. محتوایی که توسط مدل در پاسخ به پیامهای کلاینت تولید میشود. |
toolCall | فقط خروجی. از کلاینت درخواست کنید تا |
toolCallCancellation | فقط خروجی. اعلانی برای کلاینت مبنی بر اینکه |
goAway | فقط خروجی. اخطاری مبنی بر اینکه سرور به زودی قطع خواهد شد. |
sessionResumptionUpdate | فقط خروجی. بهروزرسانی وضعیت از سرگیری جلسه. |
تنظیمات محتوای BidiGenerate
پیامی که قرار است در اولین (و فقط در اولین) BidiGenerateContentClientMessage ارسال شود. حاوی پیکربندی است که در طول مدت RPC استریمینگ اعمال خواهد شد.
کلاینتها باید قبل از ارسال هرگونه پیام اضافی، منتظر پیام BidiGenerateContentSetupComplete باشند.
| فیلدها | |
|---|---|
model | الزامی. نام منبع مدل. این به عنوان شناسهای برای استفاده مدل عمل میکند. قالب: |
generationConfig | اختیاری. پیکربندی نسل. فیلدهای زیر پشتیبانی نمیشوند:
|
systemInstruction | اختیاری. کاربر دستورالعملهای سیستمی را برای مدل ارائه داده است. توجه: فقط متن باید در بخشها استفاده شود و محتوای هر بخش در یک پاراگراف جداگانه قرار گیرد. |
tools[] | اختیاری. فهرستی از |
realtimeInputConfig | اختیاری. نحوهی مدیریت ورودیهای بیدرنگ را پیکربندی میکند. |
sessionResumption | اختیاری. مکانیزم از سرگیری جلسه را پیکربندی میکند. در صورت وجود، سرور پیامهای |
contextWindowCompression | اختیاری. مکانیزم فشردهسازی پنجره زمینه را پیکربندی میکند. در صورت وجود، سرور به طور خودکار اندازه context را هنگامی که از طول پیکربندی شده تجاوز کند، کاهش میدهد. |
inputAudioTranscription | اختیاری. در صورت تنظیم، رونویسی ورودی صوتی را فعال میکند. در صورت پیکربندی، رونویسی با زبان صوتی ورودی همتراز میشود. |
outputAudioTranscription | اختیاری. در صورت تنظیم، رونویسی خروجی صدای مدل را فعال میکند. در صورت پیکربندی، رونویسی با کد زبان مشخص شده برای صدای خروجی همتراز میشود. |
proactivity | اختیاری. میزان فعالیت مدل را پیکربندی میکند. این به مدل اجازه میدهد تا به ورودیها به صورت پیشگیرانه پاسخ دهد و ورودیهای نامربوط را نادیده بگیرد. |
historyConfig | اختیاری. تبادل تاریخچه بین کلاینت و سرور را پیکربندی میکند. |
BidiGenerateContentSetupComplete
این نوع هیچ فیلدی ندارد.
در پاسخ به پیام BidiGenerateContentSetup از کلاینت ارسال شده است.
تماس با ابزار تولید محتوا (BidiGenerateContentToolCall)
از کلاینت درخواست کنید تا functionCalls را اجرا کند و پاسخها را با id منطبق s برگرداند.
| فیلدها | |
|---|---|
functionCalls[] | فقط خروجی. فراخوانی تابعی که قرار است اجرا شود. |
BidiGenerateContentToolCallCancellation
اعلانی برای کلاینت مبنی بر اینکه ToolCallMessage قبلاً صادر شده با id مشخص شده نباید اجرا میشد و باید لغو شود. اگر عوارض جانبی در آن فراخوانیهای ابزار وجود داشته باشد، کلاینتها میتوانند سعی کنند فراخوانیهای ابزار را لغو کنند. این پیام فقط در مواردی رخ میدهد که کلاینتها چرخشهای سرور را قطع میکنند.
| فیلدها | |
|---|---|
ids[] | فقط خروجی. شناسههای فراخوانیهای ابزار باید لغو شوند. |
BidiGenerateContentToolResponse
پاسخ تولید شده توسط کلاینت به یک ToolCall دریافتی از سرور. اشیاء FunctionResponse به صورت جداگانه توسط فیلد id با اشیاء FunctionCall مربوطه تطبیق داده میشوند.
توجه داشته باشید که در APIهای GenerateContent تکی و جریانی سرور، فراخوانی تابع با تبادل بخشهای Content انجام میشود، در حالی که در APIهای GenerateContent بیدی، فراخوانی تابع روی این مجموعه اختصاصی از پیامها انجام میشود.
| فیلدها | |
|---|---|
functionResponses[] | اختیاری. پاسخ به فراخوانیهای تابع. |
BidiGenerateContentTranscription
رونویسی صدا (ورودی یا خروجی).
| فیلدها | |
|---|---|
text | متن رونویسی. |
languageCode | کد زبانی BCP-47 مربوط به رونویسی. |
پیکربندی ContextWindowCompression
فشردهسازی پنجره زمینه را فعال میکند - مکانیزمی برای مدیریت پنجره زمینه مدل به طوری که از طول مشخصی تجاوز نکند.
| فیلدها | |
|---|---|
compressionMechanism فیلد یونیون. مکانیسم فشردهسازی پنجره زمینه مورد استفاده. compressionMechanism میتواند فقط یکی از موارد زیر باشد: | |
slidingWindow | مکانیزم پنجره کشویی. |
triggerTokens | تعداد توکنهایی (قبل از اجرای یک نوبت) که برای فشردهسازی پنجرهی زمینه لازم است. این میتواند برای ایجاد تعادل بین کیفیت و تأخیر استفاده شود، زیرا پنجرههای متنی کوتاهتر ممکن است منجر به پاسخهای سریعتر مدل شوند. با این حال، هرگونه عملیات فشردهسازی باعث افزایش موقت تأخیر میشود، بنابراین نباید مرتباً فعال شوند. اگر تنظیم نشود، پیشفرض ۸۰٪ از محدودیت پنجرهی زمینهی مدل است. این مقدار ۲۰٪ را برای درخواست/پاسخ مدل بعدی کاربر باقی میگذارد. |
پایان حساسیت
نحوه تشخیص پایان گفتار را تعیین میکند.
| انومها | |
|---|---|
END_SENSITIVITY_UNSPECIFIED | مقدار پیشفرض END_SENSITIVITY_HIGH است. |
END_SENSITIVITY_HIGH | تشخیص خودکار، گفتار را بیشتر قطع میکند. |
END_SENSITIVITY_LOW | تشخیص خودکار، گفتار را کمتر قطع میکند. |
برو کنار
اخطاری مبنی بر اینکه سرور به زودی قطع خواهد شد.
| فیلدها | |
|---|---|
timeLeft | زمان باقی مانده قبل از اتصال به عنوان لغو شده (ABORTED) خاتمه خواهد یافت. این مدت هرگز کمتر از حداقل مشخص شده برای مدل نخواهد بود، که همراه با محدودیتهای نرخ برای مدل مشخص میشود. |
پیکربندی تاریخچه
پیکربندی تاریخچه.
این پیام در پیکربندی جلسه با عنوان BidiGenerateContentSetup.historyConfig گنجانده شده است. تبادل پیامهای تاریخچه را پیکربندی میکند.
| فیلدها | |
|---|---|
initialHistoryInClientContent | اختیاری. اگر مقدار setupComplete درست باشد، پس از ارسال |
پرواکتیویته کانفیگ
پیکربندی برای ویژگیهای پیشگیرانه.
| فیلدها | |
|---|---|
proactiveAudio | اختیاری. در صورت فعال بودن، مدل میتواند از پاسخ دادن به آخرین درخواست خودداری کند. برای مثال، این به مدل اجازه میدهد تا سخنان خارج از متن را نادیده بگیرد یا اگر کاربر هنوز درخواستی نداده است، سکوت کند. |
پیکربندی ورودی بلادرنگ
رفتار ورودی بیدرنگ را در BidiGenerateContent پیکربندی میکند.
| فیلدها | |
|---|---|
automaticActivityDetection | اختیاری. اگر تنظیم نشود، تشخیص خودکار فعالیت به طور پیشفرض فعال است. اگر تشخیص خودکار صدا غیرفعال باشد، کلاینت باید سیگنالهای فعالیت ارسال کند. |
activityHandling | اختیاری. تعریف میکند که فعالیت چه تأثیری دارد. |
turnCoverage | اختیاری. مشخص میکند که کدام ورودی در نوبت کاربر لحاظ شود. |
پیکربندی SessionResum
پیکربندی از سرگیری جلسه.
این پیام در پیکربندی جلسه با عنوان BidiGenerateContentSetup.sessionResumption گنجانده شده است. در صورت پیکربندی، سرور پیامهای SessionResumptionUpdate ارسال خواهد کرد.
| فیلدها | |
|---|---|
handle | شناسهی جلسهی قبلی. اگر موجود نباشد، یک جلسهی جدید ایجاد میشود. شناسههای نشست (Session Handles) از مقادیر |
بهروزرسانی از سرگیری جلسه
بهروزرسانی وضعیت از سرگیری جلسه.
فقط در صورتی ارسال میشود که BidiGenerateContentSetup.sessionResumption تنظیم شده باشد.
| فیلدها | |
|---|---|
newHandle | یک شناسه جدید که نشاندهنده وضعیتی است که میتواند از سر گرفته شود. در صورت |
resumable | اگر بتوان نشست فعلی را در این مرحله از سر گرفت، صحیح است. در برخی از نقاط session، از سرگیری امکانپذیر نیست. برای مثال، هنگامی که مدل در حال اجرای فراخوانیهای تابع یا تولید است. از سرگیری session (با استفاده از توکن session قبلی) در چنین حالتی منجر به از دست رفتن مقداری از دادهها خواهد شد. در این موارد، |
پنجره کشویی
متد SlidingWindow با حذف محتوا در ابتدای پنجره context عمل میکند. context حاصل همیشه از ابتدای چرخش نقش USER شروع میشود. دستورالعملهای سیستم و هرگونه BidiGenerateContentSetup.prefixTurns همیشه در ابتدای نتیجه باقی میمانند.
| فیلدها | |
|---|---|
targetTokens | تعداد هدف توکنهایی که باید نگه داشته شوند. مقدار پیشفرض trigger_tokens/2 است. حذف بخشهایی از پنجره زمینه باعث افزایش موقت تأخیر میشود، بنابراین این مقدار باید کالیبره شود تا از عملیات فشردهسازی مکرر جلوگیری شود. |
حساسیت شروع
نحوه تشخیص شروع گفتار را تعیین میکند.
| انومها | |
|---|---|
START_SENSITIVITY_UNSPECIFIED | مقدار پیشفرض START_SENSITIVITY_HIGH است. |
START_SENSITIVITY_HIGH | تشخیص خودکار، شروع گفتار را بیشتر تشخیص میدهد. |
START_SENSITIVITY_LOW | تشخیص خودکار، شروع گفتار را کمتر تشخیص میدهد. |
پوشش نوبتی
گزینههایی در مورد اینکه کدام ورودیها در نوبت کاربر لحاظ شوند.
| انومها | |
|---|---|
TURN_COVERAGE_UNSPECIFIED | اگر مشخص نشده باشد، یک رفتار پیشفرض بر اساس مدل انتخاب میشود. به عنوان مثال، برای Gemini 2.5، پیشفرض TURN_INCLUDES_ONLY_ACTIVITY است، در حالی که برای Gemini 3.1 و بالاتر، TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO است. |
TURN_INCLUDES_ONLY_ACTIVITY | شامل فعالیت از آخرین نوبت به بعد، به استثنای عدم فعالیت (مثلاً سکوت در جریان صوتی) میشود. |
TURN_INCLUDES_ALL_INPUT | شامل تمام ورودیهای بلادرنگ از آخرین نوبت، از جمله عدم فعالیت (مثلاً سکوت در جریان صوتی) میشود. |
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO | شامل فعالیتهای صوتی و تمام ویدیوهای ضبط شده از آخرین نوبت. با تشخیص خودکار فعالیت، فعالیت صوتی شامل گفتار میشود و سکوت را شامل نمیشود. |
پیکربندی ترجمه
پیکربندی برای ویژگیهای ترجمه.
| فیلدها | |
|---|---|
targetLanguageCode | الزامی. زبان مقصد برای ترجمه. مقادیر پشتیبانیشده کدهای زبان BCP-47 هستند (مثلاً "en"، "es"، "fr"). |
echoTargetLanguage | اختیاری. اگر مقدار آن درست باشد، مدل هنگام صحبت به زبان مقصد، صدا تولید میکند، اساساً ورودی را طوطیوار تکرار میکند. اگر مقدار آن نادرست باشد، ما برای زبان مقصد صدا تولید نمیکنیم. |
فرادادهی UrlContext
فراداده مربوط به ابزار بازیابی متن url.
| فیلدها | |
|---|---|
urlMetadata[] | فهرست زمینه آدرس اینترنتی. |
کاربردفراداده
فرادادههای مربوط به پاسخ(ها).
| فیلدها | |
|---|---|
promptTokenCount | فقط خروجی. تعداد توکنهای موجود در اعلان. وقتی |
cachedContentTokenCount | تعداد توکنها در بخش ذخیرهشدهی اعلان (محتوای ذخیرهشده) |
responseTokenCount | فقط خروجی. تعداد کل توکنها در بین تمام کاندیدهای پاسخ تولید شده. |
toolUsePromptTokenCount | فقط خروجی. تعداد توکنهای موجود در اعلان(های) استفاده از ابزار. |
thoughtsTokenCount | فقط خروجی. تعداد توکنهای افکار برای مدلهای تفکر. |
totalTokenCount | فقط خروجی. تعداد کل توکنها برای درخواست تولید (نامزدهای اعلان + پاسخ). |
promptTokensDetails[] | فقط خروجی. فهرست روشهایی که در ورودی درخواست پردازش شدهاند. |
cacheTokensDetails[] | فقط خروجی. فهرستی از روشهای محتوای ذخیرهشده در ورودی درخواست. |
responseTokensDetails[] | فقط خروجی. فهرست روشهایی که در پاسخ برگردانده شدهاند. |
toolUsePromptTokensDetails[] | فقط خروجی. فهرست روشهایی که برای ورودیهای درخواست استفاده از ابزار پردازش شدهاند. |
توکنهای احراز هویت موقت
توکنهای احراز هویت موقت را میتوان با فراخوانی AuthTokenService.CreateToken به دست آورد و سپس با GenerativeService.BidiGenerateContentConstrained استفاده کرد، یا با ارسال توکن در یک پارامتر پرسوجوی access_token ، یا در یک هدر HTTP Authorization با پیشوند " Token " به آن.
درخواست ایجاد توکن احراز هویت
یک توکن احراز هویت موقت ایجاد کنید.
| فیلدها | |
|---|---|
authToken | الزامی. توکنی که باید ایجاد شود. |
توکن احراز هویت
درخواستی برای ایجاد یک توکن احراز هویت موقت.
| فیلدها | |
|---|---|
name | فقط خروجی. شناسه. خود توکن. |
expireTime | اختیاری. فقط ورودی. تغییرناپذیر. یک زمان اختیاری که پس از آن، هنگام استفاده از توکن حاصل، پیامهای موجود در جلسات BidiGenerateContent رد میشوند. (Gemini ممکن است پس از این زمان، جلسه را به صورت پیشگیرانه ببندد.) اگر تنظیم نشده باشد، این مقدار به صورت پیشفرض روی ۳۰ دقیقه در آینده تنظیم میشود. در صورت تنظیم، این مقدار باید کمتر از ۲۰ ساعت در آینده باشد. |
newSessionExpireTime | اختیاری. فقط ورودی. تغییرناپذیر. زمانی که پس از آن، جلسات جدید Live API با استفاده از توکن حاصل از این درخواست رد میشوند. اگر تنظیم نشود، پیشفرضها روی ۶۰ ثانیه در آینده تنظیم میشوند. اگر تنظیم شود، این مقدار باید کمتر از ۲۰ ساعت در آینده باشد. |
fieldMask | اختیاری. فقط ورودی. تغییرناپذیر. اگر field_mask خالی باشد و اگر field_mask خالی باشد و اگر field_mask خالی نباشد، فیلدهای مربوطه از |
config فیلد Union. پیکربندی مختص متد برای token.config config میتواند فقط یکی از موارد زیر باشد: | |
bidiGenerateContentSetup | اختیاری. فقط ورودی. تغییرناپذیر. پیکربندی مختص |
uses | اختیاری. فقط ورودی. تغییرناپذیر. تعداد دفعاتی که میتوان از توکن استفاده کرد. اگر این مقدار صفر باشد، هیچ محدودیتی اعمال نمیشود. از سرگیری یک جلسه Live API به عنوان یک استفاده حساب نمیشود. اگر مشخص نشود، پیشفرض ۱ است. |
اطلاعات بیشتر در مورد انواع رایج
برای اطلاعات بیشتر در مورد انواع منابع API رایج Blob ، Content ، FunctionCall ، FunctionResponse ، GenerationConfig ، GroundingMetadata ، ModalityTokenCount و Tool ، به بخش تولید محتوا مراجعه کنید.
Live API یک API با قابلیت Stateful است که از WebSockets استفاده میکند. در این بخش، جزئیات بیشتری در مورد WebSockets API خواهید یافت.
جلسات
یک اتصال WebSocket یک جلسه بین کلاینت و سرور Gemini برقرار میکند. پس از اینکه کلاینت یک اتصال جدید را آغاز میکند، جلسه میتواند پیامهایی را با سرور رد و بدل کند تا:
- متن، صدا یا ویدیو را به سرور جمینی ارسال کنید.
- درخواستهای صوتی، متنی یا تماس عملکردی را از سرور Gemini دریافت کنید.
اتصال وبسوکت
برای شروع یک جلسه، به این نقطه پایانی وب سوکت متصل شوید:
wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent
پیکربندی جلسه
پیام اولیهای که پس از برقراری اتصال WebSocket ارسال میشود، پیکربندی جلسه را تنظیم میکند که شامل مدل، پارامترهای تولید، دستورالعملهای سیستم و ابزارها میشود.
شما نمیتوانید پیکربندی را در حالی که اتصال باز است، بهروزرسانی کنید. با این حال، میتوانید پارامترهای پیکربندی، به جز مدل، را هنگام مکث و از سرگیری از طریق مکانیسم از سرگیری جلسه تغییر دهید.
به مثال پیکربندی زیر توجه کنید. توجه داشته باشید که نحوهی نامگذاری در SDKها ممکن است متفاوت باشد. میتوانید گزینههای پیکربندی SDK پایتون را اینجا جستجو کنید.
{
"model": string,
"generationConfig": {
"candidateCount": integer,
"maxOutputTokens": integer,
"temperature": number,
"topP": number,
"topK": integer,
"presencePenalty": number,
"frequencyPenalty": number,
"responseModalities": [string],
"speechConfig": object,
"mediaResolution": object,
"translationConfig": object
},
"systemInstruction": string,
"tools": [object]
}
برای اطلاعات بیشتر در مورد فیلد API، به generationConfig مراجعه کنید.
ارسال پیام
برای تبادل پیام از طریق اتصال WebSocket، کلاینت باید یک شیء JSON را از طریق یک اتصال WebSocket باز ارسال کند. شیء JSON باید دقیقاً یکی از فیلدهای مجموعه شیء زیر را داشته باشد:
{
"setup": BidiGenerateContentSetup,
"clientContent": BidiGenerateContentClientContent,
"realtimeInput": BidiGenerateContentRealtimeInput,
"toolResponse": BidiGenerateContentToolResponse
}
پیامهای کلاینت پشتیبانیشده
پیامهای کلاینت پشتیبانیشده را در جدول زیر مشاهده کنید:
| پیام | توضیحات |
|---|---|
BidiGenerateContentSetup | پیکربندی جلسه که باید در اولین پیام ارسال شود |
BidiGenerateContentClientContent | بهروزرسانی تدریجی محتوای مکالمه فعلی ارائه شده از طرف کلاینت |
BidiGenerateContentRealtimeInput | ورودی صوتی، تصویری یا متنی بلادرنگ |
BidiGenerateContentToolResponse | پاسخ به یک ToolCallMessage دریافت شده از سرور |
دریافت پیامها
برای دریافت پیامها از Gemini، به رویداد 'message' در WebSocket گوش دهید و سپس نتیجه را طبق تعریف پیامهای سرور پشتیبانیشده تجزیه کنید.
موارد زیر را ببینید:
async with client.aio.live.connect(model='...', config=config) as session:
await session.send(input='Hello world!', end_of_turn=True)
async for message in session.receive():
print(message)
پیامهای سرور ممکن است دارای فیلد usageMetadata باشند، اما در غیر این صورت دقیقاً یکی از فیلدهای دیگر پیام BidiGenerateContentServerMessage را شامل میشوند. (اتحادیه messageType در JSON بیان نشده است، بنابراین این فیلد در سطح بالای پیام ظاهر میشود.)
پیامها و رویدادها
پایان فعالیت
این نوع هیچ فیلدی ندارد.
پایان فعالیت کاربر را نشان میدهد.
مدیریت فعالیت
روشهای مختلف مدیریت فعالیت کاربران
| انومها | |
|---|---|
ACTIVITY_HANDLING_UNSPECIFIED | اگر مشخص نشده باشد، رفتار پیشفرض START_OF_ACTIVITY_INTERRUPTS است. |
START_OF_ACTIVITY_INTERRUPTS | اگر درست باشد، شروع فعالیت، پاسخ مدل را قطع میکند (که به آن "barge in" نیز میگویند). پاسخ فعلی مدل در لحظه وقفه قطع میشود. این رفتار پیشفرض است. |
NO_INTERRUPTION | پاسخ مدل قطع نخواهد شد. |
شروع فعالیت
این نوع هیچ فیلدی ندارد.
شروع فعالیت کاربر را نشان میدهد.
پیکربندی رونویسی صوتی
پیکربندی رونویسی صوتی.
| فیلدها | |
|---|---|
languageCodes[] | اختیاری. کدهای زبان BCP-47 نکاتی در مورد زبانهای موجود در صدا ارائه میدهند. در صورت حذف یا خالی بودن، به طور پیشفرض روی تشخیص خودکار زبان تنظیم میشود. |
customVocabulary[] | اختیاری. فهرستی از عبارات واژگانی سفارشی برای جهتدهی مدل تشخیص گفتار به سمت تشخیص اصطلاحات خاص (نام محصولات، اسمهای خاص، اصطلاحات تخصصی). |
wordTimestamp | اختیاری. تولید مهر زمانی در سطح کلمه را پیکربندی میکند. |
diarization | اختیاری. تنظیم فاصله بین بلندگوها را پیکربندی میکند. |
mode | اختیاری. حالت رونویسی را پیکربندی میکند. مقادیر پشتیبانیشده: |
حالت
حالت رونویسی.
| انومها | |
|---|---|
MODE_UNSPECIFIED | حالت رونویسی نامشخص. |
VERBATIM | حالت رونویسی کلمه به کلمه. |
SMART | حالت رونویسی هوشمند. |
تشخیص خودکار فعالیت
تشخیص خودکار فعالیت را پیکربندی میکند.
| فیلدها | |
|---|---|
disabled | اختیاری. در صورت فعال بودن (پیشفرض)، ورودیهای صوتی و متنی شناساییشده به عنوان فعالیت شمارش میشوند. در صورت غیرفعال بودن، کلاینت باید سیگنالهای فعالیت ارسال کند. |
startOfSpeechSensitivity | اختیاری. تعیین میکند که احتمال تشخیص گفتار چقدر است. |
prefixPaddingMs | اختیاری. مدت زمان مورد نیاز برای تشخیص گفتار قبل از شروع گفتار. هرچه این مقدار کمتر باشد، تشخیص شروع گفتار حساستر است و گفتار کوتاهتر قابل تشخیص است. با این حال، این امر احتمال تشخیصهای مثبت کاذب را نیز افزایش میدهد. |
endOfSpeechSensitivity | اختیاری. تعیین میکند که احتمال پایان یافتن گفتار شناساییشده چقدر است. |
silenceDurationMs | اختیاری. مدت زمان مورد نیاز برای تشخیص عدم گفتار (مثلاً سکوت) قبل از پایان گفتار. هرچه این مقدار بزرگتر باشد، میتوان فواصل گفتاری را بدون ایجاد وقفه در فعالیت کاربر طولانیتر کرد، اما این امر تأخیر مدل را افزایش میدهد. |
BidiGenerateContentClientContent
بهروزرسانی افزایشی مکالمه فعلی ارائه شده از کلاینت. تمام محتوای اینجا بدون قید و شرط به تاریخچه مکالمه اضافه میشود و به عنوان بخشی از اعلان مدل برای تولید محتوا استفاده میشود.
یک پیام در اینجا هرگونه تولید مدل فعلی را قطع میکند.
| فیلدها | |
|---|---|
turns[] | اختیاری. محتوایی که به مکالمه فعلی با مدل اضافه شده است. برای پرسوجوهای تک نوبتی، این یک نمونه واحد است. برای پرسوجوهای چند نوبتی، این یک فیلد تکراری است که شامل سابقه مکالمه و آخرین درخواست است. |
turnComplete | اختیاری. اگر درست باشد، نشان میدهد که تولید محتوای سرور باید با اعلان انباشتهشدهی فعلی شروع شود. در غیر این صورت، سرور قبل از شروع تولید منتظر پیامهای اضافی میماند. |
BidiGenerateContentRealtimeInput
ورودی کاربر که به صورت بلادرنگ ارسال میشود.
روشهای مختلف (صوت، تصویر و متن) به صورت جریانهای همزمان مدیریت میشوند. ترتیب قرارگیری در بین این جریانها تضمین شده نیست.
این از چند جهت با BidiGenerateContentClientContent متفاوت است:
- میتواند به طور مداوم و بدون وقفه برای تولید مدل ارسال شود.
- اگر نیاز به ترکیب دادههای موجود در بین
BidiGenerateContentClientContentوBidiGenerateContentRealtimeInputباشد، سرور تلاش میکند تا بهترین پاسخ را بهینهسازی کند، اما هیچ تضمینی وجود ندارد. - پایان نوبت به صراحت مشخص نشده است، بلکه از فعالیت کاربر (مثلاً پایان سخنرانی) استنباط میشود.
- حتی قبل از پایان نوبت، دادهها به صورت تدریجی پردازش میشوند تا برای شروع سریع پاسخ از مدل بهینه شوند.
| فیلدها | |
|---|---|
mediaChunks[] | اختیاری. دادههای بایت درونخطی برای ورودی رسانه. چندین منسوخ شده: به جای آن از یکی از |
audio | اختیاری. اینها جریان ورودی صوتی بلادرنگ را تشکیل میدهند. |
video | اختیاری. اینها جریان ورودی ویدیوی بلادرنگ را تشکیل میدهند. |
activityStart | اختیاری. شروع فعالیت کاربر را نشان میدهد. این فقط در صورتی قابل ارسال است که تشخیص خودکار فعالیت (یعنی سمت سرور) غیرفعال باشد. |
activityEnd | اختیاری. پایان فعالیت کاربر را نشان میدهد. این فقط در صورتی قابل ارسال است که تشخیص خودکار فعالیت (یعنی سمت سرور) غیرفعال باشد. |
mediaResolution | اختیاری. وضوح رسانهای که باید استفاده شود. اگر مشخص نشده باشد، |
audioStreamEnd | اختیاری. نشان میدهد که جریان صوتی پایان یافته است، مثلاً به دلیل خاموش شدن میکروفون. این فقط باید زمانی ارسال شود که تشخیص خودکار فعالیت فعال باشد (که پیشفرض است). کلاینت میتواند با ارسال یک پیام صوتی، استریم را دوباره باز کند. |
text | اختیاری. اینها جریان ورودی متن بلادرنگ را تشکیل میدهند. |
BidiGenerateContentServerContent
بهروزرسانی افزایشی سرور که توسط مدل در پاسخ به پیامهای کلاینت ایجاد میشود.
محتوا در سریعترین زمان ممکن تولید میشود و نه به صورت بلادرنگ. مشتریان میتوانند آن را ذخیره کرده و به صورت بلادرنگ پخش کنند.
| فیلدها | |
|---|---|
generationComplete | فقط خروجی. اگر درست باشد، نشان میدهد که مدل تولید را تمام کرده است. وقتی مدل هنگام تولید دچار وقفه شود، هیچ پیام «generation_complete» در نوبت وقفهدار وجود نخواهد داشت، و از طریق «interrupted > turn_complete» اجرا میشود. وقتی مدل فرض میکند که پخش در زمان واقعی انجام میشود، بین generation_complete و turn_complete تأخیری وجود خواهد داشت که ناشی از انتظار مدل برای پایان پخش است. |
turnComplete | فقط خروجی. اگر درست باشد، نشان میدهد که مدل نوبت خود را تکمیل کرده است. تولید فقط در پاسخ به پیامهای اضافی کلاینت شروع میشود. توجه داشته باشید که وقتی گزارش وضعیت پخش فعال است، این فقط زمانی منتشر میشود که وضعیت پخش نشان دهد که پخش انجام شده است. وضعیت پخش آینده در همان نسل نادیده گرفته میشود. |
interrupted | فقط خروجی. اگر درست باشد، نشان میدهد که یک پیام کلاینت، تولید مدل فعلی را متوقف کرده است. اگر کلاینت در حال پخش محتوا به صورت بلادرنگ است، این سیگنال خوبی برای توقف و خالی کردن صف پخش فعلی است. |
groundingMetadata | فقط خروجی. فرادادههای زمینهای برای محتوای تولید شده. |
inputTranscription | فقط خروجی. رونویسی صوتی ورودی. رونویسی مستقل از سایر پیامهای سرور ارسال میشود و هیچ ترتیب تضمینشدهای وجود ندارد. |
interimInputTranscription | فقط خروجی. رونویسی با تأخیر کم هنگام صحبت کاربر بهروزرسانی میشود. این فیلد مرتباً بهروزرسانی میشود. |
outputTranscription | Output only. Output audio transcription. These transcriptions are part of the Generation output of the server. The last output transcription of this turn is sent before either |
urlContextMetadata | |
waitingForInput | Output only. If true, indicates that the model is not generating content because it is waiting for more input from the user, eg because it expects the user to continue talking. |
speechState | Output only. DEPRECATED: Use VoiceActivity instead. Indicates the current state of speech detection on |
interactionStatus | Output only. The current activity status of the live session. Always sent alongside |
modelTurn | Output only. The content that the model has generated as part of the current conversation with the user. |
BidiGenerateContentServerMessage
Response message for the BidiGenerateContent call.
| Fields | |
|---|---|
usageMetadata | Output only. Usage metadata about the response(s). |
Union field messageType . The type of the message. messageType can be only one of the following: | |
setupComplete | Output only. Sent in response to a |
serverContent | Output only. Content generated by the model in response to client messages. |
toolCall | Output only. Request for the client to execute the |
toolCallCancellation | Output only. Notification for the client that a previously issued |
goAway | Output only. A notice that the server will soon disconnect. |
sessionResumptionUpdate | Output only. Update of the session resumption state. |
BidiGenerateContentSetup
Message to be sent in the first (and only in the first) BidiGenerateContentClientMessage . Contains configuration that will apply for the duration of the streaming RPC.
Clients should wait for a BidiGenerateContentSetupComplete message before sending any additional messages.
| Fields | |
|---|---|
model | Required. The model's resource name. This serves as an ID for the Model to use. Format: |
generationConfig | Optional. Generation config. The following fields are not supported:
|
systemInstruction | Optional. The user provided system instructions for the model. Note: Only text should be used in parts and content in each part will be in a separate paragraph. |
tools[] | Optional. A list of A |
realtimeInputConfig | Optional. Configures the handling of realtime input. |
sessionResumption | Optional. Configures session resumption mechanism. If included, the server will send |
contextWindowCompression | Optional. Configures a context window compression mechanism. If included, the server will automatically reduce the size of the context when it exceeds the configured length. |
inputAudioTranscription | Optional. If set, enables transcription of voice input. The transcription aligns with the input audio language, if configured. |
outputAudioTranscription | Optional. If set, enables transcription of the model's audio output. The transcription aligns with the language code specified for the output audio, if configured. |
proactivity | Optional. Configures the proactivity of the model. This allows the model to respond proactively to the input and to ignore irrelevant input. |
historyConfig | Optional. Configures the exchange of history between the client and the server. |
BidiGenerateContentSetupComplete
This type has no fields.
Sent in response to a BidiGenerateContentSetup message from the client.
BidiGenerateContentToolCall
Request for the client to execute the functionCalls and return the responses with the matching id s.
| Fields | |
|---|---|
functionCalls[] | Output only. The function call to be executed. |
BidiGenerateContentToolCallCancellation
Notification for the client that a previously issued ToolCallMessage with the specified id s should not have been executed and should be cancelled. If there were side-effects to those tool calls, clients may attempt to undo the tool calls. This message occurs only in cases where the clients interrupt server turns.
| Fields | |
|---|---|
ids[] | Output only. The ids of the tool calls to be cancelled. |
BidiGenerateContentToolResponse
Client generated response to a ToolCall received from the server. Individual FunctionResponse objects are matched to the respective FunctionCall objects by the id field.
Note that in the unary and server-streaming GenerateContent APIs function calling happens by exchanging the Content parts, while in the bidi GenerateContent APIs function calling happens over these dedicated set of messages.
| Fields | |
|---|---|
functionResponses[] | Optional. The response to the function calls. |
BidiGenerateContentTranscription
Transcription of audio (input or output).
| Fields | |
|---|---|
text | Transcription text. |
languageCode | The BCP-47 language code of the transcription. |
ContextWindowCompressionConfig
Enables context window compression — a mechanism for managing the model's context window so that it does not exceed a given length.
| Fields | |
|---|---|
Union field compressionMechanism . The context window compression mechanism used. compressionMechanism can be only one of the following: | |
slidingWindow | A sliding-window mechanism. |
triggerTokens | The number of tokens (before running a turn) required to trigger a context window compression. This can be used to balance quality against latency as shorter context windows may result in faster model responses. However, any compression operation will cause a temporary latency increase, so they should not be triggered frequently. If not set, the default is 80% of the model's context window limit. This leaves 20% for the next user request/model response. |
EndSensitivity
Determines how end of speech is detected.
| Enums | |
|---|---|
END_SENSITIVITY_UNSPECIFIED | The default is END_SENSITIVITY_HIGH. |
END_SENSITIVITY_HIGH | Automatic detection ends speech more often. |
END_SENSITIVITY_LOW | Automatic detection ends speech less often. |
GoAway
A notice that the server will soon disconnect.
| Fields | |
|---|---|
timeLeft | The remaining time before the connection will be terminated as ABORTED. This duration will never be less than a model-specific minimum, which will be specified together with the rate limits for the model. |
HistoryConfig
History configuration.
This message is included in the session configuration as BidiGenerateContentSetup.historyConfig . Configures the exchange of history messages.
| Fields | |
|---|---|
initialHistoryInClientContent | Optional. If true, after sending |
ProactivityConfig
Config for proactivity features.
| Fields | |
|---|---|
proactiveAudio | Optional. If enabled, the model can reject responding to the last prompt. For example, this allows the model to ignore out of context speech or to stay silent if the user did not make a request, yet. |
RealtimeInputConfig
Configures the realtime input behavior in BidiGenerateContent .
| Fields | |
|---|---|
automaticActivityDetection | Optional. If not set, automatic activity detection is enabled by default. If automatic voice detection is disabled, the client must send activity signals. |
activityHandling | Optional. Defines what effect activity has. |
turnCoverage | Optional. Defines which input is included in the user's turn. |
SessionResumptionConfig
Session resumption configuration.
This message is included in the session configuration as BidiGenerateContentSetup.sessionResumption . If configured, the server will send SessionResumptionUpdate messages.
| Fields | |
|---|---|
handle | The handle of a previous session. If not present then a new session is created. Session handles come from |
SessionResumptionUpdate
Update of the session resumption state.
Only sent if BidiGenerateContentSetup.sessionResumption was set.
| Fields | |
|---|---|
newHandle | New handle that represents a state that can be resumed. Empty if |
resumable | True if the current session can be resumed at this point. Resumption is not possible at some points in the session. For example, when the model is executing function calls or generating. Resuming the session (using a previous session token) in such a state will result in some data loss. In these cases, |
SlidingWindow
The SlidingWindow method operates by discarding content at the beginning of the context window. The resulting context will always begin at the start of a USER role turn. System instructions and any BidiGenerateContentSetup.prefixTurns will always remain at the beginning of the result.
| Fields | |
|---|---|
targetTokens | The target number of tokens to keep. The default value is trigger_tokens/2. Discarding parts of the context window causes a temporary latency increase so this value should be calibrated to avoid frequent compression operations. |
StartSensitivity
Determines how start of speech is detected.
| Enums | |
|---|---|
START_SENSITIVITY_UNSPECIFIED | The default is START_SENSITIVITY_HIGH. |
START_SENSITIVITY_HIGH | Automatic detection will detect the start of speech more often. |
START_SENSITIVITY_LOW | Automatic detection will detect the start of speech less often. |
TurnCoverage
Options about which input is included in the user's turn.
| Enums | |
|---|---|
TURN_COVERAGE_UNSPECIFIED | If unspecified, a default behavior is selected based on the model. Eg, for Gemini 2.5, the default is TURN_INCLUDES_ONLY_ACTIVITY , while for Gemini 3.1 and onwards, it's TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO . |
TURN_INCLUDES_ONLY_ACTIVITY | Includes activity since the last turn, excluding inactivity (eg silence on the audio stream). |
TURN_INCLUDES_ALL_INPUT | Includes all realtime input since the last turn, including inactivity (eg silence on the audio stream). |
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO | Includes audio activity and all video since the last turn. With automatic activity detection, audio activity means speech and excludes silence. |
TranslationConfig
Config for translation features.
| Fields | |
|---|---|
targetLanguageCode | Required. The target language for translation. Supported values are BCP-47 language codes (eg "en", "es", "fr"). |
echoTargetLanguage | Optional. If true, the model will generate audio when the target language is spoken, essentially it will parrot the input. If false, we will not produce audio for the target language. |
UrlContextMetadata
Metadata related to url context retrieval tool.
| Fields | |
|---|---|
urlMetadata[] | List of url context. |
UsageMetadata
Usage metadata about response(s).
| Fields | |
|---|---|
promptTokenCount | Output only. Number of tokens in the prompt. When |
cachedContentTokenCount | Number of tokens in the cached part of the prompt (the cached content) |
responseTokenCount | Output only. Total number of tokens across all the generated response candidates. |
toolUsePromptTokenCount | Output only. Number of tokens present in tool-use prompt(s). |
thoughtsTokenCount | Output only. Number of tokens of thoughts for thinking models. |
totalTokenCount | Output only. Total token count for the generation request (prompt + response candidates). |
promptTokensDetails[] | Output only. List of modalities that were processed in the request input. |
cacheTokensDetails[] | Output only. List of modalities of the cached content in the request input. |
responseTokensDetails[] | Output only. List of modalities that were returned in the response. |
toolUsePromptTokensDetails[] | Output only. List of modalities that were processed for tool-use request inputs. |
Ephemeral authentication tokens
Ephemeral authentication tokens can be obtained by calling AuthTokenService.CreateToken and then used with GenerativeService.BidiGenerateContentConstrained , either by passing the token in an access_token query parameter, or in an HTTP Authorization header with " Token " prefixed to it.
CreateAuthTokenRequest
Create an ephemeral authentication token.
| Fields | |
|---|---|
authToken | Required. The token to create. |
AuthToken
A request to create an ephemeral authentication token.
| Fields | |
|---|---|
name | Output only. Identifier. The token itself. |
expireTime | Optional. Input only. Immutable. An optional time after which, when using the resulting token, messages in BidiGenerateContent sessions will be rejected. (Gemini may preemptively close the session after this time.) If not set then this defaults to 30 minutes in the future. If set, this value must be less than 20 hours in the future. |
newSessionExpireTime | Optional. Input only. Immutable. The time after which new Live API sessions using the token resulting from this request will be rejected. If not set this defaults to 60 seconds in the future. If set, this value must be less than 20 hours in the future. |
fieldMask | Optional. Input only. Immutable. If field_mask is empty, and If field_mask is empty, and If field_mask is not empty, then the corresponding fields from |
Union field config . The method-specific configuration for the resulting token. config can be only one of the following: | |
bidiGenerateContentSetup | Optional. Input only. Immutable. Configuration specific to |
uses | Optional. Input only. Immutable. The number of times the token can be used. If this value is zero then no limit is applied. Resuming a Live API session does not count as a use. If unspecified, the default is 1. |
More information on common types
For more information on the commonly-used API resource types Blob , Content , FunctionCall , FunctionResponse , GenerationConfig , GroundingMetadata , ModalityTokenCount , and Tool , see Generating content .
The Live API is a stateful API that uses WebSockets . In this section, you'll find additional details regarding the WebSockets API.
جلسات
A WebSocket connection establishes a session between the client and the Gemini server. After a client initiates a new connection the session can exchange messages with the server to:
- Send text, audio, or video to the Gemini server.
- Receive audio, text, or function call requests from the Gemini server.
WebSocket connection
To start a session, connect to this websocket endpoint:
wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent
Session configuration
The initial message sent after establishing the WebSocket connection sets the session configuration, which includes the model, generation parameters, system instructions, and tools.
You cannot update the configuration while the connection is open. However, you can change the configuration parameters, except the model, when pausing and resuming via the session resumption mechanism .
See the following example configuration. Note that the name casing in SDKs may vary. You can look up the Python SDK configuration options here .
{
"model": string,
"generationConfig": {
"candidateCount": integer,
"maxOutputTokens": integer,
"temperature": number,
"topP": number,
"topK": integer,
"presencePenalty": number,
"frequencyPenalty": number,
"responseModalities": [string],
"speechConfig": object,
"mediaResolution": object,
"translationConfig": object
},
"systemInstruction": string,
"tools": [object]
}
For more information on the API field, see generationConfig .
Send messages
To exchange messages over the WebSocket connection, the client must send a JSON object over an open WebSocket connection. The JSON object must have exactly one of the fields from the following object set:
{
"setup": BidiGenerateContentSetup,
"clientContent": BidiGenerateContentClientContent,
"realtimeInput": BidiGenerateContentRealtimeInput,
"toolResponse": BidiGenerateContentToolResponse
}
Supported client messages
See the supported client messages in the following table:
| پیام | توضیحات |
|---|---|
BidiGenerateContentSetup | Session configuration to be sent in the first message |
BidiGenerateContentClientContent | Incremental content update of the current conversation delivered from the client |
BidiGenerateContentRealtimeInput | Real time audio, video, or text input |
BidiGenerateContentToolResponse | Response to a ToolCallMessage received from the server |
Receive messages
To receive messages from Gemini, listen for the WebSocket 'message' event, and then parse the result according to the definition of the supported server messages.
See the following:
async with client.aio.live.connect(model='...', config=config) as session:
await session.send(input='Hello world!', end_of_turn=True)
async for message in session.receive():
print(message)
Server messages may have a usageMetadata field but will otherwise include exactly one of the other fields from the BidiGenerateContentServerMessage message. (The messageType union is not expressed in JSON so the field will appear at the top-level of the message.)
Messages and events
ActivityEnd
This type has no fields.
Marks the end of user activity.
ActivityHandling
The different ways of handling user activity.
| Enums | |
|---|---|
ACTIVITY_HANDLING_UNSPECIFIED | If unspecified, the default behavior is START_OF_ACTIVITY_INTERRUPTS . |
START_OF_ACTIVITY_INTERRUPTS | If true, start of activity will interrupt the model's response (also called "barge in"). The model's current response will be cut-off in the moment of the interruption. This is the default behavior. |
NO_INTERRUPTION | The model's response will not be interrupted. |
ActivityStart
This type has no fields.
Marks the start of user activity.
AudioTranscriptionConfig
The audio transcription configuration.
| Fields | |
|---|---|
languageCodes[] | Optional. BCP-47 language codes providing hints about the languages present in the audio. If omitted or empty, defaults to automatic language detection. |
customVocabulary[] | Optional. A list of custom vocabulary phrases to bias the speech recognition model toward recognizing specific terms (product names, proper nouns, jargon). |
wordTimestamp | Optional. Configures word-level timestamp generation. |
diarization | Optional. Configures speaker diarization. |
mode | Optional. Configures transcription mode. Supported values: |
حالت
Transcription mode.
| Enums | |
|---|---|
MODE_UNSPECIFIED | Unspecified transcription mode. |
VERBATIM | Verbatim transcription mode. |
SMART | Smart transcription mode. |
AutomaticActivityDetection
Configures automatic detection of activity.
| Fields | |
|---|---|
disabled | Optional. If enabled (the default), detected voice and text input count as activity. If disabled, the client must send activity signals. |
startOfSpeechSensitivity | Optional. Determines how likely speech is to be detected. |
prefixPaddingMs | Optional. The required duration of detected speech before start-of-speech is committed. The lower this value, the more sensitive the start-of-speech detection is and shorter speech can be recognized. However, this also increases the probability of false positives. |
endOfSpeechSensitivity | Optional. Determines how likely detected speech is ended. |
silenceDurationMs | Optional. The required duration of detected non-speech (eg silence) before end-of-speech is committed. The larger this value, the longer speech gaps can be without interrupting the user's activity but this will increase the model's latency. |
BidiGenerateContentClientContent
Incremental update of the current conversation delivered from the client. All of the content here is unconditionally appended to the conversation history and used as part of the prompt to the model to generate content.
A message here will interrupt any current model generation.
| Fields | |
|---|---|
turns[] | Optional. The content appended to the current conversation with the model. For single-turn queries, this is a single instance. For multi-turn queries, this is a repeated field that contains conversation history and the latest request. |
turnComplete | Optional. If true, indicates that the server content generation should start with the currently accumulated prompt. Otherwise, the server awaits additional messages before starting generation. |
BidiGenerateContentRealtimeInput
User input that is sent in real time.
The different modalities (audio, video and text) are handled as concurrent streams. The ordering across these streams is not guaranteed.
This is different from BidiGenerateContentClientContent in a few ways:
- Can be sent continuously without interruption to model generation.
- If there is a need to mix data interleaved across the
BidiGenerateContentClientContentand theBidiGenerateContentRealtimeInput, the server attempts to optimize for best response, but there are no guarantees. - End of turn is not explicitly specified, but is rather derived from user activity (for example, end of speech).
- Even before the end of turn, the data is processed incrementally to optimize for a fast start of the response from the model.
| Fields | |
|---|---|
mediaChunks[] | Optional. Inlined bytes data for media input. Multiple DEPRECATED: Use one of |
audio | Optional. These form the realtime audio input stream. |
video | Optional. These form the realtime video input stream. |
activityStart | Optional. Marks the start of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled. |
activityEnd | Optional. Marks the end of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled. |
mediaResolution | Optional. The media resolution to use. If not specified, |
audioStreamEnd | Optional. Indicates that the audio stream has ended, eg because the microphone was turned off. This should only be sent when automatic activity detection is enabled (which is the default). The client can reopen the stream by sending an audio message. |
text | Optional. These form the realtime text input stream. |
BidiGenerateContentServerContent
Incremental server update generated by the model in response to client messages.
Content is generated as quickly as possible, and not in real time. Clients may choose to buffer and play it out in real time.
| Fields | |
|---|---|
generationComplete | Output only. If true, indicates that the model is done generating. When model is interrupted while generating there will be no 'generation_complete' message in interrupted turn, it will go through 'interrupted > turn_complete'. When model assumes realtime playback there will be delay between generation_complete and turn_complete that is caused by model waiting for playback to finish. |
turnComplete | Output only. If true, indicates that the model has completed its turn. Generation will only start in response to additional client messages. Note when playback status reporting is enabled, this is emitted only when the playback status indicates that the playback is done. Future playback status of the same generation will be ignored. |
interrupted | Output only. If true, indicates that a client message has interrupted current model generation. If the client is playing out the content in real time, this is a good signal to stop and empty the current playback queue. |
groundingMetadata | Output only. Grounding metadata for the generated content. |
inputTranscription | Output only. Input audio transcription. The transcription is sent independently of the other server messages and there is no guaranteed ordering. |
interimInputTranscription | Output only. Low latency transcription updated while the user is speaking. This field is subject to frequent updates. |
outputTranscription | Output only. Output audio transcription. These transcriptions are part of the Generation output of the server. The last output transcription of this turn is sent before either |
urlContextMetadata | |
waitingForInput | Output only. If true, indicates that the model is not generating content because it is waiting for more input from the user, eg because it expects the user to continue talking. |
speechState | Output only. DEPRECATED: Use VoiceActivity instead. Indicates the current state of speech detection on |
interactionStatus | Output only. The current activity status of the live session. Always sent alongside |
modelTurn | Output only. The content that the model has generated as part of the current conversation with the user. |
BidiGenerateContentServerMessage
Response message for the BidiGenerateContent call.
| Fields | |
|---|---|
usageMetadata | Output only. Usage metadata about the response(s). |
Union field messageType . The type of the message. messageType can be only one of the following: | |
setupComplete | Output only. Sent in response to a |
serverContent | Output only. Content generated by the model in response to client messages. |
toolCall | Output only. Request for the client to execute the |
toolCallCancellation | Output only. Notification for the client that a previously issued |
goAway | Output only. A notice that the server will soon disconnect. |
sessionResumptionUpdate | Output only. Update of the session resumption state. |
BidiGenerateContentSetup
Message to be sent in the first (and only in the first) BidiGenerateContentClientMessage . Contains configuration that will apply for the duration of the streaming RPC.
Clients should wait for a BidiGenerateContentSetupComplete message before sending any additional messages.
| Fields | |
|---|---|
model | Required. The model's resource name. This serves as an ID for the Model to use. Format: |
generationConfig | Optional. Generation config. The following fields are not supported:
|
systemInstruction | Optional. The user provided system instructions for the model. Note: Only text should be used in parts and content in each part will be in a separate paragraph. |
tools[] | Optional. A list of A |
realtimeInputConfig | Optional. Configures the handling of realtime input. |
sessionResumption | Optional. Configures session resumption mechanism. If included, the server will send |
contextWindowCompression | Optional. Configures a context window compression mechanism. If included, the server will automatically reduce the size of the context when it exceeds the configured length. |
inputAudioTranscription | Optional. If set, enables transcription of voice input. The transcription aligns with the input audio language, if configured. |
outputAudioTranscription | Optional. If set, enables transcription of the model's audio output. The transcription aligns with the language code specified for the output audio, if configured. |
proactivity | Optional. Configures the proactivity of the model. This allows the model to respond proactively to the input and to ignore irrelevant input. |
historyConfig | Optional. Configures the exchange of history between the client and the server. |
BidiGenerateContentSetupComplete
This type has no fields.
Sent in response to a BidiGenerateContentSetup message from the client.
BidiGenerateContentToolCall
Request for the client to execute the functionCalls and return the responses with the matching id s.
| Fields | |
|---|---|
functionCalls[] | Output only. The function call to be executed. |
BidiGenerateContentToolCallCancellation
Notification for the client that a previously issued ToolCallMessage with the specified id s should not have been executed and should be cancelled. If there were side-effects to those tool calls, clients may attempt to undo the tool calls. This message occurs only in cases where the clients interrupt server turns.
| Fields | |
|---|---|
ids[] | Output only. The ids of the tool calls to be cancelled. |
BidiGenerateContentToolResponse
Client generated response to a ToolCall received from the server. Individual FunctionResponse objects are matched to the respective FunctionCall objects by the id field.
Note that in the unary and server-streaming GenerateContent APIs function calling happens by exchanging the Content parts, while in the bidi GenerateContent APIs function calling happens over these dedicated set of messages.
| Fields | |
|---|---|
functionResponses[] | Optional. The response to the function calls. |
BidiGenerateContentTranscription
Transcription of audio (input or output).
| Fields | |
|---|---|
text | Transcription text. |
languageCode | The BCP-47 language code of the transcription. |
ContextWindowCompressionConfig
Enables context window compression — a mechanism for managing the model's context window so that it does not exceed a given length.
| Fields | |
|---|---|
Union field compressionMechanism . The context window compression mechanism used. compressionMechanism can be only one of the following: | |
slidingWindow | A sliding-window mechanism. |
triggerTokens | The number of tokens (before running a turn) required to trigger a context window compression. This can be used to balance quality against latency as shorter context windows may result in faster model responses. However, any compression operation will cause a temporary latency increase, so they should not be triggered frequently. If not set, the default is 80% of the model's context window limit. This leaves 20% for the next user request/model response. |
EndSensitivity
Determines how end of speech is detected.
| Enums | |
|---|---|
END_SENSITIVITY_UNSPECIFIED | The default is END_SENSITIVITY_HIGH. |
END_SENSITIVITY_HIGH | Automatic detection ends speech more often. |
END_SENSITIVITY_LOW | Automatic detection ends speech less often. |
GoAway
A notice that the server will soon disconnect.
| Fields | |
|---|---|
timeLeft | The remaining time before the connection will be terminated as ABORTED. This duration will never be less than a model-specific minimum, which will be specified together with the rate limits for the model. |
HistoryConfig
History configuration.
This message is included in the session configuration as BidiGenerateContentSetup.historyConfig . Configures the exchange of history messages.
| Fields | |
|---|---|
initialHistoryInClientContent | Optional. If true, after sending |
ProactivityConfig
Config for proactivity features.
| Fields | |
|---|---|
proactiveAudio | Optional. If enabled, the model can reject responding to the last prompt. For example, this allows the model to ignore out of context speech or to stay silent if the user did not make a request, yet. |
RealtimeInputConfig
Configures the realtime input behavior in BidiGenerateContent .
| Fields | |
|---|---|
automaticActivityDetection | Optional. If not set, automatic activity detection is enabled by default. If automatic voice detection is disabled, the client must send activity signals. |
activityHandling | Optional. Defines what effect activity has. |
turnCoverage | Optional. Defines which input is included in the user's turn. |
SessionResumptionConfig
Session resumption configuration.
This message is included in the session configuration as BidiGenerateContentSetup.sessionResumption . If configured, the server will send SessionResumptionUpdate messages.
| Fields | |
|---|---|
handle | The handle of a previous session. If not present then a new session is created. Session handles come from |
SessionResumptionUpdate
Update of the session resumption state.
Only sent if BidiGenerateContentSetup.sessionResumption was set.
| Fields | |
|---|---|
newHandle | New handle that represents a state that can be resumed. Empty if |
resumable | True if the current session can be resumed at this point. Resumption is not possible at some points in the session. For example, when the model is executing function calls or generating. Resuming the session (using a previous session token) in such a state will result in some data loss. In these cases, |
SlidingWindow
The SlidingWindow method operates by discarding content at the beginning of the context window. The resulting context will always begin at the start of a USER role turn. System instructions and any BidiGenerateContentSetup.prefixTurns will always remain at the beginning of the result.
| Fields | |
|---|---|
targetTokens | The target number of tokens to keep. The default value is trigger_tokens/2. Discarding parts of the context window causes a temporary latency increase so this value should be calibrated to avoid frequent compression operations. |
StartSensitivity
Determines how start of speech is detected.
| Enums | |
|---|---|
START_SENSITIVITY_UNSPECIFIED | The default is START_SENSITIVITY_HIGH. |
START_SENSITIVITY_HIGH | Automatic detection will detect the start of speech more often. |
START_SENSITIVITY_LOW | Automatic detection will detect the start of speech less often. |
TurnCoverage
Options about which input is included in the user's turn.
| Enums | |
|---|---|
TURN_COVERAGE_UNSPECIFIED | If unspecified, a default behavior is selected based on the model. Eg, for Gemini 2.5, the default is TURN_INCLUDES_ONLY_ACTIVITY , while for Gemini 3.1 and onwards, it's TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO . |
TURN_INCLUDES_ONLY_ACTIVITY | Includes activity since the last turn, excluding inactivity (eg silence on the audio stream). |
TURN_INCLUDES_ALL_INPUT | Includes all realtime input since the last turn, including inactivity (eg silence on the audio stream). |
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO | Includes audio activity and all video since the last turn. With automatic activity detection, audio activity means speech and excludes silence. |
TranslationConfig
Config for translation features.
| Fields | |
|---|---|
targetLanguageCode | Required. The target language for translation. Supported values are BCP-47 language codes (eg "en", "es", "fr"). |
echoTargetLanguage | Optional. If true, the model will generate audio when the target language is spoken, essentially it will parrot the input. If false, we will not produce audio for the target language. |
UrlContextMetadata
Metadata related to url context retrieval tool.
| Fields | |
|---|---|
urlMetadata[] | List of url context. |
UsageMetadata
Usage metadata about response(s).
| Fields | |
|---|---|
promptTokenCount | Output only. Number of tokens in the prompt. When |
cachedContentTokenCount | Number of tokens in the cached part of the prompt (the cached content) |
responseTokenCount | Output only. Total number of tokens across all the generated response candidates. |
toolUsePromptTokenCount | Output only. Number of tokens present in tool-use prompt(s). |
thoughtsTokenCount | Output only. Number of tokens of thoughts for thinking models. |
totalTokenCount | Output only. Total token count for the generation request (prompt + response candidates). |
promptTokensDetails[] | Output only. List of modalities that were processed in the request input. |
cacheTokensDetails[] | Output only. List of modalities of the cached content in the request input. |
responseTokensDetails[] | Output only. List of modalities that were returned in the response. |
toolUsePromptTokensDetails[] | Output only. List of modalities that were processed for tool-use request inputs. |
Ephemeral authentication tokens
Ephemeral authentication tokens can be obtained by calling AuthTokenService.CreateToken and then used with GenerativeService.BidiGenerateContentConstrained , either by passing the token in an access_token query parameter, or in an HTTP Authorization header with " Token " prefixed to it.
CreateAuthTokenRequest
Create an ephemeral authentication token.
| Fields | |
|---|---|
authToken | Required. The token to create. |
AuthToken
A request to create an ephemeral authentication token.
| Fields | |
|---|---|
name | Output only. Identifier. The token itself. |
expireTime | Optional. Input only. Immutable. An optional time after which, when using the resulting token, messages in BidiGenerateContent sessions will be rejected. (Gemini may preemptively close the session after this time.) If not set then this defaults to 30 minutes in the future. If set, this value must be less than 20 hours in the future. |
newSessionExpireTime | Optional. Input only. Immutable. The time after which new Live API sessions using the token resulting from this request will be rejected. If not set this defaults to 60 seconds in the future. If set, this value must be less than 20 hours in the future. |
fieldMask | Optional. Input only. Immutable. If field_mask is empty, and If field_mask is empty, and If field_mask is not empty, then the corresponding fields from |
Union field config . The method-specific configuration for the resulting token. config can be only one of the following: | |
bidiGenerateContentSetup | Optional. Input only. Immutable. Configuration specific to |
uses | Optional. Input only. Immutable. The number of times the token can be used. If this value is zero then no limit is applied. Resuming a Live API session does not count as a use. If unspecified, the default is 1. |
More information on common types
For more information on the commonly-used API resource types Blob , Content , FunctionCall , FunctionResponse , GenerationConfig , GroundingMetadata , ModalityTokenCount , and Tool , see Generating content .
The Live API is a stateful API that uses WebSockets . In this section, you'll find additional details regarding the WebSockets API.
جلسات
A WebSocket connection establishes a session between the client and the Gemini server. After a client initiates a new connection the session can exchange messages with the server to:
- Send text, audio, or video to the Gemini server.
- Receive audio, text, or function call requests from the Gemini server.
WebSocket connection
To start a session, connect to this websocket endpoint:
wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1beta.GenerativeService.BidiGenerateContent
Session configuration
The initial message sent after establishing the WebSocket connection sets the session configuration, which includes the model, generation parameters, system instructions, and tools.
You cannot update the configuration while the connection is open. However, you can change the configuration parameters, except the model, when pausing and resuming via the session resumption mechanism .
See the following example configuration. Note that the name casing in SDKs may vary. You can look up the Python SDK configuration options here .
{
"model": string,
"generationConfig": {
"candidateCount": integer,
"maxOutputTokens": integer,
"temperature": number,
"topP": number,
"topK": integer,
"presencePenalty": number,
"frequencyPenalty": number,
"responseModalities": [string],
"speechConfig": object,
"mediaResolution": object,
"translationConfig": object
},
"systemInstruction": string,
"tools": [object]
}
For more information on the API field, see generationConfig .
Send messages
To exchange messages over the WebSocket connection, the client must send a JSON object over an open WebSocket connection. The JSON object must have exactly one of the fields from the following object set:
{
"setup": BidiGenerateContentSetup,
"clientContent": BidiGenerateContentClientContent,
"realtimeInput": BidiGenerateContentRealtimeInput,
"toolResponse": BidiGenerateContentToolResponse
}
Supported client messages
See the supported client messages in the following table:
| پیام | توضیحات |
|---|---|
BidiGenerateContentSetup | Session configuration to be sent in the first message |
BidiGenerateContentClientContent | Incremental content update of the current conversation delivered from the client |
BidiGenerateContentRealtimeInput | Real time audio, video, or text input |
BidiGenerateContentToolResponse | Response to a ToolCallMessage received from the server |
Receive messages
To receive messages from Gemini, listen for the WebSocket 'message' event, and then parse the result according to the definition of the supported server messages.
See the following:
async with client.aio.live.connect(model='...', config=config) as session:
await session.send(input='Hello world!', end_of_turn=True)
async for message in session.receive():
print(message)
Server messages may have a usageMetadata field but will otherwise include exactly one of the other fields from the BidiGenerateContentServerMessage message. (The messageType union is not expressed in JSON so the field will appear at the top-level of the message.)
Messages and events
ActivityEnd
This type has no fields.
Marks the end of user activity.
ActivityHandling
The different ways of handling user activity.
| Enums | |
|---|---|
ACTIVITY_HANDLING_UNSPECIFIED | If unspecified, the default behavior is START_OF_ACTIVITY_INTERRUPTS . |
START_OF_ACTIVITY_INTERRUPTS | If true, start of activity will interrupt the model's response (also called "barge in"). The model's current response will be cut-off in the moment of the interruption. This is the default behavior. |
NO_INTERRUPTION | The model's response will not be interrupted. |
ActivityStart
This type has no fields.
Marks the start of user activity.
AudioTranscriptionConfig
The audio transcription configuration.
| Fields | |
|---|---|
languageCodes[] | Optional. BCP-47 language codes providing hints about the languages present in the audio. If omitted or empty, defaults to automatic language detection. |
customVocabulary[] | Optional. A list of custom vocabulary phrases to bias the speech recognition model toward recognizing specific terms (product names, proper nouns, jargon). |
wordTimestamp | Optional. Configures word-level timestamp generation. |
diarization | Optional. Configures speaker diarization. |
mode | Optional. Configures transcription mode. Supported values: |
حالت
Transcription mode.
| Enums | |
|---|---|
MODE_UNSPECIFIED | Unspecified transcription mode. |
VERBATIM | Verbatim transcription mode. |
SMART | Smart transcription mode. |
AutomaticActivityDetection
Configures automatic detection of activity.
| Fields | |
|---|---|
disabled | Optional. If enabled (the default), detected voice and text input count as activity. If disabled, the client must send activity signals. |
startOfSpeechSensitivity | Optional. Determines how likely speech is to be detected. |
prefixPaddingMs | Optional. The required duration of detected speech before start-of-speech is committed. The lower this value, the more sensitive the start-of-speech detection is and shorter speech can be recognized. However, this also increases the probability of false positives. |
endOfSpeechSensitivity | Optional. Determines how likely detected speech is ended. |
silenceDurationMs | Optional. The required duration of detected non-speech (eg silence) before end-of-speech is committed. The larger this value, the longer speech gaps can be without interrupting the user's activity but this will increase the model's latency. |
BidiGenerateContentClientContent
Incremental update of the current conversation delivered from the client. All of the content here is unconditionally appended to the conversation history and used as part of the prompt to the model to generate content.
A message here will interrupt any current model generation.
| Fields | |
|---|---|
turns[] | Optional. The content appended to the current conversation with the model. For single-turn queries, this is a single instance. For multi-turn queries, this is a repeated field that contains conversation history and the latest request. |
turnComplete | Optional. If true, indicates that the server content generation should start with the currently accumulated prompt. Otherwise, the server awaits additional messages before starting generation. |
BidiGenerateContentRealtimeInput
User input that is sent in real time.
The different modalities (audio, video and text) are handled as concurrent streams. The ordering across these streams is not guaranteed.
This is different from BidiGenerateContentClientContent in a few ways:
- Can be sent continuously without interruption to model generation.
- If there is a need to mix data interleaved across the
BidiGenerateContentClientContentand theBidiGenerateContentRealtimeInput, the server attempts to optimize for best response, but there are no guarantees. - End of turn is not explicitly specified, but is rather derived from user activity (for example, end of speech).
- Even before the end of turn, the data is processed incrementally to optimize for a fast start of the response from the model.
| Fields | |
|---|---|
mediaChunks[] | Optional. Inlined bytes data for media input. Multiple DEPRECATED: Use one of |
audio | Optional. These form the realtime audio input stream. |
video | Optional. These form the realtime video input stream. |
activityStart | Optional. Marks the start of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled. |
activityEnd | Optional. Marks the end of user activity. This can only be sent if automatic (ie server-side) activity detection is disabled. |
mediaResolution | Optional. The media resolution to use. If not specified, |
audioStreamEnd | Optional. Indicates that the audio stream has ended, eg because the microphone was turned off. This should only be sent when automatic activity detection is enabled (which is the default). The client can reopen the stream by sending an audio message. |
text | Optional. These form the realtime text input stream. |
BidiGenerateContentServerContent
Incremental server update generated by the model in response to client messages.
Content is generated as quickly as possible, and not in real time. Clients may choose to buffer and play it out in real time.
| Fields | |
|---|---|
generationComplete | Output only. If true, indicates that the model is done generating. When model is interrupted while generating there will be no 'generation_complete' message in interrupted turn, it will go through 'interrupted > turn_complete'. When model assumes realtime playback there will be delay between generation_complete and turn_complete that is caused by model waiting for playback to finish. |
turnComplete | Output only. If true, indicates that the model has completed its turn. Generation will only start in response to additional client messages. Note when playback status reporting is enabled, this is emitted only when the playback status indicates that the playback is done. Future playback status of the same generation will be ignored. |
interrupted | Output only. If true, indicates that a client message has interrupted current model generation. If the client is playing out the content in real time, this is a good signal to stop and empty the current playback queue. |
groundingMetadata | Output only. Grounding metadata for the generated content. |
inputTranscription | Output only. Input audio transcription. The transcription is sent independently of the other server messages and there is no guaranteed ordering. |
interimInputTranscription | Output only. Low latency transcription updated while the user is speaking. This field is subject to frequent updates. |
outputTranscription | Output only. Output audio transcription. These transcriptions are part of the Generation output of the server. The last output transcription of this turn is sent before either |
urlContextMetadata | |
waitingForInput | Output only. If true, indicates that the model is not generating content because it is waiting for more input from the user, eg because it expects the user to continue talking. |
speechState | Output only. DEPRECATED: Use VoiceActivity instead. Indicates the current state of speech detection on |
interactionStatus | Output only. The current activity status of the live session. Always sent alongside |
modelTurn | Output only. The content that the model has generated as part of the current conversation with the user. |
BidiGenerateContentServerMessage
Response message for the BidiGenerateContent call.
| Fields | |
|---|---|
usageMetadata | Output only. Usage metadata about the response(s). |
Union field messageType . The type of the message. messageType can be only one of the following: | |
setupComplete | Output only. Sent in response to a |
serverContent | Output only. Content generated by the model in response to client messages. |
toolCall | Output only. Request for the client to execute the |
toolCallCancellation | Output only. Notification for the client that a previously issued |
goAway | Output only. A notice that the server will soon disconnect. |
sessionResumptionUpdate | Output only. Update of the session resumption state. |
BidiGenerateContentSetup
Message to be sent in the first (and only in the first) BidiGenerateContentClientMessage . Contains configuration that will apply for the duration of the streaming RPC.
Clients should wait for a BidiGenerateContentSetupComplete message before sending any additional messages.
| Fields | |
|---|---|
model | Required. The model's resource name. This serves as an ID for the Model to use. Format: |
generationConfig | Optional. Generation config. The following fields are not supported:
|
systemInstruction | Optional. The user provided system instructions for the model. Note: Only text should be used in parts and content in each part will be in a separate paragraph. |
tools[] | Optional. A list of A |
realtimeInputConfig | Optional. Configures the handling of realtime input. |
sessionResumption | Optional. Configures session resumption mechanism. If included, the server will send |
contextWindowCompression | Optional. Configures a context window compression mechanism. If included, the server will automatically reduce the size of the context when it exceeds the configured length. |
inputAudioTranscription | Optional. If set, enables transcription of voice input. The transcription aligns with the input audio language, if configured. |
outputAudioTranscription | Optional. If set, enables transcription of the model's audio output. The transcription aligns with the language code specified for the output audio, if configured. |
proactivity | Optional. Configures the proactivity of the model. This allows the model to respond proactively to the input and to ignore irrelevant input. |
historyConfig | Optional. Configures the exchange of history between the client and the server. |
BidiGenerateContentSetupComplete
This type has no fields.
Sent in response to a BidiGenerateContentSetup message from the client.
BidiGenerateContentToolCall
Request for the client to execute the functionCalls and return the responses with the matching id s.
| Fields | |
|---|---|
functionCalls[] | Output only. The function call to be executed. |
BidiGenerateContentToolCallCancellation
Notification for the client that a previously issued ToolCallMessage with the specified id s should not have been executed and should be cancelled. If there were side-effects to those tool calls, clients may attempt to undo the tool calls. This message occurs only in cases where the clients interrupt server turns.
| Fields | |
|---|---|
ids[] | Output only. The ids of the tool calls to be cancelled. |
BidiGenerateContentToolResponse
Client generated response to a ToolCall received from the server. Individual FunctionResponse objects are matched to the respective FunctionCall objects by the id field.
Note that in the unary and server-streaming GenerateContent APIs function calling happens by exchanging the Content parts, while in the bidi GenerateContent APIs function calling happens over these dedicated set of messages.
| Fields | |
|---|---|
functionResponses[] | Optional. The response to the function calls. |
BidiGenerateContentTranscription
Transcription of audio (input or output).
| Fields | |
|---|---|
text | Transcription text. |
languageCode | The BCP-47 language code of the transcription. |
ContextWindowCompressionConfig
Enables context window compression — a mechanism for managing the model's context window so that it does not exceed a given length.
| Fields | |
|---|---|
Union field compressionMechanism . The context window compression mechanism used. compressionMechanism can be only one of the following: | |
slidingWindow | A sliding-window mechanism. |
triggerTokens | The number of tokens (before running a turn) required to trigger a context window compression. This can be used to balance quality against latency as shorter context windows may result in faster model responses. However, any compression operation will cause a temporary latency increase, so they should not be triggered frequently. If not set, the default is 80% of the model's context window limit. This leaves 20% for the next user request/model response. |
EndSensitivity
Determines how end of speech is detected.
| Enums | |
|---|---|
END_SENSITIVITY_UNSPECIFIED | The default is END_SENSITIVITY_HIGH. |
END_SENSITIVITY_HIGH | Automatic detection ends speech more often. |
END_SENSITIVITY_LOW | Automatic detection ends speech less often. |
GoAway
A notice that the server will soon disconnect.
| Fields | |
|---|---|
timeLeft | The remaining time before the connection will be terminated as ABORTED. This duration will never be less than a model-specific minimum, which will be specified together with the rate limits for the model. |
HistoryConfig
History configuration.
This message is included in the session configuration as BidiGenerateContentSetup.historyConfig . Configures the exchange of history messages.
| Fields | |
|---|---|
initialHistoryInClientContent | Optional. If true, after sending |
ProactivityConfig
Config for proactivity features.
| Fields | |
|---|---|
proactiveAudio | Optional. If enabled, the model can reject responding to the last prompt. For example, this allows the model to ignore out of context speech or to stay silent if the user did not make a request, yet. |
RealtimeInputConfig
Configures the realtime input behavior in BidiGenerateContent .
| Fields | |
|---|---|
automaticActivityDetection | Optional. If not set, automatic activity detection is enabled by default. If automatic voice detection is disabled, the client must send activity signals. |
activityHandling | Optional. Defines what effect activity has. |
turnCoverage | Optional. Defines which input is included in the user's turn. |
SessionResumptionConfig
Session resumption configuration.
This message is included in the session configuration as BidiGenerateContentSetup.sessionResumption . If configured, the server will send SessionResumptionUpdate messages.
| Fields | |
|---|---|
handle | The handle of a previous session. If not present then a new session is created. Session handles come from |
SessionResumptionUpdate
Update of the session resumption state.
Only sent if BidiGenerateContentSetup.sessionResumption was set.
| Fields | |
|---|---|
newHandle | New handle that represents a state that can be resumed. Empty if |
resumable | True if the current session can be resumed at this point. Resumption is not possible at some points in the session. For example, when the model is executing function calls or generating. Resuming the session (using a previous session token) in such a state will result in some data loss. In these cases, |
SlidingWindow
The SlidingWindow method operates by discarding content at the beginning of the context window. The resulting context will always begin at the start of a USER role turn. System instructions and any BidiGenerateContentSetup.prefixTurns will always remain at the beginning of the result.
| Fields | |
|---|---|
targetTokens | The target number of tokens to keep. The default value is trigger_tokens/2. Discarding parts of the context window causes a temporary latency increase so this value should be calibrated to avoid frequent compression operations. |
StartSensitivity
Determines how start of speech is detected.
| Enums | |
|---|---|
START_SENSITIVITY_UNSPECIFIED | The default is START_SENSITIVITY_HIGH. |
START_SENSITIVITY_HIGH | Automatic detection will detect the start of speech more often. |
START_SENSITIVITY_LOW | Automatic detection will detect the start of speech less often. |
TurnCoverage
Options about which input is included in the user's turn.
| Enums | |
|---|---|
TURN_COVERAGE_UNSPECIFIED | If unspecified, a default behavior is selected based on the model. Eg, for Gemini 2.5, the default is TURN_INCLUDES_ONLY_ACTIVITY , while for Gemini 3.1 and onwards, it's TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO . |
TURN_INCLUDES_ONLY_ACTIVITY | Includes activity since the last turn, excluding inactivity (eg silence on the audio stream). |
TURN_INCLUDES_ALL_INPUT | Includes all realtime input since the last turn, including inactivity (eg silence on the audio stream). |
TURN_INCLUDES_AUDIO_ACTIVITY_AND_ALL_VIDEO | Includes audio activity and all video since the last turn. With automatic activity detection, audio activity means speech and excludes silence. |
TranslationConfig
Config for translation features.
| Fields | |
|---|---|
targetLanguageCode | Required. The target language for translation. Supported values are BCP-47 language codes (eg "en", "es", "fr"). |
echoTargetLanguage | Optional. If true, the model will generate audio when the target language is spoken, essentially it will parrot the input. If false, we will not produce audio for the target language. |
UrlContextMetadata
Metadata related to url context retrieval tool.
| Fields | |
|---|---|
urlMetadata[] | List of url context. |
UsageMetadata
Usage metadata about response(s).
| Fields | |
|---|---|
promptTokenCount | Output only. Number of tokens in the prompt. When |
cachedContentTokenCount | Number of tokens in the cached part of the prompt (the cached content) |
responseTokenCount | Output only. Total number of tokens across all the generated response candidates. |
toolUsePromptTokenCount | Output only. Number of tokens present in tool-use prompt(s). |
thoughtsTokenCount | Output only. Number of tokens of thoughts for thinking models. |
totalTokenCount | Output only. Total token count for the generation request (prompt + response candidates). |
promptTokensDetails[] | Output only. List of modalities that were processed in the request input. |
cacheTokensDetails[] | Output only. List of modalities of the cached content in the request input. |
responseTokensDetails[] | Output only. List of modalities that were returned in the response. |
toolUsePromptTokensDetails[] | Output only. List of modalities that were processed for tool-use request inputs. |
Ephemeral authentication tokens
Ephemeral authentication tokens can be obtained by calling AuthTokenService.CreateToken and then used with GenerativeService.BidiGenerateContentConstrained , either by passing the token in an access_token query parameter, or in an HTTP Authorization header with " Token " prefixed to it.
CreateAuthTokenRequest
Create an ephemeral authentication token.
| Fields | |
|---|---|
authToken | Required. The token to create. |
AuthToken
A request to create an ephemeral authentication token.
| Fields | |
|---|---|
name | Output only. Identifier. The token itself. |
expireTime | Optional. Input only. Immutable. An optional time after which, when using the resulting token, messages in BidiGenerateContent sessions will be rejected. (Gemini may preemptively close the session after this time.) If not set then this defaults to 30 minutes in the future. If set, this value must be less than 20 hours in the future. |
newSessionExpireTime | Optional. Input only. Immutable. The time after which new Live API sessions using the token resulting from this request will be rejected. If not set this defaults to 60 seconds in the future. If set, this value must be less than 20 hours in the future. |
fieldMask | Optional. Input only. Immutable. If field_mask is empty, and If field_mask is empty, and If field_mask is not empty, then the corresponding fields from |
Union field config . The method-specific configuration for the resulting token. config can be only one of the following: | |
bidiGenerateContentSetup | Optional. Input only. Immutable. Configuration specific to |
uses | Optional. Input only. Immutable. The number of times the token can be used. If this value is zero then no limit is applied. Resuming a Live API session does not count as a use. If unspecified, the default is 1. |
More information on common types
For more information on the commonly-used API resource types Blob , Content , FunctionCall , FunctionResponse , GenerationConfig , GroundingMetadata , ModalityTokenCount , and Tool , see Generating content .