Gemini Interactions API

رابط برنامه‌نویسی کاربردی (API) تعاملات جمینی (Gemini Interactions API) به توسعه‌دهندگان اجازه می‌دهد تا با استفاده از مدل‌های جمینی، برنامه‌های هوش مصنوعی مولد (generative AI applications) بسازند. جمینی توانمندترین مدل ما است که از پایه برای چندوجهی بودن ساخته شده است. این مدل می‌تواند انواع مختلف اطلاعات از جمله زبان، تصاویر، صدا، ویدئو و کد را تعمیم داده و به طور یکپارچه درک کند، در میان آنها عمل کند و ترکیب کند. می‌توانید از رابط برنامه‌نویسی کاربردی جمینی برای موارد استفاده‌ای مانند استدلال در متن و تصاویر، تولید محتوا، عامل‌های گفتگو، سیستم‌های خلاصه‌سازی و طبقه‌بندی و موارد دیگر استفاده کنید.

نسخه API: v1beta v1

ایجاد تعامل

ارسال به آدرس https://generativelanguage.googleapis.com/v1beta/interactions

یک تعامل جدید ایجاد می‌کند.

پارامترهای مسیر/پرس‌وجو

رشته api_version (الزامی)

از کدام نسخه API استفاده کنیم.

درخواست بدنه

بدنه درخواست شامل داده‌هایی با ساختار زیر است:

مدل ModelOption (اختیاری)

نام «مدل» مورد استفاده برای تولید تعامل.
در صورت عدم ارائه «عامل»، الزامی است.

مدلی که اعلان شما را تکمیل می‌کند.\n\nبرای جزئیات بیشتر به [models](https://ai.google.dev/gemini-api/docs/models) مراجعه کنید.

مقادیر ممکن

  • gemini-2.5-flash

    اولین مدل استدلال ترکیبی ما که از یک پنجره زمینه ۱ میلیون توکنی پشتیبانی می‌کند و دارای بودجه‌های تفکر است.

  • gemini-2.5-pro

    مدل چندمنظوره پیشرفته ما، که در کدنویسی و کارهای استدلالی پیچیده عالی عمل می‌کند.

  • gemma-4-26b-a4b-it

    جما ۴ ۲۶ب A4ب آی تی

  • gemma-4-31b-it

    جما ۴ ۳۱ب آی‌تی

  • gemini-flash-latest

    آخرین نسخه از بازی Gemini Flash

  • gemini-flash-lite-latest

    آخرین نسخه Gemini Flash-Lite

  • gemini-pro-latest

    آخرین نسخه Gemini Pro

  • gemini-2.5-flash-lite

    کوچکترین و مقرون به صرفه ترین مدل ما، ساخته شده برای استفاده در مقیاس بزرگ.

  • gemini-2.5-flash-image

    مدل تولید تصویر بومی ما، که برای سرعت، انعطاف‌پذیری و درک متنی بهینه شده است. ورودی و خروجی متن با همان قیمت ۲.۵ فلش ارائه می‌شود.

  • gemini-3-flash-preview

    هوشمندترین مدل ما که برای سرعت ساخته شده است، هوش مرزی را با جستجو و ردیابی برتر ترکیب می‌کند.

  • gemini-3.1-pro-preview

    جدیدترین مدل استدلال SOTA ما با عمق و ظرافت بی‌سابقه و قابلیت‌های قدرتمند درک و کدنویسی چندوجهی.

  • gemini-3.1-pro-preview-customtools

    پیش‌نمایش Gemini 3.1 Pro برای استفاده از ابزارهای سفارشی بهینه شده است

  • gemini-3.1-flash-lite

    مقرون‌به‌صرفه‌ترین مدل ما، بهینه‌شده برای وظایف عامل‌محور با حجم بالا، ترجمه و پردازش داده‌های ساده.

  • gemini-3-pro-image

    تصویر Gemini 3 Pro

  • nano-banana-pro-preview

    پیش‌نمایش تصویر Gemini 3 Pro

  • gemini-3.1-flash-image

    تصویر فلش جمینی ۳.۱.

  • gemini-3.5-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

  • gemini-3.6-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

  • gemini-3.7-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

  • lyria-3-clip-preview

    مدل تولید موسیقی با تأخیر کم ما برای کلیپ‌های صوتی با کیفیت بالا و کنترل دقیق ریتمیک بهینه شده است.

  • lyria-3-pro-preview

    مدل پیشرفته و کامل ما برای تولید آهنگ با درک عمیق از آهنگسازی، بهینه شده برای کنترل ساختاری دقیق و انتقال‌های پیچیده در سبک‌های مختلف موسیقی.

  • gemini-robotics-er-1.6-preview

    پیش‌نمایش Gemini Robotics-ER 1.6

  • gemini-robotics-er-2-preview

    پیش‌نمایش Gemini Robotics Embodied Reasoning 2

گزینه عامل (اختیاری)

نام «عامل» مورد استفاده برای ایجاد تعامل.
در صورت عدم ارائه «مدل»، الزامی است.

عاملی که باید با آن تعامل داشت.

مقادیر ممکن

  • deep-research-pro-preview-12-2025

    نماینده تحقیقات عمیق جمینی

  • deep-research-preview-04-2026

    نماینده تحقیقات عمیق جمینی

  • deep-research-max-preview-04-2026

    مامور مکس تحقیقات عمیق جمینی

  • antigravity-preview-05-2026

    از عامل مدیریت‌شده‌ی Antigravity برای انجام وظایف چند مرحله‌ای که نیاز به استدلال، عملیات فایل و استفاده از ابزار دارند، استفاده کنید.

ورودی محتوا یا آرایه ( Content ) یا آرایه ( Step ) یا رشته (الزامی)

ورودی‌های تعامل (مشترک برای مدل و عامل).

رشته system_instruction (اختیاری)

دستورالعمل سیستم برای تعامل.

آرایه ابزارها ( ابزار ) (اختیاری)

فهرستی از اعلان‌های ابزار که مدل ممکن است در طول تعامل فراخوانی کند.

response_format فرمت پاسخ یا آرایه ( ResponseFormat ) (اختیاری)

تأکید می‌کند که پاسخ تولید شده یک شیء JSON است که با طرحواره JSON مشخص شده در این فیلد مطابقت دارد.

جریان بولی (اختیاری)

فقط ورودی. اینکه آیا تعامل پخش زنده خواهد شد یا خیر.

ذخیره بولی (اختیاری)

فقط ورودی. آیا پاسخ و درخواست برای بازیابی بعدی ذخیره شود یا خیر.

مقدار بولی پس‌زمینه (اختیاری)

فقط ورودی. اینکه آیا تعامل مدل در پس‌زمینه اجرا شود یا خیر.

generation_config GenerationConfig (اختیاری)

پیکربندی مدل
پارامترهای پیکربندی برای تعامل مدل.
جایگزینی برای `agent_config`. فقط زمانی قابل اجرا است که `model` تنظیم شده باشد.

پارامترهای پیکربندی برای تعاملات مدل.

فیلدها

عدد صحیح max_output_tokens (اختیاری)

حداکثر تعداد توکن‌هایی که باید در پاسخ گنجانده شوند.

عدد صحیح اولیه (اختیاری)

بذر مورد استفاده در رمزگشایی برای تکرارپذیری.

speech_config SpeakerConfig یا آرایه (SpeechConfig) (اختیاری)

اختیاری. پیکربندی گفتار و چند بلندگو.

پیکربندی برای چند گوینده و تولید گفتار.

فیلدها

آرایه بلندگوها (SpeechConfig) (اختیاری)

تنظیمات بلندگوهای جداگانه.

پیکربندی برای تعامل گفتاری.

فیلدها

رشته زبان (اختیاری)

زبان گفتار.

سیم بلندگو (اختیاری)

نام گوینده، باید با نام گوینده داده شده در سوال مطابقت داشته باشد.

رشته صدا (اختیاری)

صدای گوینده.

آرایه stop_sequences (رشته) (اختیاری)

فهرستی از توالی‌های کاراکتری که تعامل خروجی را متوقف می‌کنند.

سطح_فکریسطح_فکری ( اختیاری )

سطح توکن‌های فکری که مدل باید تولید کند.

مقادیر ممکن

  • minimal

    کم یا بدون فکر کردن.

  • low

    سطح فکری پایین.

  • medium

    سطح فکری متوسط.

  • high

    سطح فکری بالا.

خلاصه‌های تفکر ( اختیاری)

اینکه آیا خلاصه نظرات در پاسخ گنجانده شود یا خیر.

مقادیر ممکن

  • auto

    خلاصه‌های تفکر خودکار

  • none

    بدون خلاصه نویسی فکری.

tool_choice ToolChoiceConfig یا enum (رشته‌ای) (اختیاری)

پیکربندی انتخاب ابزار.

مقادیر ممکن:

  • auto

    انتخاب خودکار ابزار.

  • any

    هر انتخاب ابزاری.

  • none

    بدون انتخاب ابزار.

  • validated

    انتخاب ابزار معتبر.

transcription_config TranscriptionConfig (اختیاری)

اختیاری. پیکربندی برای تشخیص گفتار (رونویسی). در صورت وجود، ASR فعال می‌شود.

پیکربندی برای تشخیص گفتار (رونویسی).

فیلدها

آرایه custom_vocabulary (رشته‌ای) (اختیاری)

اختیاری. فهرستی از عبارات واژگانی سفارشی برای جهت‌دهی مدل تشخیص گفتار به سمت تشخیص اصطلاحات خاص.

آرایه language_codes (رشته) (اختیاری)

اختیاری. کدهای زبان BCP-47 نکاتی در مورد زبان‌های موجود در صدا ارائه می‌دهند. در صورت حذف یا خالی بودن، به طور پیش‌فرض روی تشخیص خودکار زبان تنظیم می‌شود.

حالت رونویسی یا enum (رشته) (اختیاری)

گزینه‌های حالت رونویسی متمایز یا enum.

پیکربندی برای حالت رونویسی.

انواع ممکن

حالت رونویسی هوشمند

پیکربندی برای حالت رونویسی هوشمند.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "smart" تنظیم شود.

حالت رونویسی کلمه به کلمه

پیکربندی برای حالت رونویسی کلمه به کلمه.

رشته diarization_mode (اختیاری)

اختیاری. تنظیم بلندگو را تنظیم می‌کند. مقادیر پشتیبانی شده: "speaker".

آرایه timestamp_granularities (رشته‌ای) (اختیاری)

اختیاری. جزئیات مهرهای زمانی که باید در خروجی رونویسی گنجانده شوند. مقادیر پشتیبانی شده: "word". اگر خالی باشد، هیچ مهر زمانی تولید نمی‌شود.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "verbatim" تنظیم شود.

video_config پیکربندی ویدیو (اختیاری)

پیکربندی برای تولید ویدیو.

گزینه‌های پیکربندی برای تولید ویدیو.

فیلدها

شمارش وظیفه (رشته) (اختیاری)

حالت وظیفه اختیاری برای تولید ویدیو. در صورت مشخص نکردن، مدل به طور خودکار حالت مناسب را بر اساس متن ارائه شده و رسانه ورودی تعیین می‌کند.

مقادیر ممکن:

  • text_to_video

    ویدیو را منحصراً از یک متن فوری تولید می‌کند.

  • image_to_video

    ویدئو را از یک یا دو تصویر منبع تولید می‌کند. تصویر اول فریم شروع و تصویر دوم که اختیاری است، فریم پایان را تعریف می‌کند.

  • reference_to_video

    با استفاده از رسانه‌های مرجع (مانند تصاویر، صدا یا ویدیو) ویدیو تولید می‌کند.

  • edit

    یک ویدیوی ورودی موجود را تغییر می‌دهد.

  • extend

    یک ویدیوی ورودی موجود را گسترش می‌دهد.

شیء agent_config (اختیاری)

پیکربندی عامل
پیکربندی برای عامل.
جایگزینی برای `generation_config`. فقط زمانی قابل اجرا است که `agent` تنظیم شده باشد.

انواع ممکن

تفکیک‌کننده چندریختی: type

پیکربندی عامل ضد جاذبه

پیکربندی برای زمان اجرای عامل Antigravity. کنترل سمت سرور بر محیط اجرای عامل و پیکربندی ابزار را فراهم می‌کند.

رشته‌ی max_total_tokens (اختیاری)

حداکثر مجموع توکن‌ها برای اجرای عامل.

رشته مدل (اختیاری)

مدلی که برای استدلال عامل استفاده می‌شود.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "antigravity" تنظیم شود.

پیکربندی CodeMenderAgentConfig

پیکربندی برای عامل CodeMender.

درخواست_یافتن (اختیاری )

پارامترهایی برای یافتن آسیب‌پذیری‌ها

پارامترهای درخواست مختص جلسات FIND، که برای کشف آسیب‌پذیری‌ها در یک پایگاه کد استفاده می‌شوند.

فیلدها

رشته توضیحات (اختیاری)

دستورالعمل‌های زمینه‌ای یا سفارشی اضافی که توسط کاربر برای هدایت تحلیل آسیب‌پذیری ارائه می‌شود.

رشته find_id (اختیاری)

شناسه‌ی یک یافته‌ی خاص برای تأیید. این مورد عمدتاً در حالت VERIFY برای تمرکز اعتبارسنجی مبتنی بر اجرای عامل بر روی یک آسیب‌پذیری واحد استفاده می‌شود.

حالت شمارشی (رشته) (اختیاری)

نحوه‌ی جلسه‌ی یافتن.

مقادیر ممکن:

  • scan

    اسکن سریع فقط با استفاده از طبقه‌بندی‌کننده اولیه.

  • verify

    طبقه‌بندی را انجام می‌دهد و به دنبال آن بررسی دقیقی انجام می‌دهد.

آرایه source_files (محتوای فایل) (اختیاری)

فهرستی از فایل‌های منبع که به عنوان زمینه برای اسکن ارائه می‌شوند.

محتوای یک فایل واحد در کدبیس.

فیلدها

رشته محتوا (اختیاری)

محتوای متنی فایل که با استاندارد UTF-8 کدگذاری شده است.

رشته مسیر (اختیاری)

مسیر نسبی فایل از ریشه پروژه.

درخواست رفع مشکل (اختیاری)

پارامترهای رفع آسیب‌پذیری‌ها

پارامترهای مخصوص جلسات FIX را درخواست کنید، که برای تولید و اعتبارسنجی وصله‌های امنیتی استفاده می‌شوند.

فیلدها

رشته توضیحات (اختیاری)

دستورالعمل‌های زمینه‌ای یا سفارشی اضافی که توسط کاربر برای هدایت فرآیند تولید پچ ارائه می‌شود.

رشته find_id (اختیاری)

شناسه‌ی یافته‌ی امنیتی خاصی که باید اصلاح شود. این شناسه به یک آسیب‌پذیری کشف‌شده‌ی قبلی اشاره دارد.

آرایه source_files (محتوای فایل) (اختیاری)

فهرستی از فایل‌های منبع که زمینه را برای اصلاح فراهم می‌کنند. این فایل‌ها معمولاً فایل‌هایی هستند که حاوی آسیب‌پذیری شناسایی‌شده هستند.

محتوای یک فایل واحد در کدبیس.

فیلدها

رشته محتوا (اختیاری)

محتوای متنی فایل که با استاندارد UTF-8 کدگذاری شده است.

رشته مسیر (اختیاری)

مسیر نسبی فایل از ریشه پروژه.

رشته مدل (اختیاری)

نام مدلی که برای عامل CodeMender استفاده می‌شود. در هر جلسه CodeMender فقط از یک مدل استفاده خواهد شد.

session_config پیکربندی جلسه (اختیاری)

پیکربندی‌های اختیاری مختص جلسه برای لغو رفتار پیش‌فرض عامل.

پیکربندی جلسات CodeMender.

فیلدها

عدد صحیح max_rounds (اختیاری)

حداکثر تعداد دورهای تعاملی که عامل مجاز است قبل از رسیدن به مهلت زمانی مشخص شده انجام دهد.

رشته session_id (اختیاری)

پارامتری برای گروه‌بندی چندین تعامل که متعلق به یک جلسه CodeMender هستند.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "code-mender" تنظیم شود.

پیکربندی DeepResearchAgent

پیکربندی برای عامل تحقیقات عمیق.

نوع بولی collaboration_planning (اختیاری)

برنامه‌ریزی انسان در حلقه را برای عامل تحقیقات عمیق فعال می‌کند. اگر روی درست تنظیم شود، عامل تحقیقات عمیق در پاسخ خود یک طرح تحقیقاتی ارائه می‌دهد. سپس عامل تنها در صورتی ادامه می‌دهد که کاربر طرح را در نوبت بعدی تأیید کند.

enable_bigquery_tool مقدار بولی (اختیاری)

ابزار bigquery را برای عامل Deep Research فعال می‌کند.

خلاصه‌های تفکر ( اختیاری)

اینکه آیا خلاصه نظرات در پاسخ گنجانده شود یا خیر.

مقادیر ممکن

  • auto

    خلاصه‌های تفکر خودکار

  • none

    بدون خلاصه نویسی فکری.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "deep-research" تنظیم شود.

تجسم enum (رشته) (اختیاری)

اینکه آیا باید از تصاویر در پاسخ استفاده کرد یا خیر.

مقادیر ممکن:

  • off

    تجسمات را لحاظ نکنید.

  • auto

    به طور خودکار تجسم‌ها را شامل می‌شود.

پیکربندی DynamicAgent

پیکربندی برای عامل‌های پویا

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "dynamic" تنظیم شود.

environmentConfig یا رشته (اختیاری)

پیکربندی محیط برای تعامل. می‌تواند یک شیء باشد که منابع محیط از راه دور را مشخص می‌کند یا رشته‌ای باشد که به یک شناسه محیط موجود اشاره می‌کند.

شیء برچسب‌ها (اختیاری)

برچسب‌هایی با فراداده‌های تعریف‌شده توسط کاربر برای درخواست.

رشته‌ی previous_interaction_id (اختیاری)

شناسه‌ی تعامل قبلی، در صورت وجود.

آرایه safety_settings (SafetySetting) (اختیاری)

تنظیمات ایمنی برای تعامل.

یک تنظیم ایمنی که بر رفتار مسدود کردن ایمنی تأثیر می‌گذارد. یک تنظیم ایمنی شامل یک دسته آسیب و یک آستانه برای آن دسته است.

فیلدها

متد enum (رشته) (اختیاری)

اختیاری. روش مسدود کردن محتوا. اگر مشخص نشده باشد، رفتار پیش‌فرض استفاده از امتیاز احتمال است.

مقادیر ممکن:

  • severity

    روش بلوک آسیب از هر دو امتیاز احتمال و شدت استفاده می‌کند.

  • probability

    روش بلوک آسیب از امتیاز احتمال استفاده می‌کند.

شمارش آستانه (رشته) (اختیاری)

الزامی. آستانه‌ی مسدود کردن محتوا. اگر احتمال آسیب از این آستانه بیشتر شود، محتوا مسدود خواهد شد.

مقادیر ممکن:

  • block_low_and_above

    مسدود کردن محتوا با احتمال آسیب کم یا بیشتر.

  • block_medium_and_above

    محتوایی را که احتمال آسیب‌رسانی آن متوسط ​​یا بالاتر است، مسدود کنید.

  • block_only_high

    مسدود کردن محتوا با احتمال آسیب بالا.

  • block_none

    صرف نظر از احتمال آسیب‌رسانی آن، هیچ محتوایی را مسدود نکنید.

  • off

    فیلتر ایمنی را کاملاً خاموش کنید.

نوع آسیب (اختیاری)

الزامی. نوع دسته‌بندی آسیبی که باید مسدود شود.

مقادیر ممکن

  • hate_speech

    محتوایی که خشونت را ترویج می‌دهد یا بر اساس ویژگی‌های خاص، نفرت علیه افراد یا گروه‌ها را برمی‌انگیزد.

  • dangerous_content

    محتوایی که فعالیت‌های خطرناک را ترویج، تسهیل یا امکان‌پذیر می‌کند.

  • harassment

    محتوای توهین‌آمیز، تهدیدآمیز یا با هدف قلدری، شکنجه یا تمسخر.

  • sexually_explicit

    محتوایی که حاوی مطالب جنسی و غیراخلاقی باشد.

  • civic_integrity

    منسوخ شده: فیلتر انتخابات دیگر پشتیبانی نمی‌شود. دسته‌بندی آسیب، سلامت مدنی است.

  • image_hate

    تصاویری که حاوی نفرت‌پراکنی هستند.

  • image_dangerous_content

    تصاویری که حاوی محتوای خطرناک هستند.

  • image_harassment

    تصاویری که حاوی آزار و اذیت هستند.

  • image_sexually_explicit

    تصاویری که حاوی محتوای جنسی هستند.

  • jailbreak

    اعلان‌هایی که برای دور زدن فیلترهای ایمنی طراحی شده‌اند.

service_tier لایه سرویس (اختیاری)

لایه سرویس برای تعامل.

مقادیر ممکن

  • flex

    سطح خدمات انعطاف‌پذیر.

  • standard

    سطح خدمات استاندارد.

  • priority

    ردیف خدمات اولویت‌دار.

  • deferred

    ردیف خدمات معوق.

webhook_config پیکربندی وب هوک (اختیاری)

اختیاری. پیکربندی وب‌هوک برای دریافت اعلان‌ها پس از اتمام تعامل.

پیام مربوط به پیکربندی رویدادهای وب‌هوک برای یک درخواست.

فیلدها

آرایه uris (رشته) (اختیاری)

اختیاری. در صورت تنظیم، این URLهای وب‌هوک به جای وب‌هوک‌های ثبت‌شده، برای رویدادهای وب‌هوک استفاده خواهند شد.

شیء user_metadata (اختیاری)

اختیاری. فراداده کاربر که در هر انتشار رویداد به وب‌هوک‌ها بازگردانده می‌شود.

پاسخ

یک منبع تعامل (Interaction) را برمی‌گرداند.

درخواست ساده

پاسخ نمونه

{
  "created": "2025-11-26T12:25:15Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "Hello! I'm functioning perfectly and ready to assist you.\n\nHow are you doing today?"
        }
      ]
    }
  ],
  "updated": "2025-11-26T12:25:15Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 7
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 7,
    "total_output_tokens": 20,
    "total_thought_tokens": 22,
    "total_tokens": 49,
    "total_tool_use_tokens": 0
  }
}

چند نوبتی

پاسخ نمونه

{
  "created": "2025-11-26T12:22:47Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "The capital of France is Paris."
        }
      ]
    }
  ],
  "updated": "2025-11-26T12:22:47Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 50
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 50,
    "total_output_tokens": 10,
    "total_thought_tokens": 0,
    "total_tokens": 60,
    "total_tool_use_tokens": 0
  }
}

ورودی تصویر

پاسخ نمونه

{
  "created": "2025-11-26T12:22:47Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "A white humanoid robot with glowing blue eyes stands holding a red skateboard."
        }
      ]
    }
  ],
  "updated": "2025-11-26T12:22:47Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 10
      },
      {
        "modality": "image",
        "tokens": 258
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 268,
    "total_output_tokens": 20,
    "total_thought_tokens": 0,
    "total_tokens": 288,
    "total_tool_use_tokens": 0
  }
}

فراخوانی تابع

پاسخ نمونه

{
  "created": "2025-11-26T12:22:47Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "requires_action",
  "steps": [
    {
      "name": "get_weather",
      "type": "function_call",
      "arguments": {
        "location": "Boston, MA"
      },
      "id": "gth23981"
    }
  ],
  "updated": "2025-11-26T12:22:47Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 100
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 100,
    "total_output_tokens": 25,
    "total_thought_tokens": 0,
    "total_tokens": 125,
    "total_tool_use_tokens": 50
  }
}

تحقیقات عمیق

پاسخ نمونه

{
  "agent": "deep-research-pro-preview-12-2025",
  "created": "2025-11-26T12:22:47Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "Here is a comprehensive research report on the current state of cancer research..."
        }
      ]
    }
  ],
  "updated": "2025-11-26T12:22:47Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 20
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 20,
    "total_output_tokens": 1000,
    "total_thought_tokens": 500,
    "total_tokens": 1520,
    "total_tool_use_tokens": 0
  }
}

عامل ضد جاذبه

پاسخ نمونه

{
  "agent": "antigravity-preview-05-2026",
  "created": "2025-11-26T12:22:47Z",
  "environment_id": "env_abc123",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "I've summarized the top 5 Hacker News stories and saved the results to /workspace/summary.md."
        }
      ]
    }
  ],
  "updated": "2025-11-26T12:22:47Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 50
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 50,
    "total_output_tokens": 500,
    "total_thought_tokens": 200,
    "total_tokens": 750,
    "total_tool_use_tokens": 0
  }
}

محیط استفاده مجدد

پاسخ نمونه

{
  "agent": "antigravity-preview-05-2026",
  "created": "2025-11-26T12:23:00Z",
  "environment_id": "env_abc123",
  "id": "v1_Chd2ZTJhYmNkZWZnaGlqa2xtbm9wcXJzdHV2d3h5ejAxMjM0NTY3ODkwMTIzNDU2Nzg",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "I've updated /workspace/hello.py to accept a name argument and greet the user."
        }
      ]
    }
  ],
  "updated": "2025-11-26T12:23:00Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 80
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 80,
    "total_output_tokens": 200,
    "total_thought_tokens": 100,
    "total_tokens": 380,
    "total_tool_use_tokens": 0
  }
}

با منابع

نماینده سفارشی

لغو یک تعامل

ارسال https://generativelanguage.googleapis.com/v1beta/interactions/{id}/cancel

یک تعامل را بر اساس شناسه لغو می‌کند. این فقط برای تعاملات پس‌زمینه‌ای که هنوز در حال اجرا هستند، اعمال می‌شود.

پارامترهای مسیر/پرس‌وجو

رشته api_version (الزامی)

از کدام نسخه API استفاده کنیم.

رشته شناسه (الزامی)

شناسه منحصر به فرد تعاملی که باید لغو شود.

پاسخ

یک منبع تعامل (Interaction) را برمی‌گرداند.

لغو تعامل

پاسخ نمونه

{
  "agent": "deep-research-pro-preview-12-2025",
  "created": "2026-06-22T04:55:47Z",
  "id": "v1_ChdVc0E0YXJTYk1zYlV6N0lQcXRXVG1BYxIXVXNBNGFyU2JNc2JVejdJUHF0V1RtQWM",
  "status": "cancelled",
  "steps": [
    {
      "type": "user_input",
      "content": [
        {
          "type": "text",
          "text": "Research the history of the Google TPUs with a focus on 2025 specs."
        }
      ]
    }
  ],
  "updated": "2026-06-22T04:55:47Z"
}

بازیابی یک تعامل

دریافت کنید https://generativelanguage.googleapis.com/v1beta/interactions/{id}

جزئیات کامل یک تعامل واحد را بر اساس `Interaction.id` آن بازیابی می‌کند.

پارامترهای مسیر/پرس‌وجو

رشته api_version (الزامی)

از کدام نسخه API استفاده کنیم.

رشته شناسه (الزامی)

شناسه منحصر به فرد تعاملی که قرار است بازیابی شود.

رشته last_event_id (اختیاری)

اختیاری. در صورت تنظیم، جریان تعامل را از بخش بعدی پس از رویداد مشخص شده توسط شناسه رویداد از سر می‌گیرد. فقط در صورتی قابل استفاده است که `stream` برابر با true باشد.

جریان بولی (اختیاری)

اگر روی درست تنظیم شود، محتوای تولید شده به صورت تدریجی پخش می‌شود.

پیش‌فرض: False

پاسخ

یک منبع تعامل (Interaction) را برمی‌گرداند.

تعامل دریافت کنید

پاسخ نمونه

{
  "created": "2025-11-26T12:25:15Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "I'm doing great, thank you for asking! How can I help you today?"
        }
      ]
    }
  ],
  "updated": "2025-11-26T12:25:15Z"
}

حذف یک تعامل

https://generativelanguage.googleapis.com/v1beta/interactions/{id} را حذف کنید

تعامل را بر اساس شناسه حذف می‌کند.

پارامترهای مسیر/پرس‌وجو

رشته api_version (الزامی)

از کدام نسخه API استفاده کنیم.

رشته شناسه (الزامی)

شناسه منحصر به فرد تعاملی که باید حذف شود.

پاسخ

در صورت موفقیت، پاسخ خالی است.

حذف

منابع

تعامل

منبع تعامل.

فیلدها

گزینه عامل (اختیاری)

نام «عامل» مورد استفاده برای ایجاد تعامل.

عاملی که باید با آن تعامل داشت.

مقادیر ممکن

  • deep-research-pro-preview-12-2025

    نماینده تحقیقات عمیق جمینی

  • deep-research-preview-04-2026

    نماینده تحقیقات عمیق جمینی

  • deep-research-max-preview-04-2026

    مامور مکس تحقیقات عمیق جمینی

  • antigravity-preview-05-2026

    از عامل مدیریت‌شده‌ی Antigravity برای انجام وظایف چند مرحله‌ای که نیاز به استدلال، عملیات فایل و استفاده از ابزار دارند، استفاده کنید.

شیء agent_config (اختیاری)

پارامترهای پیکربندی برای تعامل عامل.

انواع ممکن

تفکیک‌کننده چندریختی: type

پیکربندی عامل ضد جاذبه

پیکربندی برای زمان اجرای عامل Antigravity. کنترل سمت سرور بر محیط اجرای عامل و پیکربندی ابزار را فراهم می‌کند.

رشته‌ی max_total_tokens (اختیاری)

حداکثر مجموع توکن‌ها برای اجرای عامل.

رشته مدل (اختیاری)

مدلی که برای استدلال عامل استفاده می‌شود.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "antigravity" تنظیم شود.

پیکربندی CodeMenderAgentConfig

پیکربندی برای عامل CodeMender.

درخواست_یافتن (اختیاری )

پارامترهایی برای یافتن آسیب‌پذیری‌ها

پارامترهای درخواست مختص جلسات FIND، که برای کشف آسیب‌پذیری‌ها در یک پایگاه کد استفاده می‌شوند.

فیلدها

رشته توضیحات (اختیاری)

دستورالعمل‌های زمینه‌ای یا سفارشی اضافی که توسط کاربر برای هدایت تحلیل آسیب‌پذیری ارائه می‌شود.

رشته find_id (اختیاری)

شناسه‌ی یک یافته‌ی خاص برای تأیید. این مورد عمدتاً در حالت VERIFY برای تمرکز اعتبارسنجی مبتنی بر اجرای عامل بر روی یک آسیب‌پذیری واحد استفاده می‌شود.

حالت شمارشی (رشته) (اختیاری)

نحوه‌ی جلسه‌ی یافتن.

مقادیر ممکن:

  • scan

    اسکن سریع فقط با استفاده از طبقه‌بندی‌کننده اولیه.

  • verify

    طبقه‌بندی را انجام می‌دهد و به دنبال آن بررسی دقیقی انجام می‌دهد.

آرایه source_files (محتوای فایل) (اختیاری)

فهرستی از فایل‌های منبع که به عنوان زمینه برای اسکن ارائه می‌شوند.

محتوای یک فایل واحد در کدبیس.

فیلدها

رشته محتوا (اختیاری)

محتوای متنی فایل که با استاندارد UTF-8 کدگذاری شده است.

رشته مسیر (اختیاری)

مسیر نسبی فایل از ریشه پروژه.

درخواست رفع مشکل (اختیاری)

پارامترهای رفع آسیب‌پذیری‌ها

پارامترهای مخصوص جلسات FIX را درخواست کنید، که برای تولید و اعتبارسنجی وصله‌های امنیتی استفاده می‌شوند.

فیلدها

رشته توضیحات (اختیاری)

دستورالعمل‌های زمینه‌ای یا سفارشی اضافی که توسط کاربر برای هدایت فرآیند تولید پچ ارائه می‌شود.

رشته find_id (اختیاری)

شناسه‌ی یافته‌ی امنیتی خاصی که باید اصلاح شود. این شناسه به یک آسیب‌پذیری کشف‌شده‌ی قبلی اشاره دارد.

آرایه source_files (محتوای فایل) (اختیاری)

فهرستی از فایل‌های منبع که زمینه را برای اصلاح فراهم می‌کنند. این فایل‌ها معمولاً فایل‌هایی هستند که حاوی آسیب‌پذیری شناسایی‌شده هستند.

محتوای یک فایل واحد در کدبیس.

فیلدها

رشته محتوا (اختیاری)

محتوای متنی فایل که با استاندارد UTF-8 کدگذاری شده است.

رشته مسیر (اختیاری)

مسیر نسبی فایل از ریشه پروژه.

رشته مدل (اختیاری)

نام مدلی که برای عامل CodeMender استفاده می‌شود. در هر جلسه CodeMender فقط از یک مدل استفاده خواهد شد.

session_config پیکربندی جلسه (اختیاری)

پیکربندی‌های اختیاری مختص جلسه برای لغو رفتار پیش‌فرض عامل.

پیکربندی جلسات CodeMender.

فیلدها

عدد صحیح max_rounds (اختیاری)

حداکثر تعداد دورهای تعاملی که عامل مجاز است قبل از رسیدن به مهلت زمانی مشخص شده انجام دهد.

رشته session_id (اختیاری)

پارامتری برای گروه‌بندی چندین تعامل که متعلق به یک جلسه CodeMender هستند.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "code-mender" تنظیم شود.

پیکربندی DeepResearchAgent

پیکربندی برای عامل تحقیقات عمیق.

نوع بولی collaboration_planning (اختیاری)

برنامه‌ریزی انسان در حلقه را برای عامل تحقیقات عمیق فعال می‌کند. اگر روی درست تنظیم شود، عامل تحقیقات عمیق در پاسخ خود یک طرح تحقیقاتی ارائه می‌دهد. سپس عامل تنها در صورتی ادامه می‌دهد که کاربر طرح را در نوبت بعدی تأیید کند.

enable_bigquery_tool مقدار بولی (اختیاری)

ابزار bigquery را برای عامل Deep Research فعال می‌کند.

خلاصه‌های تفکر ( اختیاری)

اینکه آیا خلاصه نظرات در پاسخ گنجانده شود یا خیر.

مقادیر ممکن

  • auto

    خلاصه‌های تفکر خودکار

  • none

    بدون خلاصه نویسی فکری.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "deep-research" تنظیم شود.

تجسم enum (رشته) (اختیاری)

اینکه آیا باید از تصاویر در پاسخ استفاده کرد یا خیر.

مقادیر ممکن:

  • off

    تجسمات را لحاظ نکنید.

  • auto

    به طور خودکار تجسم‌ها را شامل می‌شود.

پیکربندی DynamicAgent

پیکربندی برای عامل‌های پویا

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "dynamic" تنظیم شود.

رشته ایجاد شده (اختیاری)

فقط خروجی. زمانی که پاسخ در قالب ISO 8601 (YYYY-MM-DDThh:mm:ssZ) ایجاد شده است.

environmentConfig یا رشته (اختیاری)

پیکربندی محیط برای تعامل. می‌تواند یک شیء باشد که منابع محیط از راه دور را مشخص می‌کند یا رشته‌ای باشد که به یک شناسه محیط موجود اشاره می‌کند.

رشته‌ی environment_id (اختیاری)

فقط خروجی. شناسه محیط برای تعامل. فقط در صورتی که پیکربندی محیط در درخواست تنظیم شده باشد، پر می‌شود.

آرایه خطاها (Error) (اختیاری)

فقط خروجی. خطاهای تشخیصی / خطاهای پلتفرم که در تعامل ثبت شده‌اند.

پیام خطا از یک تعامل.

فیلدها

رشته کد (اختیاری)

یک URI که نوع خطا را مشخص می‌کند.

رشته پیام (اختیاری)

یک پیام خطا که برای انسان قابل خواندن باشد.

رشته شناسه (اختیاری)

الزامی. فقط خروجی. یک شناسه منحصر به فرد برای تکمیل تعامل.

پیش‌فرض‌ها به:

ورودی محتوا یا آرایه ( Content ) یا آرایه ( Step ) یا رشته (اختیاری)

ورودی برای تعامل.

شیء برچسب‌ها (اختیاری)

برچسب‌هایی با فراداده‌های تعریف‌شده توسط کاربر برای درخواست.

مدل ModelOption (اختیاری)

نام «مدل» مورد استفاده برای تولید تعامل.

مدلی که اعلان شما را تکمیل می‌کند.\n\nبرای جزئیات بیشتر به [models](https://ai.google.dev/gemini-api/docs/models) مراجعه کنید.

مقادیر ممکن

  • gemini-2.5-flash

    اولین مدل استدلال ترکیبی ما که از یک پنجره زمینه ۱ میلیون توکنی پشتیبانی می‌کند و دارای بودجه‌های تفکر است.

  • gemini-2.5-pro

    مدل چندمنظوره پیشرفته ما، که در کدنویسی و کارهای استدلالی پیچیده عالی عمل می‌کند.

  • gemma-4-26b-a4b-it

    جما ۴ ۲۶ب A4ب آی تی

  • gemma-4-31b-it

    جما ۴ ۳۱ب آی‌تی

  • gemini-flash-latest

    آخرین نسخه از بازی Gemini Flash

  • gemini-flash-lite-latest

    آخرین نسخه Gemini Flash-Lite

  • gemini-pro-latest

    آخرین نسخه Gemini Pro

  • gemini-2.5-flash-lite

    کوچکترین و مقرون به صرفه ترین مدل ما، ساخته شده برای استفاده در مقیاس بزرگ.

  • gemini-2.5-flash-image

    مدل تولید تصویر بومی ما، که برای سرعت، انعطاف‌پذیری و درک متنی بهینه شده است. ورودی و خروجی متن با همان قیمت ۲.۵ فلش ارائه می‌شود.

  • gemini-3-flash-preview

    هوشمندترین مدل ما که برای سرعت ساخته شده است، هوش مرزی را با جستجو و ردیابی برتر ترکیب می‌کند.

  • gemini-3.1-pro-preview

    جدیدترین مدل استدلال SOTA ما با عمق و ظرافت بی‌سابقه و قابلیت‌های قدرتمند درک و کدنویسی چندوجهی.

  • gemini-3.1-pro-preview-customtools

    پیش‌نمایش Gemini 3.1 Pro برای استفاده از ابزارهای سفارشی بهینه شده است

  • gemini-3.1-flash-lite

    مقرون‌به‌صرفه‌ترین مدل ما، بهینه‌شده برای وظایف عامل‌محور با حجم بالا، ترجمه و پردازش داده‌های ساده.

  • gemini-3-pro-image

    تصویر Gemini 3 Pro

  • nano-banana-pro-preview

    پیش‌نمایش تصویر Gemini 3 Pro

  • gemini-3.1-flash-image

    تصویر فلش جمینی ۳.۱.

  • gemini-3.5-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

  • gemini-3.6-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

  • gemini-3.7-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

  • lyria-3-clip-preview

    مدل تولید موسیقی با تأخیر کم ما برای کلیپ‌های صوتی با کیفیت بالا و کنترل دقیق ریتمیک بهینه شده است.

  • lyria-3-pro-preview

    مدل پیشرفته و کامل ما برای تولید آهنگ با درک عمیق از آهنگسازی، بهینه شده برای کنترل ساختاری دقیق و انتقال‌های پیچیده در سبک‌های مختلف موسیقی.

  • gemini-robotics-er-1.6-preview

    پیش‌نمایش Gemini Robotics-ER 1.6

  • gemini-robotics-er-2-preview

    پیش‌نمایش Gemini Robotics Embodied Reasoning 2

محتوای صوتی output_audio (اختیاری)

آخرین صدای تولید شده توسط مدل در پاسخ به درخواست فعلی. توجه: این توسط SDK اضافه شده است.

یک بلوک محتوای صوتی.

فیلدها

عدد صحیح کانال‌ها (اختیاری)

تعداد کانال‌های صوتی

رشته داده (اختیاری)

محتوای صوتی.

mime_type enum (رشته) (اختیاری)

نوع مایم صدا.

مقادیر ممکن:

  • audio/wav

    فرمت صوتی WAV

  • audio/mp3

    فرمت صوتی MP3

  • audio/aiff

    فرمت صوتی AIFF

  • audio/aac

    فرمت صوتی AAC

  • audio/ogg

    فرمت صوتی OGG

  • audio/flac

    فرمت صوتی FLAC

  • audio/mpeg

    فرمت صوتی MPEG

  • audio/m4a

    فرمت صوتی M4A

  • audio/l16

    فرمت صوتی L16

  • audio/opus

    فرمت صوتی OPUS

  • audio/alaw

    فرمت صوتی ALAW

  • audio/mulaw

    فرمت صوتی MULAW

عدد صحیح sample_rate (اختیاری)

نرخ نمونه‌برداری صدا.

نوع شیء (اختیاری)

هیچ توضیحی ارائه نشده است.

همیشه روی "audio" تنظیم شود.

رشته uri (اختیاری)

آدرس اینترنتی (URI) فایل صوتی.

محتوای تصویر خروجی (اختیاری)

آخرین تصویری که توسط مدل در پاسخ به درخواست فعلی تولید شده است. توجه: این تصویر توسط SDK اضافه شده است.

رشته‌ی output_text (اختیاری)

متن به هم پیوسته از آخرین خروجی مدل در پاسخ به درخواست فعلی. توجه: این توسط SDK اضافه شده است.

خروجی_ویدئو محتوای ویدیویی (اختیاری)

آخرین ویدیویی که توسط مدل در پاسخ به درخواست فعلی تولید شده است. توجه: این توسط SDK اضافه شده است.

یک بلوک محتوای ویدیویی.

فیلدها

رشته داده (اختیاری)

محتوای ویدیویی.

mime_type enum (رشته) (اختیاری)

نوع میم (شبیه‌سازی) ویدیو.

مقادیر ممکن:

  • video/mp4

    فرمت ویدیویی MP4

  • video/mpeg

    فرمت ویدیویی MPEG

  • video/mpg

    فرمت ویدیویی MPG

  • video/mov

    فرمت ویدیویی MOV

  • video/avi

    فرمت ویدیویی AVI

  • video/x-flv

    فرمت ویدیویی FLV

  • video/webm

    فرمت ویدیویی وب‌ام

  • video/wmv

    فرمت ویدیویی WMV

  • video/3gpp

    فرمت ویدیویی 3GPP

پردازش MediaProcessing یا enum (رشته‌ای) (اختیاری)

چگونه مدل این ویدیو را برای درک پردازش می‌کند.

وضوح تصویر MediaResolution (اختیاری)

قطعنامه رسانه‌ها.

مقادیر ممکن

  • low

    وضوح پایین.

  • medium

    وضوح متوسط.

  • high

    وضوح بالا.

  • ultra_high

    وضوح فوق العاده بالا.

نوع شیء (اختیاری)

هیچ توضیحی ارائه نشده است.

همیشه روی "video" تنظیم شود.

رشته uri (اختیاری)

آدرس اینترنتی (URI) ویدیو.

رشته‌ی previous_interaction_id (اختیاری)

شناسه‌ی تعامل قبلی، در صورت وجود.

response_format فرمت پاسخ یا آرایه ( ResponseFormat ) (اختیاری)

تأکید می‌کند که پاسخ تولید شده یک شیء JSON است که با طرحواره JSON مشخص شده در این فیلد مطابقت دارد.

آرایه safety_settings (SafetySetting) (اختیاری)

تنظیمات ایمنی برای تعامل.

یک تنظیم ایمنی که بر رفتار مسدود کردن ایمنی تأثیر می‌گذارد. یک تنظیم ایمنی شامل یک دسته آسیب و یک آستانه برای آن دسته است.

فیلدها

متد enum (رشته) (اختیاری)

اختیاری. روش مسدود کردن محتوا. اگر مشخص نشده باشد، رفتار پیش‌فرض استفاده از امتیاز احتمال است.

مقادیر ممکن:

  • severity

    روش بلوک آسیب از هر دو امتیاز احتمال و شدت استفاده می‌کند.

  • probability

    روش بلوک آسیب از امتیاز احتمال استفاده می‌کند.

شمارش آستانه (رشته) (اختیاری)

الزامی. آستانه‌ی مسدود کردن محتوا. اگر احتمال آسیب از این آستانه بیشتر شود، محتوا مسدود خواهد شد.

مقادیر ممکن:

  • block_low_and_above

    مسدود کردن محتوا با احتمال آسیب کم یا بیشتر.

  • block_medium_and_above

    محتوایی را که احتمال آسیب‌رسانی آن متوسط ​​یا بالاتر است، مسدود کنید.

  • block_only_high

    مسدود کردن محتوا با احتمال آسیب بالا.

  • block_none

    صرف نظر از احتمال آسیب‌رسانی آن، هیچ محتوایی را مسدود نکنید.

  • off

    فیلتر ایمنی را کاملاً خاموش کنید.

نوع آسیب (اختیاری)

الزامی. نوع دسته‌بندی آسیبی که باید مسدود شود.

مقادیر ممکن

  • hate_speech

    محتوایی که خشونت را ترویج می‌دهد یا بر اساس ویژگی‌های خاص، نفرت علیه افراد یا گروه‌ها را برمی‌انگیزد.

  • dangerous_content

    محتوایی که فعالیت‌های خطرناک را ترویج، تسهیل یا امکان‌پذیر می‌کند.

  • harassment

    محتوای توهین‌آمیز، تهدیدآمیز یا با هدف قلدری، شکنجه یا تمسخر.

  • sexually_explicit

    محتوایی که حاوی مطالب جنسی و غیراخلاقی باشد.

  • civic_integrity

    منسوخ شده: فیلتر انتخابات دیگر پشتیبانی نمی‌شود. دسته‌بندی آسیب، سلامت مدنی است.

  • image_hate

    تصاویری که حاوی نفرت‌پراکنی هستند.

  • image_dangerous_content

    تصاویری که حاوی محتوای خطرناک هستند.

  • image_harassment

    تصاویری که حاوی آزار و اذیت هستند.

  • image_sexually_explicit

    تصاویری که حاوی محتوای جنسی هستند.

  • jailbreak

    اعلان‌هایی که برای دور زدن فیلترهای ایمنی طراحی شده‌اند.

service_tier لایه سرویس (اختیاری)

لایه سرویس برای تعامل.

مقادیر ممکن

  • flex

    سطح خدمات انعطاف‌پذیر.

  • standard

    سطح خدمات استاندارد.

  • priority

    ردیف خدمات اولویت‌دار.

  • deferred

    ردیف خدمات معوق.

شمارش وضعیت (رشته) (اختیاری)

الزامی. فقط خروجی. وضعیت تعامل.

مقادیر ممکن:

  • in_progress

    تعامل در حال انجام است.

  • requires_action

    این تعامل نیاز به اقدام/ورودی از سوی کاربر دارد.

  • completed

    تعامل تکمیل شده است.

  • failed

    تعامل شکست خورد.

  • cancelled

    تعامل لغو شد.

  • incomplete

    تعامل تکمیل شده است، اما شامل نتایج ناقص است (مثلاً رسیدن به max_tokens).

  • budget_exceeded

    تعامل متوقف شد زیرا بودجه توکن از حد مجاز فراتر رفته بود.

  • queued

    تعامل در صف انتظار پردازش قرار می‌گیرد.

آرایه گام‌ها ( Step ) (اختیاری)

فقط خروجی. مراحلی که تعامل را تشکیل می‌دهند، زمانی که در پاسخ گنجانده شوند.

رشته system_instruction (اختیاری)

دستورالعمل سیستم برای تعامل.

آرایه ابزارها ( ابزار ) (اختیاری)

فهرستی از اعلان‌های ابزار که مدل ممکن است در طول تعامل فراخوانی کند.

رشته به‌روزرسانی‌شده (اختیاری)

فقط خروجی. زمانی که پاسخ آخرین بار در قالب ISO 8601 (YYYY-MM-DDThh:mm:ssZ) به‌روزرسانی شده است.

کاربرد (اختیاری )

فقط خروجی. آمار مربوط به میزان استفاده از توکن درخواست تعامل.

آمار مربوط به میزان استفاده از توکن درخواست تعامل.

فیلدها

آرایه cached_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک میزان استفاده از توکن‌های ذخیره‌شده بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه grounding_tool_count (GroundingToolCount) (اختیاری)

تعداد ابزار اتصال به زمین

تعداد ابزار اتصال به زمین مهم است.

فیلدها

شمارش عدد صحیح (اختیاری)

تعداد ابزار اتصال به زمین مهم است.

نوع enum (رشته) (اختیاری)

نوع ابزار اتصال زمین مرتبط با شمارش.

مقادیر ممکن:

  • google_search

    اتصال به زمین با جستجوی وب و جستجوی تصویر گوگل، و اتصال به زمین وب برای سازمان‌ها.

  • google_maps

    اتصال به زمین با نقشه‌های گوگل.

  • retrieval

    پایه گذاری با داده های مشتری، به عنوان مثال، VertexAISearch.

آرایه input_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک استفاده از توکن ورودی بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه output_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک استفاده از توکن خروجی بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه tool_use_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک میزان استفاده از توکن‌های ابزار بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

عدد صحیح total_cached_tokens (اختیاری)

تعداد توکن‌ها در بخش ذخیره‌شده‌ی اعلان (محتوای ذخیره‌شده).

عدد صحیح total_input_tokens (اختیاری)

تعداد توکن‌ها در اعلان (زمینه).

total_output_tokens عدد صحیح (اختیاری)

تعداد کل توکن‌ها در تمام پاسخ‌های تولید شده.

total_thought_tokens عدد صحیح (اختیاری)

تعداد توکن‌های افکار برای مدل‌های تفکر.

عدد صحیح total_tokens (اختیاری)

تعداد کل توکن‌ها برای درخواست تعامل (درخواست + پاسخ‌ها + سایر توکن‌های داخلی).

عدد صحیح total_tool_use_tokens (اختیاری)

تعداد توکن‌های موجود در اعلان(های) استفاده از ابزار.

webhook_config پیکربندی وب هوک (اختیاری)

اختیاری. پیکربندی وب‌هوک برای دریافت اعلان‌ها پس از اتمام تعامل.

پیام مربوط به پیکربندی رویدادهای وب‌هوک برای یک درخواست.

فیلدها

آرایه uris (رشته) (اختیاری)

اختیاری. در صورت تنظیم، این URLهای وب‌هوک به جای وب‌هوک‌های ثبت‌شده، برای رویدادهای وب‌هوک استفاده می‌شوند.

شیء user_metadata (اختیاری)

اختیاری. فراداده کاربر که در هر انتشار رویداد به وب‌هوک‌ها بازگردانده می‌شود.

مثال‌ها

مثال

{
  "created": "2025-12-04T15:01:45Z",
  "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "Hello! I'm doing well, functioning as expected. Thank you for asking! How are you doing today?"
        }
      ]
    }
  ],
  "updated": "2025-12-04T15:01:45Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 7
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 7,
    "total_output_tokens": 23,
    "total_thought_tokens": 49,
    "total_tokens": 79,
    "total_tool_use_tokens": 0
  }
}

مدل‌های داده

محتوا

محتوای پاسخ.

انواع ممکن

محتوای صوتی

یک بلوک محتوای صوتی.

عدد صحیح کانال‌ها (اختیاری)

The number of audio channels.

data string (optional)

The audio content.

mime_type enum (string) (optional)

The mime type of the audio.

Possible values:

  • audio/wav

    WAV audio format

  • audio/mp3

    MP3 audio format

  • audio/aiff

    AIFF audio format

  • audio/aac

    AAC audio format

  • audio/ogg

    OGG audio format

  • audio/flac

    FLAC audio format

  • audio/mpeg

    MPEG audio format

  • audio/m4a

    M4A audio format

  • audio/l16

    L16 audio format

  • audio/opus

    OPUS audio format

  • audio/alaw

    ALAW audio format

  • audio/mulaw

    MULAW audio format

sample_rate integer (optional)

The sample rate of the audio.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "audio" .

uri string (optional)

The URI of the audio.

DocumentContent

A document content block.

data string (optional)

The document content.

mime_type enum (string) (optional)

The mime type of the document.

Possible values:

  • application/pdf

    PDF document format

  • text/csv

    CSV document format

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "document" .

uri string (optional)

The URI of the document.

ImageContent

An image content block.

data string (optional)

The image content.

mime_type enum (string) (optional)

The mime type of the image.

Possible values:

  • image/png

    PNG image format

  • image/jpeg

    JPEG image format

  • image/webp

    WebP image format

  • image/heic

    HEIC image format

  • image/heif

    HEIF image format

  • image/gif

    GIF image format

  • image/bmp

    BMP image format

  • image/tiff

    TIFF image format

resolution MediaResolution (optional)

The resolution of the media.

Possible values

  • low

    Low resolution.

  • medium

    Medium resolution.

  • high

    High resolution.

  • ultra_high

    Ultra high resolution.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "image" .

uri string (optional)

The URI of the image.

TextContent

A text content block.

annotations array (Annotation) (optional)

Citation information for model-generated content.

Citation information for model-generated content.

Possible Types

FileCitation

A file citation annotation.

custom_metadata object (optional)

User provided metadata about the retrieved context.

document_uri string (optional)

The URI of the file.

end_index integer (optional)

End of the attributed segment, exclusive.

file_name string (optional)

The name of the file.

media_id string (optional)

Media ID in-case of image citations, if applicable.

page_number integer (optional)

Page number of the cited document, if applicable.

source string (optional)

Source attributed for a portion of the text.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "file_citation" .

PlaceCitation

A place citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

name string (optional)

Title of the place.

place_id string (optional)

The ID of the place, in `places/{place_id}` format.

review_snippets array (ReviewSnippet) (optional)

Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.

Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.

فیلدها

review_id string (optional)

The ID of the review snippet.

title string (optional)

Title of the review.

url string (optional)

A link that corresponds to the user review on Google Maps.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "place_citation" .

url string (optional)

URI reference of the place.

UrlCitation

A URL citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

title string (optional)

The title of the URL.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "url_citation" .

url string (optional)

The URL.

WordInfo

Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.

end_index integer (optional)

End of the attributed segment, exclusive.

end_offset string (optional)

End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".

speaker string (optional)

Optional. Speaker label for this word (eg "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

start_offset string (optional)

Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".

text string (optional)

The transcribed word.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "word_info" .

text string (required)

Required. The text content.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "text" .

VideoContent

A video content block.

data string (optional)

The video content.

mime_type enum (string) (optional)

The mime type of the video.

Possible values:

  • video/mp4

    MP4 video format

  • video/mpeg

    MPEG video format

  • video/mpg

    MPG video format

  • video/mov

    MOV video format

  • video/avi

    AVI video format

  • video/x-flv

    FLV video format

  • video/webm

    WebM video format

  • video/wmv

    WMV video format

  • video/3gpp

    3GPP video format

processing MediaProcessing or enum (string) (optional)

How the model processes this video for understanding.

resolution MediaResolution (optional)

The resolution of the media.

Possible values

  • low

    Low resolution.

  • medium

    Medium resolution.

  • high

    High resolution.

  • ultra_high

    Ultra high resolution.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "video" .

uri string (optional)

The URI of the video.

مثال‌ها

صوتی

{
  "type": "audio",
  "data": "BASE64_ENCODED_AUDIO",
  "mime_type": "audio/wav"
}

سند

{
  "type": "document",
  "data": "BASE64_ENCODED_DOCUMENT",
  "mime_type": "application/pdf"
}

تصویر

{
  "type": "image",
  "data": "BASE64_ENCODED_IMAGE",
  "mime_type": "image/png"
}

متن

{
  "type": "text",
  "text": "Hello, how are you?"
}

ویدئو

{
  "type": "video",
  "uri": "https://www.youtube.com/watch?v=9hE5-98ZeCg"
}

ابزار

A tool that can be used by the model.

Possible Types

CodeExecution

A tool that can be used by the model to execute code.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "code_execution" .

ComputerUse

A tool that can be used by the model to interact with the computer.

disabled_safety_policies array (enum (string)) (optional)

Optional. Disabled safety policies for computer use.

Possible values:

  • financial_transactions

    Safety policy for financial transactions.

  • sensitive_data_modification

    Safety policy for sensitive data modification.

  • communication_tool

    Safety policy for communication tools (eg Gmail, Chat, Meet).

  • account_creation

    Safety policy for account creation.

  • data_modification

    Safety policy for data modification.

  • user_consent_management

    Safety policy for user consent management.

  • legal_terms_and_agreements

    Safety policy for legal terms and agreements.

enable_prompt_injection_detection boolean (optional)

Whether enable the prompt injection detection check on computer-use request.

environment enum (string) (optional)

The environment being operated.

Possible values:

  • browser

    Operates in a web browser.

  • mobile

    Operates in a mobile environment.

  • desktop

    Operates in a desktop environment.

excluded_predefined_functions array (string) (optional)

The list of predefined functions that are excluded from the model call.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "computer_use" .

FileSearch

A tool that can be used by the model to search files.

file_search_store_names array (string) (optional)

The file search store names to search.

metadata_filter string (optional)

Metadata filter to apply to the semantic retrieval documents and chunks.

top_k integer (optional)

The number of semantic retrieval chunks to retrieve.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "file_search" .

عملکرد

A tool that can be used by the model.

description string (optional)

A description of the function.

name string (optional)

The name of the function.

parameters object (optional)

The JSON Schema for the function's parameters.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "function" .

GoogleMaps

A tool that can be used by the model to call Google Maps.

enable_widget boolean (optional)

Whether to return a widget context token in the tool call result of the response.

latitude number (optional)

The latitude of the user's location.

longitude number (optional)

The longitude of the user's location.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "google_maps" .

GoogleSearch

A tool that can be used by the model to search Google.

search_types array (enum (string)) (optional)

The types of search grounding to enable.

Possible values:

  • web_search

    Setting this field enables web search. Only text results are returned.

  • image_search

    Setting this field enables image search. Image bytes are returned.

  • enterprise_web_search

    Setting this field enables enterprise web search.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "google_search" .

McpServer

A MCPServer is a server that can be called by the model to perform actions.

allowed_tools array (AllowedTools) (optional)

The allowed tools.

The configuration for allowed tools.

فیلدها

mode enum (string) (optional)

The mode of the tool choice.

Possible values:

  • auto

    Auto tool choice.

  • any

    Any tool choice.

  • none

    No tool choice.

  • validated

    Validated tool choice.

tools array (string) (optional)

The names of the allowed tools.

headers object (optional)

Optional: Fields for authentication headers, timeouts, etc., if needed.

name string (optional)

The name of the MCPServer.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "mcp_server" .

url string (optional)

The full URL for the MCPServer endpoint. Example: "https://api.example.com/mcp"

بازیابی

A tool that can be used by the model to retrieve files.

exa_ai_search_config ExaAISearchConfig (optional)

Used to specify configuration for ExaAISearch.

Used to specify configuration for ExaAISearch.

فیلدها

api_key string (optional)

Required. The API key for ExaAiSearch.

custom_config object (optional)

Optional. This field can be used to pass any parameter from the Exa.ai Search API.

parallel_ai_search_config ParallelAISearchConfig (optional)

Used to specify configuration for ParallelAISearch.

Used to specify configuration for ParallelAISearch.

فیلدها

api_key string (optional)

Optional. The API key for ParallelAiSearch.

custom_config object (optional)

Optional. Custom configs for ParallelAiSearch.

rag_store_config RagStoreConfig (optional)

Used to specify configuration for RagStore.

Use to specify configuration for RAG Store.

فیلدها

rag_resources array (RagResource) (optional)

Optional. The representation of the rag source.

The definition of the Rag resource.

فیلدها

rag_corpus string (optional)

Optional. RagCorpora resource name.

rag_file_ids array (string) (optional)

Optional. rag_file_id. The files should be in the same rag_corpus set in rag_corpus field.

rag_retrieval_config RagRetrievalConfig (optional)

Optional. The retrieval config for the Rag query.

Specifies the context retrieval config.

فیلدها

filter Filter (optional)

Optional. Config for filters.

Config for filters.

فیلدها

metadata_filter string (optional)

Optional. String for metadata filtering.

vector_distance_threshold number (optional)

Optional. Only returns contexts with vector distance smaller than the threshold.

vector_similarity_threshold number (optional)

Optional. Only returns contexts with vector similarity larger than the threshold.

hybrid_search HybridSearch (optional)

Optional. Config for Hybrid Search.

Config for Hybrid Search.

فیلدها

alpha number (optional)

Optional. Alpha value controls the weight between dense and sparse vector search results.

ranking Ranking (optional)

Optional. Config for ranking and reranking.

Config for ranking and reranking.

top_k integer (optional)

Optional. The number of contexts to retrieve.

retrieval_types array (enum (string)) (optional)

The types of file retrieval to enable.

Possible values:

  • rag_store
  • exa_ai_search
  • parallel_ai_search
type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "retrieval" .

UrlContext

A tool that can be used by the model to fetch URL context.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "url_context" .

مثال‌ها

CodeExecution

ComputerUse

FileSearch

عملکرد

GoogleMaps

GoogleSearch

McpServer

بازیابی

No examples available for this type.

UrlContext

InteractionSseEvent

Possible Types

Polymorphic discriminator: event_type

ErrorEvent

error Error (optional)

هیچ توضیحی ارائه نشده است.

Error message from an interaction.

فیلدها

code string (optional)

A URI that identifies the error type.

message string (optional)

A human-readable error message.

event_id string (optional)

The event_id token to be used to resume the interaction stream, from this event.

event_type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "error" .

InteractionCompletedEvent

event_id string (optional)

The event_id token to be used to resume the interaction stream, from this event.

event_type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "interaction.completed" .

interaction InteractionSseEventInteraction (required)

Partial completed interaction resource emitted at the end of the stream.

Partial interaction resource emitted by interaction lifecycle SSE events. Streaming lifecycle payloads may omit fields that are only available on full non-streaming Interaction responses.

فیلدها

agent string (optional)

The agent to interact with.

created string (optional)

Output only. The time at which the response was created in ISO 8601 format.

id string (optional)

Required. Output only. A unique identifier for the interaction completion.

model string (optional)

The model that will complete your prompt.

object string (optional)

Output only. The resource type.

service_tier ServiceTier (optional)

The service tier for the interaction.

Possible values

  • flex

    Flex service tier.

  • standard

    Standard service tier.

  • priority

    Priority service tier.

  • deferred

    Deferred service tier.

status enum (string) (optional)

Required. Output only. The status of the interaction.

Possible values:

  • in_progress

    The interaction is in progress.

  • requires_action

    The interaction requires action/input from the user.

  • completed

    The interaction is completed.

  • failed

    The interaction failed.

  • cancelled

    The interaction was cancelled.

  • incomplete

    The interaction is completed, but contains incomplete results (eg hitting max_tokens).

steps array ( Step ) (optional)

Output only. The steps that make up the interaction, if included in this event.

updated string (optional)

Output only. The time at which the response was last updated in ISO 8601 format.

usage Usage (optional)

Output only. Statistics on the interaction request's token usage.

Statistics on the interaction request's token usage.

فیلدها

cached_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of cached token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

grounding_tool_count array (GroundingToolCount) (optional)

Grounding tool count.

The number of grounding tool counts.

فیلدها

count integer (optional)

The number of grounding tool counts.

type enum (string) (optional)

The grounding tool type associated with the count.

Possible values:

  • google_search

    Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.

  • google_maps

    Grounding with Google Maps.

  • retrieval

    Grounding with customer's data, for example, VertexAISearch.

input_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of input token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

output_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of output token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

tool_use_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of tool-use token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

total_cached_tokens integer (optional)

Number of tokens in the cached part of the prompt (the cached content).

total_input_tokens integer (optional)

Number of tokens in the prompt (context).

total_output_tokens integer (optional)

Total number of tokens across all the generated responses.

total_thought_tokens integer (optional)

Number of tokens of thoughts for thinking models.

total_tokens integer (optional)

Total token count for the interaction request (prompt + responses + other internal tokens).

total_tool_use_tokens integer (optional)

Number of tokens present in tool-use prompt(s).

InteractionCreatedEvent

event_id string (optional)

The event_id token to be used to resume the interaction stream, from this event.

event_type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "interaction.created" .

interaction InteractionSseEventInteraction (required)

Partial interaction resource emitted when the stream is created.

Partial interaction resource emitted by interaction lifecycle SSE events. Streaming lifecycle payloads may omit fields that are only available on full non-streaming Interaction responses.

فیلدها

agent string (optional)

The agent to interact with.

created string (optional)

Output only. The time at which the response was created in ISO 8601 format.

id string (optional)

Required. Output only. A unique identifier for the interaction completion.

model string (optional)

The model that will complete your prompt.

object string (optional)

Output only. The resource type.

service_tier ServiceTier (optional)

The service tier for the interaction.

Possible values

  • flex

    Flex service tier.

  • standard

    Standard service tier.

  • priority

    Priority service tier.

  • deferred

    Deferred service tier.

status enum (string) (optional)

Required. Output only. The status of the interaction.

Possible values:

  • in_progress

    The interaction is in progress.

  • requires_action

    The interaction requires action/input from the user.

  • completed

    The interaction is completed.

  • failed

    The interaction failed.

  • cancelled

    The interaction was cancelled.

  • incomplete

    The interaction is completed, but contains incomplete results (eg hitting max_tokens).

steps array ( Step ) (optional)

Output only. The steps that make up the interaction, if included in this event.

updated string (optional)

Output only. The time at which the response was last updated in ISO 8601 format.

usage Usage (optional)

Output only. Statistics on the interaction request's token usage.

Statistics on the interaction request's token usage.

فیلدها

cached_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of cached token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

grounding_tool_count array (GroundingToolCount) (optional)

Grounding tool count.

The number of grounding tool counts.

فیلدها

count integer (optional)

The number of grounding tool counts.

type enum (string) (optional)

The grounding tool type associated with the count.

Possible values:

  • google_search

    Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.

  • google_maps

    Grounding with Google Maps.

  • retrieval

    Grounding with customer's data, for example, VertexAISearch.

input_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of input token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

output_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of output token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

tool_use_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of tool-use token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

total_cached_tokens integer (optional)

Number of tokens in the cached part of the prompt (the cached content).

total_input_tokens integer (optional)

Number of tokens in the prompt (context).

total_output_tokens integer (optional)

Total number of tokens across all the generated responses.

total_thought_tokens integer (optional)

Number of tokens of thoughts for thinking models.

total_tokens integer (optional)

Total token count for the interaction request (prompt + responses + other internal tokens).

total_tool_use_tokens integer (optional)

Number of tokens present in tool-use prompt(s).

InteractionStatusUpdate

event_id string (optional)

The event_id token to be used to resume the interaction stream, from this event.

event_type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "interaction.status_update" .

interaction_id string (required)

هیچ توضیحی ارائه نشده است.

status enum (string) (required)

هیچ توضیحی ارائه نشده است.

Possible values:

  • in_progress

    The interaction is in progress.

  • requires_action

    The interaction requires action/input from the user.

  • completed

    The interaction is completed.

  • failed

    The interaction failed.

  • cancelled

    The interaction was cancelled.

  • incomplete

    The interaction is completed, but contains incomplete results (eg hitting max_tokens).

  • budget_exceeded

    The interaction was halted because the token budget was exceeded.

  • queued

    The interaction is queued, waiting for processing (eg waiting for off-peak capacity).

StepDelta

delta StepDeltaData (required)

هیچ توضیحی ارائه نشده است.

Possible Types

ArgumentsDelta

arguments string (optional)

هیچ توضیحی ارائه نشده است.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "arguments_delta" .

AudioDelta

channels integer (optional)

The number of audio channels.

data string (optional)

هیچ توضیحی ارائه نشده است.

mime_type enum (string) (optional)

هیچ توضیحی ارائه نشده است.

Possible values:

  • audio/wav

    WAV audio format

  • audio/mp3

    MP3 audio format

  • audio/aiff

    AIFF audio format

  • audio/aac

    AAC audio format

  • audio/ogg

    OGG audio format

  • audio/flac

    FLAC audio format

  • audio/mpeg

    MPEG audio format

  • audio/m4a

    M4A audio format

  • audio/l16

    L16 audio format

  • audio/opus

    OPUS audio format

  • audio/alaw

    ALAW audio format

  • audio/mulaw

    MULAW audio format

sample_rate integer (optional)

The sample rate of the audio.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "audio" .

uri string (optional)

هیچ توضیحی ارائه نشده است.

CodeExecutionCallDelta

arguments CodeExecutionCallArguments (required)

هیچ توضیحی ارائه نشده است.

The arguments to pass to the code execution.

فیلدها

code string (optional)

The code to be executed.

language enum (string) (optional)

Programming language of the `code`.

Possible values:

  • python

    Python >= 3.10, with numpy and simpy available.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "code_execution_call" .

CodeExecutionResultDelta

is_error boolean (optional)

هیچ توضیحی ارائه نشده است.

result string (required)

هیچ توضیحی ارائه نشده است.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "code_execution_result" .

DocumentDelta

data string (optional)

هیچ توضیحی ارائه نشده است.

mime_type enum (string) (optional)

هیچ توضیحی ارائه نشده است.

Possible values:

  • application/pdf

    PDF document format

  • text/csv

    CSV document format

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "document" .

uri string (optional)

هیچ توضیحی ارائه نشده است.

FileSearchCallDelta

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "file_search_call" .

FileSearchResultDelta

result array (FileSearchResult) (required)

هیچ توضیحی ارائه نشده است.

The result of the File Search.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "file_search_result" .

FunctionResultDelta

call_id string (required)

Required. ID to match the ID from the function call block.

is_error boolean (optional)

هیچ توضیحی ارائه نشده است.

name string (optional)

هیچ توضیحی ارائه نشده است.

result array ( ImageContent or TextContent ) or object or string (required)

هیچ توضیحی ارائه نشده است.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "function_result" .

GoogleMapsCallDelta

arguments GoogleMapsCallArguments (optional)

The arguments to pass to the Google Maps tool.

The arguments to pass to the Google Maps tool.

فیلدها

queries array (string) (optional)

The queries to be executed.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "google_maps_call" .

GoogleMapsResultDelta

result array (GoogleMapsResult) (optional)

The results of the Google Maps.

The result of the Google Maps.

فیلدها

places array (Places) (optional)

The places that were found.

فیلدها

name string (optional)

Title of the place.

place_id string (optional)

The ID of the place, in `places/{place_id}` format.

review_snippets array (ReviewSnippet) (optional)

Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.

Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.

فیلدها

review_id string (optional)

The ID of the review snippet.

title string (optional)

Title of the review.

url string (optional)

A link that corresponds to the user review on Google Maps.

url string (optional)

URI reference of the place.

widget_context_token string (optional)

Resource name of the Google Maps widget context token.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "google_maps_result" .

GoogleSearchCallDelta

arguments GoogleSearchCallArguments (required)

هیچ توضیحی ارائه نشده است.

The arguments to pass to Google Search.

فیلدها

queries array (string) (optional)

Web search queries for the following-up web search.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "google_search_call" .

GoogleSearchResultDelta

is_error boolean (optional)

هیچ توضیحی ارائه نشده است.

result array (GoogleSearchResult) (required)

هیچ توضیحی ارائه نشده است.

The result of the Google Search.

فیلدها

search_suggestions string (optional)

Web content snippet that can be embedded in a web page or an app webview.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "google_search_result" .

ImageDelta

data string (optional)

هیچ توضیحی ارائه نشده است.

mime_type enum (string) (optional)

هیچ توضیحی ارائه نشده است.

Possible values:

  • image/png

    PNG image format

  • image/jpeg

    JPEG image format

  • image/webp

    WebP image format

  • image/heic

    HEIC image format

  • image/heif

    HEIF image format

  • image/gif

    GIF image format

  • image/bmp

    BMP image format

  • image/tiff

    TIFF image format

resolution MediaResolution (optional)

The resolution of the media.

Possible values

  • low

    Low resolution.

  • medium

    Medium resolution.

  • high

    High resolution.

  • ultra_high

    Ultra high resolution.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "image" .

uri string (optional)

هیچ توضیحی ارائه نشده است.

McpServerToolCallDelta

arguments object (required)

هیچ توضیحی ارائه نشده است.

name string (required)

هیچ توضیحی ارائه نشده است.

server_name string (required)

هیچ توضیحی ارائه نشده است.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "mcp_server_tool_call" .

McpServerToolResultDelta

name string (optional)

هیچ توضیحی ارائه نشده است.

result array ( ImageContent or TextContent ) or object or string (required)

هیچ توضیحی ارائه نشده است.

server_name string (optional)

هیچ توضیحی ارائه نشده است.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "mcp_server_tool_result" .

RetrievalCallDelta

Used by Vertex Retrieval tools such as Parallel AI, Exa AI, Vertex AI Search, etc. RetrievalType decides which tool is used.

arguments RetrievalStepArguments (required)

Required. The arguments to pass to the Retrieval tool.

The arguments to pass to Retrieval tools.

فیلدها

queries array (string) (optional)

Queries for Retrieval information.

retrieval_type enum (string) (optional)

The type of retrieval tools.

Possible values:

  • rag_store

    The type of retrieval tools.

  • exa_ai_search

    The type of retrieval tools.

  • parallel_ai_search

    The type of retrieval tools.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "retrieval_call" .

RetrievalResultDelta

Used by Vertex Retrieval tools such as Parallel AI, Exa AI, Vertex AI Search, etc. ToolResultDelta.type

is_error boolean (optional)

Whether the retrieval resulted in an error.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "retrieval_result" .

TextAnnotationDelta

annotations array (Annotation) (optional)

Citation information for model-generated content.

Citation information for model-generated content.

Possible Types

FileCitation

A file citation annotation.

custom_metadata object (optional)

User provided metadata about the retrieved context.

document_uri string (optional)

The URI of the file.

end_index integer (optional)

End of the attributed segment, exclusive.

file_name string (optional)

The name of the file.

media_id string (optional)

Media ID in-case of image citations, if applicable.

page_number integer (optional)

Page number of the cited document, if applicable.

source string (optional)

Source attributed for a portion of the text.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "file_citation" .

PlaceCitation

A place citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

name string (optional)

Title of the place.

place_id string (optional)

The ID of the place, in `places/{place_id}` format.

review_snippets array (ReviewSnippet) (optional)

Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.

Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.

فیلدها

review_id string (optional)

The ID of the review snippet.

title string (optional)

Title of the review.

url string (optional)

A link that corresponds to the user review on Google Maps.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "place_citation" .

url string (optional)

URI reference of the place.

UrlCitation

A URL citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

title string (optional)

The title of the URL.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "url_citation" .

url string (optional)

The URL.

WordInfo

Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.

end_index integer (optional)

End of the attributed segment, exclusive.

end_offset string (optional)

End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".

speaker string (optional)

Optional. Speaker label for this word (eg "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

start_offset string (optional)

Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".

text string (optional)

The transcribed word.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "word_info" .

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "text_annotation_delta" .

TextDelta

text string (required)

هیچ توضیحی ارائه نشده است.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "text" .

ThoughtSignatureDelta

signature string (optional)

Signature to match the backend source to be part of the generation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "thought_signature" .

ThoughtSummaryDelta

content Content (optional)

A new summary item to be added to the thought.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "thought_summary" .

UrlContextCallDelta

arguments UrlContextCallArguments (required)

هیچ توضیحی ارائه نشده است.

The arguments to pass to the URL context.

فیلدها

urls array (string) (optional)

The URLs to fetch.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "url_context_call" .

UrlContextResultDelta

is_error boolean (optional)

هیچ توضیحی ارائه نشده است.

result array (UrlContextResult) (required)

هیچ توضیحی ارائه نشده است.

The result of the URL context.

فیلدها

status enum (string) (optional)

The status of the URL retrieval.

Possible values:

  • success

    Url retrieval is successful.

  • error

    Url retrieval is failed due to error.

  • paywall

    Url retrieval is failed because the content is behind paywall.

  • unsafe

    Url retrieval is failed because the content is unsafe.

url string (optional)

The URL that was fetched.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "url_context_result" .

VideoDelta

data string (optional)

هیچ توضیحی ارائه نشده است.

mime_type enum (string) (optional)

هیچ توضیحی ارائه نشده است.

Possible values:

  • video/mp4

    MP4 video format

  • video/mpeg

    MPEG video format

  • video/mpg

    MPG video format

  • video/mov

    MOV video format

  • video/avi

    AVI video format

  • video/x-flv

    FLV video format

  • video/webm

    WebM video format

  • video/wmv

    WMV video format

  • video/3gpp

    3GPP video format

  • video/jpeg2000

    JPEG 2000 video format

resolution MediaResolution (optional)

The resolution of the media.

Possible values

  • low

    Low resolution.

  • medium

    Medium resolution.

  • high

    High resolution.

  • ultra_high

    Ultra high resolution.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "video" .

uri string (optional)

هیچ توضیحی ارائه نشده است.

event_id string (optional)

The event_id token to be used to resume the interaction stream, from this event.

event_type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "step.delta" .

index integer (required)

هیچ توضیحی ارائه نشده است.

metadata StepDeltaMetadata (optional)

هیچ توضیحی ارائه نشده است.

Optional metadata accompanying ANY streamed event.

فیلدها

total_usage Usage (optional)

Statistics on the interaction request's token usage.

Statistics on the interaction request's token usage.

فیلدها

cached_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of cached token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

grounding_tool_count array (GroundingToolCount) (optional)

Grounding tool count.

The number of grounding tool counts.

فیلدها

count integer (optional)

The number of grounding tool counts.

type enum (string) (optional)

The grounding tool type associated with the count.

Possible values:

  • google_search

    Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.

  • google_maps

    Grounding with Google Maps.

  • retrieval

    Grounding with customer's data, for example, VertexAISearch.

input_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of input token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

output_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of output token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

tool_use_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of tool-use token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

total_cached_tokens integer (optional)

Number of tokens in the cached part of the prompt (the cached content).

total_input_tokens integer (optional)

Number of tokens in the prompt (context).

total_output_tokens integer (optional)

Total number of tokens across all the generated responses.

total_thought_tokens integer (optional)

Number of tokens of thoughts for thinking models.

total_tokens integer (optional)

Total token count for the interaction request (prompt + responses + other internal tokens).

total_tool_use_tokens integer (optional)

Number of tokens present in tool-use prompt(s).

StepStart

event_id string (optional)

The event_id token to be used to resume the interaction stream, from this event.

event_type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "step.start" .

index integer (required)

هیچ توضیحی ارائه نشده است.

step Step (required)

هیچ توضیحی ارائه نشده است.

StepStop

event_id string (optional)

The event_id token to be used to resume the interaction stream, from this event.

event_type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "step.stop" .

index integer (required)

هیچ توضیحی ارائه نشده است.

step_usage Usage (optional)

Model usage stats for this specific step.

Statistics on the interaction request's token usage.

فیلدها

cached_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of cached token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

grounding_tool_count array (GroundingToolCount) (optional)

Grounding tool count.

The number of grounding tool counts.

فیلدها

count integer (optional)

The number of grounding tool counts.

type enum (string) (optional)

The grounding tool type associated with the count.

Possible values:

  • google_search

    Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.

  • google_maps

    Grounding with Google Maps.

  • retrieval

    Grounding with customer's data, for example, VertexAISearch.

input_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of input token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

output_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of output token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

tool_use_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of tool-use token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

total_cached_tokens integer (optional)

Number of tokens in the cached part of the prompt (the cached content).

total_input_tokens integer (optional)

Number of tokens in the prompt (context).

total_output_tokens integer (optional)

Total number of tokens across all the generated responses.

total_thought_tokens integer (optional)

Number of tokens of thoughts for thinking models.

total_tokens integer (optional)

Total token count for the interaction request (prompt + responses + other internal tokens).

total_tool_use_tokens integer (optional)

Number of tokens present in tool-use prompt(s).

usage Usage (optional)

Cumulative model usage stats from the start of the session.

Statistics on the interaction request's token usage.

فیلدها

cached_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of cached token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

grounding_tool_count array (GroundingToolCount) (optional)

Grounding tool count.

The number of grounding tool counts.

فیلدها

count integer (optional)

The number of grounding tool counts.

type enum (string) (optional)

The grounding tool type associated with the count.

Possible values:

  • google_search

    Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.

  • google_maps

    Grounding with Google Maps.

  • retrieval

    Grounding with customer's data, for example, VertexAISearch.

input_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of input token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

output_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of output token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

tool_use_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of tool-use token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

total_cached_tokens integer (optional)

Number of tokens in the cached part of the prompt (the cached content).

total_input_tokens integer (optional)

Number of tokens in the prompt (context).

total_output_tokens integer (optional)

Total number of tokens across all the generated responses.

total_thought_tokens integer (optional)

Number of tokens of thoughts for thinking models.

total_tokens integer (optional)

Total token count for the interaction request (prompt + responses + other internal tokens).

total_tool_use_tokens integer (optional)

Number of tokens present in tool-use prompt(s).

مثال‌ها

Error Event

{
  "error": {
    "code": "not_found",
    "message": "Failed to get completed interaction: Result not found."
  },
  "event_type": "error"
}

Interaction Completed

{
  "event_id": "evt_123",
  "event_type": "interaction.completed",
  "interaction": {
    "created": "2025-12-04T15:01:45Z",
    "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg",
    "model": "gemini-3.6-flash",
    "status": "completed",
    "updated": "2025-12-04T15:01:45Z"
  }
}

Interaction Completed

{
  "event_id": "evt_123",
  "event_type": "interaction.completed",
  "interaction": {
    "created": "2025-12-04T15:01:45Z",
    "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg",
    "model": "gemini-3-flash-preview",
    "object": "interaction",
    "status": "completed",
    "updated": "2025-12-04T15:01:45Z"
  }
}

Interaction Created

{
  "event_id": "evt_123",
  "event_type": "interaction.created",
  "interaction": {
    "created": "2025-12-04T15:01:45Z",
    "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg",
    "model": "gemini-3.6-flash",
    "status": "in_progress",
    "updated": "2025-12-04T15:01:45Z"
  }
}

Interaction Created

{
  "event_id": "evt_123",
  "event_type": "interaction.created",
  "interaction": {
    "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg",
    "model": "gemini-3-flash-preview",
    "object": "interaction",
    "status": "in_progress"
  }
}

Interaction Status Update

{
  "event_type": "interaction.status_update",
  "interaction_id": "v1_ChdTMjQ0YWJ5TUF1TzcxZThQdjRpcnFRcxIXUzI0NGFieU1BdU83MWU4UHY0aXJxUXM",
  "status": "in_progress"
}

Step Delta

{
  "delta": {
    "type": "text",
    "text": "Hello"
  },
  "event_type": "step.delta",
  "index": 0
}

Step Start

{
  "event_type": "step.start",
  "index": 0,
  "step": {
    "type": "model_output"
  }
}

Step Stop

{
  "event_type": "step.stop",
  "index": 0
}

ResponseFormat

Possible Types

AudioResponseFormat

Configuration for audio output format.

bit_rate integer (optional)

Bit rate in bits per second (bps). Only applicable for compressed formats (MP3, Opus).

delivery enum (string) (optional)

The delivery mode for the audio output.

Possible values:

  • inline

    Audio data is returned inline in the response.

  • uri

    Audio data is returned as a URI.

mime_type enum (string) (optional)

The MIME type of the audio output.

Possible values:

  • audio/mp3

    MP3 audio format.

  • audio/ogg_opus

    OGG Opus audio format.

  • audio/l16

    Raw PCM (L16) audio format.

  • audio/wav

    WAV audio format.

  • audio/alaw

    A-law audio format.

  • audio/mulaw

    Mu-law audio format.

sample_rate integer (optional)

Sample rate in Hz.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "audio" .

ImageResponseFormat

Configuration for image output format.

aspect_ratio enum (string) (optional)

The aspect ratio for the image output.

Possible values:

  • 1:1

    1:1 aspect ratio.

  • 2:3

    2:3 aspect ratio.

  • 3:2

    3:2 aspect ratio.

  • 3:4

    3:4 aspect ratio.

  • 4:3

    4:3 aspect ratio.

  • 4:5

    4:5 aspect ratio.

  • 5:4

    5:4 aspect ratio.

  • 9:16

    9:16 aspect ratio.

  • 16:9

    16:9 aspect ratio.

  • 21:9

    21:9 aspect ratio.

  • 1:8

    1:8 aspect ratio.

  • 8:1

    8:1 aspect ratio.

  • 1:4

    1:4 aspect ratio.

  • 4:1

    4:1 aspect ratio.

delivery enum (string) (optional)

The delivery mode for the image output.

Possible values:

  • inline

    Image data is returned inline in the response.

  • uri

    Image data is returned as a URI.

image_size enum (string) (optional)

The size of the image output.

Possible values:

  • 512

    512px image size.

  • 1K

    1K image size.

  • 2K

    2K image size.

  • 4K

    4K image size.

mime_type enum (string) (optional)

The MIME type of the image output.

Possible values:

  • image/jpeg

    JPEG image format.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "image" .

TextResponseFormat

Configuration for text output format.

mime_type enum (string) (optional)

The MIME type of the text output.

Possible values:

  • application/json

    JSON output format.

  • text/plain

    Plain text output format.

schema object (optional)

The JSON schema that the output should conform to. Only applicable when mime_type is application/json.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "text" .

VideoResponseFormat

Configuration for video output format.

aspect_ratio enum (string) (optional)

The aspect ratio for the video output.

Possible values:

  • 16:9

    16:9 aspect ratio.

  • 9:16

    9:16 aspect ratio.

delivery enum (string) (optional)

The delivery mode for the video output.

Possible values:

  • inline

    Video data is returned inline in the response.

  • uri

    Video data is returned as a URI.

duration string (optional)

The duration for the video output.

gcs_uri string (optional)

The Cloud Storage URI to store the video output. Required for Vertex if delivery mode is URI.

resolution enum (string) (optional)

The video output resolution. Defaults to 720p.

Possible values:

  • 360p

    360p resolution.

  • 720p

    720p resolution.

  • 1080p

    1080p resolution.

  • 4k

    4K resolution.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "video" .

مثال‌ها

خروجی صدا

{
  "type": "audio",
  "sample_rate": 24000
}

Image Output

{
  "type": "image",
  "aspect_ratio": "16:9",
  "image_size": "1K",
  "mime_type": "image/jpeg"
}

Text Output (JSON Schema)

{
  "type": "text",
  "mime_type": "application/json",
  "schema": {
    "type": "object",
    "properties": {
      "ingredients": {
        "type": "array",
        "items": {
          "type": "string"
        }
      },
      "recipe_name": {
        "type": "string"
      }
    },
    "required": [
      "ingredients",
      "recipe_name"
    ]
  }
}

خروجی ویدئو

{
  "type": "video",
  "aspect_ratio": "16:9",
  "delivery": "inline"
}

قدم

A step in the interaction.

Possible Types

CodeExecutionCallStep

Code execution call step.

arguments CodeExecutionCallStepArguments (required)

Required. The arguments to pass to the code execution.

The arguments to pass to the code execution.

فیلدها

code string (optional)

The code to be executed.

language enum (string) (optional)

Programming language of the `code`.

Possible values:

  • python

    Python >= 3.10, with numpy and simpy available.

id string (required)

Required. A unique ID for this specific tool call.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "code_execution_call" .

CodeExecutionResultStep

Code execution result step.

call_id string (required)

Required. ID to match the ID from the function call block.

is_error boolean (optional)

Whether the code execution resulted in an error.

result string (required)

Required. The output of the code execution.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "code_execution_result" .

FileSearchCallStep

File Search call step.

id string (required)

Required. A unique ID for this specific tool call.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "file_search_call" .

FileSearchResultStep

File Search result step.

call_id string (required)

Required. ID to match the ID from the function call block.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "file_search_result" .

FunctionCallStep

A function tool call step.

arguments object (required)

Required. The arguments to pass to the function.

id string (required)

Required. A unique ID for this specific tool call.

name string (required)

Required. The name of the tool to call.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "function_call" .

FunctionResultStep

Result of a function tool call.

call_id string (required)

Required. ID to match the ID from the function call block.

is_error boolean (optional)

Whether the tool call resulted in an error.

name string (optional)

The name of the tool that was called.

result array ( ImageContent or TextContent ) or object or string (required)

Required. The result of the tool call.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "function_result" .

GoogleMapsCallStep

Google Maps call step.

arguments GoogleMapsCallStepArguments (optional)

The arguments to pass to the Google Maps tool.

The arguments to pass to the Google Maps tool.

فیلدها

queries array (string) (optional)

The queries to be executed.

id string (required)

Required. A unique ID for this specific tool call.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "google_maps_call" .

GoogleMapsResultStep

Google Maps result step.

call_id string (required)

Required. ID to match the ID from the function call block.

result array (GoogleMapsResultItem) (required)

هیچ توضیحی ارائه نشده است.

The result of the Google Maps.

فیلدها

places array (GoogleMapsResultPlaces) (optional)

هیچ توضیحی ارائه نشده است.

فیلدها

name string (optional)

هیچ توضیحی ارائه نشده است.

place_id string (optional)

هیچ توضیحی ارائه نشده است.

review_snippets array (ReviewSnippet) (optional)

هیچ توضیحی ارائه نشده است.

Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.

فیلدها

review_id string (optional)

The ID of the review snippet.

title string (optional)

Title of the review.

url string (optional)

A link that corresponds to the user review on Google Maps.

url string (optional)

هیچ توضیحی ارائه نشده است.

widget_context_token string (optional)

هیچ توضیحی ارائه نشده است.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "google_maps_result" .

GoogleSearchCallStep

Google Search call step.

arguments GoogleSearchCallStepArguments (required)

Required. The arguments to pass to Google Search.

The arguments to pass to Google Search.

فیلدها

queries array (string) (optional)

Web search queries for the following-up web search.

id string (required)

Required. A unique ID for this specific tool call.

search_type enum (string) (optional)

The type of search grounding enabled.

Possible values:

  • web_search

    Setting this field enables web search. Only text results are returned.

  • image_search

    Setting this field enables image search. Image bytes are returned.

  • enterprise_web_search

    Setting this field enables enterprise web search.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "google_search_call" .

GoogleSearchResultStep

Google Search result step.

call_id string (required)

Required. ID to match the ID from the function call block.

is_error boolean (optional)

Whether the Google Search resulted in an error.

result array (GoogleSearchResultItem) (required)

Required. The results of the Google Search.

The result of the Google Search.

فیلدها

search_suggestions string (optional)

Web content snippet that can be embedded in a web page or an app webview.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "google_search_result" .

McpServerToolCallStep

MCPServer tool call step.

arguments object (required)

Required. The JSON object of arguments for the function.

id string (required)

Required. A unique ID for this specific tool call.

name string (required)

Required. The name of the tool which was called.

server_name string (required)

Required. The name of the used MCP server.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "mcp_server_tool_call" .

McpServerToolResultStep

MCPServer tool result step.

call_id string (required)

Required. ID to match the ID from the function call block.

name string (optional)

Name of the tool which is called for this specific tool call.

result array ( ImageContent or TextContent ) or object or string (required)

Required. The output from the MCP server call. Can be simple text or rich content.

server_name string (optional)

The name of the used MCP server.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "mcp_server_tool_result" .

ModelOutputStep

Output generated by the model.

content array ( Content ) (optional)

هیچ توضیحی ارائه نشده است.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "model_output" .

ThoughtStep

A thought step.

signature string (optional)

A signature hash for backend validation.

summary array (ThoughtSummaryContent) (optional)

A summary of the thought.

Possible Types

ImageContent

An image content block.

data string (optional)

The image content.

mime_type enum (string) (optional)

The mime type of the image.

Possible values:

  • image/png

    PNG image format

  • image/jpeg

    JPEG image format

  • image/webp

    WebP image format

  • image/heic

    HEIC image format

  • image/heif

    HEIF image format

  • image/gif

    GIF image format

  • image/bmp

    BMP image format

  • image/tiff

    TIFF image format

resolution MediaResolution (optional)

The resolution of the media.

Possible values

  • low

    Low resolution.

  • medium

    Medium resolution.

  • high

    High resolution.

  • ultra_high

    Ultra high resolution.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "image" .

uri string (optional)

The URI of the image.

TextContent

A text content block.

annotations array (Annotation) (optional)

Citation information for model-generated content.

Citation information for model-generated content.

Possible Types

FileCitation

A file citation annotation.

custom_metadata object (optional)

User provided metadata about the retrieved context.

document_uri string (optional)

The URI of the file.

end_index integer (optional)

End of the attributed segment, exclusive.

file_name string (optional)

The name of the file.

media_id string (optional)

Media ID in-case of image citations, if applicable.

page_number integer (optional)

Page number of the cited document, if applicable.

source string (optional)

Source attributed for a portion of the text.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "file_citation" .

PlaceCitation

A place citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

name string (optional)

Title of the place.

place_id string (optional)

The ID of the place, in `places/{place_id}` format.

review_snippets array (ReviewSnippet) (optional)

Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.

Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.

فیلدها

review_id string (optional)

The ID of the review snippet.

title string (optional)

Title of the review.

url string (optional)

A link that corresponds to the user review on Google Maps.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

No description provided.

Always set to "place_citation" .

url string (optional)

URI reference of the place.

UrlCitation

A URL citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

title string (optional)

The title of the URL.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "url_citation" .

url string (optional)

The URL.

WordInfo

Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.

end_index integer (optional)

End of the attributed segment, exclusive.

end_offset string (optional)

End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".

speaker string (optional)

Optional. Speaker label for this word (eg "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

start_offset string (optional)

Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".

text string (optional)

The transcribed word.

type object (required)

No description provided.

Always set to "word_info" .

text string (required)

Required. The text content.

type object (required)

No description provided.

Always set to "text" .

type object (required)

No description provided.

Always set to "thought" .

UrlContextCallStep

URL context call step.

arguments UrlContextCallArguments (required)

Required. The arguments to pass to the URL context.

The arguments to pass to the URL context.

فیلدها

urls array (string) (optional)

The URLs to fetch.

id string (required)

Required. A unique ID for this specific tool call.

signature string (optional)

A signature hash for backend validation.

type object (required)

هیچ توضیحی ارائه نشده است.

Always set to "url_context_call" .

UrlContextResultStep

URL context result step.

call_id string (required)

Required. ID to match the ID from the function call block.

is_error boolean (optional)

Whether the URL context resulted in an error.

result array (UrlContextResult) (required)

Required. The results of the URL context.

The result of the URL context.

فیلدها

status enum (string) (optional)

The status of the URL retrieval.

Possible values:

  • success

    Url retrieval is successful.

  • error

    Url retrieval is failed due to error.

  • paywall

    Url retrieval is failed because the content is behind paywall.

  • unsafe

    Url retrieval is failed because the content is unsafe.

url string (optional)

The URL that was fetched.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "url_context_result" .

UserInputStep

Input provided by the user.

content array ( Content ) (optional)

No description provided.

type object (required)

No description provided.

Always set to "user_input" .

مثال‌ها

CodeExecutionCallStep

{
  "type": "code_execution_call",
  "arguments": {
    "code": "print(sum(range(1, 11)))"
  },
  "id": "code_call_71021"
}

CodeExecutionResultStep

{
  "type": "code_execution_result",
  "call_id": "code_call_71021",
  "result": "55\n"
}

FileSearchCallStep

{
  "type": "file_search_call",
  "id": "file_call_88192"
}

FileSearchResultStep

{
  "type": "file_search_result",
  "call_id": "file_call_88192"
}

FunctionCallStep

{
  "name": "get_weather",
  "type": "function_call",
  "arguments": {
    "location": "Boston, MA"
  },
  "id": "call_98231"
}

FunctionResultStep

{
  "name": "get_weather",
  "type": "function_result",
  "call_id": "call_98231",
  "result": [
    {
      "type": "text",
      "text": "{\"weather\":\"sunny\"}"
    }
  ]
}

GoogleMapsCallStep

{
  "type": "google_maps_call",
  "arguments": {
    "latitude": 37.7749,
    "longitude": -122.4194
  },
  "id": "maps_call_39201"
}

GoogleMapsResultStep

{
  "type": "google_maps_result",
  "call_id": "maps_call_39201",
  "result": [
    {
      "name": "Golden Gate Park",
      "place_id": "ChIJIQBpAG2ahYAR9R7bNdTLg8M",
      "rating": 4.8
    }
  ]
}

GoogleSearchCallStep

{
  "type": "google_search_call",
  "arguments": {
    "query": "Who won the men's 100m in Paris 2024?"
  },
  "id": "search_call_19201"
}

GoogleSearchResultStep

{
  "type": "google_search_result",
  "call_id": "search_call_19201",
  "result": [
    {
      "title": "Paris 2024 Olympics: Noah Lyles wins men's 100m gold",
      "url": "https://olympics.com/en/news/paris-2024-noah-lyles-wins-mens-100m-gold",
      "snippet": "American Noah Lyles won the Olympic men's 100m gold medal in a photo finish."
    }
  ]
}

McpServerToolCallStep

{
  "name": "calculate_tax",
  "type": "mcp_server_tool_call",
  "arguments": {
    "income": 120000,
    "state": "CA"
  },
  "id": "mcp_call_29012",
  "server_name": "financial_mcp_server"
}

McpServerToolResultStep

{
  "type": "mcp_server_tool_result",
  "call_id": "mcp_call_29012",
  "result": {
    "tax_due": 32400
  }
}

ModelOutputStep

{
  "type": "model_output",
  "content": [
    {
      "type": "text",
      "text": "The capital of France is Paris."
    }
  ]
}

ThoughtStep

{
  "type": "thought",
  "signature": "thought_sig_abcd1234",
  "summary": [
    {
      "type": "text",
      "text": "The model is searching Google for the capital of France."
    }
  ]
}

UrlContextCallStep

{
  "type": "url_context_call",
  "arguments": {
    "urls": [
      "https://www.example.com"
    ]
  },
  "id": "url_call_10219"
}

UrlContextResultStep

{
  "type": "url_context_result",
  "call_id": "url_call_10219",
  "result": [
    {
      "title": "Example Domain",
      "url": "https://www.example.com",
      "snippet": "This domain is for use in illustrative examples in documents."
    }
  ]
}

UserInputStep

{
  "type": "user_input",
  "content": [
    {
      "type": "text",
      "text": "What is the capital of France?"
    }
  ]
}

EnvironmentConfig

Configuration for a custom environment.

فیلدها

environment_id string (optional)

Optional. The environment ID for the interaction. If specified, the request will update the existing environment instead of creating a new one.

network EnvironmentNetworkEgressAllowlist or enum (string) (optional)

Network configuration for the environment.

Possible values:

  • disabled

    Turns all network off.

sources array (Source) (optional)

No description provided.

A source to be mounted into the environment.

فیلدها

content string (optional)

The inline content if `type` is `INLINE`.

encoding string (optional)

Optional encoding for inline content (eg `base64`).

source string (optional)

The source of the environment. For Cloud Storage, this is the Cloud Storage path. For GitHub, this is the GitHub path.

target string (optional)

Where the source should appear in the environment.

type enum (string) (optional)

No description provided.

Possible values:

  • gcs

    A Cloud Storage bucket.

  • inline

    Inline content.

  • repository

    A generic repository. The protocol prefix in the source URL identifies the provider (eg, github://, gcs://).

  • skill_registry

    A skill resource from the Skill Registry Service. Skill: projects/{project}/locations/{location}/skills/{skill} SkillRevision: projects/{project}/locations/{location}/skills/{skill}/revisions/{revision} Support mounting all skills under a project: projects/{project}/locations/{location}/skills.

type object (optional)

No description provided.

Always set to "remote" .

مثال‌ها

Inline Sources

{
  "type": "remote",
  "sources": [
    {
      "type": "inline",
      "content": "You are a data analyst. Always include visualizations and export results as PDF.",
      "target": ".agents/AGENTS.md"
    },
    {
      "type": "inline",
      "content": "---\nname: slide-maker\ndescription: Create HTML slide decks\n---\n# Slide Maker\n\nWhen asked to create a presentation:\n1. Analyze the input data\n2. Create an HTML slide deck with reveal.js\n3. Save to /workspace/output/slides.html",
      "target": ".agents/skills/slide-maker/SKILL.md"
    }
  ]
}

External Sources

{
  "type": "remote",
  "sources": [
    {
      "type": "repository",
      "source": "https://github.com/my-org/my-skills.git",
      "target": ".agents/skills"
    },
    {
      "type": "gcs",
      "source": "gs://my-bucket/my-folder",
      "target": "/workspace/data"
    }
  ]
}

Network Allowlist

{
  "type": "remote",
  "network": {
    "allowlist": [
      {
        "domain": "pypi.org"
      },
      {
        "domain": "*.github.com"
      }
    ]
  }
}

Proxy Credentials

{
  "type": "remote",
  "network": {
    "allowlist": [
      {
        "domain": "api.github.com",
        "transform": {
          "Authorization": "Bearer YOUR_GITHUB_TOKEN"
        }
      }
    ]
  }
}

EnvironmentNetworkEgressAllowlist

Outbound networking configuration for the sandbox. Accepts an object with an 'allowlist' array to restrict traffic, or the string 'disabled' to turn off all network access. Omit entirely to allow all outbound traffic with no header injection.

Possible Types

شیء

Outbound networking configuration for the sandbox. When specified, restricts which external domains the sandbox can reach. Omit entirely to allow all outbound traffic with no header injection.

allowlist array (AllowlistEntry) (optional)

List of allowed outbound domains. Only requests to listed domains are permitted. Use [{'domain': '*'}] to allow all domains while still injecting headers on specific ones.

A single domain allowlist rule with optional header injection.

فیلدها

domain string (optional)

Domain to allow outbound requests to. Supports wildcards (eg '*.googleapis.com'). Use '*' to allow all domains.

transform array (object) or object (optional)

Headers to inject on all outbound requests matching this domain. Accepts a single dict or a list of dicts. The egress proxy injects these automatically.

رشته

Turns all network off.

Possible values

  • disabled

    Turns all network off.

مثال‌ها

مثال

{
  "allowlist": [
    {
      "domain": "github.com",
      "transform": [
        {
          "Authorization": "Bearer your-token"
        }
      ]
    },
    {
      "domain": "*.googleapis.com"
    }
  ]
}

ToolChoiceConfig

The tool choice configuration containing allowed tools.

فیلدها

allowed_tools AllowedTools (optional)

The allowed tools.

The configuration for allowed tools.

فیلدها

mode enum (string) (optional)

The mode of the tool choice.

Possible values:

  • auto

    Auto tool choice.

  • any

    Any tool choice.

  • none

    No tool choice.

  • validated

    Validated tool choice.

tools array (string) (optional)

The names of the allowed tools.

مثال‌ها

مثال

{
  "allowed_tools": {
    "mode": "any",
    "tools": [
      "my_tool"
    ]
  }
}

ImageContent

An image content block.

فیلدها

data string (optional)

The image content.

mime_type enum (string) (optional)

The mime type of the image.

Possible values:

  • image/png

    PNG image format

  • image/jpeg

    JPEG image format

  • image/webp

    WebP image format

  • image/heic

    HEIC image format

  • image/heif

    HEIF image format

  • image/gif

    GIF image format

  • image/bmp

    BMP image format

  • image/tiff

    TIFF image format

resolution MediaResolution (optional)

The resolution of the media.

Possible values

  • low

    Low resolution.

  • medium

    Medium resolution.

  • high

    High resolution.

  • ultra_high

    Ultra high resolution.

type object (optional)

No description provided.

Always set to "image" .

uri string (optional)

The URI of the image.

مثال‌ها

تصویر

{
  "type": "image",
  "data": "BASE64_ENCODED_IMAGE",
  "mime_type": "image/png"
}

TextContent

A text content block.

فیلدها

annotations array (Annotation) (optional)

Citation information for model-generated content.

Citation information for model-generated content.

Possible Types

FileCitation

A file citation annotation.

custom_metadata object (optional)

User provided metadata about the retrieved context.

document_uri string (optional)

The URI of the file.

end_index integer (optional)

End of the attributed segment, exclusive.

file_name string (optional)

The name of the file.

media_id string (optional)

Media ID in-case of image citations, if applicable.

page_number integer (optional)

Page number of the cited document, if applicable.

source string (optional)

Source attributed for a portion of the text.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

No description provided.

Always set to "file_citation" .

PlaceCitation

A place citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

name string (optional)

Title of the place.

place_id string (optional)

The ID of the place, in `places/{place_id}` format.

review_snippets array (ReviewSnippet) (optional)

Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.

Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.

فیلدها

review_id string (optional)

The ID of the review snippet.

title string (optional)

Title of the review.

url string (optional)

A link that corresponds to the user review on Google Maps.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

No description provided.

Always set to "place_citation" .

url string (optional)

URI reference of the place.

UrlCitation

A URL citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

title string (optional)

The title of the URL.

type object (required)

No description provided.

Always set to "url_citation" .

url string (optional)

The URL.

WordInfo

Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.

end_index integer (optional)

End of the attributed segment, exclusive.

end_offset string (optional)

End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".

speaker string (optional)

Optional. Speaker label for this word (eg "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

start_offset string (optional)

Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".

text string (optional)

The transcribed word.

type object (required)

No description provided.

Always set to "word_info" .

text string (optional)

Required. The text content.

type object (optional)

No description provided.

Always set to "text" .

مثال‌ها

متن

{
  "type": "text",
  "text": "Hello, how are you?"
}