رابط برنامهنویسی کاربردی (API) تعاملات جمینی (Gemini Interactions API) به توسعهدهندگان اجازه میدهد تا با استفاده از مدلهای جمینی، برنامههای هوش مصنوعی مولد (generative AI applications) بسازند. جمینی توانمندترین مدل ما است که از پایه برای چندوجهی بودن ساخته شده است. این مدل میتواند انواع مختلف اطلاعات از جمله زبان، تصاویر، صدا، ویدئو و کد را تعمیم داده و به طور یکپارچه درک کند، در میان آنها عمل کند و ترکیب کند. میتوانید از رابط برنامهنویسی کاربردی جمینی برای موارد استفادهای مانند استدلال در متن و تصاویر، تولید محتوا، عاملهای گفتگو، سیستمهای خلاصهسازی و طبقهبندی و موارد دیگر استفاده کنید.
ایجاد تعامل
یک تعامل جدید ایجاد میکند.
پارامترهای مسیر/پرسوجو
از کدام نسخه API استفاده کنیم.
درخواست بدنه
بدنه درخواست شامل دادههایی با ساختار زیر است:
مدل ModelOption (اختیاری)
نام «مدل» مورد استفاده برای تولید تعامل.
در صورت عدم ارائه «عامل»، الزامی است.
مقادیر ممکن
-
gemini-2.5-flashاولین مدل استدلال ترکیبی ما که از یک پنجره زمینه ۱ میلیون توکنی پشتیبانی میکند و دارای بودجههای تفکر است.
-
gemini-2.5-proمدل چندمنظوره پیشرفته ما، که در کدنویسی و کارهای استدلالی پیچیده عالی عمل میکند.
-
gemma-4-26b-a4b-itجما ۴ ۲۶ب A4ب آی تی
-
gemma-4-31b-itجما ۴ ۳۱ب آیتی
-
gemini-flash-latestآخرین نسخه از بازی Gemini Flash
-
gemini-flash-lite-latestآخرین نسخه Gemini Flash-Lite
-
gemini-pro-latestآخرین نسخه Gemini Pro
-
gemini-2.5-flash-liteکوچکترین و مقرون به صرفه ترین مدل ما، ساخته شده برای استفاده در مقیاس بزرگ.
-
gemini-2.5-flash-imageمدل تولید تصویر بومی ما، که برای سرعت، انعطافپذیری و درک متنی بهینه شده است. ورودی و خروجی متن با همان قیمت ۲.۵ فلش ارائه میشود.
-
gemini-3-flash-previewهوشمندترین مدل ما که برای سرعت ساخته شده است، هوش مرزی را با جستجو و ردیابی برتر ترکیب میکند.
-
gemini-3.1-pro-previewجدیدترین مدل استدلال SOTA ما با عمق و ظرافت بیسابقه و قابلیتهای قدرتمند درک و کدنویسی چندوجهی.
-
gemini-3.1-pro-preview-customtoolsپیشنمایش Gemini 3.1 Pro برای استفاده از ابزارهای سفارشی بهینه شده است
-
gemini-3.1-flash-liteمقرونبهصرفهترین مدل ما، بهینهشده برای وظایف عاملمحور با حجم بالا، ترجمه و پردازش دادههای ساده.
-
gemini-3-pro-imageتصویر Gemini 3 Pro
-
nano-banana-pro-previewپیشنمایش تصویر Gemini 3 Pro
-
gemini-3.1-flash-imageتصویر فلش جمینی ۳.۱.
-
gemini-3.5-flashهوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عاملدار و کدنویسی.
-
gemini-3.6-flashهوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عاملدار و کدنویسی.
-
gemini-3.7-flashهوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عاملدار و کدنویسی.
-
lyria-3-clip-previewمدل تولید موسیقی با تأخیر کم ما برای کلیپهای صوتی با کیفیت بالا و کنترل دقیق ریتمیک بهینه شده است.
-
lyria-3-pro-previewمدل پیشرفته و کامل ما برای تولید آهنگ با درک عمیق از آهنگسازی، بهینه شده برای کنترل ساختاری دقیق و انتقالهای پیچیده در سبکهای مختلف موسیقی.
-
gemini-robotics-er-1.6-previewپیشنمایش Gemini Robotics-ER 1.6
-
gemini-robotics-er-2-previewپیشنمایش Gemini Robotics Embodied Reasoning 2
گزینه عامل (اختیاری)
نام «عامل» مورد استفاده برای ایجاد تعامل.
در صورت عدم ارائه «مدل»، الزامی است.
مقادیر ممکن
-
deep-research-pro-preview-12-2025نماینده تحقیقات عمیق جمینی
-
deep-research-preview-04-2026نماینده تحقیقات عمیق جمینی
-
deep-research-max-preview-04-2026مامور مکس تحقیقات عمیق جمینی
-
antigravity-preview-05-2026از عامل مدیریتشدهی Antigravity برای انجام وظایف چند مرحلهای که نیاز به استدلال، عملیات فایل و استفاده از ابزار دارند، استفاده کنید.
ورودیهای تعامل (مشترک برای مدل و عامل).
دستورالعمل سیستم برای تعامل.
فهرستی از اعلانهای ابزار که مدل ممکن است در طول تعامل فراخوانی کند.
تأکید میکند که پاسخ تولید شده یک شیء JSON است که با طرحواره JSON مشخص شده در این فیلد مطابقت دارد.
فقط ورودی. اینکه آیا تعامل پخش زنده خواهد شد یا خیر.
فقط ورودی. آیا پاسخ و درخواست برای بازیابی بعدی ذخیره شود یا خیر.
فقط ورودی. اینکه آیا تعامل مدل در پسزمینه اجرا شود یا خیر.
generation_config GenerationConfig (اختیاری)
پیکربندی مدل
پارامترهای پیکربندی برای تعامل مدل.
جایگزینی برای `agent_config`. فقط زمانی قابل اجرا است که `model` تنظیم شده باشد.
فیلدها
حداکثر تعداد توکنهایی که باید در پاسخ گنجانده شوند.
بذر مورد استفاده در رمزگشایی برای تکرارپذیری.
speech_config SpeakerConfig یا آرایه (SpeechConfig) (اختیاری)
اختیاری. پیکربندی گفتار و چند بلندگو.
فیلدها
آرایه بلندگوها (SpeechConfig) (اختیاری)
تنظیمات بلندگوهای جداگانه.
فیلدها
زبان گفتار.
نام گوینده، باید با نام گوینده داده شده در سوال مطابقت داشته باشد.
صدای گوینده.
فهرستی از توالیهای کاراکتری که تعامل خروجی را متوقف میکنند.
سطح_فکریسطح_فکری ( اختیاری )
سطح توکنهای فکری که مدل باید تولید کند.
مقادیر ممکن
-
minimalکم یا بدون فکر کردن.
-
lowسطح فکری پایین.
-
mediumسطح فکری متوسط.
-
highسطح فکری بالا.
خلاصههای تفکر ( اختیاری)
اینکه آیا خلاصه نظرات در پاسخ گنجانده شود یا خیر.
مقادیر ممکن
-
autoخلاصههای تفکر خودکار
-
noneبدون خلاصه نویسی فکری.
پیکربندی انتخاب ابزار.
مقادیر ممکن:
-
autoانتخاب خودکار ابزار.
-
anyهر انتخاب ابزاری.
-
noneبدون انتخاب ابزار.
-
validatedانتخاب ابزار معتبر.
transcription_config TranscriptionConfig (اختیاری)
اختیاری. پیکربندی برای تشخیص گفتار (رونویسی). در صورت وجود، ASR فعال میشود.
فیلدها
اختیاری. فهرستی از عبارات واژگانی سفارشی برای جهتدهی مدل تشخیص گفتار به سمت تشخیص اصطلاحات خاص.
اختیاری. کدهای زبان BCP-47 نکاتی در مورد زبانهای موجود در صدا ارائه میدهند. در صورت حذف یا خالی بودن، به طور پیشفرض روی تشخیص خودکار زبان تنظیم میشود.
حالت رونویسی یا enum (رشته) (اختیاری)
گزینههای حالت رونویسی متمایز یا enum.
انواع ممکن
حالت رونویسی هوشمند
پیکربندی برای حالت رونویسی هوشمند.
هیچ توضیحی ارائه نشده است.
همیشه روی "smart" تنظیم شود.
حالت رونویسی کلمه به کلمه
پیکربندی برای حالت رونویسی کلمه به کلمه.
اختیاری. تنظیم بلندگو را تنظیم میکند. مقادیر پشتیبانی شده: "speaker".
اختیاری. جزئیات مهرهای زمانی که باید در خروجی رونویسی گنجانده شوند. مقادیر پشتیبانی شده: "word". اگر خالی باشد، هیچ مهر زمانی تولید نمیشود.
هیچ توضیحی ارائه نشده است.
همیشه روی "verbatim" تنظیم شود.
video_config پیکربندی ویدیو (اختیاری)
پیکربندی برای تولید ویدیو.
فیلدها
حالت وظیفه اختیاری برای تولید ویدیو. در صورت مشخص نکردن، مدل به طور خودکار حالت مناسب را بر اساس متن ارائه شده و رسانه ورودی تعیین میکند.
مقادیر ممکن:
-
text_to_videoویدیو را منحصراً از یک متن فوری تولید میکند.
-
image_to_videoویدئو را از یک یا دو تصویر منبع تولید میکند. تصویر اول فریم شروع و تصویر دوم که اختیاری است، فریم پایان را تعریف میکند.
-
reference_to_videoبا استفاده از رسانههای مرجع (مانند تصاویر، صدا یا ویدیو) ویدیو تولید میکند.
-
editیک ویدیوی ورودی موجود را تغییر میدهد.
-
extendیک ویدیوی ورودی موجود را گسترش میدهد.
شیء agent_config (اختیاری)
پیکربندی عامل
پیکربندی برای عامل.
جایگزینی برای `generation_config`. فقط زمانی قابل اجرا است که `agent` تنظیم شده باشد.
انواع ممکن
تفکیککننده چندریختی: type
پیکربندی عامل ضد جاذبه
پیکربندی برای زمان اجرای عامل Antigravity. کنترل سمت سرور بر محیط اجرای عامل و پیکربندی ابزار را فراهم میکند.
حداکثر مجموع توکنها برای اجرای عامل.
مدلی که برای استدلال عامل استفاده میشود.
هیچ توضیحی ارائه نشده است.
همیشه روی "antigravity" تنظیم شود.
پیکربندی CodeMenderAgentConfig
پیکربندی برای عامل CodeMender.
درخواست_یافتن (اختیاری )
پارامترهایی برای یافتن آسیبپذیریها
فیلدها
دستورالعملهای زمینهای یا سفارشی اضافی که توسط کاربر برای هدایت تحلیل آسیبپذیری ارائه میشود.
شناسهی یک یافتهی خاص برای تأیید. این مورد عمدتاً در حالت VERIFY برای تمرکز اعتبارسنجی مبتنی بر اجرای عامل بر روی یک آسیبپذیری واحد استفاده میشود.
نحوهی جلسهی یافتن.
مقادیر ممکن:
-
scanاسکن سریع فقط با استفاده از طبقهبندیکننده اولیه.
-
verifyطبقهبندی را انجام میدهد و به دنبال آن بررسی دقیقی انجام میدهد.
آرایه source_files (محتوای فایل) (اختیاری)
فهرستی از فایلهای منبع که به عنوان زمینه برای اسکن ارائه میشوند.
فیلدها
محتوای متنی فایل که با استاندارد UTF-8 کدگذاری شده است.
مسیر نسبی فایل از ریشه پروژه.
درخواست رفع مشکل (اختیاری)
پارامترهای رفع آسیبپذیریها
فیلدها
دستورالعملهای زمینهای یا سفارشی اضافی که توسط کاربر برای هدایت فرآیند تولید پچ ارائه میشود.
شناسهی یافتهی امنیتی خاصی که باید اصلاح شود. این شناسه به یک آسیبپذیری کشفشدهی قبلی اشاره دارد.
آرایه source_files (محتوای فایل) (اختیاری)
فهرستی از فایلهای منبع که زمینه را برای اصلاح فراهم میکنند. این فایلها معمولاً فایلهایی هستند که حاوی آسیبپذیری شناساییشده هستند.
فیلدها
محتوای متنی فایل که با استاندارد UTF-8 کدگذاری شده است.
مسیر نسبی فایل از ریشه پروژه.
نام مدلی که برای عامل CodeMender استفاده میشود. در هر جلسه CodeMender فقط از یک مدل استفاده خواهد شد.
session_config پیکربندی جلسه (اختیاری)
پیکربندیهای اختیاری مختص جلسه برای لغو رفتار پیشفرض عامل.
فیلدها
حداکثر تعداد دورهای تعاملی که عامل مجاز است قبل از رسیدن به مهلت زمانی مشخص شده انجام دهد.
پارامتری برای گروهبندی چندین تعامل که متعلق به یک جلسه CodeMender هستند.
هیچ توضیحی ارائه نشده است.
همیشه روی "code-mender" تنظیم شود.
پیکربندی DeepResearchAgent
پیکربندی برای عامل تحقیقات عمیق.
برنامهریزی انسان در حلقه را برای عامل تحقیقات عمیق فعال میکند. اگر روی درست تنظیم شود، عامل تحقیقات عمیق در پاسخ خود یک طرح تحقیقاتی ارائه میدهد. سپس عامل تنها در صورتی ادامه میدهد که کاربر طرح را در نوبت بعدی تأیید کند.
ابزار bigquery را برای عامل Deep Research فعال میکند.
خلاصههای تفکر ( اختیاری)
اینکه آیا خلاصه نظرات در پاسخ گنجانده شود یا خیر.
مقادیر ممکن
-
autoخلاصههای تفکر خودکار
-
noneبدون خلاصه نویسی فکری.
هیچ توضیحی ارائه نشده است.
همیشه روی "deep-research" تنظیم شود.
اینکه آیا باید از تصاویر در پاسخ استفاده کرد یا خیر.
مقادیر ممکن:
-
offتجسمات را لحاظ نکنید.
-
autoبه طور خودکار تجسمها را شامل میشود.
پیکربندی DynamicAgent
پیکربندی برای عاملهای پویا
هیچ توضیحی ارائه نشده است.
همیشه روی "dynamic" تنظیم شود.
پیکربندی محیط برای تعامل. میتواند یک شیء باشد که منابع محیط از راه دور را مشخص میکند یا رشتهای باشد که به یک شناسه محیط موجود اشاره میکند.
برچسبهایی با فرادادههای تعریفشده توسط کاربر برای درخواست.
شناسهی تعامل قبلی، در صورت وجود.
آرایه safety_settings (SafetySetting) (اختیاری)
تنظیمات ایمنی برای تعامل.
فیلدها
اختیاری. روش مسدود کردن محتوا. اگر مشخص نشده باشد، رفتار پیشفرض استفاده از امتیاز احتمال است.
مقادیر ممکن:
-
severityروش بلوک آسیب از هر دو امتیاز احتمال و شدت استفاده میکند.
-
probabilityروش بلوک آسیب از امتیاز احتمال استفاده میکند.
الزامی. آستانهی مسدود کردن محتوا. اگر احتمال آسیب از این آستانه بیشتر شود، محتوا مسدود خواهد شد.
مقادیر ممکن:
-
block_low_and_aboveمسدود کردن محتوا با احتمال آسیب کم یا بیشتر.
-
block_medium_and_aboveمحتوایی را که احتمال آسیبرسانی آن متوسط یا بالاتر است، مسدود کنید.
-
block_only_highمسدود کردن محتوا با احتمال آسیب بالا.
-
block_noneصرف نظر از احتمال آسیبرسانی آن، هیچ محتوایی را مسدود نکنید.
-
offفیلتر ایمنی را کاملاً خاموش کنید.
نوع آسیب (اختیاری)
الزامی. نوع دستهبندی آسیبی که باید مسدود شود.
مقادیر ممکن
-
hate_speechمحتوایی که خشونت را ترویج میدهد یا بر اساس ویژگیهای خاص، نفرت علیه افراد یا گروهها را برمیانگیزد.
-
dangerous_contentمحتوایی که فعالیتهای خطرناک را ترویج، تسهیل یا امکانپذیر میکند.
-
harassmentمحتوای توهینآمیز، تهدیدآمیز یا با هدف قلدری، شکنجه یا تمسخر.
-
sexually_explicitمحتوایی که حاوی مطالب جنسی و غیراخلاقی باشد.
-
civic_integrityمنسوخ شده: فیلتر انتخابات دیگر پشتیبانی نمیشود. دستهبندی آسیب، سلامت مدنی است.
-
image_hateتصاویری که حاوی نفرتپراکنی هستند.
-
image_dangerous_contentتصاویری که حاوی محتوای خطرناک هستند.
-
image_harassmentتصاویری که حاوی آزار و اذیت هستند.
-
image_sexually_explicitتصاویری که حاوی محتوای جنسی هستند.
-
jailbreakاعلانهایی که برای دور زدن فیلترهای ایمنی طراحی شدهاند.
service_tier لایه سرویس (اختیاری)
لایه سرویس برای تعامل.
مقادیر ممکن
-
flexسطح خدمات انعطافپذیر.
-
standardسطح خدمات استاندارد.
-
priorityردیف خدمات اولویتدار.
-
deferredردیف خدمات معوق.
webhook_config پیکربندی وب هوک (اختیاری)
اختیاری. پیکربندی وبهوک برای دریافت اعلانها پس از اتمام تعامل.
فیلدها
اختیاری. در صورت تنظیم، این URLهای وبهوک به جای وبهوکهای ثبتشده، برای رویدادهای وبهوک استفاده خواهند شد.
اختیاری. فراداده کاربر که در هر انتشار رویداد به وبهوکها بازگردانده میشود.
پاسخ
یک منبع تعامل (Interaction) را برمیگرداند.
درخواست ساده
پاسخ نمونه
{ "created": "2025-11-26T12:25:15Z", "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg", "model": "gemini-3.6-flash", "object": "interaction", "status": "completed", "steps": [ { "type": "model_output", "content": [ { "type": "text", "text": "Hello! I'm functioning perfectly and ready to assist you.\n\nHow are you doing today?" } ] } ], "updated": "2025-11-26T12:25:15Z", "usage": { "input_tokens_by_modality": [ { "modality": "text", "tokens": 7 } ], "total_cached_tokens": 0, "total_input_tokens": 7, "total_output_tokens": 20, "total_thought_tokens": 22, "total_tokens": 49, "total_tool_use_tokens": 0 } }
چند نوبتی
پاسخ نمونه
{ "created": "2025-11-26T12:22:47Z", "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg", "model": "gemini-3.6-flash", "object": "interaction", "status": "completed", "steps": [ { "type": "model_output", "content": [ { "type": "text", "text": "The capital of France is Paris." } ] } ], "updated": "2025-11-26T12:22:47Z", "usage": { "input_tokens_by_modality": [ { "modality": "text", "tokens": 50 } ], "total_cached_tokens": 0, "total_input_tokens": 50, "total_output_tokens": 10, "total_thought_tokens": 0, "total_tokens": 60, "total_tool_use_tokens": 0 } }
ورودی تصویر
پاسخ نمونه
{ "created": "2025-11-26T12:22:47Z", "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg", "model": "gemini-3.6-flash", "object": "interaction", "status": "completed", "steps": [ { "type": "model_output", "content": [ { "type": "text", "text": "A white humanoid robot with glowing blue eyes stands holding a red skateboard." } ] } ], "updated": "2025-11-26T12:22:47Z", "usage": { "input_tokens_by_modality": [ { "modality": "text", "tokens": 10 }, { "modality": "image", "tokens": 258 } ], "total_cached_tokens": 0, "total_input_tokens": 268, "total_output_tokens": 20, "total_thought_tokens": 0, "total_tokens": 288, "total_tool_use_tokens": 0 } }
فراخوانی تابع
پاسخ نمونه
{ "created": "2025-11-26T12:22:47Z", "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg", "model": "gemini-3.6-flash", "object": "interaction", "status": "requires_action", "steps": [ { "name": "get_weather", "type": "function_call", "arguments": { "location": "Boston, MA" }, "id": "gth23981" } ], "updated": "2025-11-26T12:22:47Z", "usage": { "input_tokens_by_modality": [ { "modality": "text", "tokens": 100 } ], "total_cached_tokens": 0, "total_input_tokens": 100, "total_output_tokens": 25, "total_thought_tokens": 0, "total_tokens": 125, "total_tool_use_tokens": 50 } }
تحقیقات عمیق
پاسخ نمونه
{ "agent": "deep-research-pro-preview-12-2025", "created": "2025-11-26T12:22:47Z", "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg", "object": "interaction", "status": "completed", "steps": [ { "type": "model_output", "content": [ { "type": "text", "text": "Here is a comprehensive research report on the current state of cancer research..." } ] } ], "updated": "2025-11-26T12:22:47Z", "usage": { "input_tokens_by_modality": [ { "modality": "text", "tokens": 20 } ], "total_cached_tokens": 0, "total_input_tokens": 20, "total_output_tokens": 1000, "total_thought_tokens": 500, "total_tokens": 1520, "total_tool_use_tokens": 0 } }
عامل ضد جاذبه
پاسخ نمونه
{ "agent": "antigravity-preview-05-2026", "created": "2025-11-26T12:22:47Z", "environment_id": "env_abc123", "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg", "object": "interaction", "status": "completed", "steps": [ { "type": "model_output", "content": [ { "type": "text", "text": "I've summarized the top 5 Hacker News stories and saved the results to /workspace/summary.md." } ] } ], "updated": "2025-11-26T12:22:47Z", "usage": { "input_tokens_by_modality": [ { "modality": "text", "tokens": 50 } ], "total_cached_tokens": 0, "total_input_tokens": 50, "total_output_tokens": 500, "total_thought_tokens": 200, "total_tokens": 750, "total_tool_use_tokens": 0 } }
محیط استفاده مجدد
پاسخ نمونه
{ "agent": "antigravity-preview-05-2026", "created": "2025-11-26T12:23:00Z", "environment_id": "env_abc123", "id": "v1_Chd2ZTJhYmNkZWZnaGlqa2xtbm9wcXJzdHV2d3h5ejAxMjM0NTY3ODkwMTIzNDU2Nzg", "object": "interaction", "status": "completed", "steps": [ { "type": "model_output", "content": [ { "type": "text", "text": "I've updated /workspace/hello.py to accept a name argument and greet the user." } ] } ], "updated": "2025-11-26T12:23:00Z", "usage": { "input_tokens_by_modality": [ { "modality": "text", "tokens": 80 } ], "total_cached_tokens": 0, "total_input_tokens": 80, "total_output_tokens": 200, "total_thought_tokens": 100, "total_tokens": 380, "total_tool_use_tokens": 0 } }
با منابع
نماینده سفارشی
لغو یک تعامل
یک تعامل را بر اساس شناسه لغو میکند. این فقط برای تعاملات پسزمینهای که هنوز در حال اجرا هستند، اعمال میشود.
پارامترهای مسیر/پرسوجو
از کدام نسخه API استفاده کنیم.
شناسه منحصر به فرد تعاملی که باید لغو شود.
پاسخ
یک منبع تعامل (Interaction) را برمیگرداند.
لغو تعامل
پاسخ نمونه
{ "agent": "deep-research-pro-preview-12-2025", "created": "2026-06-22T04:55:47Z", "id": "v1_ChdVc0E0YXJTYk1zYlV6N0lQcXRXVG1BYxIXVXNBNGFyU2JNc2JVejdJUHF0V1RtQWM", "status": "cancelled", "steps": [ { "type": "user_input", "content": [ { "type": "text", "text": "Research the history of the Google TPUs with a focus on 2025 specs." } ] } ], "updated": "2026-06-22T04:55:47Z" }
بازیابی یک تعامل
جزئیات کامل یک تعامل واحد را بر اساس `Interaction.id` آن بازیابی میکند.
پارامترهای مسیر/پرسوجو
از کدام نسخه API استفاده کنیم.
شناسه منحصر به فرد تعاملی که قرار است بازیابی شود.
اختیاری. در صورت تنظیم، جریان تعامل را از بخش بعدی پس از رویداد مشخص شده توسط شناسه رویداد از سر میگیرد. فقط در صورتی قابل استفاده است که `stream` برابر با true باشد.
اگر روی درست تنظیم شود، محتوای تولید شده به صورت تدریجی پخش میشود.
پیشفرض: False
پاسخ
یک منبع تعامل (Interaction) را برمیگرداند.
تعامل دریافت کنید
پاسخ نمونه
{ "created": "2025-11-26T12:25:15Z", "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg", "model": "gemini-3.6-flash", "object": "interaction", "status": "completed", "steps": [ { "type": "model_output", "content": [ { "type": "text", "text": "I'm doing great, thank you for asking! How can I help you today?" } ] } ], "updated": "2025-11-26T12:25:15Z" }
حذف یک تعامل
تعامل را بر اساس شناسه حذف میکند.
پارامترهای مسیر/پرسوجو
از کدام نسخه API استفاده کنیم.
شناسه منحصر به فرد تعاملی که باید حذف شود.
پاسخ
در صورت موفقیت، پاسخ خالی است.
حذف
منابع
تعامل
منبع تعامل.
فیلدها
گزینه عامل (اختیاری)
نام «عامل» مورد استفاده برای ایجاد تعامل.
مقادیر ممکن
-
deep-research-pro-preview-12-2025نماینده تحقیقات عمیق جمینی
-
deep-research-preview-04-2026نماینده تحقیقات عمیق جمینی
-
deep-research-max-preview-04-2026مامور مکس تحقیقات عمیق جمینی
-
antigravity-preview-05-2026از عامل مدیریتشدهی Antigravity برای انجام وظایف چند مرحلهای که نیاز به استدلال، عملیات فایل و استفاده از ابزار دارند، استفاده کنید.
شیء agent_config (اختیاری)
پارامترهای پیکربندی برای تعامل عامل.
انواع ممکن
تفکیککننده چندریختی: type
پیکربندی عامل ضد جاذبه
پیکربندی برای زمان اجرای عامل Antigravity. کنترل سمت سرور بر محیط اجرای عامل و پیکربندی ابزار را فراهم میکند.
حداکثر مجموع توکنها برای اجرای عامل.
مدلی که برای استدلال عامل استفاده میشود.
هیچ توضیحی ارائه نشده است.
همیشه روی "antigravity" تنظیم شود.
پیکربندی CodeMenderAgentConfig
پیکربندی برای عامل CodeMender.
درخواست_یافتن (اختیاری )
پارامترهایی برای یافتن آسیبپذیریها
فیلدها
دستورالعملهای زمینهای یا سفارشی اضافی که توسط کاربر برای هدایت تحلیل آسیبپذیری ارائه میشود.
شناسهی یک یافتهی خاص برای تأیید. این مورد عمدتاً در حالت VERIFY برای تمرکز اعتبارسنجی مبتنی بر اجرای عامل بر روی یک آسیبپذیری واحد استفاده میشود.
نحوهی جلسهی یافتن.
مقادیر ممکن:
-
scanاسکن سریع فقط با استفاده از طبقهبندیکننده اولیه.
-
verifyطبقهبندی را انجام میدهد و به دنبال آن بررسی دقیقی انجام میدهد.
آرایه source_files (محتوای فایل) (اختیاری)
فهرستی از فایلهای منبع که به عنوان زمینه برای اسکن ارائه میشوند.
فیلدها
محتوای متنی فایل که با استاندارد UTF-8 کدگذاری شده است.
مسیر نسبی فایل از ریشه پروژه.
درخواست رفع مشکل (اختیاری)
پارامترهای رفع آسیبپذیریها
فیلدها
دستورالعملهای زمینهای یا سفارشی اضافی که توسط کاربر برای هدایت فرآیند تولید پچ ارائه میشود.
شناسهی یافتهی امنیتی خاصی که باید اصلاح شود. این شناسه به یک آسیبپذیری کشفشدهی قبلی اشاره دارد.
آرایه source_files (محتوای فایل) (اختیاری)
فهرستی از فایلهای منبع که زمینه را برای اصلاح فراهم میکنند. این فایلها معمولاً فایلهایی هستند که حاوی آسیبپذیری شناساییشده هستند.
فیلدها
محتوای متنی فایل که با استاندارد UTF-8 کدگذاری شده است.
مسیر نسبی فایل از ریشه پروژه.
نام مدلی که برای عامل CodeMender استفاده میشود. در هر جلسه CodeMender فقط از یک مدل استفاده خواهد شد.
session_config پیکربندی جلسه (اختیاری)
پیکربندیهای اختیاری مختص جلسه برای لغو رفتار پیشفرض عامل.
فیلدها
حداکثر تعداد دورهای تعاملی که عامل مجاز است قبل از رسیدن به مهلت زمانی مشخص شده انجام دهد.
پارامتری برای گروهبندی چندین تعامل که متعلق به یک جلسه CodeMender هستند.
هیچ توضیحی ارائه نشده است.
همیشه روی "code-mender" تنظیم شود.
پیکربندی DeepResearchAgent
پیکربندی برای عامل تحقیقات عمیق.
برنامهریزی انسان در حلقه را برای عامل تحقیقات عمیق فعال میکند. اگر روی درست تنظیم شود، عامل تحقیقات عمیق در پاسخ خود یک طرح تحقیقاتی ارائه میدهد. سپس عامل تنها در صورتی ادامه میدهد که کاربر طرح را در نوبت بعدی تأیید کند.
ابزار bigquery را برای عامل Deep Research فعال میکند.
خلاصههای تفکر ( اختیاری)
اینکه آیا خلاصه نظرات در پاسخ گنجانده شود یا خیر.
مقادیر ممکن
-
autoخلاصههای تفکر خودکار
-
noneبدون خلاصه نویسی فکری.
هیچ توضیحی ارائه نشده است.
همیشه روی "deep-research" تنظیم شود.
اینکه آیا باید از تصاویر در پاسخ استفاده کرد یا خیر.
مقادیر ممکن:
-
offتجسمات را لحاظ نکنید.
-
autoبه طور خودکار تجسمها را شامل میشود.
پیکربندی DynamicAgent
پیکربندی برای عاملهای پویا
هیچ توضیحی ارائه نشده است.
همیشه روی "dynamic" تنظیم شود.
فقط خروجی. زمانی که پاسخ در قالب ISO 8601 (YYYY-MM-DDThh:mm:ssZ) ایجاد شده است.
پیکربندی محیط برای تعامل. میتواند یک شیء باشد که منابع محیط از راه دور را مشخص میکند یا رشتهای باشد که به یک شناسه محیط موجود اشاره میکند.
فقط خروجی. شناسه محیط برای تعامل. فقط در صورتی که پیکربندی محیط در درخواست تنظیم شده باشد، پر میشود.
آرایه خطاها (Error) (اختیاری)
فقط خروجی. خطاهای تشخیصی / خطاهای پلتفرم که در تعامل ثبت شدهاند.
فیلدها
یک URI که نوع خطا را مشخص میکند.
یک پیام خطا که برای انسان قابل خواندن باشد.
الزامی. فقط خروجی. یک شناسه منحصر به فرد برای تکمیل تعامل.
پیشفرضها به:
برچسبهایی با فرادادههای تعریفشده توسط کاربر برای درخواست.
مدل ModelOption (اختیاری)
نام «مدل» مورد استفاده برای تولید تعامل.
مقادیر ممکن
-
gemini-2.5-flashاولین مدل استدلال ترکیبی ما که از یک پنجره زمینه ۱ میلیون توکنی پشتیبانی میکند و دارای بودجههای تفکر است.
-
gemini-2.5-proمدل چندمنظوره پیشرفته ما، که در کدنویسی و کارهای استدلالی پیچیده عالی عمل میکند.
-
gemma-4-26b-a4b-itجما ۴ ۲۶ب A4ب آی تی
-
gemma-4-31b-itجما ۴ ۳۱ب آیتی
-
gemini-flash-latestآخرین نسخه از بازی Gemini Flash
-
gemini-flash-lite-latestآخرین نسخه Gemini Flash-Lite
-
gemini-pro-latestآخرین نسخه Gemini Pro
-
gemini-2.5-flash-liteکوچکترین و مقرون به صرفه ترین مدل ما، ساخته شده برای استفاده در مقیاس بزرگ.
-
gemini-2.5-flash-imageمدل تولید تصویر بومی ما، که برای سرعت، انعطافپذیری و درک متنی بهینه شده است. ورودی و خروجی متن با همان قیمت ۲.۵ فلش ارائه میشود.
-
gemini-3-flash-previewهوشمندترین مدل ما که برای سرعت ساخته شده است، هوش مرزی را با جستجو و ردیابی برتر ترکیب میکند.
-
gemini-3.1-pro-previewجدیدترین مدل استدلال SOTA ما با عمق و ظرافت بیسابقه و قابلیتهای قدرتمند درک و کدنویسی چندوجهی.
-
gemini-3.1-pro-preview-customtoolsپیشنمایش Gemini 3.1 Pro برای استفاده از ابزارهای سفارشی بهینه شده است
-
gemini-3.1-flash-liteمقرونبهصرفهترین مدل ما، بهینهشده برای وظایف عاملمحور با حجم بالا، ترجمه و پردازش دادههای ساده.
-
gemini-3-pro-imageتصویر Gemini 3 Pro
-
nano-banana-pro-previewپیشنمایش تصویر Gemini 3 Pro
-
gemini-3.1-flash-imageتصویر فلش جمینی ۳.۱.
-
gemini-3.5-flashهوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عاملدار و کدنویسی.
-
gemini-3.6-flashهوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عاملدار و کدنویسی.
-
gemini-3.7-flashهوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عاملدار و کدنویسی.
-
lyria-3-clip-previewمدل تولید موسیقی با تأخیر کم ما برای کلیپهای صوتی با کیفیت بالا و کنترل دقیق ریتمیک بهینه شده است.
-
lyria-3-pro-previewمدل پیشرفته و کامل ما برای تولید آهنگ با درک عمیق از آهنگسازی، بهینه شده برای کنترل ساختاری دقیق و انتقالهای پیچیده در سبکهای مختلف موسیقی.
-
gemini-robotics-er-1.6-previewپیشنمایش Gemini Robotics-ER 1.6
-
gemini-robotics-er-2-previewپیشنمایش Gemini Robotics Embodied Reasoning 2
محتوای صوتی output_audio (اختیاری)
آخرین صدای تولید شده توسط مدل در پاسخ به درخواست فعلی. توجه: این توسط SDK اضافه شده است.
فیلدها
تعداد کانالهای صوتی
محتوای صوتی.
نوع مایم صدا.
مقادیر ممکن:
-
audio/wavفرمت صوتی WAV
-
audio/mp3فرمت صوتی MP3
-
audio/aiffفرمت صوتی AIFF
-
audio/aacفرمت صوتی AAC
-
audio/oggفرمت صوتی OGG
-
audio/flacفرمت صوتی FLAC
-
audio/mpegفرمت صوتی MPEG
-
audio/m4aفرمت صوتی M4A
-
audio/l16فرمت صوتی L16
-
audio/opusفرمت صوتی OPUS
-
audio/alawفرمت صوتی ALAW
-
audio/mulawفرمت صوتی MULAW
نرخ نمونهبرداری صدا.
هیچ توضیحی ارائه نشده است.
همیشه روی "audio" تنظیم شود.
آدرس اینترنتی (URI) فایل صوتی.
آخرین تصویری که توسط مدل در پاسخ به درخواست فعلی تولید شده است. توجه: این تصویر توسط SDK اضافه شده است.
متن به هم پیوسته از آخرین خروجی مدل در پاسخ به درخواست فعلی. توجه: این توسط SDK اضافه شده است.
خروجی_ویدئو محتوای ویدیویی (اختیاری)
آخرین ویدیویی که توسط مدل در پاسخ به درخواست فعلی تولید شده است. توجه: این توسط SDK اضافه شده است.
فیلدها
محتوای ویدیویی.
نوع میم (شبیهسازی) ویدیو.
مقادیر ممکن:
-
video/mp4فرمت ویدیویی MP4
-
video/mpegفرمت ویدیویی MPEG
-
video/mpgفرمت ویدیویی MPG
-
video/movفرمت ویدیویی MOV
-
video/aviفرمت ویدیویی AVI
-
video/x-flvفرمت ویدیویی FLV
-
video/webmفرمت ویدیویی وبام
-
video/wmvفرمت ویدیویی WMV
-
video/3gppفرمت ویدیویی 3GPP
پردازش MediaProcessing یا enum (رشتهای) (اختیاری)
چگونه مدل این ویدیو را برای درک پردازش میکند.
وضوح تصویر MediaResolution (اختیاری)
قطعنامه رسانهها.
مقادیر ممکن
-
lowوضوح پایین.
-
mediumوضوح متوسط.
-
highوضوح بالا.
-
ultra_highوضوح فوق العاده بالا.
هیچ توضیحی ارائه نشده است.
همیشه روی "video" تنظیم شود.
آدرس اینترنتی (URI) ویدیو.
شناسهی تعامل قبلی، در صورت وجود.
تأکید میکند که پاسخ تولید شده یک شیء JSON است که با طرحواره JSON مشخص شده در این فیلد مطابقت دارد.
آرایه safety_settings (SafetySetting) (اختیاری)
تنظیمات ایمنی برای تعامل.
فیلدها
اختیاری. روش مسدود کردن محتوا. اگر مشخص نشده باشد، رفتار پیشفرض استفاده از امتیاز احتمال است.
مقادیر ممکن:
-
severityروش بلوک آسیب از هر دو امتیاز احتمال و شدت استفاده میکند.
-
probabilityروش بلوک آسیب از امتیاز احتمال استفاده میکند.
الزامی. آستانهی مسدود کردن محتوا. اگر احتمال آسیب از این آستانه بیشتر شود، محتوا مسدود خواهد شد.
مقادیر ممکن:
-
block_low_and_aboveمسدود کردن محتوا با احتمال آسیب کم یا بیشتر.
-
block_medium_and_aboveمحتوایی را که احتمال آسیبرسانی آن متوسط یا بالاتر است، مسدود کنید.
-
block_only_highمسدود کردن محتوا با احتمال آسیب بالا.
-
block_noneصرف نظر از احتمال آسیبرسانی آن، هیچ محتوایی را مسدود نکنید.
-
offفیلتر ایمنی را کاملاً خاموش کنید.
نوع آسیب (اختیاری)
الزامی. نوع دستهبندی آسیبی که باید مسدود شود.
مقادیر ممکن
-
hate_speechمحتوایی که خشونت را ترویج میدهد یا بر اساس ویژگیهای خاص، نفرت علیه افراد یا گروهها را برمیانگیزد.
-
dangerous_contentمحتوایی که فعالیتهای خطرناک را ترویج، تسهیل یا امکانپذیر میکند.
-
harassmentمحتوای توهینآمیز، تهدیدآمیز یا با هدف قلدری، شکنجه یا تمسخر.
-
sexually_explicitمحتوایی که حاوی مطالب جنسی و غیراخلاقی باشد.
-
civic_integrityمنسوخ شده: فیلتر انتخابات دیگر پشتیبانی نمیشود. دستهبندی آسیب، سلامت مدنی است.
-
image_hateتصاویری که حاوی نفرتپراکنی هستند.
-
image_dangerous_contentتصاویری که حاوی محتوای خطرناک هستند.
-
image_harassmentتصاویری که حاوی آزار و اذیت هستند.
-
image_sexually_explicitتصاویری که حاوی محتوای جنسی هستند.
-
jailbreakاعلانهایی که برای دور زدن فیلترهای ایمنی طراحی شدهاند.
service_tier لایه سرویس (اختیاری)
لایه سرویس برای تعامل.
مقادیر ممکن
-
flexسطح خدمات انعطافپذیر.
-
standardسطح خدمات استاندارد.
-
priorityردیف خدمات اولویتدار.
-
deferredردیف خدمات معوق.
الزامی. فقط خروجی. وضعیت تعامل.
مقادیر ممکن:
-
in_progressتعامل در حال انجام است.
-
requires_actionاین تعامل نیاز به اقدام/ورودی از سوی کاربر دارد.
-
completedتعامل تکمیل شده است.
-
failedتعامل شکست خورد.
-
cancelledتعامل لغو شد.
-
incompleteتعامل تکمیل شده است، اما شامل نتایج ناقص است (مثلاً رسیدن به max_tokens).
-
budget_exceededتعامل متوقف شد زیرا بودجه توکن از حد مجاز فراتر رفته بود.
-
queuedتعامل در صف انتظار پردازش قرار میگیرد.
فقط خروجی. مراحلی که تعامل را تشکیل میدهند، زمانی که در پاسخ گنجانده شوند.
دستورالعمل سیستم برای تعامل.
فهرستی از اعلانهای ابزار که مدل ممکن است در طول تعامل فراخوانی کند.
فقط خروجی. زمانی که پاسخ آخرین بار در قالب ISO 8601 (YYYY-MM-DDThh:mm:ssZ) بهروزرسانی شده است.
کاربرد (اختیاری )
فقط خروجی. آمار مربوط به میزان استفاده از توکن درخواست تعامل.
فیلدها
آرایه cached_tokens_by_modality (ModalityTokens) (اختیاری)
تفکیک میزان استفاده از توکنهای ذخیرهشده بر اساس روش.
فیلدها
روش پاسخ (اختیاری)
روش مرتبط با شمارش توکنها.
مقادیر ممکن
-
textنشان میدهد که مدل باید متن را برگرداند.
-
imageنشان میدهد که مدل باید تصاویر را برگرداند.
-
audioنشان میدهد که مدل باید صدا را برگرداند.
-
videoنشان میدهد که مدل باید ویدیو برگرداند.
-
documentنشان میدهد که مدل باید اسناد را برگرداند.
تعداد توکنها برای روش.
آرایه grounding_tool_count (GroundingToolCount) (اختیاری)
تعداد ابزار اتصال به زمین
فیلدها
تعداد ابزار اتصال به زمین مهم است.
نوع ابزار اتصال زمین مرتبط با شمارش.
مقادیر ممکن:
-
google_searchاتصال به زمین با جستجوی وب و جستجوی تصویر گوگل، و اتصال به زمین وب برای سازمانها.
-
google_mapsاتصال به زمین با نقشههای گوگل.
-
retrievalپایه گذاری با داده های مشتری، به عنوان مثال، VertexAISearch.
آرایه input_tokens_by_modality (ModalityTokens) (اختیاری)
تفکیک استفاده از توکن ورودی بر اساس روش.
فیلدها
روش پاسخ (اختیاری)
روش مرتبط با شمارش توکنها.
مقادیر ممکن
-
textنشان میدهد که مدل باید متن را برگرداند.
-
imageنشان میدهد که مدل باید تصاویر را برگرداند.
-
audioنشان میدهد که مدل باید صدا را برگرداند.
-
videoنشان میدهد که مدل باید ویدیو برگرداند.
-
documentنشان میدهد که مدل باید اسناد را برگرداند.
تعداد توکنها برای روش.
آرایه output_tokens_by_modality (ModalityTokens) (اختیاری)
تفکیک استفاده از توکن خروجی بر اساس روش.
فیلدها
روش پاسخ (اختیاری)
روش مرتبط با شمارش توکنها.
مقادیر ممکن
-
textنشان میدهد که مدل باید متن را برگرداند.
-
imageنشان میدهد که مدل باید تصاویر را برگرداند.
-
audioنشان میدهد که مدل باید صدا را برگرداند.
-
videoنشان میدهد که مدل باید ویدیو برگرداند.
-
documentنشان میدهد که مدل باید اسناد را برگرداند.
تعداد توکنها برای روش.
آرایه tool_use_tokens_by_modality (ModalityTokens) (اختیاری)
تفکیک میزان استفاده از توکنهای ابزار بر اساس روش.
فیلدها
روش پاسخ (اختیاری)
روش مرتبط با شمارش توکنها.
مقادیر ممکن
-
textنشان میدهد که مدل باید متن را برگرداند.
-
imageنشان میدهد که مدل باید تصاویر را برگرداند.
-
audioنشان میدهد که مدل باید صدا را برگرداند.
-
videoنشان میدهد که مدل باید ویدیو برگرداند.
-
documentنشان میدهد که مدل باید اسناد را برگرداند.
تعداد توکنها برای روش.
تعداد توکنها در بخش ذخیرهشدهی اعلان (محتوای ذخیرهشده).
تعداد توکنها در اعلان (زمینه).
تعداد کل توکنها در تمام پاسخهای تولید شده.
تعداد توکنهای افکار برای مدلهای تفکر.
تعداد کل توکنها برای درخواست تعامل (درخواست + پاسخها + سایر توکنهای داخلی).
تعداد توکنهای موجود در اعلان(های) استفاده از ابزار.
webhook_config پیکربندی وب هوک (اختیاری)
اختیاری. پیکربندی وبهوک برای دریافت اعلانها پس از اتمام تعامل.
فیلدها
اختیاری. در صورت تنظیم، این URLهای وبهوک به جای وبهوکهای ثبتشده، برای رویدادهای وبهوک استفاده میشوند.
اختیاری. فراداده کاربر که در هر انتشار رویداد به وبهوکها بازگردانده میشود.
مثالها
مثال
{ "created": "2025-12-04T15:01:45Z", "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg", "model": "gemini-3.6-flash", "object": "interaction", "status": "completed", "steps": [ { "type": "model_output", "content": [ { "type": "text", "text": "Hello! I'm doing well, functioning as expected. Thank you for asking! How are you doing today?" } ] } ], "updated": "2025-12-04T15:01:45Z", "usage": { "input_tokens_by_modality": [ { "modality": "text", "tokens": 7 } ], "total_cached_tokens": 0, "total_input_tokens": 7, "total_output_tokens": 23, "total_thought_tokens": 49, "total_tokens": 79, "total_tool_use_tokens": 0 } }
مدلهای داده
محتوا
محتوای پاسخ.
انواع ممکن
محتوای صوتی
یک بلوک محتوای صوتی.
The number of audio channels.
The audio content.
The mime type of the audio.
Possible values:
-
audio/wavWAV audio format
-
audio/mp3MP3 audio format
-
audio/aiffAIFF audio format
-
audio/aacAAC audio format
-
audio/oggOGG audio format
-
audio/flacFLAC audio format
-
audio/mpegMPEG audio format
-
audio/m4aM4A audio format
-
audio/l16L16 audio format
-
audio/opusOPUS audio format
-
audio/alawALAW audio format
-
audio/mulawMULAW audio format
The sample rate of the audio.
هیچ توضیحی ارائه نشده است.
Always set to "audio" .
The URI of the audio.
DocumentContent
A document content block.
The document content.
The mime type of the document.
Possible values:
-
application/pdfPDF document format
-
text/csvCSV document format
هیچ توضیحی ارائه نشده است.
Always set to "document" .
The URI of the document.
ImageContent
An image content block.
The image content.
The mime type of the image.
Possible values:
-
image/pngPNG image format
-
image/jpegJPEG image format
-
image/webpWebP image format
-
image/heicHEIC image format
-
image/heifHEIF image format
-
image/gifGIF image format
-
image/bmpBMP image format
-
image/tiffTIFF image format
resolution MediaResolution (optional)
The resolution of the media.
Possible values
-
lowLow resolution.
-
mediumMedium resolution.
-
highHigh resolution.
-
ultra_highUltra high resolution.
هیچ توضیحی ارائه نشده است.
Always set to "image" .
The URI of the image.
TextContent
A text content block.
annotations array (Annotation) (optional)
Citation information for model-generated content.
Possible Types
FileCitation
A file citation annotation.
User provided metadata about the retrieved context.
The URI of the file.
End of the attributed segment, exclusive.
The name of the file.
Media ID in-case of image citations, if applicable.
Page number of the cited document, if applicable.
Source attributed for a portion of the text.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
هیچ توضیحی ارائه نشده است.
Always set to "file_citation" .
PlaceCitation
A place citation annotation.
End of the attributed segment, exclusive.
Title of the place.
The ID of the place, in `places/{place_id}` format.
review_snippets array (ReviewSnippet) (optional)
Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.
فیلدها
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
هیچ توضیحی ارائه نشده است.
Always set to "place_citation" .
URI reference of the place.
UrlCitation
A URL citation annotation.
End of the attributed segment, exclusive.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
The title of the URL.
هیچ توضیحی ارائه نشده است.
Always set to "url_citation" .
The URL.
WordInfo
Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.
End of the attributed segment, exclusive.
End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
Optional. Speaker label for this word (eg "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
The transcribed word.
هیچ توضیحی ارائه نشده است.
Always set to "word_info" .
Required. The text content.
هیچ توضیحی ارائه نشده است.
Always set to "text" .
VideoContent
A video content block.
The video content.
The mime type of the video.
Possible values:
-
video/mp4MP4 video format
-
video/mpegMPEG video format
-
video/mpgMPG video format
-
video/movMOV video format
-
video/aviAVI video format
-
video/x-flvFLV video format
-
video/webmWebM video format
-
video/wmvWMV video format
-
video/3gpp3GPP video format
processing MediaProcessing or enum (string) (optional)
How the model processes this video for understanding.
resolution MediaResolution (optional)
The resolution of the media.
Possible values
-
lowLow resolution.
-
mediumMedium resolution.
-
highHigh resolution.
-
ultra_highUltra high resolution.
هیچ توضیحی ارائه نشده است.
Always set to "video" .
The URI of the video.
مثالها
صوتی
{ "type": "audio", "data": "BASE64_ENCODED_AUDIO", "mime_type": "audio/wav" }
سند
{ "type": "document", "data": "BASE64_ENCODED_DOCUMENT", "mime_type": "application/pdf" }
تصویر
{ "type": "image", "data": "BASE64_ENCODED_IMAGE", "mime_type": "image/png" }
متن
{ "type": "text", "text": "Hello, how are you?" }
ویدئو
{ "type": "video", "uri": "https://www.youtube.com/watch?v=9hE5-98ZeCg" }
ابزار
A tool that can be used by the model.
Possible Types
CodeExecution
A tool that can be used by the model to execute code.
هیچ توضیحی ارائه نشده است.
Always set to "code_execution" .
ComputerUse
A tool that can be used by the model to interact with the computer.
Optional. Disabled safety policies for computer use.
Possible values:
-
financial_transactionsSafety policy for financial transactions.
-
sensitive_data_modificationSafety policy for sensitive data modification.
-
communication_toolSafety policy for communication tools (eg Gmail, Chat, Meet).
-
account_creationSafety policy for account creation.
-
data_modificationSafety policy for data modification.
-
user_consent_managementSafety policy for user consent management.
-
legal_terms_and_agreementsSafety policy for legal terms and agreements.
Whether enable the prompt injection detection check on computer-use request.
The environment being operated.
Possible values:
-
browserOperates in a web browser.
-
mobileOperates in a mobile environment.
-
desktopOperates in a desktop environment.
The list of predefined functions that are excluded from the model call.
هیچ توضیحی ارائه نشده است.
Always set to "computer_use" .
FileSearch
A tool that can be used by the model to search files.
The file search store names to search.
Metadata filter to apply to the semantic retrieval documents and chunks.
The number of semantic retrieval chunks to retrieve.
هیچ توضیحی ارائه نشده است.
Always set to "file_search" .
عملکرد
A tool that can be used by the model.
A description of the function.
The name of the function.
The JSON Schema for the function's parameters.
هیچ توضیحی ارائه نشده است.
Always set to "function" .
GoogleMaps
A tool that can be used by the model to call Google Maps.
Whether to return a widget context token in the tool call result of the response.
The latitude of the user's location.
The longitude of the user's location.
هیچ توضیحی ارائه نشده است.
Always set to "google_maps" .
GoogleSearch
A tool that can be used by the model to search Google.
The types of search grounding to enable.
Possible values:
-
web_searchSetting this field enables web search. Only text results are returned.
-
image_searchSetting this field enables image search. Image bytes are returned.
-
enterprise_web_searchSetting this field enables enterprise web search.
هیچ توضیحی ارائه نشده است.
Always set to "google_search" .
McpServer
A MCPServer is a server that can be called by the model to perform actions.
allowed_tools array (AllowedTools) (optional)
The allowed tools.
فیلدها
The mode of the tool choice.
Possible values:
-
autoAuto tool choice.
-
anyAny tool choice.
-
noneNo tool choice.
-
validatedValidated tool choice.
The names of the allowed tools.
Optional: Fields for authentication headers, timeouts, etc., if needed.
The name of the MCPServer.
هیچ توضیحی ارائه نشده است.
Always set to "mcp_server" .
The full URL for the MCPServer endpoint. Example: "https://api.example.com/mcp"
بازیابی
A tool that can be used by the model to retrieve files.
exa_ai_search_config ExaAISearchConfig (optional)
Used to specify configuration for ExaAISearch.
فیلدها
Required. The API key for ExaAiSearch.
Optional. This field can be used to pass any parameter from the Exa.ai Search API.
parallel_ai_search_config ParallelAISearchConfig (optional)
Used to specify configuration for ParallelAISearch.
فیلدها
Optional. The API key for ParallelAiSearch.
Optional. Custom configs for ParallelAiSearch.
rag_store_config RagStoreConfig (optional)
Used to specify configuration for RagStore.
فیلدها
rag_resources array (RagResource) (optional)
Optional. The representation of the rag source.
فیلدها
Optional. RagCorpora resource name.
Optional. rag_file_id. The files should be in the same rag_corpus set in rag_corpus field.
rag_retrieval_config RagRetrievalConfig (optional)
Optional. The retrieval config for the Rag query.
فیلدها
filter Filter (optional)
Optional. Config for filters.
فیلدها
Optional. String for metadata filtering.
Optional. Only returns contexts with vector distance smaller than the threshold.
Optional. Only returns contexts with vector similarity larger than the threshold.
hybrid_search HybridSearch (optional)
Optional. Config for Hybrid Search.
فیلدها
Optional. Alpha value controls the weight between dense and sparse vector search results.
ranking Ranking (optional)
Optional. Config for ranking and reranking.
Optional. The number of contexts to retrieve.
The types of file retrieval to enable.
Possible values:
-
rag_store -
exa_ai_search -
parallel_ai_search
هیچ توضیحی ارائه نشده است.
Always set to "retrieval" .
UrlContext
A tool that can be used by the model to fetch URL context.
هیچ توضیحی ارائه نشده است.
Always set to "url_context" .
مثالها
CodeExecution
ComputerUse
FileSearch
عملکرد
GoogleMaps
GoogleSearch
McpServer
بازیابی
No examples available for this type.
UrlContext
InteractionSseEvent
Possible Types
Polymorphic discriminator: event_type
ErrorEvent
error Error (optional)
هیچ توضیحی ارائه نشده است.
فیلدها
A URI that identifies the error type.
A human-readable error message.
The event_id token to be used to resume the interaction stream, from this event.
هیچ توضیحی ارائه نشده است.
Always set to "error" .
InteractionCompletedEvent
The event_id token to be used to resume the interaction stream, from this event.
هیچ توضیحی ارائه نشده است.
Always set to "interaction.completed" .
interaction InteractionSseEventInteraction (required)
Partial completed interaction resource emitted at the end of the stream.
فیلدها
The agent to interact with.
Output only. The time at which the response was created in ISO 8601 format.
Required. Output only. A unique identifier for the interaction completion.
The model that will complete your prompt.
Output only. The resource type.
service_tier ServiceTier (optional)
The service tier for the interaction.
Possible values
-
flexFlex service tier.
-
standardStandard service tier.
-
priorityPriority service tier.
-
deferredDeferred service tier.
Required. Output only. The status of the interaction.
Possible values:
-
in_progressThe interaction is in progress.
-
requires_actionThe interaction requires action/input from the user.
-
completedThe interaction is completed.
-
failedThe interaction failed.
-
cancelledThe interaction was cancelled.
-
incompleteThe interaction is completed, but contains incomplete results (eg hitting max_tokens).
Output only. The steps that make up the interaction, if included in this event.
Output only. The time at which the response was last updated in ISO 8601 format.
usage Usage (optional)
Output only. Statistics on the interaction request's token usage.
فیلدها
cached_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of cached token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
grounding_tool_count array (GroundingToolCount) (optional)
Grounding tool count.
فیلدها
The number of grounding tool counts.
The grounding tool type associated with the count.
Possible values:
-
google_searchGrounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
-
google_mapsGrounding with Google Maps.
-
retrievalGrounding with customer's data, for example, VertexAISearch.
input_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of input token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
output_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of output token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
tool_use_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of tool-use token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
Number of tokens in the cached part of the prompt (the cached content).
Number of tokens in the prompt (context).
Total number of tokens across all the generated responses.
Number of tokens of thoughts for thinking models.
Total token count for the interaction request (prompt + responses + other internal tokens).
Number of tokens present in tool-use prompt(s).
InteractionCreatedEvent
The event_id token to be used to resume the interaction stream, from this event.
هیچ توضیحی ارائه نشده است.
Always set to "interaction.created" .
interaction InteractionSseEventInteraction (required)
Partial interaction resource emitted when the stream is created.
فیلدها
The agent to interact with.
Output only. The time at which the response was created in ISO 8601 format.
Required. Output only. A unique identifier for the interaction completion.
The model that will complete your prompt.
Output only. The resource type.
service_tier ServiceTier (optional)
The service tier for the interaction.
Possible values
-
flexFlex service tier.
-
standardStandard service tier.
-
priorityPriority service tier.
-
deferredDeferred service tier.
Required. Output only. The status of the interaction.
Possible values:
-
in_progressThe interaction is in progress.
-
requires_actionThe interaction requires action/input from the user.
-
completedThe interaction is completed.
-
failedThe interaction failed.
-
cancelledThe interaction was cancelled.
-
incompleteThe interaction is completed, but contains incomplete results (eg hitting max_tokens).
Output only. The steps that make up the interaction, if included in this event.
Output only. The time at which the response was last updated in ISO 8601 format.
usage Usage (optional)
Output only. Statistics on the interaction request's token usage.
فیلدها
cached_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of cached token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
grounding_tool_count array (GroundingToolCount) (optional)
Grounding tool count.
فیلدها
The number of grounding tool counts.
The grounding tool type associated with the count.
Possible values:
-
google_searchGrounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
-
google_mapsGrounding with Google Maps.
-
retrievalGrounding with customer's data, for example, VertexAISearch.
input_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of input token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
output_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of output token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
tool_use_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of tool-use token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
Number of tokens in the cached part of the prompt (the cached content).
Number of tokens in the prompt (context).
Total number of tokens across all the generated responses.
Number of tokens of thoughts for thinking models.
Total token count for the interaction request (prompt + responses + other internal tokens).
Number of tokens present in tool-use prompt(s).
InteractionStatusUpdate
The event_id token to be used to resume the interaction stream, from this event.
هیچ توضیحی ارائه نشده است.
Always set to "interaction.status_update" .
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
Possible values:
-
in_progressThe interaction is in progress.
-
requires_actionThe interaction requires action/input from the user.
-
completedThe interaction is completed.
-
failedThe interaction failed.
-
cancelledThe interaction was cancelled.
-
incompleteThe interaction is completed, but contains incomplete results (eg hitting max_tokens).
-
budget_exceededThe interaction was halted because the token budget was exceeded.
-
queuedThe interaction is queued, waiting for processing (eg waiting for off-peak capacity).
StepDelta
delta StepDeltaData (required)
هیچ توضیحی ارائه نشده است.
Possible Types
ArgumentsDelta
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
Always set to "arguments_delta" .
AudioDelta
The number of audio channels.
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
Possible values:
-
audio/wavWAV audio format
-
audio/mp3MP3 audio format
-
audio/aiffAIFF audio format
-
audio/aacAAC audio format
-
audio/oggOGG audio format
-
audio/flacFLAC audio format
-
audio/mpegMPEG audio format
-
audio/m4aM4A audio format
-
audio/l16L16 audio format
-
audio/opusOPUS audio format
-
audio/alawALAW audio format
-
audio/mulawMULAW audio format
The sample rate of the audio.
هیچ توضیحی ارائه نشده است.
Always set to "audio" .
هیچ توضیحی ارائه نشده است.
CodeExecutionCallDelta
arguments CodeExecutionCallArguments (required)
هیچ توضیحی ارائه نشده است.
فیلدها
The code to be executed.
Programming language of the `code`.
Possible values:
-
pythonPython >= 3.10, with numpy and simpy available.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "code_execution_call" .
CodeExecutionResultDelta
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "code_execution_result" .
DocumentDelta
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
Possible values:
-
application/pdfPDF document format
-
text/csvCSV document format
هیچ توضیحی ارائه نشده است.
Always set to "document" .
هیچ توضیحی ارائه نشده است.
FileSearchCallDelta
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "file_search_call" .
FileSearchResultDelta
result array (FileSearchResult) (required)
هیچ توضیحی ارائه نشده است.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "file_search_result" .
FunctionResultDelta
Required. ID to match the ID from the function call block.
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
Always set to "function_result" .
GoogleMapsCallDelta
arguments GoogleMapsCallArguments (optional)
The arguments to pass to the Google Maps tool.
فیلدها
The queries to be executed.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "google_maps_call" .
GoogleMapsResultDelta
result array (GoogleMapsResult) (optional)
The results of the Google Maps.
فیلدها
places array (Places) (optional)
The places that were found.
فیلدها
Title of the place.
The ID of the place, in `places/{place_id}` format.
review_snippets array (ReviewSnippet) (optional)
Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.
فیلدها
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
URI reference of the place.
Resource name of the Google Maps widget context token.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "google_maps_result" .
GoogleSearchCallDelta
arguments GoogleSearchCallArguments (required)
هیچ توضیحی ارائه نشده است.
فیلدها
Web search queries for the following-up web search.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "google_search_call" .
GoogleSearchResultDelta
هیچ توضیحی ارائه نشده است.
result array (GoogleSearchResult) (required)
هیچ توضیحی ارائه نشده است.
فیلدها
Web content snippet that can be embedded in a web page or an app webview.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "google_search_result" .
ImageDelta
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
Possible values:
-
image/pngPNG image format
-
image/jpegJPEG image format
-
image/webpWebP image format
-
image/heicHEIC image format
-
image/heifHEIF image format
-
image/gifGIF image format
-
image/bmpBMP image format
-
image/tiffTIFF image format
resolution MediaResolution (optional)
The resolution of the media.
Possible values
-
lowLow resolution.
-
mediumMedium resolution.
-
highHigh resolution.
-
ultra_highUltra high resolution.
هیچ توضیحی ارائه نشده است.
Always set to "image" .
هیچ توضیحی ارائه نشده است.
McpServerToolCallDelta
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
Always set to "mcp_server_tool_call" .
McpServerToolResultDelta
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
Always set to "mcp_server_tool_result" .
RetrievalCallDelta
Used by Vertex Retrieval tools such as Parallel AI, Exa AI, Vertex AI Search, etc. RetrievalType decides which tool is used.
arguments RetrievalStepArguments (required)
Required. The arguments to pass to the Retrieval tool.
فیلدها
Queries for Retrieval information.
The type of retrieval tools.
Possible values:
-
rag_storeThe type of retrieval tools.
-
exa_ai_searchThe type of retrieval tools.
-
parallel_ai_searchThe type of retrieval tools.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "retrieval_call" .
RetrievalResultDelta
Used by Vertex Retrieval tools such as Parallel AI, Exa AI, Vertex AI Search, etc. ToolResultDelta.type
Whether the retrieval resulted in an error.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "retrieval_result" .
TextAnnotationDelta
annotations array (Annotation) (optional)
Citation information for model-generated content.
Possible Types
FileCitation
A file citation annotation.
User provided metadata about the retrieved context.
The URI of the file.
End of the attributed segment, exclusive.
The name of the file.
Media ID in-case of image citations, if applicable.
Page number of the cited document, if applicable.
Source attributed for a portion of the text.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
هیچ توضیحی ارائه نشده است.
Always set to "file_citation" .
PlaceCitation
A place citation annotation.
End of the attributed segment, exclusive.
Title of the place.
The ID of the place, in `places/{place_id}` format.
review_snippets array (ReviewSnippet) (optional)
Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.
فیلدها
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
هیچ توضیحی ارائه نشده است.
Always set to "place_citation" .
URI reference of the place.
UrlCitation
A URL citation annotation.
End of the attributed segment, exclusive.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
The title of the URL.
هیچ توضیحی ارائه نشده است.
Always set to "url_citation" .
The URL.
WordInfo
Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.
End of the attributed segment, exclusive.
End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
Optional. Speaker label for this word (eg "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
The transcribed word.
هیچ توضیحی ارائه نشده است.
Always set to "word_info" .
هیچ توضیحی ارائه نشده است.
Always set to "text_annotation_delta" .
TextDelta
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
Always set to "text" .
ThoughtSignatureDelta
Signature to match the backend source to be part of the generation.
هیچ توضیحی ارائه نشده است.
Always set to "thought_signature" .
ThoughtSummaryDelta
A new summary item to be added to the thought.
هیچ توضیحی ارائه نشده است.
Always set to "thought_summary" .
UrlContextCallDelta
arguments UrlContextCallArguments (required)
هیچ توضیحی ارائه نشده است.
فیلدها
The URLs to fetch.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "url_context_call" .
UrlContextResultDelta
هیچ توضیحی ارائه نشده است.
result array (UrlContextResult) (required)
هیچ توضیحی ارائه نشده است.
فیلدها
The status of the URL retrieval.
Possible values:
-
successUrl retrieval is successful.
-
errorUrl retrieval is failed due to error.
-
paywallUrl retrieval is failed because the content is behind paywall.
-
unsafeUrl retrieval is failed because the content is unsafe.
The URL that was fetched.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "url_context_result" .
VideoDelta
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
Possible values:
-
video/mp4MP4 video format
-
video/mpegMPEG video format
-
video/mpgMPG video format
-
video/movMOV video format
-
video/aviAVI video format
-
video/x-flvFLV video format
-
video/webmWebM video format
-
video/wmvWMV video format
-
video/3gpp3GPP video format
-
video/jpeg2000JPEG 2000 video format
resolution MediaResolution (optional)
The resolution of the media.
Possible values
-
lowLow resolution.
-
mediumMedium resolution.
-
highHigh resolution.
-
ultra_highUltra high resolution.
هیچ توضیحی ارائه نشده است.
Always set to "video" .
هیچ توضیحی ارائه نشده است.
The event_id token to be used to resume the interaction stream, from this event.
هیچ توضیحی ارائه نشده است.
Always set to "step.delta" .
هیچ توضیحی ارائه نشده است.
metadata StepDeltaMetadata (optional)
هیچ توضیحی ارائه نشده است.
فیلدها
total_usage Usage (optional)
Statistics on the interaction request's token usage.
فیلدها
cached_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of cached token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
grounding_tool_count array (GroundingToolCount) (optional)
Grounding tool count.
فیلدها
The number of grounding tool counts.
The grounding tool type associated with the count.
Possible values:
-
google_searchGrounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
-
google_mapsGrounding with Google Maps.
-
retrievalGrounding with customer's data, for example, VertexAISearch.
input_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of input token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
output_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of output token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
tool_use_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of tool-use token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
Number of tokens in the cached part of the prompt (the cached content).
Number of tokens in the prompt (context).
Total number of tokens across all the generated responses.
Number of tokens of thoughts for thinking models.
Total token count for the interaction request (prompt + responses + other internal tokens).
Number of tokens present in tool-use prompt(s).
StepStart
The event_id token to be used to resume the interaction stream, from this event.
هیچ توضیحی ارائه نشده است.
Always set to "step.start" .
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
StepStop
The event_id token to be used to resume the interaction stream, from this event.
هیچ توضیحی ارائه نشده است.
Always set to "step.stop" .
هیچ توضیحی ارائه نشده است.
step_usage Usage (optional)
Model usage stats for this specific step.
فیلدها
cached_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of cached token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
grounding_tool_count array (GroundingToolCount) (optional)
Grounding tool count.
فیلدها
The number of grounding tool counts.
The grounding tool type associated with the count.
Possible values:
-
google_searchGrounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
-
google_mapsGrounding with Google Maps.
-
retrievalGrounding with customer's data, for example, VertexAISearch.
input_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of input token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
output_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of output token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
tool_use_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of tool-use token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
Number of tokens in the cached part of the prompt (the cached content).
Number of tokens in the prompt (context).
Total number of tokens across all the generated responses.
Number of tokens of thoughts for thinking models.
Total token count for the interaction request (prompt + responses + other internal tokens).
Number of tokens present in tool-use prompt(s).
usage Usage (optional)
Cumulative model usage stats from the start of the session.
فیلدها
cached_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of cached token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
grounding_tool_count array (GroundingToolCount) (optional)
Grounding tool count.
فیلدها
The number of grounding tool counts.
The grounding tool type associated with the count.
Possible values:
-
google_searchGrounding with Google Web Search and Image Search, & Web Grounding for Enterprise.
-
google_mapsGrounding with Google Maps.
-
retrievalGrounding with customer's data, for example, VertexAISearch.
input_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of input token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
output_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of output token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
tool_use_tokens_by_modality array (ModalityTokens) (optional)
A breakdown of tool-use token usage by modality.
فیلدها
modality ResponseModality (optional)
The modality associated with the token count.
Possible values
-
textIndicates the model should return text.
-
imageIndicates the model should return images.
-
audioIndicates the model should return audio.
-
videoIndicates the model should return video.
-
documentIndicates the model should return documents.
Number of tokens for the modality.
Number of tokens in the cached part of the prompt (the cached content).
Number of tokens in the prompt (context).
Total number of tokens across all the generated responses.
Number of tokens of thoughts for thinking models.
Total token count for the interaction request (prompt + responses + other internal tokens).
Number of tokens present in tool-use prompt(s).
مثالها
Error Event
{ "error": { "code": "not_found", "message": "Failed to get completed interaction: Result not found." }, "event_type": "error" }
Interaction Completed
{ "event_id": "evt_123", "event_type": "interaction.completed", "interaction": { "created": "2025-12-04T15:01:45Z", "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg", "model": "gemini-3.6-flash", "status": "completed", "updated": "2025-12-04T15:01:45Z" } }
Interaction Completed
{ "event_id": "evt_123", "event_type": "interaction.completed", "interaction": { "created": "2025-12-04T15:01:45Z", "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg", "model": "gemini-3-flash-preview", "object": "interaction", "status": "completed", "updated": "2025-12-04T15:01:45Z" } }
Interaction Created
{ "event_id": "evt_123", "event_type": "interaction.created", "interaction": { "created": "2025-12-04T15:01:45Z", "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg", "model": "gemini-3.6-flash", "status": "in_progress", "updated": "2025-12-04T15:01:45Z" } }
Interaction Created
{ "event_id": "evt_123", "event_type": "interaction.created", "interaction": { "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg", "model": "gemini-3-flash-preview", "object": "interaction", "status": "in_progress" } }
Interaction Status Update
{ "event_type": "interaction.status_update", "interaction_id": "v1_ChdTMjQ0YWJ5TUF1TzcxZThQdjRpcnFRcxIXUzI0NGFieU1BdU83MWU4UHY0aXJxUXM", "status": "in_progress" }
Step Delta
{ "delta": { "type": "text", "text": "Hello" }, "event_type": "step.delta", "index": 0 }
Step Start
{ "event_type": "step.start", "index": 0, "step": { "type": "model_output" } }
Step Stop
{ "event_type": "step.stop", "index": 0 }
ResponseFormat
Possible Types
AudioResponseFormat
Configuration for audio output format.
Bit rate in bits per second (bps). Only applicable for compressed formats (MP3, Opus).
The delivery mode for the audio output.
Possible values:
-
inlineAudio data is returned inline in the response.
-
uriAudio data is returned as a URI.
The MIME type of the audio output.
Possible values:
-
audio/mp3MP3 audio format.
-
audio/ogg_opusOGG Opus audio format.
-
audio/l16Raw PCM (L16) audio format.
-
audio/wavWAV audio format.
-
audio/alawA-law audio format.
-
audio/mulawMu-law audio format.
Sample rate in Hz.
هیچ توضیحی ارائه نشده است.
Always set to "audio" .
ImageResponseFormat
Configuration for image output format.
The aspect ratio for the image output.
Possible values:
-
1:11:1 aspect ratio.
-
2:32:3 aspect ratio.
-
3:23:2 aspect ratio.
-
3:43:4 aspect ratio.
-
4:34:3 aspect ratio.
-
4:54:5 aspect ratio.
-
5:45:4 aspect ratio.
-
9:169:16 aspect ratio.
-
16:916:9 aspect ratio.
-
21:921:9 aspect ratio.
-
1:81:8 aspect ratio.
-
8:18:1 aspect ratio.
-
1:41:4 aspect ratio.
-
4:14:1 aspect ratio.
The delivery mode for the image output.
Possible values:
-
inlineImage data is returned inline in the response.
-
uriImage data is returned as a URI.
The size of the image output.
Possible values:
-
512512px image size.
-
1K1K image size.
-
2K2K image size.
-
4K4K image size.
The MIME type of the image output.
Possible values:
-
image/jpegJPEG image format.
هیچ توضیحی ارائه نشده است.
Always set to "image" .
TextResponseFormat
Configuration for text output format.
The MIME type of the text output.
Possible values:
-
application/jsonJSON output format.
-
text/plainPlain text output format.
The JSON schema that the output should conform to. Only applicable when mime_type is application/json.
هیچ توضیحی ارائه نشده است.
Always set to "text" .
VideoResponseFormat
Configuration for video output format.
The aspect ratio for the video output.
Possible values:
-
16:916:9 aspect ratio.
-
9:169:16 aspect ratio.
The delivery mode for the video output.
Possible values:
-
inlineVideo data is returned inline in the response.
-
uriVideo data is returned as a URI.
The duration for the video output.
The Cloud Storage URI to store the video output. Required for Vertex if delivery mode is URI.
The video output resolution. Defaults to 720p.
Possible values:
-
360p360p resolution.
-
720p720p resolution.
-
1080p1080p resolution.
-
4k4K resolution.
هیچ توضیحی ارائه نشده است.
Always set to "video" .
مثالها
خروجی صدا
{ "type": "audio", "sample_rate": 24000 }
Image Output
{ "type": "image", "aspect_ratio": "16:9", "image_size": "1K", "mime_type": "image/jpeg" }
Text Output (JSON Schema)
{ "type": "text", "mime_type": "application/json", "schema": { "type": "object", "properties": { "ingredients": { "type": "array", "items": { "type": "string" } }, "recipe_name": { "type": "string" } }, "required": [ "ingredients", "recipe_name" ] } }
خروجی ویدئو
{ "type": "video", "aspect_ratio": "16:9", "delivery": "inline" }
قدم
A step in the interaction.
Possible Types
CodeExecutionCallStep
Code execution call step.
arguments CodeExecutionCallStepArguments (required)
Required. The arguments to pass to the code execution.
فیلدها
The code to be executed.
Programming language of the `code`.
Possible values:
-
pythonPython >= 3.10, with numpy and simpy available.
Required. A unique ID for this specific tool call.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "code_execution_call" .
CodeExecutionResultStep
Code execution result step.
Required. ID to match the ID from the function call block.
Whether the code execution resulted in an error.
Required. The output of the code execution.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "code_execution_result" .
FileSearchCallStep
File Search call step.
Required. A unique ID for this specific tool call.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "file_search_call" .
FileSearchResultStep
File Search result step.
Required. ID to match the ID from the function call block.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "file_search_result" .
FunctionCallStep
A function tool call step.
Required. The arguments to pass to the function.
Required. A unique ID for this specific tool call.
Required. The name of the tool to call.
هیچ توضیحی ارائه نشده است.
Always set to "function_call" .
FunctionResultStep
Result of a function tool call.
Required. ID to match the ID from the function call block.
Whether the tool call resulted in an error.
The name of the tool that was called.
Required. The result of the tool call.
هیچ توضیحی ارائه نشده است.
Always set to "function_result" .
GoogleMapsCallStep
Google Maps call step.
arguments GoogleMapsCallStepArguments (optional)
The arguments to pass to the Google Maps tool.
فیلدها
The queries to be executed.
Required. A unique ID for this specific tool call.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "google_maps_call" .
GoogleMapsResultStep
Google Maps result step.
Required. ID to match the ID from the function call block.
result array (GoogleMapsResultItem) (required)
هیچ توضیحی ارائه نشده است.
فیلدها
places array (GoogleMapsResultPlaces) (optional)
هیچ توضیحی ارائه نشده است.
فیلدها
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
review_snippets array (ReviewSnippet) (optional)
هیچ توضیحی ارائه نشده است.
فیلدها
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "google_maps_result" .
GoogleSearchCallStep
Google Search call step.
arguments GoogleSearchCallStepArguments (required)
Required. The arguments to pass to Google Search.
فیلدها
Web search queries for the following-up web search.
Required. A unique ID for this specific tool call.
The type of search grounding enabled.
Possible values:
-
web_searchSetting this field enables web search. Only text results are returned.
-
image_searchSetting this field enables image search. Image bytes are returned.
-
enterprise_web_searchSetting this field enables enterprise web search.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "google_search_call" .
GoogleSearchResultStep
Google Search result step.
Required. ID to match the ID from the function call block.
Whether the Google Search resulted in an error.
result array (GoogleSearchResultItem) (required)
Required. The results of the Google Search.
فیلدها
Web content snippet that can be embedded in a web page or an app webview.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "google_search_result" .
McpServerToolCallStep
MCPServer tool call step.
Required. The JSON object of arguments for the function.
Required. A unique ID for this specific tool call.
Required. The name of the tool which was called.
Required. The name of the used MCP server.
هیچ توضیحی ارائه نشده است.
Always set to "mcp_server_tool_call" .
McpServerToolResultStep
MCPServer tool result step.
Required. ID to match the ID from the function call block.
Name of the tool which is called for this specific tool call.
Required. The output from the MCP server call. Can be simple text or rich content.
The name of the used MCP server.
هیچ توضیحی ارائه نشده است.
Always set to "mcp_server_tool_result" .
ModelOutputStep
Output generated by the model.
هیچ توضیحی ارائه نشده است.
هیچ توضیحی ارائه نشده است.
Always set to "model_output" .
ThoughtStep
A thought step.
A signature hash for backend validation.
summary array (ThoughtSummaryContent) (optional)
A summary of the thought.
Possible Types
ImageContent
An image content block.
The image content.
The mime type of the image.
Possible values:
-
image/pngPNG image format
-
image/jpegJPEG image format
-
image/webpWebP image format
-
image/heicHEIC image format
-
image/heifHEIF image format
-
image/gifGIF image format
-
image/bmpBMP image format
-
image/tiffTIFF image format
resolution MediaResolution (optional)
The resolution of the media.
Possible values
-
lowLow resolution.
-
mediumMedium resolution.
-
highHigh resolution.
-
ultra_highUltra high resolution.
هیچ توضیحی ارائه نشده است.
Always set to "image" .
The URI of the image.
TextContent
A text content block.
annotations array (Annotation) (optional)
Citation information for model-generated content.
Possible Types
FileCitation
A file citation annotation.
User provided metadata about the retrieved context.
The URI of the file.
End of the attributed segment, exclusive.
The name of the file.
Media ID in-case of image citations, if applicable.
Page number of the cited document, if applicable.
Source attributed for a portion of the text.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
هیچ توضیحی ارائه نشده است.
Always set to "file_citation" .
PlaceCitation
A place citation annotation.
End of the attributed segment, exclusive.
Title of the place.
The ID of the place, in `places/{place_id}` format.
review_snippets array (ReviewSnippet) (optional)
Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.
فیلدها
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
No description provided.
Always set to "place_citation" .
URI reference of the place.
UrlCitation
A URL citation annotation.
End of the attributed segment, exclusive.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
The title of the URL.
هیچ توضیحی ارائه نشده است.
Always set to "url_citation" .
The URL.
WordInfo
Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.
End of the attributed segment, exclusive.
End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
Optional. Speaker label for this word (eg "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
The transcribed word.
No description provided.
Always set to "word_info" .
Required. The text content.
No description provided.
Always set to "text" .
No description provided.
Always set to "thought" .
UrlContextCallStep
URL context call step.
arguments UrlContextCallArguments (required)
Required. The arguments to pass to the URL context.
فیلدها
The URLs to fetch.
Required. A unique ID for this specific tool call.
A signature hash for backend validation.
هیچ توضیحی ارائه نشده است.
Always set to "url_context_call" .
UrlContextResultStep
URL context result step.
Required. ID to match the ID from the function call block.
Whether the URL context resulted in an error.
result array (UrlContextResult) (required)
Required. The results of the URL context.
فیلدها
The status of the URL retrieval.
Possible values:
-
successUrl retrieval is successful.
-
errorUrl retrieval is failed due to error.
-
paywallUrl retrieval is failed because the content is behind paywall.
-
unsafeUrl retrieval is failed because the content is unsafe.
The URL that was fetched.
A signature hash for backend validation.
No description provided.
Always set to "url_context_result" .
UserInputStep
Input provided by the user.
No description provided.
No description provided.
Always set to "user_input" .
مثالها
CodeExecutionCallStep
{ "type": "code_execution_call", "arguments": { "code": "print(sum(range(1, 11)))" }, "id": "code_call_71021" }
CodeExecutionResultStep
{ "type": "code_execution_result", "call_id": "code_call_71021", "result": "55\n" }
FileSearchCallStep
{ "type": "file_search_call", "id": "file_call_88192" }
FileSearchResultStep
{ "type": "file_search_result", "call_id": "file_call_88192" }
FunctionCallStep
{ "name": "get_weather", "type": "function_call", "arguments": { "location": "Boston, MA" }, "id": "call_98231" }
FunctionResultStep
{ "name": "get_weather", "type": "function_result", "call_id": "call_98231", "result": [ { "type": "text", "text": "{\"weather\":\"sunny\"}" } ] }
GoogleMapsCallStep
{ "type": "google_maps_call", "arguments": { "latitude": 37.7749, "longitude": -122.4194 }, "id": "maps_call_39201" }
GoogleMapsResultStep
{ "type": "google_maps_result", "call_id": "maps_call_39201", "result": [ { "name": "Golden Gate Park", "place_id": "ChIJIQBpAG2ahYAR9R7bNdTLg8M", "rating": 4.8 } ] }
GoogleSearchCallStep
{ "type": "google_search_call", "arguments": { "query": "Who won the men's 100m in Paris 2024?" }, "id": "search_call_19201" }
GoogleSearchResultStep
{ "type": "google_search_result", "call_id": "search_call_19201", "result": [ { "title": "Paris 2024 Olympics: Noah Lyles wins men's 100m gold", "url": "https://olympics.com/en/news/paris-2024-noah-lyles-wins-mens-100m-gold", "snippet": "American Noah Lyles won the Olympic men's 100m gold medal in a photo finish." } ] }
McpServerToolCallStep
{ "name": "calculate_tax", "type": "mcp_server_tool_call", "arguments": { "income": 120000, "state": "CA" }, "id": "mcp_call_29012", "server_name": "financial_mcp_server" }
McpServerToolResultStep
{ "type": "mcp_server_tool_result", "call_id": "mcp_call_29012", "result": { "tax_due": 32400 } }
ModelOutputStep
{ "type": "model_output", "content": [ { "type": "text", "text": "The capital of France is Paris." } ] }
ThoughtStep
{ "type": "thought", "signature": "thought_sig_abcd1234", "summary": [ { "type": "text", "text": "The model is searching Google for the capital of France." } ] }
UrlContextCallStep
{ "type": "url_context_call", "arguments": { "urls": [ "https://www.example.com" ] }, "id": "url_call_10219" }
UrlContextResultStep
{ "type": "url_context_result", "call_id": "url_call_10219", "result": [ { "title": "Example Domain", "url": "https://www.example.com", "snippet": "This domain is for use in illustrative examples in documents." } ] }
UserInputStep
{ "type": "user_input", "content": [ { "type": "text", "text": "What is the capital of France?" } ] }
EnvironmentConfig
Configuration for a custom environment.
فیلدها
Optional. The environment ID for the interaction. If specified, the request will update the existing environment instead of creating a new one.
Network configuration for the environment.
Possible values:
-
disabledTurns all network off.
sources array (Source) (optional)
No description provided.
فیلدها
The inline content if `type` is `INLINE`.
Optional encoding for inline content (eg `base64`).
The source of the environment. For Cloud Storage, this is the Cloud Storage path. For GitHub, this is the GitHub path.
Where the source should appear in the environment.
No description provided.
Possible values:
-
gcsA Cloud Storage bucket.
-
inlineInline content.
-
repositoryA generic repository. The protocol prefix in the source URL identifies the provider (eg, github://, gcs://).
-
skill_registryA skill resource from the Skill Registry Service. Skill: projects/{project}/locations/{location}/skills/{skill} SkillRevision: projects/{project}/locations/{location}/skills/{skill}/revisions/{revision} Support mounting all skills under a project: projects/{project}/locations/{location}/skills.
No description provided.
Always set to "remote" .
مثالها
Inline Sources
{ "type": "remote", "sources": [ { "type": "inline", "content": "You are a data analyst. Always include visualizations and export results as PDF.", "target": ".agents/AGENTS.md" }, { "type": "inline", "content": "---\nname: slide-maker\ndescription: Create HTML slide decks\n---\n# Slide Maker\n\nWhen asked to create a presentation:\n1. Analyze the input data\n2. Create an HTML slide deck with reveal.js\n3. Save to /workspace/output/slides.html", "target": ".agents/skills/slide-maker/SKILL.md" } ] }
External Sources
{ "type": "remote", "sources": [ { "type": "repository", "source": "https://github.com/my-org/my-skills.git", "target": ".agents/skills" }, { "type": "gcs", "source": "gs://my-bucket/my-folder", "target": "/workspace/data" } ] }
Network Allowlist
{ "type": "remote", "network": { "allowlist": [ { "domain": "pypi.org" }, { "domain": "*.github.com" } ] } }
Proxy Credentials
{ "type": "remote", "network": { "allowlist": [ { "domain": "api.github.com", "transform": { "Authorization": "Bearer YOUR_GITHUB_TOKEN" } } ] } }
EnvironmentNetworkEgressAllowlist
Outbound networking configuration for the sandbox. Accepts an object with an 'allowlist' array to restrict traffic, or the string 'disabled' to turn off all network access. Omit entirely to allow all outbound traffic with no header injection.
Possible Types
شیء
Outbound networking configuration for the sandbox. When specified, restricts which external domains the sandbox can reach. Omit entirely to allow all outbound traffic with no header injection.
allowlist array (AllowlistEntry) (optional)
List of allowed outbound domains. Only requests to listed domains are permitted. Use [{'domain': '*'}] to allow all domains while still injecting headers on specific ones.
فیلدها
Domain to allow outbound requests to. Supports wildcards (eg '*.googleapis.com'). Use '*' to allow all domains.
Headers to inject on all outbound requests matching this domain. Accepts a single dict or a list of dicts. The egress proxy injects these automatically.
رشته
Turns all network off.
Possible values
-
disabledTurns all network off.
مثالها
مثال
{ "allowlist": [ { "domain": "github.com", "transform": [ { "Authorization": "Bearer your-token" } ] }, { "domain": "*.googleapis.com" } ] }
ToolChoiceConfig
The tool choice configuration containing allowed tools.
فیلدها
allowed_tools AllowedTools (optional)
The allowed tools.
فیلدها
The mode of the tool choice.
Possible values:
-
autoAuto tool choice.
-
anyAny tool choice.
-
noneNo tool choice.
-
validatedValidated tool choice.
The names of the allowed tools.
مثالها
مثال
{ "allowed_tools": { "mode": "any", "tools": [ "my_tool" ] } }
ImageContent
An image content block.
فیلدها
The image content.
The mime type of the image.
Possible values:
-
image/pngPNG image format
-
image/jpegJPEG image format
-
image/webpWebP image format
-
image/heicHEIC image format
-
image/heifHEIF image format
-
image/gifGIF image format
-
image/bmpBMP image format
-
image/tiffTIFF image format
resolution MediaResolution (optional)
The resolution of the media.
Possible values
-
lowLow resolution.
-
mediumMedium resolution.
-
highHigh resolution.
-
ultra_highUltra high resolution.
No description provided.
Always set to "image" .
The URI of the image.
مثالها
تصویر
{ "type": "image", "data": "BASE64_ENCODED_IMAGE", "mime_type": "image/png" }
TextContent
A text content block.
فیلدها
annotations array (Annotation) (optional)
Citation information for model-generated content.
Possible Types
FileCitation
A file citation annotation.
User provided metadata about the retrieved context.
The URI of the file.
End of the attributed segment, exclusive.
The name of the file.
Media ID in-case of image citations, if applicable.
Page number of the cited document, if applicable.
Source attributed for a portion of the text.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
No description provided.
Always set to "file_citation" .
PlaceCitation
A place citation annotation.
End of the attributed segment, exclusive.
Title of the place.
The ID of the place, in `places/{place_id}` format.
review_snippets array (ReviewSnippet) (optional)
Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.
فیلدها
The ID of the review snippet.
Title of the review.
A link that corresponds to the user review on Google Maps.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
No description provided.
Always set to "place_citation" .
URI reference of the place.
UrlCitation
A URL citation annotation.
End of the attributed segment, exclusive.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
The title of the URL.
No description provided.
Always set to "url_citation" .
The URL.
WordInfo
Word-level ASR annotation for transcription output. Carries the word text, optional timing, and optional speaker attribution.
End of the attributed segment, exclusive.
End offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
Optional. Speaker label for this word (eg "spk_1", "spk_2"). Present when diarization_mode is set in TranscriptionConfig.
Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.
Start offset in time of the word relative to the start of the audio. Present when timestamp_granularities contains "word".
The transcribed word.
No description provided.
Always set to "word_info" .
Required. The text content.
No description provided.
Always set to "text" .
مثالها
متن
{ "type": "text", "text": "Hello, how are you?" }