Gemini Interactions API

رابط برنامه‌نویسی کاربردی (API) تعاملات جمینی (Gemini Interactions API) به توسعه‌دهندگان اجازه می‌دهد تا با استفاده از مدل‌های جمینی، برنامه‌های هوش مصنوعی مولد (generative AI applications) بسازند. جمینی توانمندترین مدل ما است که از پایه برای چندوجهی بودن ساخته شده است. این مدل می‌تواند انواع مختلف اطلاعات از جمله زبان، تصاویر، صدا، ویدئو و کد را تعمیم داده و به طور یکپارچه درک کند، در میان آنها عمل کند و ترکیب کند. می‌توانید از رابط برنامه‌نویسی کاربردی جمینی برای موارد استفاده‌ای مانند استدلال در متن و تصاویر، تولید محتوا، عامل‌های گفتگو، سیستم‌های خلاصه‌سازی و طبقه‌بندی و موارد دیگر استفاده کنید.

نسخه API: v1beta v1

ایجاد تعامل

ارسال به آدرس https://generativelanguage.googleapis.com/v1beta/interactions

یک تعامل جدید ایجاد می‌کند.

پارامترهای مسیر/پرس‌وجو

رشته api_version (الزامی)

از کدام نسخه API استفاده کنیم.

درخواست بدنه

بدنه درخواست شامل داده‌هایی با ساختار زیر است:

مدل ModelOption (اختیاری)

نام «مدل» مورد استفاده برای تولید تعامل.
در صورت عدم ارائه «عامل»، الزامی است.

مدلی که اعلان شما را تکمیل می‌کند.\n\nبرای جزئیات بیشتر به [models](https://ai.google.dev/gemini-api/docs/models) مراجعه کنید.

مقادیر ممکن

  • models/gemini-2.5-flash-lite

    کوچکترین و مقرون به صرفه ترین مدل ما، ساخته شده برای استفاده در مقیاس بزرگ.

  • models/gemini-2.5-flash-image

    مدل تولید تصویر بومی ما، که برای سرعت، انعطاف‌پذیری و درک متنی بهینه شده است. ورودی و خروجی متن با همان قیمت ۲.۵ فلش ارائه می‌شود.

  • models/gemini-3.1-flash-lite

    مقرون‌به‌صرفه‌ترین مدل ما، بهینه‌شده برای وظایف عامل‌محور با حجم بالا، ترجمه و پردازش داده‌های ساده.

  • models/gemini-3.1-flash-image

    هوش بصری حرفه‌ای با کارایی فوق‌العاده و قابلیت‌های تولید محتوای مبتنی بر واقعیت.

  • models/gemini-3.5-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

  • models/gemini-3.6-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

  • models/gemini-3.7-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

گزینه عامل (اختیاری)

نام «عامل» مورد استفاده برای ایجاد تعامل.
در صورت عدم ارائه «مدل»، الزامی است.

عاملی که باید با آن تعامل داشت.

مقادیر ممکن

  • deep-research-pro-preview-12-2025

    نماینده تحقیقات عمیق جمینی

  • deep-research-preview-04-2026

    نماینده تحقیقات عمیق جمینی

  • deep-research-max-preview-04-2026

    مامور مکس تحقیقات عمیق جمینی

  • antigravity-preview-05-2026

    از عامل مدیریت‌شده‌ی Antigravity برای انجام وظایف چند مرحله‌ای که نیاز به استدلال، عملیات فایل و استفاده از ابزار دارند، استفاده کنید.

ورودی محتوا یا آرایه ( Content ) یا آرایه ( Step ) یا رشته (الزامی)

ورودی‌های تعامل (مشترک برای مدل و عامل).

رشته system_instruction (اختیاری)

دستورالعمل سیستم برای تعامل.

آرایه ابزارها ( ابزار ) (اختیاری)

فهرستی از اعلان‌های ابزار که مدل ممکن است در طول تعامل فراخوانی کند.

response_format فرمت پاسخ یا آرایه ( ResponseFormat ) (اختیاری)

تأکید می‌کند که پاسخ تولید شده یک شیء JSON است که با طرحواره JSON مشخص شده در این فیلد مطابقت دارد.

جریان بولی (اختیاری)

فقط ورودی. اینکه آیا تعامل پخش زنده خواهد شد یا خیر.

ذخیره بولی (اختیاری)

فقط ورودی. آیا پاسخ و درخواست برای بازیابی بعدی ذخیره شود یا خیر.

مقدار بولی پس‌زمینه (اختیاری)

فقط ورودی. اینکه آیا تعامل مدل در پس‌زمینه اجرا شود یا خیر.

generation_config GenerationConfig (اختیاری)

پیکربندی مدل
پارامترهای پیکربندی برای تعامل مدل.
جایگزینی برای `agent_config`. فقط زمانی قابل اجرا است که `model` تنظیم شده باشد.

پارامترهای پیکربندی برای تعاملات مدل.

فیلدها

عدد صحیح max_output_tokens (اختیاری)

حداکثر تعداد توکن‌هایی که باید در پاسخ گنجانده شوند.

عدد صحیح اولیه (اختیاری)

بذر مورد استفاده در رمزگشایی برای تکرارپذیری.

speech_config SpeakerConfig یا آرایه (SpeechConfig) (اختیاری)

اختیاری. پیکربندی گفتار و چند بلندگو.

آرایه stop_sequences (رشته) (اختیاری)

فهرستی از توالی‌های کاراکتری که تعامل خروجی را متوقف می‌کنند.

سطح_فکریسطح_فکری ( اختیاری )

سطح توکن‌های فکری که مدل باید تولید کند.

مقادیر ممکن

  • minimal

    کم یا بدون فکر کردن.

  • low

    سطح فکری پایین.

  • medium

    سطح فکری متوسط.

  • high

    سطح فکری بالا.

خلاصه‌های تفکر ( اختیاری)

اینکه آیا خلاصه نظرات در پاسخ گنجانده شود یا خیر.

مقادیر ممکن

  • auto

    خلاصه‌های تفکر خودکار

  • none

    بدون خلاصه نویسی فکری.

tool_choice ToolChoiceConfig یا enum (رشته‌ای) (اختیاری)

پیکربندی انتخاب ابزار.

مقادیر ممکن:

  • auto

    انتخاب خودکار ابزار.

  • any

    هر انتخاب ابزاری.

  • none

    بدون انتخاب ابزار.

  • validated

    انتخاب ابزار معتبر.

پیکربندی عامل پویا (اختیاری)

پیکربندی عامل
پیکربندی برای عامل.
جایگزینی برای `generation_config`. فقط زمانی قابل اجرا است که `agent` تنظیم شده باشد.

پیکربندی برای عامل‌های پویا

فیلدها

نوع شیء (اختیاری)

هیچ توضیحی ارائه نشده است.

همیشه روی "dynamic" تنظیم شود.

شیء برچسب‌ها (اختیاری)

برچسب‌هایی با فراداده‌های تعریف‌شده توسط کاربر برای درخواست.

رشته‌ی max_total_tokens (اختیاری)

حداکثر مجموع توکن‌ها برای اجرای عامل.

رشته‌ی previous_interaction_id (اختیاری)

شناسه‌ی تعامل قبلی، در صورت وجود.

آرایه safety_settings (SafetySetting) (اختیاری)

تنظیمات ایمنی برای تعامل.

پاسخ

یک منبع تعامل (Interaction) را برمی‌گرداند.

درخواست ساده

پاسخ نمونه

{
  "created": "2025-11-26T12:25:15Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "Hello! I'm functioning perfectly and ready to assist you.\n\nHow are you doing today?"
        }
      ]
    }
  ],
  "updated": "2025-11-26T12:25:15Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 7
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 7,
    "total_output_tokens": 20,
    "total_thought_tokens": 22,
    "total_tokens": 49,
    "total_tool_use_tokens": 0
  }
}

چند نوبتی

پاسخ نمونه

{
  "created": "2025-11-26T12:22:47Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "The capital of France is Paris."
        }
      ]
    }
  ],
  "updated": "2025-11-26T12:22:47Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 50
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 50,
    "total_output_tokens": 10,
    "total_thought_tokens": 0,
    "total_tokens": 60,
    "total_tool_use_tokens": 0
  }
}

ورودی تصویر

پاسخ نمونه

{
  "created": "2025-11-26T12:22:47Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "A white humanoid robot with glowing blue eyes stands holding a red skateboard."
        }
      ]
    }
  ],
  "updated": "2025-11-26T12:22:47Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 10
      },
      {
        "modality": "image",
        "tokens": 258
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 268,
    "total_output_tokens": 20,
    "total_thought_tokens": 0,
    "total_tokens": 288,
    "total_tool_use_tokens": 0
  }
}

فراخوانی تابع

پاسخ نمونه

{
  "created": "2025-11-26T12:22:47Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "requires_action",
  "steps": [
    {
      "name": "get_weather",
      "type": "function_call",
      "arguments": {
        "location": "Boston, MA"
      },
      "id": "gth23981"
    }
  ],
  "updated": "2025-11-26T12:22:47Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 100
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 100,
    "total_output_tokens": 25,
    "total_thought_tokens": 0,
    "total_tokens": 125,
    "total_tool_use_tokens": 50
  }
}

لغو یک تعامل

ارسال https://generativelanguage.googleapis.com/v1beta/interactions/{id}/cancel

یک تعامل را بر اساس شناسه لغو می‌کند. این فقط برای تعاملات پس‌زمینه‌ای که هنوز در حال اجرا هستند، اعمال می‌شود.

پارامترهای مسیر/پرس‌وجو

رشته api_version (الزامی)

از کدام نسخه API استفاده کنیم.

رشته شناسه (الزامی)

شناسه منحصر به فرد تعاملی که باید لغو شود.

پاسخ

یک منبع تعامل (Interaction) را برمی‌گرداند.

لغو تعامل

پاسخ نمونه

{
  "created": "2025-11-26T12:25:15Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "cancelled",
  "updated": "2025-11-26T12:25:15Z"
}

بازیابی یک تعامل

دریافت کنید https://generativelanguage.googleapis.com/v1beta/interactions/{id}

جزئیات کامل یک تعامل واحد را بر اساس `Interaction.id` آن بازیابی می‌کند.

پارامترهای مسیر/پرس‌وجو

رشته api_version (الزامی)

از کدام نسخه API استفاده کنیم.

رشته شناسه (الزامی)

شناسه منحصر به فرد تعاملی که قرار است بازیابی شود.

رشته last_event_id (اختیاری)

اختیاری. در صورت تنظیم، جریان تعامل را از بخش بعدی پس از رویداد مشخص شده توسط شناسه رویداد از سر می‌گیرد. فقط در صورتی قابل استفاده است که `stream` برابر با true باشد.

جریان بولی (اختیاری)

اگر روی درست تنظیم شود، محتوای تولید شده به صورت تدریجی پخش می‌شود.

پیش‌فرض: False

پاسخ

یک منبع تعامل (Interaction) را برمی‌گرداند.

تعامل دریافت کنید

پاسخ نمونه

{
  "created": "2025-11-26T12:25:15Z",
  "id": "v1_ChdPU0F4YWFtNkFwS2kxZThQZ05lbXdROBIXT1NBeGFhbTZBcEtpMWU4UGdOZW13UTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "I'm doing great, thank you for asking! How can I help you today?"
        }
      ]
    }
  ],
  "updated": "2025-11-26T12:25:15Z"
}

حذف یک تعامل

https://generativelanguage.googleapis.com/v1beta/interactions/{id} را حذف کنید

تعامل را بر اساس شناسه حذف می‌کند.

پارامترهای مسیر/پرس‌وجو

رشته api_version (الزامی)

از کدام نسخه API استفاده کنیم.

رشته شناسه (الزامی)

شناسه منحصر به فرد تعاملی که باید حذف شود.

پاسخ

در صورت موفقیت، پاسخ خالی است.

حذف

منابع

تعامل

منبع تعامل.

فیلدها

گزینه عامل (اختیاری)

نام «عامل» مورد استفاده برای ایجاد تعامل.

عاملی که باید با آن تعامل داشت.

مقادیر ممکن

  • deep-research-pro-preview-12-2025

    نماینده تحقیقات عمیق جمینی

  • deep-research-preview-04-2026

    نماینده تحقیقات عمیق جمینی

  • deep-research-max-preview-04-2026

    مامور مکس تحقیقات عمیق جمینی

  • antigravity-preview-05-2026

    از عامل مدیریت‌شده‌ی Antigravity برای انجام وظایف چند مرحله‌ای که نیاز به استدلال، عملیات فایل و استفاده از ابزار دارند، استفاده کنید.

پیکربندی عامل پویا (اختیاری)

پارامترهای پیکربندی برای تعامل عامل.

پیکربندی برای عامل‌های پویا

فیلدها

نوع شیء (اختیاری)

هیچ توضیحی ارائه نشده است.

همیشه روی "dynamic" تنظیم شود.

رشته ایجاد شده (اختیاری)

فقط خروجی. زمانی که پاسخ در قالب ISO 8601 (YYYY-MM-DDThh:mm:ssZ) ایجاد شده است.

آرایه خطاها (Error) (اختیاری)

فقط خروجی. خطاهای تشخیصی / خطاهای پلتفرم که در تعامل ثبت شده‌اند.

پیام خطا از یک تعامل.

فیلدها

رشته کد (اختیاری)

یک URI که نوع خطا را مشخص می‌کند.

رشته پیام (اختیاری)

یک پیام خطا که برای انسان قابل خواندن باشد.

رشته شناسه (اختیاری)

الزامی. فقط خروجی. یک شناسه منحصر به فرد برای تکمیل تعامل.

پیش‌فرض‌ها به:

ورودی محتوا یا آرایه ( Content ) یا آرایه ( Step ) یا رشته (اختیاری)

ورودی برای تعامل.

شیء برچسب‌ها (اختیاری)

برچسب‌هایی با فراداده‌های تعریف‌شده توسط کاربر برای درخواست.

رشته‌ی max_total_tokens (اختیاری)

حداکثر مجموع توکن‌ها برای اجرای عامل.

مدل ModelOption (اختیاری)

نام «مدل» مورد استفاده برای تولید تعامل.

مدلی که اعلان شما را تکمیل می‌کند.\n\nبرای جزئیات بیشتر به [models](https://ai.google.dev/gemini-api/docs/models) مراجعه کنید.

مقادیر ممکن

  • models/gemini-2.5-flash-lite

    کوچکترین و مقرون به صرفه ترین مدل ما، ساخته شده برای استفاده در مقیاس بزرگ.

  • models/gemini-2.5-flash-image

    مدل تولید تصویر بومی ما، که برای سرعت، انعطاف‌پذیری و درک متنی بهینه شده است. ورودی و خروجی متن با همان قیمت ۲.۵ فلش ارائه می‌شود.

  • models/gemini-3.1-flash-lite

    مقرون‌به‌صرفه‌ترین مدل ما، بهینه‌شده برای وظایف عامل‌محور با حجم بالا، ترجمه و پردازش داده‌های ساده.

  • models/gemini-3.1-flash-image

    هوش بصری حرفه‌ای با کارایی فوق‌العاده و قابلیت‌های تولید محتوای مبتنی بر واقعیت.

  • models/gemini-3.5-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

  • models/gemini-3.6-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

  • models/gemini-3.7-flash

    هوشمندترین مدل ما برای عملکرد مرزی پایدار در وظایف عامل‌دار و کدنویسی.

محتوای صوتی output_audio (اختیاری)

آخرین صدای تولید شده توسط مدل در پاسخ به درخواست فعلی. توجه: این توسط SDK اضافه شده است.

یک بلوک محتوای صوتی.

فیلدها

عدد صحیح کانال‌ها (اختیاری)

تعداد کانال‌های صوتی

رشته داده (اختیاری)

محتوای صوتی.

mime_type enum (رشته) (اختیاری)

نوع مایم صدا.

مقادیر ممکن:

  • audio/wav

    فرمت صوتی WAV

  • audio/mp3

    فرمت صوتی MP3

  • audio/aiff

    فرمت صوتی AIFF

  • audio/aac

    فرمت صوتی AAC

  • audio/ogg

    فرمت صوتی OGG

  • audio/flac

    فرمت صوتی FLAC

  • audio/mpeg

    فرمت صوتی MPEG

  • audio/m4a

    فرمت صوتی M4A

  • audio/l16

    فرمت صوتی L16

  • audio/opus

    فرمت صوتی OPUS

  • audio/alaw

    فرمت صوتی ALAW

  • audio/mulaw

    فرمت صوتی MULAW

عدد صحیح sample_rate (اختیاری)

نرخ نمونه‌برداری صدا.

نوع شیء (اختیاری)

هیچ توضیحی ارائه نشده است.

همیشه روی "audio" تنظیم شود.

رشته uri (اختیاری)

آدرس اینترنتی (URI) فایل صوتی.

محتوای تصویر خروجی (اختیاری)

آخرین تصویری که توسط مدل در پاسخ به درخواست فعلی تولید شده است. توجه: این تصویر توسط SDK اضافه شده است.

رشته‌ی output_text (اختیاری)

متن به هم پیوسته از آخرین خروجی مدل در پاسخ به درخواست فعلی. توجه: این توسط SDK اضافه شده است.

رشته‌ی previous_interaction_id (اختیاری)

شناسه‌ی تعامل قبلی، در صورت وجود.

response_format فرمت پاسخ یا آرایه ( ResponseFormat ) (اختیاری)

تأکید می‌کند که پاسخ تولید شده یک شیء JSON است که با طرحواره JSON مشخص شده در این فیلد مطابقت دارد.

آرایه safety_settings (SafetySetting) (اختیاری)

تنظیمات ایمنی برای تعامل.

شمارش وضعیت (رشته) (اختیاری)

الزامی. فقط خروجی. وضعیت تعامل.

مقادیر ممکن:

  • in_progress

    تعامل در حال انجام است.

  • requires_action

    این تعامل نیاز به اقدام/ورودی از سوی کاربر دارد.

  • completed

    تعامل تکمیل شده است.

  • failed

    تعامل شکست خورد.

  • cancelled

    تعامل لغو شد.

  • incomplete

    تعامل تکمیل شده است، اما شامل نتایج ناقص است (مثلاً رسیدن به max_tokens).

آرایه گام‌ها ( Step ) (اختیاری)

فقط خروجی. مراحلی که تعامل را تشکیل می‌دهند، زمانی که در پاسخ گنجانده شوند.

رشته system_instruction (اختیاری)

دستورالعمل سیستم برای تعامل.

آرایه ابزارها ( ابزار ) (اختیاری)

فهرستی از اعلان‌های ابزار که مدل ممکن است در طول تعامل فراخوانی کند.

رشته به‌روزرسانی‌شده (اختیاری)

فقط خروجی. زمانی که پاسخ آخرین بار در قالب ISO 8601 (YYYY-MM-DDThh:mm:ssZ) به‌روزرسانی شده است.

کاربرد (اختیاری )

فقط خروجی. آمار مربوط به میزان استفاده از توکن درخواست تعامل.

آمار مربوط به میزان استفاده از توکن درخواست تعامل.

فیلدها

آرایه cached_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک میزان استفاده از توکن‌های ذخیره‌شده بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه grounding_tool_count (GroundingToolCount) (اختیاری)

تعداد ابزار اتصال به زمین

تعداد ابزار اتصال به زمین مهم است.

فیلدها

شمارش عدد صحیح (اختیاری)

تعداد ابزار اتصال به زمین مهم است.

نوع enum (رشته) (اختیاری)

نوع ابزار اتصال زمین مرتبط با شمارش.

مقادیر ممکن:

  • google_search

    اتصال به زمین با جستجوی وب و جستجوی تصویر گوگل، و اتصال به زمین وب برای سازمان‌ها.

  • google_maps

    اتصال به زمین با نقشه‌های گوگل.

آرایه input_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک استفاده از توکن ورودی بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه output_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک استفاده از توکن خروجی بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه tool_use_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک میزان استفاده از توکن‌های ابزار بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

عدد صحیح total_cached_tokens (اختیاری)

تعداد توکن‌ها در بخش ذخیره‌شده‌ی اعلان (محتوای ذخیره‌شده).

عدد صحیح total_input_tokens (اختیاری)

تعداد توکن‌ها در اعلان (زمینه).

total_output_tokens عدد صحیح (اختیاری)

تعداد کل توکن‌ها در تمام پاسخ‌های تولید شده.

total_thought_tokens عدد صحیح (اختیاری)

تعداد توکن‌های افکار برای مدل‌های تفکر.

عدد صحیح total_tokens (اختیاری)

تعداد کل توکن‌ها برای درخواست تعامل (درخواست + پاسخ‌ها + سایر توکن‌های داخلی).

عدد صحیح total_tool_use_tokens (اختیاری)

تعداد توکن‌های موجود در اعلان(های) استفاده از ابزار.

مثال‌ها

مثال

{
  "created": "2025-12-04T15:01:45Z",
  "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg",
  "model": "gemini-3.6-flash",
  "object": "interaction",
  "status": "completed",
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "Hello! I'm doing well, functioning as expected. Thank you for asking! How are you doing today?"
        }
      ]
    }
  ],
  "updated": "2025-12-04T15:01:45Z",
  "usage": {
    "input_tokens_by_modality": [
      {
        "modality": "text",
        "tokens": 7
      }
    ],
    "total_cached_tokens": 0,
    "total_input_tokens": 7,
    "total_output_tokens": 23,
    "total_thought_tokens": 49,
    "total_tokens": 79,
    "total_tool_use_tokens": 0
  }
}

مدل‌های داده

محتوا

محتوای پاسخ.

انواع ممکن

محتوای صوتی

یک بلوک محتوای صوتی.

عدد صحیح کانال‌ها (اختیاری)

تعداد کانال‌های صوتی

رشته داده (اختیاری)

محتوای صوتی.

mime_type enum (رشته) (اختیاری)

نوع مایم صدا.

مقادیر ممکن:

  • audio/wav

    فرمت صوتی WAV

  • audio/mp3

    فرمت صوتی MP3

  • audio/aiff

    فرمت صوتی AIFF

  • audio/aac

    فرمت صوتی AAC

  • audio/ogg

    فرمت صوتی OGG

  • audio/flac

    فرمت صوتی FLAC

  • audio/mpeg

    فرمت صوتی MPEG

  • audio/m4a

    فرمت صوتی M4A

  • audio/l16

    فرمت صوتی L16

  • audio/opus

    فرمت صوتی OPUS

  • audio/alaw

    فرمت صوتی ALAW

  • audio/mulaw

    فرمت صوتی MULAW

عدد صحیح sample_rate (اختیاری)

نرخ نمونه‌برداری صدا.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "audio" تنظیم شود.

رشته uri (اختیاری)

آدرس اینترنتی (URI) فایل صوتی.

محتوای سند

یک بلوک محتوای سند.

رشته داده (اختیاری)

محتوای سند.

mime_type enum (رشته) (اختیاری)

نوع MIME سند.

مقادیر ممکن:

  • application/pdf

    فرمت سند PDF

  • text/csv

    فرمت سند CSV

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "document" تنظیم شود.

رشته uri (اختیاری)

آدرس اینترنتی (URI) سند.

محتوای تصویر

یک بلوک محتوای تصویر.

رشته داده (اختیاری)

محتوای تصویر.

mime_type enum (رشته) (اختیاری)

نوع مایم تصویر.

مقادیر ممکن:

  • image/png

    فرمت تصویر PNG

  • image/jpeg

    فرمت تصویر JPEG

  • image/webp

    فرمت تصویر WebP

  • image/heic

    فرمت تصویر HEIC

  • image/heif

    فرمت تصویر HEIF

  • image/gif

    فرمت تصویر GIF

  • image/bmp

    فرمت تصویر BMP

  • image/tiff

    فرمت تصویر TIFF

وضوح تصویر MediaResolution (اختیاری)

قطعنامه رسانه‌ها.

مقادیر ممکن

  • low

    وضوح پایین.

  • medium

    وضوح متوسط.

  • high

    وضوح بالا.

  • ultra_high

    وضوح فوق العاده بالا.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "image" تنظیم شود.

رشته uri (اختیاری)

آدرس اینترنتی (URI) تصویر.

محتوای متن

یک بلوک محتوای متنی.

آرایه حاشیه‌نویسی‌ها (Annotation) (اختیاری)

اطلاعات استناد برای محتوای تولید شده توسط مدل.

اطلاعات استناد برای محتوای تولید شده توسط مدل.

انواع ممکن

استناد به فایل

حاشیه‌نویسی استناد به فایل.

شیء custom_metadata (اختیاری)

فراداده‌های ارائه شده توسط کاربر در مورد متن بازیابی شده.

رشته document_uri (اختیاری)

آدرس اینترنتی (URI) فایل.

عدد صحیح end_index (اختیاری)

پایان بخش منسوب، منحصر به فرد.

رشته نام فایل (اختیاری)

نام فایل.

رشته media_id (اختیاری)

شناسه رسانه در صورت استناد به تصویر، در صورت لزوم.

عدد صحیح شماره صفحه (اختیاری)

شماره صفحه سند ذکر شده، در صورت وجود.

رشته منبع (اختیاری)

منبع برای بخشی از متن ذکر شده است.

عدد صحیح start_index (اختیاری)

شروع بخش پاسخی که به این منبع نسبت داده می‌شود. اندیس، شروع بخش را نشان می‌دهد که بر حسب بایت اندازه‌گیری می‌شود.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "file_citation" تنظیم شود.

استناد به مکان

حاشیه‌نویسی برای استناد به مکان.

عدد صحیح end_index (اختیاری)

پایان بخش منسوب، منحصر به فرد.

رشته نام (اختیاری)

عنوان مکان.

رشته place_id (اختیاری)

شناسه مکان، با فرمت `places/{place_id}`.

آرایه review_snippets (ReviewSnippet) (اختیاری)

گزیده‌هایی از نظرات که برای تولید پاسخ در مورد ویژگی‌های یک مکان مشخص در نقشه‌های گوگل استفاده می‌شوند.

بخشی از نقد کاربر را که به سوالی در مورد ویژگی‌های یک مکان خاص در نقشه‌های گوگل پاسخ می‌دهد، در بر می‌گیرد.

فیلدها

رشته review_id (اختیاری)

شناسه‌ی قطعه نقد و بررسی.

رشته عنوان (اختیاری)

عنوان نقد.

رشته آدرس اینترنتی (اختیاری)

لینکی که مربوط به نظر کاربر در نقشه گوگل باشد.

عدد صحیح start_index (اختیاری)

شروع بخش پاسخی که به این منبع نسبت داده می‌شود. اندیس، شروع بخش را نشان می‌دهد که بر حسب بایت اندازه‌گیری می‌شود.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "place_citation" تنظیم شود.

رشته آدرس اینترنتی (اختیاری)

مرجع URI آن مکان.

استناد به آدرس اینترنتی

حاشیه‌نویسی استناد URL.

عدد صحیح end_index (اختیاری)

پایان بخش منسوب، منحصر به فرد.

عدد صحیح start_index (اختیاری)

شروع بخش پاسخی که به این منبع نسبت داده می‌شود. اندیس، شروع بخش را نشان می‌دهد که بر حسب بایت اندازه‌گیری می‌شود.

رشته عنوان (اختیاری)

عنوان URL.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "url_citation" تنظیم شود.

رشته آدرس اینترنتی (اختیاری)

آدرس اینترنتی (URL).

رشته متن (الزامی)

محتوای متن الزامی است.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "text" تنظیم شود.

مثال‌ها

صوتی

{
  "type": "audio",
  "data": "BASE64_ENCODED_AUDIO",
  "mime_type": "audio/wav"
}

سند

{
  "type": "document",
  "data": "BASE64_ENCODED_DOCUMENT",
  "mime_type": "application/pdf"
}

تصویر

{
  "type": "image",
  "data": "BASE64_ENCODED_IMAGE",
  "mime_type": "image/png"
}

متن

{
  "type": "text",
  "text": "Hello, how are you?"
}

ابزار

ابزاری که می‌تواند توسط مدل مورد استفاده قرار گیرد.

انواع ممکن

اجرای کد

ابزاری که می‌تواند توسط مدل برای اجرای کد استفاده شود.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "code_execution" تنظیم شود.

جستجوی فایل

ابزاری که می‌تواند توسط مدل برای جستجوی فایل‌ها استفاده شود.

آرایه file_search_store_names (رشته‌ای) (اختیاری)

جستجوی فایل، نام‌های فروشگاه را برای جستجو ذخیره می‌کند.

رشته metadata_filter (اختیاری)

فیلتر فراداده برای اعمال روی اسناد و تکه‌های بازیابی معنایی.

عدد صحیح top_k (اختیاری)

تعداد تکه‌های بازیابی معنایی که باید بازیابی شوند.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "file_search" تنظیم شود.

عملکرد

ابزاری که می‌تواند توسط مدل مورد استفاده قرار گیرد.

رشته توضیحات (اختیاری)

شرحی از تابع.

رشته نام (اختیاری)

نام تابع.

پارامتر شیء (اختیاری)

طرحواره JSON برای پارامترهای تابع.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "function" تنظیم شود.

گوگل مپ

ابزاری که می‌تواند توسط مدل برای فراخوانی نقشه‌های گوگل استفاده شود.

enable_widget بولی (اختیاری)

اینکه آیا در نتیجه فراخوانی ابزار، توکن زمینه ویجت برگردانده شود یا خیر.

شماره عرض جغرافیایی (اختیاری)

عرض جغرافیایی محل کاربر.

شماره طول جغرافیایی (اختیاری)

طول جغرافیایی موقعیت مکانی کاربر.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "google_maps" تنظیم شود.

جستجوی گوگل

ابزاری که می‌تواند توسط مدل برای جستجو در گوگل استفاده شود.

آرایه search_types (enum (رشته)) (اختیاری)

انواع زمینه‌سازی جستجو برای فعال‌سازی.

مقادیر ممکن:

  • web_search

    تنظیم این فیلد، جستجوی وب را فعال می‌کند. فقط نتایج متنی برگردانده می‌شوند.

  • image_search

    تنظیم این فیلد، جستجوی تصویر را فعال می‌کند. بایت‌های تصویر برگردانده می‌شوند.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "google_search" تنظیم شود.

متن آدرس

ابزاری که می‌تواند توسط مدل برای دریافت متن URL استفاده شود.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "url_context" تنظیم شود.

مثال‌ها

اجرای کد

جستجوی فایل

عملکرد

گوگل مپ

جستجوی گوگل

متن آدرس

رویداد تعامل

انواع ممکن

تفکیک‌کننده چندریختی: event_type

رویداد خطا

خطا ( اختیاری )

هیچ توضیحی ارائه نشده است.

پیام خطا از یک تعامل.

فیلدها

رشته کد (اختیاری)

یک URI که نوع خطا را مشخص می‌کند.

رشته پیام (اختیاری)

یک پیام خطا که برای انسان قابل خواندن باشد.

رشته event_id (اختیاری)

توکن event_id که برای از سرگیری جریان تعامل، از این رویداد، استفاده می‌شود.

شیء event_type (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "error" تنظیم شود.

رویداد تکمیل‌شده‌ی تعامل

رشته event_id (اختیاری)

توکن event_id که برای از سرگیری جریان تعامل، از این رویداد، استفاده می‌شود.

شیء event_type (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "interaction.completed" تنظیم شود.

تعامل InteractionSseEventInteraction (الزامی)

منبع تعاملی که تا حدی تکمیل شده و در انتهای جریان منتشر شده است.

منبع تعامل جزئی که توسط رویدادهای SSE چرخه حیات تعامل منتشر می‌شود. بارهای چرخه حیات جریان ممکن است فیلدهایی را که فقط در پاسخ‌های تعامل کامل غیر جریانی در دسترس هستند، حذف کنند.

فیلدها

رشته عامل (اختیاری)

عاملی که باید با آن تعامل داشت.

رشته ایجاد شده (اختیاری)

فقط خروجی. زمانی که پاسخ در قالب ISO 8601 ایجاد شده است.

رشته شناسه (اختیاری)

الزامی. فقط خروجی. یک شناسه منحصر به فرد برای تکمیل تعامل.

رشته مدل (اختیاری)

مدلی که درخواست شما را تکمیل می‌کند.

رشته شیء (اختیاری)

فقط خروجی. نوع منبع.

شمارش وضعیت (رشته) (اختیاری)

الزامی. فقط خروجی. وضعیت تعامل.

مقادیر ممکن:

  • in_progress

    تعامل در حال انجام است.

  • requires_action

    این تعامل نیاز به اقدام/ورودی از سوی کاربر دارد.

  • completed

    تعامل تکمیل شده است.

  • failed

    تعامل شکست خورد.

  • cancelled

    تعامل لغو شد.

  • incomplete

    تعامل تکمیل شده است، اما شامل نتایج ناقص است (مثلاً رسیدن به max_tokens).

آرایه گام‌ها ( Step ) (اختیاری)

فقط خروجی. مراحلی که تعامل را تشکیل می‌دهند، در صورتی که در این رویداد گنجانده شده باشند.

رشته به‌روزرسانی‌شده (اختیاری)

فقط خروجی. زمانی که پاسخ آخرین بار در قالب ISO 8601 به‌روزرسانی شده است.

کاربرد (اختیاری )

فقط خروجی. آمار مربوط به میزان استفاده از توکن درخواست تعامل.

آمار مربوط به میزان استفاده از توکن درخواست تعامل.

فیلدها

آرایه cached_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک میزان استفاده از توکن‌های ذخیره‌شده بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه grounding_tool_count (GroundingToolCount) (اختیاری)

تعداد ابزار اتصال به زمین

تعداد ابزار اتصال به زمین مهم است.

فیلدها

شمارش عدد صحیح (اختیاری)

تعداد ابزار اتصال به زمین مهم است.

نوع enum (رشته) (اختیاری)

نوع ابزار اتصال زمین مرتبط با شمارش.

مقادیر ممکن:

  • google_search

    اتصال به زمین با جستجوی وب و جستجوی تصویر گوگل، و اتصال به زمین وب برای سازمان‌ها.

  • google_maps

    اتصال به زمین با نقشه‌های گوگل.

آرایه input_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک استفاده از توکن ورودی بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه output_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک استفاده از توکن خروجی بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه tool_use_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک میزان استفاده از توکن‌های ابزار بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

عدد صحیح total_cached_tokens (اختیاری)

تعداد توکن‌ها در بخش ذخیره‌شده‌ی اعلان (محتوای ذخیره‌شده).

عدد صحیح total_input_tokens (اختیاری)

تعداد توکن‌ها در اعلان (زمینه).

total_output_tokens عدد صحیح (اختیاری)

تعداد کل توکن‌ها در تمام پاسخ‌های تولید شده.

total_thought_tokens عدد صحیح (اختیاری)

تعداد توکن‌های افکار برای مدل‌های تفکر.

عدد صحیح total_tokens (اختیاری)

تعداد کل توکن‌ها برای درخواست تعامل (درخواست + پاسخ‌ها + سایر توکن‌های داخلی).

عدد صحیح total_tool_use_tokens (اختیاری)

تعداد توکن‌های موجود در اعلان(های) استفاده از ابزار.

رویداد ایجاد شده توسط تعامل

رشته event_id (اختیاری)

توکن event_id که برای از سرگیری جریان تعامل، از این رویداد، استفاده می‌شود.

شیء event_type (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "interaction.created" تنظیم شود.

تعامل InteractionSseEventInteraction (الزامی)

منبع تعامل جزئی هنگام ایجاد جریان منتشر می‌شود.

منبع تعامل جزئی که توسط رویدادهای SSE چرخه حیات تعامل منتشر می‌شود. بارهای چرخه حیات جریان ممکن است فیلدهایی را که فقط در پاسخ‌های تعامل کامل غیر جریانی در دسترس هستند، حذف کنند.

فیلدها

رشته عامل (اختیاری)

عاملی که باید با آن تعامل داشت.

رشته ایجاد شده (اختیاری)

فقط خروجی. زمانی که پاسخ در قالب ISO 8601 ایجاد شده است.

رشته شناسه (اختیاری)

الزامی. فقط خروجی. یک شناسه منحصر به فرد برای تکمیل تعامل.

رشته مدل (اختیاری)

مدلی که درخواست شما را تکمیل می‌کند.

رشته شیء (اختیاری)

فقط خروجی. نوع منبع.

شمارش وضعیت (رشته) (اختیاری)

الزامی. فقط خروجی. وضعیت تعامل.

مقادیر ممکن:

  • in_progress

    تعامل در حال انجام است.

  • requires_action

    این تعامل نیاز به اقدام/ورودی از سوی کاربر دارد.

  • completed

    تعامل تکمیل شده است.

  • failed

    تعامل شکست خورد.

  • cancelled

    تعامل لغو شد.

  • incomplete

    تعامل تکمیل شده است، اما شامل نتایج ناقص است (مثلاً رسیدن به max_tokens).

آرایه گام‌ها ( Step ) (اختیاری)

فقط خروجی. مراحلی که تعامل را تشکیل می‌دهند، در صورتی که در این رویداد گنجانده شده باشند.

رشته به‌روزرسانی‌شده (اختیاری)

فقط خروجی. زمانی که پاسخ آخرین بار در قالب ISO 8601 به‌روزرسانی شده است.

کاربرد (اختیاری )

فقط خروجی. آمار مربوط به میزان استفاده از توکن درخواست تعامل.

آمار مربوط به میزان استفاده از توکن درخواست تعامل.

فیلدها

آرایه cached_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک میزان استفاده از توکن‌های ذخیره‌شده بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه grounding_tool_count (GroundingToolCount) (اختیاری)

تعداد ابزار اتصال به زمین

تعداد ابزار اتصال به زمین مهم است.

فیلدها

شمارش عدد صحیح (اختیاری)

تعداد ابزار اتصال به زمین مهم است.

نوع enum (رشته) (اختیاری)

نوع ابزار اتصال زمین مرتبط با شمارش.

مقادیر ممکن:

  • google_search

    اتصال به زمین با جستجوی وب و جستجوی تصویر گوگل، و اتصال به زمین وب برای سازمان‌ها.

  • google_maps

    اتصال به زمین با نقشه‌های گوگل.

آرایه input_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک استفاده از توکن ورودی بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه output_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک استفاده از توکن خروجی بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

آرایه tool_use_tokens_by_modality (ModalityTokens) (اختیاری)

تفکیک میزان استفاده از توکن‌های ابزار بر اساس روش.

تعداد توکن‌ها برای یک روش پاسخ واحد.

فیلدها

روش پاسخ (اختیاری)

روش مرتبط با شمارش توکن‌ها.

مقادیر ممکن

  • text

    نشان می‌دهد که مدل باید متن را برگرداند.

  • image

    نشان می‌دهد که مدل باید تصاویر را برگرداند.

  • audio

    نشان می‌دهد که مدل باید صدا را برگرداند.

  • video

    نشان می‌دهد که مدل باید ویدیو برگرداند.

  • document

    نشان می‌دهد که مدل باید اسناد را برگرداند.

عدد صحیح توکن (اختیاری)

تعداد توکن‌ها برای روش.

عدد صحیح total_cached_tokens (اختیاری)

تعداد توکن‌ها در بخش ذخیره‌شده‌ی اعلان (محتوای ذخیره‌شده).

عدد صحیح total_input_tokens (اختیاری)

تعداد توکن‌ها در اعلان (زمینه).

total_output_tokens عدد صحیح (اختیاری)

تعداد کل توکن‌ها در تمام پاسخ‌های تولید شده.

total_thought_tokens عدد صحیح (اختیاری)

تعداد توکن‌های افکار برای مدل‌های تفکر.

عدد صحیح total_tokens (اختیاری)

تعداد کل توکن‌ها برای درخواست تعامل (درخواست + پاسخ‌ها + سایر توکن‌های داخلی).

عدد صحیح total_tool_use_tokens (اختیاری)

تعداد توکن‌های موجود در اعلان(های) استفاده از ابزار.

به‌روزرسانی وضعیت تعامل

رشته event_id (اختیاری)

توکن event_id که برای از سرگیری جریان تعامل، از این رویداد، استفاده می‌شود.

شیء event_type (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "interaction.status_update" تنظیم شود.

رشته‌ی interaction_id (الزامی)

هیچ توضیحی ارائه نشده است.

شمارش وضعیت (رشته) (الزامی)

هیچ توضیحی ارائه نشده است.

مقادیر ممکن:

  • in_progress

    تعامل در حال انجام است.

  • requires_action

    این تعامل نیاز به اقدام/ورودی از سوی کاربر دارد.

  • completed

    تعامل تکمیل شده است.

  • failed

    تعامل شکست خورد.

  • cancelled

    تعامل لغو شد.

  • incomplete

    تعامل تکمیل شده است، اما شامل نتایج ناقص است (مثلاً رسیدن به max_tokens).

استپ دلتا

دلتا گام دلتا داده (الزامی)

هیچ توضیحی ارائه نشده است.

انواع ممکن

آرگومان‌هادلتا

رشته آرگومان‌ها (اختیاری)

هیچ توضیحی ارائه نشده است.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "arguments_delta" تنظیم شود.

آدیودلتا

عدد صحیح کانال‌ها (اختیاری)

تعداد کانال‌های صوتی

رشته داده (اختیاری)

هیچ توضیحی ارائه نشده است.

mime_type enum (رشته) (اختیاری)

هیچ توضیحی ارائه نشده است.

مقادیر ممکن:

  • audio/wav

    فرمت صوتی WAV

  • audio/mp3

    فرمت صوتی MP3

  • audio/aiff

    فرمت صوتی AIFF

  • audio/aac

    فرمت صوتی AAC

  • audio/ogg

    فرمت صوتی OGG

  • audio/flac

    فرمت صوتی FLAC

  • audio/mpeg

    فرمت صوتی MPEG

  • audio/m4a

    فرمت صوتی M4A

  • audio/l16

    فرمت صوتی L16

  • audio/opus

    فرمت صوتی OPUS

  • audio/alaw

    فرمت صوتی ALAW

  • audio/mulaw

    فرمت صوتی MULAW

عدد صحیح sample_rate (اختیاری)

نرخ نمونه‌برداری صدا.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "audio" تنظیم شود.

رشته uri (اختیاری)

هیچ توضیحی ارائه نشده است.

اجرای کدCallDelta

آرگومان‌های CodeExecutionCallArguments (الزامی)

هیچ توضیحی ارائه نشده است.

آرگومان‌هایی که باید به اجرای کد ارسال شوند.

فیلدها

رشته کد (اختیاری)

کدی که قرار است اجرا شود.

شمارش زبان (رشته) (اختیاری)

زبان برنامه‌نویسی «کد».

مقادیر ممکن:

  • python

    پایتون >= 3.10، با numpy و simpy در دسترس است.

رشته امضا (اختیاری)

یک هش امضا برای اعتبارسنجی backend.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "code_execution_call" تنظیم شود.

دلتای نتیجه اجرای کد

is_error نوع داده بولی (اختیاری)

هیچ توضیحی ارائه نشده است.

رشته نتیجه (الزامی)

هیچ توضیحی ارائه نشده است.

رشته امضا (اختیاری)

یک هش امضا برای اعتبارسنجی backend.

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

همیشه روی "code_execution_result" تنظیم شود.

سند دلتا

رشته داده (اختیاری)

هیچ توضیحی ارائه نشده است.

mime_type enum (رشته) (اختیاری)

هیچ توضیحی ارائه نشده است.

مقادیر ممکن:

  • application/pdf

    فرمت سند PDF

  • text/csv

    فرمت سند CSV

نوع شیء (الزامی)

هیچ توضیحی ارائه نشده است.

Always set to "document" .

uri string (optional)

No description provided.

FileSearchCallDelta

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "file_search_call" .

FileSearchResultDelta

result array (FileSearchResult) (required)

No description provided.

The result of the File Search.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "file_search_result" .

FunctionResultDelta

call_id string (required)

Required. ID to match the ID from the function call block.

is_error boolean (optional)

No description provided.

name string (optional)

No description provided.

result array ( ImageContent or TextContent ) or object or string (required)

No description provided.

type object (required)

No description provided.

Always set to "function_result" .

GoogleMapsCallDelta

arguments GoogleMapsCallArguments (optional)

The arguments to pass to the Google Maps tool.

The arguments to pass to the Google Maps tool.

فیلدها

queries array (string) (optional)

The queries to be executed.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "google_maps_call" .

GoogleMapsResultDelta

result array (GoogleMapsResult) (optional)

The results of the Google Maps.

The result of the Google Maps.

فیلدها

places array (Places) (optional)

The places that were found.

فیلدها

name string (optional)

Title of the place.

place_id string (optional)

The ID of the place, in `places/{place_id}` format.

review_snippets array (ReviewSnippet) (optional)

Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.

Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.

فیلدها

review_id string (optional)

The ID of the review snippet.

title string (optional)

Title of the review.

url string (optional)

A link that corresponds to the user review on Google Maps.

url string (optional)

URI reference of the place.

widget_context_token string (optional)

Resource name of the Google Maps widget context token.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "google_maps_result" .

GoogleSearchCallDelta

arguments GoogleSearchCallArguments (required)

No description provided.

The arguments to pass to Google Search.

فیلدها

queries array (string) (optional)

Web search queries for the following-up web search.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "google_search_call" .

GoogleSearchResultDelta

is_error boolean (optional)

No description provided.

result array (GoogleSearchResult) (required)

No description provided.

The result of the Google Search.

فیلدها

search_suggestions string (optional)

Web content snippet that can be embedded in a web page or an app webview.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "google_search_result" .

ImageDelta

data string (optional)

No description provided.

mime_type enum (string) (optional)

No description provided.

Possible values:

  • image/png

    PNG image format

  • image/jpeg

    JPEG image format

  • image/webp

    WebP image format

  • image/heic

    HEIC image format

  • image/heif

    HEIF image format

  • image/gif

    GIF image format

  • image/bmp

    BMP image format

  • image/tiff

    TIFF image format

resolution MediaResolution (optional)

The resolution of the media.

Possible values

  • low

    Low resolution.

  • medium

    Medium resolution.

  • high

    High resolution.

  • ultra_high

    Ultra high resolution.

type object (required)

No description provided.

Always set to "image" .

uri string (optional)

No description provided.

TextAnnotationDelta

annotations array (Annotation) (optional)

Citation information for model-generated content.

Citation information for model-generated content.

Possible Types

FileCitation

A file citation annotation.

custom_metadata object (optional)

User provided metadata about the retrieved context.

document_uri string (optional)

The URI of the file.

end_index integer (optional)

End of the attributed segment, exclusive.

file_name string (optional)

The name of the file.

media_id string (optional)

Media ID in-case of image citations, if applicable.

page_number integer (optional)

Page number of the cited document, if applicable.

source string (optional)

Source attributed for a portion of the text.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

No description provided.

Always set to "file_citation" .

PlaceCitation

A place citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

name string (optional)

Title of the place.

place_id string (optional)

The ID of the place, in `places/{place_id}` format.

review_snippets array (ReviewSnippet) (optional)

Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.

Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.

فیلدها

review_id string (optional)

The ID of the review snippet.

title string (optional)

Title of the review.

url string (optional)

A link that corresponds to the user review on Google Maps.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

No description provided.

Always set to "place_citation" .

url string (optional)

URI reference of the place.

UrlCitation

A URL citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

title string (optional)

The title of the URL.

type object (required)

No description provided.

Always set to "url_citation" .

url string (optional)

The URL.

type object (required)

No description provided.

Always set to "text_annotation_delta" .

TextDelta

text string (required)

No description provided.

type object (required)

No description provided.

Always set to "text" .

ThoughtSignatureDelta

signature string (optional)

Signature to match the backend source to be part of the generation.

type object (required)

No description provided.

Always set to "thought_signature" .

ThoughtSummaryDelta

content Content (optional)

A new summary item to be added to the thought.

type object (required)

No description provided.

Always set to "thought_summary" .

UrlContextCallDelta

arguments UrlContextCallArguments (required)

No description provided.

The arguments to pass to the URL context.

فیلدها

urls array (string) (optional)

The URLs to fetch.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "url_context_call" .

UrlContextResultDelta

is_error boolean (optional)

No description provided.

result array (UrlContextResult) (required)

No description provided.

The result of the URL context.

فیلدها

status enum (string) (optional)

The status of the URL retrieval.

Possible values:

  • success

    Url retrieval is successful.

  • error

    Url retrieval is failed due to error.

  • paywall

    Url retrieval is failed because the content is behind paywall.

  • unsafe

    Url retrieval is failed because the content is unsafe.

url string (optional)

The URL that was fetched.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "url_context_result" .

event_id string (optional)

The event_id token to be used to resume the interaction stream, from this event.

event_type object (required)

No description provided.

Always set to "step.delta" .

index integer (required)

No description provided.

metadata StepDeltaMetadata (optional)

No description provided.

Optional metadata accompanying ANY streamed event.

فیلدها

total_usage Usage (optional)

Statistics on the interaction request's token usage.

Statistics on the interaction request's token usage.

فیلدها

cached_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of cached token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

grounding_tool_count array (GroundingToolCount) (optional)

Grounding tool count.

The number of grounding tool counts.

فیلدها

count integer (optional)

The number of grounding tool counts.

type enum (string) (optional)

The grounding tool type associated with the count.

Possible values:

  • google_search

    Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.

  • google_maps

    Grounding with Google Maps.

input_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of input token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

output_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of output token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

tool_use_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of tool-use token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

total_cached_tokens integer (optional)

Number of tokens in the cached part of the prompt (the cached content).

total_input_tokens integer (optional)

Number of tokens in the prompt (context).

total_output_tokens integer (optional)

Total number of tokens across all the generated responses.

total_thought_tokens integer (optional)

Number of tokens of thoughts for thinking models.

total_tokens integer (optional)

Total token count for the interaction request (prompt + responses + other internal tokens).

total_tool_use_tokens integer (optional)

Number of tokens present in tool-use prompt(s).

StepStart

event_id string (optional)

The event_id token to be used to resume the interaction stream, from this event.

event_type object (required)

No description provided.

Always set to "step.start" .

index integer (required)

No description provided.

step Step (required)

No description provided.

StepStop

event_id string (optional)

The event_id token to be used to resume the interaction stream, from this event.

event_type object (required)

No description provided.

Always set to "step.stop" .

index integer (required)

No description provided.

step_usage Usage (optional)

Model usage stats for this specific step.

Statistics on the interaction request's token usage.

فیلدها

cached_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of cached token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

grounding_tool_count array (GroundingToolCount) (optional)

Grounding tool count.

The number of grounding tool counts.

فیلدها

count integer (optional)

The number of grounding tool counts.

type enum (string) (optional)

The grounding tool type associated with the count.

Possible values:

  • google_search

    Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.

  • google_maps

    Grounding with Google Maps.

input_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of input token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

output_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of output token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

tool_use_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of tool-use token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

total_cached_tokens integer (optional)

Number of tokens in the cached part of the prompt (the cached content).

total_input_tokens integer (optional)

Number of tokens in the prompt (context).

total_output_tokens integer (optional)

Total number of tokens across all the generated responses.

total_thought_tokens integer (optional)

Number of tokens of thoughts for thinking models.

total_tokens integer (optional)

Total token count for the interaction request (prompt + responses + other internal tokens).

total_tool_use_tokens integer (optional)

Number of tokens present in tool-use prompt(s).

usage Usage (optional)

Cumulative model usage stats from the start of the session.

Statistics on the interaction request's token usage.

فیلدها

cached_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of cached token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

grounding_tool_count array (GroundingToolCount) (optional)

Grounding tool count.

The number of grounding tool counts.

فیلدها

count integer (optional)

The number of grounding tool counts.

type enum (string) (optional)

The grounding tool type associated with the count.

Possible values:

  • google_search

    Grounding with Google Web Search and Image Search, & Web Grounding for Enterprise.

  • google_maps

    Grounding with Google Maps.

input_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of input token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

output_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of output token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

tool_use_tokens_by_modality array (ModalityTokens) (optional)

A breakdown of tool-use token usage by modality.

The token count for a single response modality.

فیلدها

modality ResponseModality (optional)

The modality associated with the token count.

Possible values

  • text

    Indicates the model should return text.

  • image

    Indicates the model should return images.

  • audio

    Indicates the model should return audio.

  • video

    Indicates the model should return video.

  • document

    Indicates the model should return documents.

tokens integer (optional)

Number of tokens for the modality.

total_cached_tokens integer (optional)

Number of tokens in the cached part of the prompt (the cached content).

total_input_tokens integer (optional)

Number of tokens in the prompt (context).

total_output_tokens integer (optional)

Total number of tokens across all the generated responses.

total_thought_tokens integer (optional)

Number of tokens of thoughts for thinking models.

total_tokens integer (optional)

Total token count for the interaction request (prompt + responses + other internal tokens).

total_tool_use_tokens integer (optional)

Number of tokens present in tool-use prompt(s).

مثال‌ها

Error Event

{
  "error": {
    "code": "not_found",
    "message": "Failed to get completed interaction: Result not found."
  },
  "event_type": "error"
}

Interaction Completed

{
  "event_id": "evt_123",
  "event_type": "interaction.completed",
  "interaction": {
    "created": "2025-12-04T15:01:45Z",
    "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg",
    "model": "gemini-3.6-flash",
    "status": "completed",
    "updated": "2025-12-04T15:01:45Z"
  }
}

Interaction Completed

{
  "event_id": "evt_123",
  "event_type": "interaction.completed",
  "interaction": {
    "created": "2025-12-04T15:01:45Z",
    "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg",
    "model": "gemini-3-flash-preview",
    "object": "interaction",
    "status": "completed",
    "updated": "2025-12-04T15:01:45Z"
  }
}

Interaction Created

{
  "event_id": "evt_123",
  "event_type": "interaction.created",
  "interaction": {
    "created": "2025-12-04T15:01:45Z",
    "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg",
    "model": "gemini-3.6-flash",
    "status": "in_progress",
    "updated": "2025-12-04T15:01:45Z"
  }
}

Interaction Created

{
  "event_id": "evt_123",
  "event_type": "interaction.created",
  "interaction": {
    "id": "v1_ChdXS0l4YWZXTk9xbk0xZThQczhEcmlROBIXV0tJeGFmV05PcW5NMWU4UHM4RHJpUTg",
    "model": "gemini-3-flash-preview",
    "object": "interaction",
    "status": "in_progress"
  }
}

Interaction Status Update

{
  "event_type": "interaction.status_update",
  "interaction_id": "v1_ChdTMjQ0YWJ5TUF1TzcxZThQdjRpcnFRcxIXUzI0NGFieU1BdU83MWU4UHY0aXJxUXM",
  "status": "in_progress"
}

Step Delta

{
  "delta": {
    "type": "text",
    "text": "Hello"
  },
  "event_type": "step.delta",
  "index": 0
}

Step Start

{
  "event_type": "step.start",
  "index": 0,
  "step": {
    "type": "model_output"
  }
}

Step Stop

{
  "event_type": "step.stop",
  "index": 0
}

ResponseFormat

Possible Types

AudioResponseFormat

Configuration for audio output format.

bit_rate integer (optional)

Bit rate in bits per second (bps). Only applicable for compressed formats (MP3, Opus).

delivery enum (string) (optional)

The delivery mode for the audio output.

Possible values:

  • inline

    Audio data is returned inline in the response.

  • uri

    Audio data is returned as a URI.

mime_type enum (string) (optional)

The MIME type of the audio output.

Possible values:

  • audio/mp3

    MP3 audio format.

  • audio/ogg_opus

    OGG Opus audio format.

  • audio/l16

    Raw PCM (L16) audio format.

  • audio/wav

    WAV audio format.

  • audio/alaw

    A-law audio format.

  • audio/mulaw

    Mu-law audio format.

sample_rate integer (optional)

Sample rate in Hz.

type object (required)

No description provided.

Always set to "audio" .

ImageResponseFormat

Configuration for image output format.

aspect_ratio enum (string) (optional)

The aspect ratio for the image output.

Possible values:

  • 1:1

    نسبت تصویر ۱:۱.

  • 2:3

    2:3 aspect ratio.

  • 3:2

    3:2 aspect ratio.

  • 3:4

    3:4 aspect ratio.

  • 4:3

    نسبت تصویر ۴:۳.

  • 4:5

    4:5 aspect ratio.

  • 5:4

    5:4 aspect ratio.

  • 9:16

    نسبت تصویر ۹:۱۶

  • 16:9

    نسبت تصویر ۱۶:۹.

  • 21:9

    21:9 aspect ratio.

  • 1:8

    1:8 aspect ratio.

  • 8:1

    8:1 aspect ratio.

  • 1:4

    1:4 aspect ratio.

  • 4:1

    4:1 aspect ratio.

delivery enum (string) (optional)

The delivery mode for the image output.

Possible values:

  • inline

    Image data is returned inline in the response.

  • uri

    Image data is returned as a URI.

image_size enum (string) (optional)

The size of the image output.

Possible values:

  • 512

    512px image size.

  • 1K

    1K image size.

  • 2K

    2K image size.

  • 4K

    4K image size.

mime_type enum (string) (optional)

The MIME type of the image output.

Possible values:

  • image/jpeg

    JPEG image format.

type object (required)

No description provided.

Always set to "image" .

TextResponseFormat

Configuration for text output format.

mime_type enum (string) (optional)

The MIME type of the text output.

Possible values:

  • application/json

    JSON output format.

  • text/plain

    Plain text output format.

schema object (optional)

The JSON schema that the output should conform to. Only applicable when mime_type is application/json.

type object (required)

No description provided.

Always set to "text" .

VideoResponseFormat

Configuration for video output format.

aspect_ratio enum (string) (optional)

The aspect ratio for the video output.

Possible values:

  • 16:9

    نسبت تصویر ۱۶:۹.

  • 9:16

    نسبت تصویر ۹:۱۶

delivery enum (string) (optional)

The delivery mode for the video output.

Possible values:

  • inline

    Video data is returned inline in the response.

  • uri

    Video data is returned as a URI.

duration string (optional)

The duration for the video output.

resolution enum (string) (optional)

The video output resolution. Defaults to 720p.

Possible values:

  • 360p

    360p resolution.

  • 720p

    720p resolution.

  • 1080p

    1080p resolution.

  • 4k

    وضوح تصویر 4K.

type object (required)

No description provided.

Always set to "video" .

مثال‌ها

Audio Output

{
  "type": "audio",
  "sample_rate": 24000
}

Image Output

{
  "type": "image",
  "aspect_ratio": "16:9",
  "image_size": "1K",
  "mime_type": "image/jpeg"
}

Text Output (JSON Schema)

{
  "type": "text",
  "mime_type": "application/json",
  "schema": {
    "type": "object",
    "properties": {
      "ingredients": {
        "type": "array",
        "items": {
          "type": "string"
        }
      },
      "recipe_name": {
        "type": "string"
      }
    },
    "required": [
      "ingredients",
      "recipe_name"
    ]
  }
}

VideoResponseFormat

No examples available for this type.

قدم

A step in the interaction.

Possible Types

CodeExecutionCallStep

Code execution call step.

arguments CodeExecutionCallStepArguments (optional)

The arguments to pass to the code execution.

The arguments to pass to the code execution.

فیلدها

code string (optional)

The code to be executed.

language enum (string) (optional)

Programming language of the `code`.

Possible values:

  • python

    Python >= 3.10, with numpy and simpy available.

id string (required)

Required. A unique ID for this specific tool call.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "code_execution_call" .

CodeExecutionResultStep

Code execution result step.

call_id string (required)

Required. ID to match the ID from the function call block.

is_error boolean (optional)

Whether the code execution resulted in an error.

result string (optional)

The output of the code execution.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "code_execution_result" .

FileSearchCallStep

File Search call step.

id string (required)

Required. A unique ID for this specific tool call.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "file_search_call" .

FileSearchResultStep

File Search result step.

call_id string (required)

Required. ID to match the ID from the function call block.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "file_search_result" .

FunctionCallStep

A function tool call step.

arguments object (required)

Required. The arguments to pass to the function.

id string (required)

Required. A unique ID for this specific tool call.

name string (required)

Required. The name of the tool to call.

type object (required)

No description provided.

Always set to "function_call" .

FunctionResultStep

Result of a function tool call.

call_id string (required)

Required. ID to match the ID from the function call block.

is_error boolean (optional)

Whether the tool call resulted in an error.

name string (optional)

The name of the tool that was called.

result array ( ImageContent or TextContent ) or object or string (required)

Required. The result of the tool call.

type object (required)

No description provided.

Always set to "function_result" .

GoogleMapsCallStep

Google Maps call step.

arguments GoogleMapsCallStepArguments (optional)

The arguments to pass to the Google Maps tool.

The arguments to pass to the Google Maps tool.

فیلدها

queries array (string) (optional)

The queries to be executed.

id string (required)

Required. A unique ID for this specific tool call.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "google_maps_call" .

GoogleMapsResultStep

Google Maps result step.

call_id string (required)

Required. ID to match the ID from the function call block.

result array (GoogleMapsResultItem) (optional)

No description provided.

The result of the Google Maps.

فیلدها

places array (GoogleMapsResultPlaces) (optional)

No description provided.

فیلدها

name string (optional)

No description provided.

place_id string (optional)

No description provided.

review_snippets array (ReviewSnippet) (optional)

No description provided.

Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.

فیلدها

review_id string (optional)

The ID of the review snippet.

title string (optional)

Title of the review.

url string (optional)

A link that corresponds to the user review on Google Maps.

url string (optional)

No description provided.

widget_context_token string (optional)

No description provided.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "google_maps_result" .

GoogleSearchCallStep

Google Search call step.

arguments GoogleSearchCallStepArguments (optional)

The arguments to pass to Google Search.

The arguments to pass to Google Search.

فیلدها

queries array (string) (optional)

Web search queries for the following-up web search.

id string (required)

Required. A unique ID for this specific tool call.

search_type enum (string) (optional)

The type of search grounding enabled.

Possible values:

  • web_search

    Setting this field enables web search. Only text results are returned.

  • image_search

    Setting this field enables image search. Image bytes are returned.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "google_search_call" .

GoogleSearchResultStep

Google Search result step.

call_id string (required)

Required. ID to match the ID from the function call block.

is_error boolean (optional)

Whether the Google Search resulted in an error.

result array (GoogleSearchResultItem) (optional)

The results of the Google Search.

The result of the Google Search.

فیلدها

search_suggestions string (optional)

Web content snippet that can be embedded in a web page or an app webview.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "google_search_result" .

ModelOutputStep

Output generated by the model.

content array ( Content ) (optional)

No description provided.

type object (required)

No description provided.

Always set to "model_output" .

ThoughtStep

A thought step.

signature string (optional)

A signature hash for backend validation.

summary array ( Content ) (optional)

A summary of the thought.

type object (required)

No description provided.

Always set to "thought" .

UrlContextCallStep

URL context call step.

arguments UrlContextCallArguments (optional)

The arguments to pass to the URL context.

The arguments to pass to the URL context.

فیلدها

urls array (string) (optional)

The URLs to fetch.

id string (required)

Required. A unique ID for this specific tool call.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "url_context_call" .

UrlContextResultStep

URL context result step.

call_id string (required)

Required. ID to match the ID from the function call block.

is_error boolean (optional)

Whether the URL context resulted in an error.

result array (UrlContextResult) (optional)

The results of the URL context.

The result of the URL context.

فیلدها

status enum (string) (optional)

The status of the URL retrieval.

Possible values:

  • success

    Url retrieval is successful.

  • error

    Url retrieval is failed due to error.

  • paywall

    Url retrieval is failed because the content is behind paywall.

  • unsafe

    Url retrieval is failed because the content is unsafe.

url string (optional)

The URL that was fetched.

signature string (optional)

A signature hash for backend validation.

type object (required)

No description provided.

Always set to "url_context_result" .

UserInputStep

Input provided by the user.

content array ( Content ) (optional)

No description provided.

type object (required)

No description provided.

Always set to "user_input" .

مثال‌ها

CodeExecutionCallStep

{
  "type": "code_execution_call",
  "arguments": {
    "code": "print(sum(range(1, 11)))"
  },
  "id": "code_call_71021"
}

CodeExecutionResultStep

{
  "type": "code_execution_result",
  "call_id": "code_call_71021",
  "result": "55\n"
}

FileSearchCallStep

{
  "type": "file_search_call",
  "id": "file_call_88192"
}

FileSearchResultStep

{
  "type": "file_search_result",
  "call_id": "file_call_88192"
}

FunctionCallStep

{
  "name": "get_weather",
  "type": "function_call",
  "arguments": {
    "location": "Boston, MA"
  },
  "id": "call_98231"
}

FunctionResultStep

{
  "name": "get_weather",
  "type": "function_result",
  "call_id": "call_98231",
  "result": [
    {
      "type": "text",
      "text": "{\"weather\":\"sunny\"}"
    }
  ]
}

GoogleMapsCallStep

{
  "type": "google_maps_call",
  "arguments": {
    "latitude": 37.7749,
    "longitude": -122.4194
  },
  "id": "maps_call_39201"
}

GoogleMapsResultStep

{
  "type": "google_maps_result",
  "call_id": "maps_call_39201",
  "result": [
    {
      "name": "Golden Gate Park",
      "place_id": "ChIJIQBpAG2ahYAR9R7bNdTLg8M",
      "rating": 4.8
    }
  ]
}

GoogleSearchCallStep

{
  "type": "google_search_call",
  "arguments": {
    "query": "Who won the men's 100m in Paris 2024?"
  },
  "id": "search_call_19201"
}

GoogleSearchResultStep

{
  "type": "google_search_result",
  "call_id": "search_call_19201",
  "result": [
    {
      "title": "Paris 2024 Olympics: Noah Lyles wins men's 100m gold",
      "url": "https://olympics.com/en/news/paris-2024-noah-lyles-wins-mens-100m-gold",
      "snippet": "American Noah Lyles won the Olympic men's 100m gold medal in a photo finish."
    }
  ]
}

ModelOutputStep

{
  "type": "model_output",
  "content": [
    {
      "type": "text",
      "text": "The capital of France is Paris."
    }
  ]
}

ThoughtStep

{
  "type": "thought",
  "signature": "thought_sig_abcd1234",
  "summary": [
    {
      "type": "text",
      "text": "The model is searching Google for the capital of France."
    }
  ]
}

UrlContextCallStep

{
  "type": "url_context_call",
  "arguments": {
    "urls": [
      "https://www.example.com"
    ]
  },
  "id": "url_call_10219"
}

UrlContextResultStep

{
  "type": "url_context_result",
  "call_id": "url_call_10219",
  "result": [
    {
      "title": "Example Domain",
      "url": "https://www.example.com",
      "snippet": "This domain is for use in illustrative examples in documents."
    }
  ]
}

UserInputStep

{
  "type": "user_input",
  "content": [
    {
      "type": "text",
      "text": "What is the capital of France?"
    }
  ]
}

ToolChoiceConfig

The tool choice configuration containing allowed tools.

فیلدها

allowed_tools AllowedTools (optional)

The allowed tools.

The configuration for allowed tools.

فیلدها

mode enum (string) (optional)

The mode of the tool choice.

Possible values:

  • auto

    Auto tool choice.

  • any

    Any tool choice.

  • none

    No tool choice.

  • validated

    Validated tool choice.

tools array (string) (optional)

The names of the allowed tools.

مثال‌ها

مثال

{
  "allowed_tools": {
    "mode": "any",
    "tools": [
      "my_tool"
    ]
  }
}

ImageContent

An image content block.

فیلدها

data string (optional)

The image content.

mime_type enum (string) (optional)

The mime type of the image.

Possible values:

  • image/png

    PNG image format

  • image/jpeg

    JPEG image format

  • image/webp

    WebP image format

  • image/heic

    HEIC image format

  • image/heif

    HEIF image format

  • image/gif

    GIF image format

  • image/bmp

    BMP image format

  • image/tiff

    TIFF image format

resolution MediaResolution (optional)

The resolution of the media.

Possible values

  • low

    Low resolution.

  • medium

    Medium resolution.

  • high

    High resolution.

  • ultra_high

    Ultra high resolution.

type object (optional)

No description provided.

Always set to "image" .

uri string (optional)

The URI of the image.

مثال‌ها

تصویر

{
  "type": "image",
  "data": "BASE64_ENCODED_IMAGE",
  "mime_type": "image/png"
}

TextContent

A text content block.

فیلدها

annotations array (Annotation) (optional)

Citation information for model-generated content.

Citation information for model-generated content.

Possible Types

FileCitation

A file citation annotation.

custom_metadata object (optional)

User provided metadata about the retrieved context.

document_uri string (optional)

The URI of the file.

end_index integer (optional)

End of the attributed segment, exclusive.

file_name string (optional)

The name of the file.

media_id string (optional)

Media ID in-case of image citations, if applicable.

page_number integer (optional)

Page number of the cited document, if applicable.

source string (optional)

Source attributed for a portion of the text.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

No description provided.

Always set to "file_citation" .

PlaceCitation

A place citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

name string (optional)

Title of the place.

place_id string (optional)

The ID of the place, in `places/{place_id}` format.

review_snippets array (ReviewSnippet) (optional)

Snippets of reviews that are used to generate answers about the features of a given place in Google Maps.

Encapsulates a snippet of a user review that answers a question about the features of a specific place in Google Maps.

فیلدها

review_id string (optional)

The ID of the review snippet.

title string (optional)

Title of the review.

url string (optional)

A link that corresponds to the user review on Google Maps.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

type object (required)

No description provided.

Always set to "place_citation" .

url string (optional)

URI reference of the place.

UrlCitation

A URL citation annotation.

end_index integer (optional)

End of the attributed segment, exclusive.

start_index integer (optional)

Start of segment of the response that is attributed to this source. Index indicates the start of the segment, measured in bytes.

title string (optional)

The title of the URL.

type object (required)

No description provided.

Always set to "url_citation" .

url string (optional)

The URL.

text string (optional)

Required. The text content.

type object (optional)

No description provided.

Always set to "text" .

مثال‌ها

متن

{
  "type": "text",
  "text": "Hello, how are you?"
}