ابزار «استفاده از رایانه» به شما امکان میدهد عاملهای کنترل مرورگر، تلفن همراه، و رایانه رومیزی بسازید که با تکالیف تعامل برقرار میکنند و آنها را خودکارسازی میکنند. بااستفاده از نماگرفتها، مدل میتواند صفحه رایانه را «ببیند» و با تولید کنشهای خاص میانای کاربر مثل کلیکهای موشواره و ورودیهای صفحهکلید «عمل کند». مشابه با فراخوانی تابع، باید محیط اجرای سمت مشتری را برای دریافت و اجرای کنشهای «استفاده از رایانه» پیادهسازی کنید.
برای فهرست مدلهای پشتیبانیشده، نسخههای مدل را ببینید. مدلهای Gemini 3.x از چندین قابلیت پیشرفته پشتیبانی میکنند:
- پشتیبانی از چند محیط: برای محیطهای مرورگر، تلفن همراه، و رایانه عامل بسازید.
- کنشهای سادهشده با هدفها: کنشها شامل فیلد
intentاست که استدلال مدل را برای هر مرحله توضیح میدهد. - خطمشیهای ایمنی پیکربندیپذیر: رفتار ایمنی را با دستهبندیهای خطمشی داخلی و ملغیسازیها تنظیم دقیق کنید.
- تشخیص تزریق پیامواره: موافقت با اسکن نماگرفت برای تشخیص دستورالعملهای خصمانه پنهان.
با «استفاده از رایانه» میتوانید کارگزارانی بسازید که:
- ورود دادههای تکراری یا تکمیل فرم در وبسایتها را خودکارسازی کنید.
- انجام آزمایش خودکار برنامههای وب و جریانهای کاربر
- انجام پژوهش در وبسایتهای مختلف (برای نمونه، جمعآوری اطلاعات محصول، قیمتها، و مرورها از سایتهای تجارت الکترونیک برای اطلاعرسانی درباره خرید)
در اینجا نمونهای حداقلی از مقداردهی اولیه کلاینت و ارسال پیامواره به مدل با فعال بودن ابزار computer_use برای محیط مرورگر ارائه شده است:
Python
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Search for 'Gemini API' on Google.",
tools=[{"type": "computer_use", "environment": "browser"}]
)
print(interaction)
JavaScript
import { GoogleGenAI } from '@google/genai';
const ai = new GoogleGenAI();
const interaction = await ai.interactions.create({
model: 'gemini-3.8-flash',
input: "Search for 'Gemini API' on Google.",
tools: [{ type: "computer_use", environment: "browser" }]
});
console.log(interaction);
جاوا
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model("gemini-3.8-flash")
.input(InteractionsInput.of("Search for 'Gemini API' on Google."))
.tools(
Arrays.asList(
ComputerUse.builder().environment(EnvironmentEnum.BROWSER).build()))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println(interaction);
رفتن
package main
import (
"context"
"fmt"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput("Search for 'Gemini API' on Google."),
Tools: []interactions.Tool{
interactions.NewTool(interactions.ComputerUse{
Environment: interactions.EnvironmentEnumBrowser.ToPointer(),
}),
},
}),
})
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Interaction)
}
نحوه عملکرد «استفاده از رایانه»
برای ساختن عامل با مدل «استفاده از رایانه»، باید حلقه پیوستهای بین برنامهتان و میانای برنامهسازی کاربردی راهاندازی کنید. در اینجا آنچه کد شما در هر مرحله انجام خواهد داد آمده است:
- ارسال درخواست به مدل
- برنامه شما درخواست میانای برنامهسازی کاربردی حاوی ابزار «استفاده از رایانه»، تنظیمات پیکربندی شما (مثل محیط هدف)، پیامواره کاربر، و نماگرفتی از صفحه فعلی ارسال میکند.
- دریافت پاسخ مدل
- مدل صفحه و پیامواره را تجزیهوتحلیل میکند و پاسخی برمیگرداند که شامل
function_callپیشنهادی است که کنش رابط کاربری را نشان میدهد (مثل کلیک، پیمایش، یا ضربه کلید). - برای مدلهای Gemini 3.x، پاسخ همچنین شامل استدلال
intentتوضیح میدهد که چرا مدل آن کنش را انتخاب کرده است. - پاسخ ممکن است شامل
safety_decisionاز سیستم ایمنی داخلی باشد که کنش را بهعنوان عادی/مجاز،require_confirmation(نیازمند تأیید کاربر)، یا مسدود طبقهبندی میکند.
- مدل صفحه و پیامواره را تجزیهوتحلیل میکند و پاسخی برمیگرداند که شامل
- اجرای کنش دریافتی
- اگر کنش مجاز باشد (یا کاربر آن را تأیید کند)، کد سمت مشتری شما
function_callرا تجزیه میکند، مختصات عادیسازیشده را مقیاسبندی میکند تا با درگاه دید شما مطابقت داشته باشد، و کنش را در محیط هدف بااستفاده از ابزارهای خودکارسازی (مثل Playwright) اجرا میکند. اگر کنش مسدود شد، کارخواه شما باید اجرا را متوقف کند یا وقفه را مدیریت کند.
- اگر کنش مجاز باشد (یا کاربر آن را تأیید کند)، کد سمت مشتری شما
- ثبت وضعیت محیط جدید
- پساز اینکه کنش اجرا شد، برنامه شما نماگرفت جدیدی میگیرد و آن را در
function_resultبه مدل برمیگرداند تا مرحله بعدی را درخواست کند.
- پساز اینکه کنش اجرا شد، برنامه شما نماگرفت جدیدی میگیرد و آن را در
این فرایند سپس از مرحله ۲ تکرار میشود و بهطور مداوم کنش بعدی را از مدل درخواست میکند تا زمانی که تکلیف تکمیل یا خاتمه یابد.

نحوه پیادهسازی «استفاده از رایانه»
قبلاز ساختن با ابزار «استفاده از رایانه»، باید موارد زیر را راهاندازی کنید:
- محیط اجرای امن: نمایندهتان را در ماشین مجازی یا محفظه جعبه شنی اجرا کنید تا آن را از سیستم میزبان خود جدا کنید و تأثیر بالقوه آن را محدود کنید. پیادهسازی مرجع شامل یک جعبه ایمنی مبتنی بر Docker آماده استفاده است که میتوانید از آن بهعنوان نقطه شروع استفاده کنید.
- مدیر کنش سمت کارخواه: منطق سمت کارخواه را برای اجرای مختصات، تایپ نوشتار، و گرفتن نماگرفت پیادهسازی کنید.
نمونههای زیر از مرورگر وب بهعنوان محیط اجرا و Playwright بهعنوان مدیریتکننده سمت مشتری استفاده میکنند.
۰. راهاندازی Playwright
ابتدا بستههای موردنیاز را نصب کنید:
pip install google-genai playwright
playwright install chromium
سپس، نمونه مرورگر Playwright را برای استفاده در اجرا مقداردهی اولیه کنید:
from playwright.sync_api import sync_playwright
# 1. Configure screen dimensions for the target environment
SCREEN_WIDTH = 1440
SCREEN_HEIGHT = 900
# 2. Start the Playwright browser
# In production, utilize a sandboxed environment.
playwright = sync_playwright().start()
# Set headless=False to see the actions performed on your screen
browser = playwright.chromium.launch(headless=False)
# 3. Create a context and page with the specified dimensions
context = browser.new_context(
viewport={"width": SCREEN_WIDTH, "height": SCREEN_HEIGHT}
)
page = context.new_page()
# 4. Navigate to an initial page to start the task
page.goto("https://www.google.com")
# The 'page', 'SCREEN_WIDTH', and 'SCREEN_HEIGHT' variables
# will be used in the steps below.
۱. ارسال درخواست به مدل
کتابخانه کارخواه را مقداردهی اولیه کنید و ابزار «استفاده از رایانه» را پیکربندی کنید. توجه داشته باشید که هنگام صدور درخواست نیازی به تعیین اندازه نمایشگر نیست؛ مدل مختصات پیکسلی را که با ارتفاع و عرض صفحه مقیاسبندی شده است پیشبینی میکند.
Python
برای پیکربندی درخواست هدفیابی محیط مرورگر، از «کیت توسعه نرمافزار google-genai Python» (نسخه 2.7.0 یا بالاتر) استفاده کنید:
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model='gemini-3.8-flash',
input="Find a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th",
tools=[
{
"type": "computer_use",
"environment": "browser",
"enable_prompt_injection_detection": True
}
]
)
print(interaction)
JavaScript
از کیت توسعه نرمافزار @google/genai Node.js برای پیکربندی درخواست هدفیابی محیط مرورگر استفاده کنید:
import { GoogleGenAI } from '@google/genai';
const ai = new GoogleGenAI();
const interaction = await ai.interactions.create({
model: 'gemini-3.8-flash',
input: "Find a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th",
tools: [
{
type: "computer_use",
environment: "browser",
enable_prompt_injection_detection: true
}
]
});
console.log(interaction);
جاوا
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model("gemini-3.8-flash")
.input(
InteractionsInput.of(
"Find a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th"))
.tools(
Arrays.asList(
ComputerUse.builder()
.environment(EnvironmentEnum.BROWSER)
.enablePromptInjectionDetection(true)
.build()))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println(interaction);
رفتن
package main
import (
"context"
"fmt"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput("Find a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th"),
Tools: []interactions.Tool{
interactions.NewTool(interactions.ComputerUse{
Environment: interactions.EnvironmentEnumBrowser.ToPointer(),
EnablePromptInjectionDetection: genai.Ptr(true),
}),
},
}),
})
if err != nil {
log.Fatal(err)
}
fmt.Println(res.Interaction)
}
REST
برای ارسال درخواست از curl استفاده کنید:
curl -X POST \
"https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.8-flash",
"input": "Find me a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th. Start by navigating directly to flights.google.com",
"tools": [
{
"type": "computer_use",
"environment": "browser",
"enable_prompt_injection_detection": true
}
]
}'
۲. دریافت پاسخ مدل
پاسخ مدل فراخوانی تابعی را پیشنهاد میدهد که حاوی مختصات و هدف استدلال سفارشی برای توضیح کنش است:
{
"steps": [
{
"type": "function_call",
"name": "click",
"arguments": {
"x": 450,
"y": 120,
"intent": "Click the search box to type the destination."
}
}
]
}
۳. اجرای کنشهای دریافتی
برنامه شما باید مختصات پاسخ را تجزیه کند، آنها را از مختصات نرمالشده ۱۰۰۰x۱۰۰۰ مقیاسبندی کند، و کنش را اجرا کند:
Python
from typing import Any, List, Tuple
import time
def denormalize_x(x: int, screen_width: int) -> int:
"""Convert normalized x coordinate (0-1000) to actual pixel coordinate."""
return int(x / 1000 * screen_width)
def denormalize_y(y: int, screen_height: int) -> int:
"""Convert normalized y coordinate (0-1000) to actual pixel coordinate."""
return int(y / 1000 * screen_height)
def execute_function_calls(interaction, page, screen_width, screen_height):
results = []
function_calls = [
step for step in interaction.steps if step.type == "function_call"
]
for function_call in function_calls:
action_result = {}
fname = function_call.name
args = function_call.arguments
print(f" -> Executing: {fname} (Intent: {args.get('intent', 'N/A')})")
try:
if fname == "open_app":
pass # Handled / already open
elif fname in ("click", "double_click", "triple_click", "middle_click", "right_click", "move", "long_press"):
actual_x = denormalize_x(args["x"], screen_width)
actual_y = denormalize_y(args["y"], screen_height)
if fname == "click":
page.mouse.click(actual_x, actual_y)
elif fname == "double_click":
page.mouse.dblclick(actual_x, actual_y)
elif fname == "right_click":
page.mouse.click(actual_x, actual_y, button="right")
elif fname == "middle_click":
page.mouse.click(actual_x, actual_y, button="middle")
elif fname == "move":
page.mouse.move(actual_x, actual_y)
elif fname == "type":
actual_x = denormalize_x(args["x"], screen_width) if "x" in args else None
actual_y = denormalize_y(args["y"], screen_height) if "y" in args else None
text = args["text"]
press_enter = args.get("press_enter", False)
if actual_x is not None and actual_y is not None:
page.mouse.click(actual_x, actual_y)
# Clear field first
page.keyboard.press("Meta+A")
page.keyboard.press("Backspace")
page.keyboard.type(text)
if press_enter:
page.keyboard.press("Enter")
elif fname == "navigate":
page.goto(args["url"])
elif fname == "go_back":
page.go_back()
elif fname == "go_forward":
page.go_forward()
elif fname == "wait":
time.sleep(args.get("seconds", 1))
else:
print(f"Warning: Custom or unhandled function {fname}")
page.wait_for_load_state(timeout=5000)
time.sleep(1)
except Exception as e:
print(f"Error executing {fname}: {e}")
action_result = {"error": str(e)}
results.append((fname, function_call.id, action_result))
return results
JavaScript
function denormalizeX(x, screenWidth) {
// Convert normalized x coordinate (0-1000) to actual pixel coordinate.
return Math.floor((x / 1000) * screenWidth);
}
function denormalizeY(y, screenHeight) {
// Convert normalized y coordinate (0-1000) to actual pixel coordinate.
return Math.floor((y / 1000) * screenHeight);
}
async function executeFunctionCalls(interaction, page, screenWidth, screenHeight) {
const results = [];
const functionCalls = interaction.steps.filter(step => step.type === "function_call");
for (const functionCall of functionCalls) {
const actionResult = {};
const fname = functionCall.name;
const args = functionCall.arguments;
console.log(` -> Executing: ${fname} (Intent: ${args.intent || 'N/A'})`);
try {
if (fname === "open_app") {
// Handled / already open
} else if (["click", "double_click", "triple_click", "middle_click", "right_click", "move", "long_press"].includes(fname)) {
const actualX = denormalizeX(args.x, screenWidth);
const actualY = denormalizeY(args.y, screenHeight);
if (fname === "click") {
await page.mouse.click(actualX, actualY);
} else if (fname === "double_click") {
await page.mouse.dblclick(actualX, actualY);
} else if (fname === "right_click") {
await page.mouse.click(actualX, actualY, { button: "right" });
} else if (fname === "middle_click") {
await page.mouse.click(actualX, actualY, { button: "middle" });
} else if (fname === "move") {
await page.mouse.move(actualX, actualY);
}
} else if (fname === "type") {
const actualX = args.x !== undefined ? denormalizeX(args.x, screenWidth) : null;
const actualY = args.y !== undefined ? denormalizeY(args.y, screenHeight) : null;
const text = args.text;
const pressEnter = args.press_enter || false;
if (actualX !== null && actualY !== null) {
await page.mouse.click(actualX, actualY);
}
// Clear field first
await page.keyboard.press("Meta+A");
await page.keyboard.press("Backspace");
await page.keyboard.type(text);
if (pressEnter) {
await page.keyboard.press("Enter");
}
} else if (fname === "navigate") {
await page.goto(args.url);
} else if (fname === "go_back") {
await page.goBack();
} else if (fname === "go_forward") {
await page.goForward();
} else if (fname === "wait") {
await new Promise(resolve => setTimeout(resolve, (args.seconds || 1) * 1000));
} else {
console.log(`Warning: Custom or unhandled function ${fname}`);
}
await page.waitForLoadState('load', { timeout: 5000 }).catch(() => {});
await new Promise(resolve => setTimeout(resolve, 1000));
} catch (e) {
console.log(`Error executing ${fname}: ${e}`);
actionResult.error = e.message;
}
results.push([fname, functionCall.id, actionResult]);
}
return results;
}
جاوا
import com.google.genai.gaos.models.interactions.FunctionCallStep;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.Step;
import java.util.ArrayList;
import java.util.Collections;
import java.util.HashMap;
import java.util.List;
import java.util.Map;
class ActionExecutor {
int denormalizeX(int x, int screenWidth) {
return (int) (x / 1000.0 * screenWidth);
}
int denormalizeY(int y, int screenHeight) {
return (int) (y / 1000.0 * screenHeight);
}
List<Map<String, Object>> executeFunctionCalls(
Interaction interaction, int screenWidth, int screenHeight) {
List<Map<String, Object>> results = new ArrayList<>();
for (Step step : interaction.steps().orElse(Collections.emptyList())) {
if (step instanceof FunctionCallStep) {
FunctionCallStep functionCall = (FunctionCallStep) step;
String fname = functionCall.name().orElse("");
Map<String, Object> args = functionCall.arguments().orElse(Collections.emptyMap());
Map<String, Object> actionResult = new HashMap<>();
System.out.println(
" -> Executing: " + fname + " (Intent: " + args.getOrDefault("intent", "N/A") + ")");
try {
if (fname.equals("click")) {
int actualX = denormalizeX(((Number) args.get("x")).intValue(), screenWidth);
int actualY = denormalizeY(((Number) args.get("y")).intValue(), screenHeight);
// Perform mouse click at (actualX, actualY) using your browser automation library
} else if (fname.equals("type")) {
String text = (String) args.get("text");
// Type text into active element using your browser automation library
} else if (fname.equals("navigate")) {
String url = (String) args.get("url");
// Navigate browser to url
}
} catch (Exception e) {
actionResult.put("error", e.getMessage());
}
Map<String, Object> entry = new HashMap<>();
entry.put("name", fname);
entry.put("callId", functionCall.id().orElse(""));
entry.put("result", actionResult);
results.add(entry);
}
}
return results;
}
}
رفتن
package main
import (
"fmt"
"google.golang.org/genai/interactions/models/interactions"
)
func denormalizeX(x, screenWidth int) int {
return int(float64(x) / 1000.0 * float64(screenWidth))
}
func denormalizeY(y, screenHeight int) int {
return int(float64(y) / 1000.0 * float64(screenHeight))
}
func executeFunctionCalls(interaction *interactions.Interaction, screenWidth, screenHeight int) []map[string]any {
var results []map[string]any
for _, step := range interaction.Steps {
if functionCall := step.FunctionCallStep; functionCall != nil {
fname := functionCall.Name
args := functionCall.Arguments
actionResult := map[string]any{}
intent := args["intent"]
if intent == nil {
intent = "N/A"
}
fmt.Printf(" -> Executing: %s (Intent: %v)\n", fname, intent)
switch fname {
case "click":
xVal, _ := args["x"].(float64)
yVal, _ := args["y"].(float64)
actualX := denormalizeX(int(xVal), screenWidth)
actualY := denormalizeY(int(yVal), screenHeight)
_ = actualX
_ = actualY
// Perform mouse click at (actualX, actualY) using your browser automation library
case "type":
text, _ := args["text"].(string)
_ = text
// Type text into active element using your browser automation library
case "navigate":
url, _ := args["url"].(string)
_ = url
// Navigate browser to url
}
results = append(results, map[string]any{
"name": fname,
"callId": functionCall.ID,
"result": actionResult,
})
}
}
return results
}
func main() {
// Example helper usage with an Interaction response
}
۴. وضعیت محیط جدید را ضبط کنید
پساز اجرای کنشها، نتیجه اجرای تابع را به مدل برگردانید تا بتواند از این اطلاعات برای تولید کنش بعدی استفاده کند. اگر
چندین کنش (تماسهای موازی) اجرا شده است، باید
function_result برای هریک از آنها در نوبت کاربر بعدی ارسال کنید.
Python
import json
import base64
def get_function_responses(page, results):
screenshot_bytes = page.screenshot(type="png")
current_url = page.url
function_responses = []
for name, call_id, result in results:
function_responses.append({
"type": "function_result",
"name": name,
"call_id": call_id,
"result": [
{
"type": "text",
"text": json.dumps({"url": current_url, **result})
},
{
"type": "image",
"data": base64.b64encode(screenshot_bytes).decode("utf-8"),
"mime_type": "image/png"
}
]
})
return function_responses
JavaScript
async function getFunctionResponses(page, results) {
const screenshotBuffer = await page.screenshot({ type: 'png' });
const screenshotBase64 = screenshotBuffer.toString('base64');
const currentUrl = page.url();
const functionResponses = [];
for (const [name, callId, result] of results) {
functionResponses.push({
type: "function_result",
name: name,
call_id: callId,
result: [
{
type: "text",
text: JSON.stringify({ url: currentUrl, ...result })
},
{
type: "image",
data: screenshotBase64,
mime_type: "image/png"
}
]
});
}
return functionResponses;
}
جاوا
import com.google.genai.gaos.models.interactions.FunctionResultStep;
import com.google.genai.gaos.models.interactions.FunctionResultStepResultUnion;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Step;
import com.google.genai.gaos.models.interactions.TextContent;
import java.util.ArrayList;
import java.util.Arrays;
import java.util.Base64;
import java.util.List;
import java.util.Map;
class StateCapturer {
List<Step> getFunctionResponses(
byte[] screenshotBytes, String currentUrl, List<Map<String, Object>> results) {
List<Step> functionResponses = new ArrayList<>();
String base64Screenshot = Base64.getEncoder().encodeToString(screenshotBytes);
for (Map<String, Object> entry : results) {
String name = (String) entry.get("name");
String callId = (String) entry.get("callId");
String jsonResult = String.format("{\"url\": \"%s\"}", currentUrl);
FunctionResultStep responseStep =
FunctionResultStep.builder()
.name(name)
.callId(callId)
.result(
FunctionResultStepResultUnion.of(
Arrays.asList(
TextContent.builder().text(jsonResult).build(),
ImageContent.builder()
.data(base64Screenshot)
.mimeType(ImageContentMimeType.IMAGE_PNG)
.build())))
.build();
functionResponses.add(responseStep);
}
return functionResponses;
}
}
رفتن
package main
import (
"encoding/base64"
"fmt"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
)
func getFunctionResponses(screenshotBytes []byte, currentURL string, results []map[string]any) []interactions.Step {
var functionResponses []interactions.Step
base64Screenshot := base64.StdEncoding.EncodeToString(screenshotBytes)
for _, entry := range results {
name, _ := entry["name"].(string)
callID, _ := entry["callId"].(string)
jsonResult := fmt.Sprintf(`{"url": "%s"}`, currentURL)
responseStep := interactions.NewStep(interactions.FunctionResultStep{
Name: genai.Ptr(name),
CallID: callID,
Result: interactions.NewFunctionResultStepResultUnion([]interactions.FunctionResultSubcontent{
interactions.NewFunctionResultSubcontent(interactions.TextContent{
Text: jsonResult,
}),
interactions.NewFunctionResultSubcontent(interactions.ImageContent{
Data: genai.Ptr(base64Screenshot),
MimeType: interactions.ImageContentMimeType("image/png").ToPointer(),
}),
}),
})
functionResponses = append(functionResponses, responseStep)
}
return functionResponses
}
func main() {
// Example helper usage to build FunctionResultStep responses
}
پساز اینکه تعریف کردید وضعیت محیط چگونه ضبط و قالببندی شود، میتوانید همه این مراحل را در یک حلقه اجرای پیوسته ترکیب کنید.
ساختن حلقه عامل
برای فعال کردن تعاملهای چندمرحلهای، چهار مرحله بخش نحوه پیادهسازی استفاده از رایانه را در یک حلقه واحد ترکیب کنید. این حلقه تا زمانی که تکلیف تکمیل شود، به درخواست کنشها و برگرداندن نتایج به مدل ادامه میدهد.
بهخاطر داشته باشید که سابقه مکالمه را بهدرستی مدیریت کنید و در هر مرحله، هم پاسخهای مدل و هم پاسخهای تابع خود را به سابقه اضافه کنید.
Python
import time
from typing import Any, List, Tuple
from playwright.sync_api import sync_playwright
from google import genai
client = genai.Client()
# Constants for screen dimensions
SCREEN_WIDTH = 1440
SCREEN_HEIGHT = 900
# Setup Playwright
print("Initializing browser...")
playwright = sync_playwright().start()
browser = playwright.chromium.launch(headless=False)
context = browser.new_context(viewport={"width": SCREEN_WIDTH, "height": SCREEN_HEIGHT})
page = context.new_page()
# Define helper functions. Copy/paste from steps 3 and 4
# def denormalize_x(...)
# def denormalize_y(...)
# def execute_function_calls(...)
# def get_function_responses(...)
try:
# Go to initial page
page.goto("https://ai.google.dev/gemini-api/docs")
# Take initial screenshot
initial_screenshot = page.screenshot(type="png")
USER_PROMPT = "Go to ai.google.dev/gemini-api/docs and search for pricing."
print(f"Goal: {USER_PROMPT}")
# First interaction
interaction = client.interactions.create(
model='gemini-3.8-flash',
input=[
{"type": "text", "text": USER_PROMPT},
{"type": "image", "data": base64.b64encode(initial_screenshot).decode("utf-8"), "mime_type": "image/png"}
],
tools=[{
"type": "computer_use",
"environment": "browser",
"enable_prompt_injection_detection": True
}]
)
# Agent Loop
turn_limit = 5
for i in range(turn_limit):
print(f"\n--- Turn {i+1} ---")
has_function_calls = any(
step.type == "function_call"
for step in interaction.steps
)
if not has_function_calls:
text_response = " ".join([
content_block.text for step in interaction.steps if step.type == "model_output"
for content_block in step.content if content_block.type == "text"
])
print("Agent finished:", text_response)
break
print("Executing actions...")
results = execute_function_calls(interaction, page, SCREEN_WIDTH, SCREEN_HEIGHT)
print("Capturing state...")
function_responses = get_function_responses(page, results)
# Continue conversation with function responses
interaction = client.interactions.create(
model='gemini-3.8-flash',
previous_interaction_id=interaction.id,
input=function_responses,
tools=[{
"type": "computer_use",
"environment": "browser",
"enable_prompt_injection_detection": True
}]
)
finally:
# Cleanup
print("\nClosing browser...")
browser.close()
playwright.stop()
JavaScript
import { chromium } from 'playwright';
import { GoogleGenAI } from '@google/genai';
const ai = new GoogleGenAI();
// Constants for screen dimensions
const SCREEN_WIDTH = 1440;
const SCREEN_HEIGHT = 900;
console.log("Initializing browser...");
const browser = await chromium.launch({ headless: false });
const context = await browser.newContext({
viewport: { width: SCREEN_WIDTH, height: SCREEN_HEIGHT }
});
const page = await context.newPage();
// Define helper functions. Copy/paste from steps 3 and 4:
// function denormalizeX(...)
// function denormalizeY(...)
// async function executeFunctionCalls(...)
// async function getFunctionResponses(...)
try {
// Go to initial page
await page.goto("https://ai.google.dev/gemini-api/docs");
// Take initial screenshot
const initialScreenshotBuffer = await page.screenshot({ type: 'png' });
const initialScreenshotBase64 = initialScreenshotBuffer.toString('base64');
const USER_PROMPT = "Go to ai.google.dev/gemini-api/docs and search for pricing.";
console.log(`Goal: ${USER_PROMPT}`);
// First interaction
let interaction = await ai.interactions.create({
model: 'gemini-3.8-flash',
input: [
{ type: 'text', text: USER_PROMPT },
{ type: 'image', data: initialScreenshotBase64, mime_type: 'image/png' }
],
tools: [{
type: 'computer_use',
environment: 'browser',
enable_prompt_injection_detection: true
}]
});
// Agent Loop
const turnLimit = 5;
for (let i = 0; i < turnLimit; i++) {
console.log(`\n--- Turn ${i + 1} ---`);
const hasFunctionCalls = interaction.steps.some(step => step.type === "function_call");
if (!hasFunctionCalls) {
const textResponses = [];
for (const step of interaction.steps) {
if (step.type === "model_output") {
for (const contentBlock of step.content || []) {
if (contentBlock.type === "text") {
textResponses.push(contentBlock.text);
}
}
}
}
console.log("Agent finished:", textResponses.join(" "));
break;
}
console.log("Executing actions...");
const results = await executeFunctionCalls(interaction, page, SCREEN_WIDTH, SCREEN_HEIGHT);
console.log("Capturing state...");
const functionResponses = await getFunctionResponses(page, results);
// Continue conversation with function responses
interaction = await ai.interactions.create({
model: 'gemini-3.8-flash',
previous_interaction_id: interaction.id,
input: functionResponses,
tools: [{
type: 'computer_use',
environment: 'browser',
enable_prompt_injection_detection: true
}]
});
}
} finally {
// Cleanup
console.log("\nClosing browser...");
await browser.close();
}
جاوا
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.Content;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.FunctionCallStep;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.ModelOutputStep;
import com.google.genai.gaos.models.interactions.Step;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.ArrayList;
import java.util.Arrays;
import java.util.Base64;
import java.util.Collections;
import java.util.List;
Client client = new Client();
// Constants for screen dimensions
int screenWidth = 1440;
int screenHeight = 900;
// Capture initial screenshot from browser driver (e.g. Playwright)
byte[] initialScreenshot = new byte[0];
String base64Screenshot = Base64.getEncoder().encodeToString(initialScreenshot);
String userPrompt = "Go to ai.google.dev/gemini-api/docs and search for pricing.";
System.out.println("Goal: " + userPrompt);
ComputerUse computerUseTool =
ComputerUse.builder()
.environment(EnvironmentEnum.BROWSER)
.enablePromptInjectionDetection(true)
.build();
CreateModelInteraction initialParams =
CreateModelInteraction.builder()
.model("gemini-3.8-flash")
.input(
InteractionsInput.ofContent(
Arrays.asList(
TextContent.builder().text(userPrompt).build(),
ImageContent.builder()
.data(base64Screenshot)
.mimeType(ImageContentMimeType.IMAGE_PNG)
.build())))
.tools(Arrays.asList(computerUseTool))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(initialParams)).interaction().get();
int turnLimit = 5;
for (int i = 0; i < turnLimit; i++) {
System.out.println("\n--- Turn " + (i + 1) + " ---");
boolean hasFunctionCalls =
interaction.steps().orElse(Collections.emptyList()).stream()
.anyMatch(step -> step instanceof FunctionCallStep);
if (!hasFunctionCalls) {
StringBuilder textResponse = new StringBuilder();
for (Step step : interaction.steps().orElse(Collections.emptyList())) {
if (step instanceof ModelOutputStep) {
for (Content contentBlock :
((ModelOutputStep) step).content().orElse(Collections.emptyList())) {
if (contentBlock instanceof TextContent) {
textResponse.append(((TextContent) contentBlock).text().orElse("")).append(" ");
}
}
}
}
System.out.println("Agent finished: " + textResponse.toString().trim());
break;
}
System.out.println("Executing actions and capturing state...");
// Execute function calls against browser driver and capture List<Step> functionResponses
List<Step> functionResponses = new ArrayList<>();
CreateModelInteraction nextParams =
CreateModelInteraction.builder()
.model("gemini-3.8-flash")
.previousInteractionId(interaction.id().get())
.input(InteractionsInput.ofStep(functionResponses))
.tools(Arrays.asList(computerUseTool))
.build();
interaction =
client.interactions.create(CreateInteractionRequestBody.of(nextParams)).interaction().get();
}
رفتن
package main
import (
"context"
"encoding/base64"
"fmt"
"log"
"strings"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
// Constants for screen dimensions
screenWidth := 1440
screenHeight := 900
_ = screenWidth
_ = screenHeight
// Capture initial screenshot from browser driver (e.g. Playwright)
initialScreenshot := []byte{}
base64Screenshot := base64.StdEncoding.EncodeToString(initialScreenshot)
userPrompt := "Go to ai.google.dev/gemini-api/docs and search for pricing."
fmt.Println("Goal:", userPrompt)
computerUseTool := interactions.NewTool(interactions.ComputerUse{
Environment: interactions.EnvironmentEnumBrowser.ToPointer(),
EnablePromptInjectionDetection: genai.Ptr(true),
})
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput([]interactions.Content{
interactions.NewContent(interactions.TextContent{Text: userPrompt}),
interactions.NewContent(interactions.ImageContent{
Data: genai.Ptr(base64Screenshot),
MimeType: interactions.ImageContentMimeType("image/png").ToPointer(),
}),
}),
Tools: []interactions.Tool{computerUseTool},
}),
})
if err != nil {
log.Fatal(err)
}
interaction := res.Interaction
turnLimit := 5
for i := 0; i < turnLimit; i++ {
fmt.Printf("\n--- Turn %d ---\n", i+1)
hasFunctionCalls := false
for _, step := range interaction.Steps {
if step.FunctionCallStep != nil {
hasFunctionCalls = true
break
}
}
if !hasFunctionCalls {
var parts []string
for _, step := range interaction.Steps {
if outStep := step.ModelOutputStep; outStep != nil {
for _, contentBlock := range outStep.Content {
if textContent := contentBlock.TextContent; textContent != nil {
parts = append(parts, textContent.GetText())
}
}
}
}
fmt.Println("Agent finished:", strings.TrimSpace(strings.Join(parts, " ")))
break
}
fmt.Println("Executing actions and capturing state...")
// Execute function calls against browser driver and capture []interactions.Step functionResponses
var functionResponses []interactions.Step
nextRes, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
PreviousInteractionID: interaction.ID,
Input: interactions.NewInteractionsInput(functionResponses),
Tools: []interactions.Tool{computerUseTool},
}),
})
if err != nil {
log.Fatal(err)
}
interaction = nextRes.Interaction
}
}
محیطهای پشتیبانیشده
مدلهای Gemini 3.x از سه محیط مشخصشده در computer_use
پیکربندیها پشتیبانی میکنند:
محیط مرورگر (ENVIRONMENT_BROWSER)
کنشهای دردسترس در ابزار مرورگر:
| نام فرمان | شرح | متغیرهای مستقل (در فراخوانی تابع) |
|---|---|---|
| کلیک کنید | در مختصات کلیک چپ میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| double_click | در مختصات دوکلیک میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| triple_click | سهکلیک در مختصات. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| middle_click | در مختصات، کلیک میانی انجام میدهد. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| right_click | در مختصات موردنظر کلیک راست میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| mouse_down | دکمه موشواره را در مختصات فشار میدهد و نگه میدارد. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| mouse_up | دکمه موشواره را در مختصات رها میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| انتقال | مکاننما را به موقعیت مشخصشده منتقل میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| نوع | نوشتار را تایپ میکند. | text: strpress_enter: bool (اختیاری، پیشفرض false)intent: str |
| drag_and_drop | موردی را از مختصات شروع به مختصات پایان میکشاند. | start_y: int (0-999)start_x: int (0-999)end_y: int (0-999)end_x: int (0-999)intent: str |
| انتظار | اجرا را برای تعداد مشخصی ثانیه موقتاً متوقف میکند. | seconds: عدد صحیح (اختیاری، پیشفرض 1)intent: رشته |
| press_key | کلید مشخصشده را فشار میدهد و آن را رها میکند. | key: strintent: str |
| key_down | کلید مشخصشده را فشار میدهد و نگه میدارد. | key: strintent: str |
| key_up | کلید مشخصشده را رها میکند. | key: strintent: str |
| کلید میانبر | ترکیب کلید مشخصشده را فشار میدهد. | keys: List[str]intent: str |
| take_screenshot | نماگرفتی از صفحه فعلی برمیگرداند. | intent: str |
| پیمایش | در مختصاتی با فاصله پیکسلی به بالا، پایین، چپ، یا راست پیمایش میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)direction: رشته ("up"، "down"، "left"، "right")magnitude_in_pixels: عدد صحیح (۰ تا ۹۹۹، اختیاری، پیشفرض 300)intent: رشته |
| go_back | به صفحه وب قبلی در سابقه مرورگر برمیگردد. | intent: str |
| پیمایش | مستقیماً به نشانی وب مشخصشدهای پیمایش میکند. | url: strintent: str |
| go_forward | به صفحه وب بعدی در سابقه مرورگر پیمایش میکند. | intent: str |
محیط تلفن همراه (ENVIRONMENT_MOBILE)
کنشهای محیط بهینهسازیشده برای Android:
| نام فرمان | شرح | متغیرهای مستقل (در فراخوانی تابع) |
|---|---|---|
| open_app | برنامهای را با نامش باز میکند. | app_name: strintent: str |
| کلیک کنید | در مختصات کلیک چپ میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| list_apps | برنامههای دردسترس در دستگاه را فهرست میکند و نام و نام بسته آنها را برمیگرداند. | intent: str |
| انتظار | اجرا را برای تعداد مشخصی ثانیه موقتاً متوقف میکند. | seconds: عدد صحیح (اختیاری، پیشفرض 1)intent: رشته |
| go_back | به صفحه یا صفحه وب قبلی برمیگردد. | intent: str |
| نوع | نوشتار را تایپ میکند. | text: strpress_enter: bool (اختیاری، پیشفرض false)intent: str |
| drag_and_drop | موردی را از مختصات شروع به مختصات پایان میکشاند. | start_y: int (0-999)start_x: int (0-999)end_y: int (0-999)end_x: int (0-999)intent: str |
| long_press | فشار طولانی در مختصات روی صفحهنمایش انجام میدهد. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)seconds: عدد صحیح (اختیاری، پیشفرض 2)intent: رشته |
| press_key | کلید مشخصشده را فشار میدهد و آن را رها میکند. | key: strintent: str |
| take_screenshot | نماگرفتی از صفحه فعلی برمیگرداند. | intent: str |
محیط رایانه رومیزی (ENVIRONMENT_DESKTOP)
فرمانهای مکاننمای سطح سیستمعامل محیطهای میز کار:
| نام فرمان | شرح | متغیرهای مستقل (در فراخوانی تابع) |
|---|---|---|
| کلیک کنید | در مختصات کلیک چپ میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| double_click | در مختصات دوکلیک میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| triple_click | سهکلیک در مختصات. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| middle_click | در مختصات، کلیک میانی انجام میدهد. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| right_click | در مختصات موردنظر کلیک راست میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| mouse_down | دکمه موشواره را در مختصات فشار میدهد و نگه میدارد. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| mouse_up | دکمه موشواره را در مختصات رها میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| انتقال | مکاننما را به موقعیت مشخصشده منتقل میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)intent: رشته |
| نوع | نوشتار را تایپ میکند. | text: strpress_enter: bool (اختیاری، پیشفرض false)intent: str |
| drag_and_drop | موردی را از مختصات شروع به مختصات پایان میکشاند. | start_y: int (0-999)start_x: int (0-999)end_y: int (0-999)end_x: int (0-999)intent: str |
| انتظار | اجرا را برای تعداد مشخصی ثانیه موقتاً متوقف میکند. | seconds: عدد صحیح (اختیاری، پیشفرض 1)intent: رشته |
| press_key | کلید مشخصشده را فشار میدهد و آن را رها میکند. | key: strintent: str |
| key_down | کلید مشخصشده را فشار میدهد و نگه میدارد. | key: strintent: str |
| key_up | کلید مشخصشده را رها میکند. | key: strintent: str |
| کلید میانبر | ترکیب کلید مشخصشده را فشار میدهد. | keys: List[str]intent: str |
| take_screenshot | نماگرفتی از صفحه فعلی برمیگرداند. | intent: str |
| پیمایش | در مختصاتی با فاصله پیکسلی به بالا، پایین، چپ، یا راست پیمایش میکند. | y: عدد صحیح (۰ تا ۹۹۹)x: عدد صحیح (۰ تا ۹۹۹)direction: رشته ("up"، "down"، "left"، "right")magnitude_in_pixels: عدد صحیح (۰ تا ۹۹۹، اختیاری، پیشفرض 300)intent: رشته |
توابع سفارشی تعریفشده توسط کاربر
با افزودن توابع سفارشی تعریفشده توسط کاربر میتوانید عملکرد مدل را گسترش دهید. برای مثال، در سناریوهای انسان در حلقه (HITL) میتوانید کنشهای پیشفرض ازپیش تعریفشده را کنار بگذارید و کنشهای سفارشی را ثبت کنید.
Python
کنشهای مرورگر ازپیش تعریفشده استاندارد (مثل click) را مستثنا کنید و ابزار سفارشی yield_to_user را ثبت کنید:
from google import genai
client = genai.Client()
yield_to_user_tool = {
"type": "function",
"name": "yield_to_user",
"description": "Yields control back to the user for assistance or verification when an automated action is unsafe or ambiguous.",
"parameters": {
"type": "object",
"properties": {
"reason": {
"type": "string",
"description": "The reason why the agent is yielding control to the human."
}
},
"required": ["reason"]
}
}
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Click the submit button. If you need a second factor authentication code, ask me.",
tools=[
{
"type": "computer_use",
"environment": "mobile",
"excluded_predefined_functions": ["click"]
},
yield_to_user_tool
]
)
JavaScript
کنشهای مرورگر ازپیش تعریفشده استاندارد (مثل click) را مستثنا کنید و ابزار سفارشی yield_to_user را ثبت کنید:
import { GoogleGenAI } from '@google/genai';
const ai = new GoogleGenAI();
const yieldToUserTool = {
type: "function",
name: "yield_to_user",
description: "Yields control back to the user for assistance or verification when an automated action is unsafe or ambiguous.",
parameters: {
type: "object",
properties: {
reason: {
type: "string",
description: "The reason why the agent is yielding control to the human."
}
},
required: ["reason"]
}
};
const interaction = await ai.interactions.create({
model: "gemini-3.8-flash",
input: "Click the submit button. If you need a second factor authentication code, ask me.",
tools: [
{
type: "computer_use",
environment: "mobile",
excluded_predefined_functions: ["click"]
},
yieldToUserTool
]
});
جاوا
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Function;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;
import java.util.Collections;
import java.util.HashMap;
import java.util.Map;
Client client = new Client();
Map<String, Object> reasonProp = new HashMap<>();
reasonProp.put("type", "string");
reasonProp.put("description", "The reason why the agent is yielding control to the human.");
Map<String, Object> properties = new HashMap<>();
properties.put("reason", reasonProp);
Map<String, Object> parameters = new HashMap<>();
parameters.put("type", "object");
parameters.put("properties", properties);
parameters.put("required", Collections.singletonList("reason"));
Function yieldToUserTool =
Function.builder()
.name("yield_to_user")
.description(
"Yields control back to the user for assistance or verification when an automated action is unsafe or ambiguous.")
.parameters(parameters)
.build();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model("gemini-3.8-flash")
.input(
InteractionsInput.of(
"Click the submit button. If you need a second factor authentication code, ask me."))
.tools(
Arrays.asList(
ComputerUse.builder()
.environment(EnvironmentEnum.MOBILE)
.excludedPredefinedFunctions(Arrays.asList("click"))
.build(),
yieldToUserTool))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
رفتن
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
yieldToUserTool := interactions.NewTool(interactions.Function{
Name: genai.Ptr("yield_to_user"),
Description: genai.Ptr("Yields control back to the user for assistance or verification when an automated action is unsafe or ambiguous."),
Parameters: map[string]any{
"type": "object",
"properties": map[string]any{
"reason": map[string]any{
"type": "string",
"description": "The reason why the agent is yielding control to the human.",
},
},
"required": []string{"reason"},
},
})
_, err = client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput("Click the submit button. If you need a second factor authentication code, ask me."),
Tools: []interactions.Tool{
interactions.NewTool(interactions.ComputerUse{
Environment: interactions.EnvironmentEnumMobile.ToPointer(),
ExcludedPredefinedFunctions: []string{"click"},
}),
yieldToUserTool,
},
}),
})
if err != nil {
log.Fatal(err)
}
}
مدیریت سطوح اندیشیدن
برای کارگزاران استفاده از رایانه، میتوانید سطوح تفکر مختلفی را پیکربندی کنید تا بین کیفیت کنش و سرعت اجرا تعادل برقرار کنید. سطوح پایینتر اندیشیدن معمولاً برای وظایف خودکارسازی استاندارد تعادل خوبی ایجاد میکنند.
ایمنی و امنیت
درحال پیکربندی خطمشیهای ایمنی
مدلهای Gemini 3.x شامل دستههای سرویس ایمنی داخلی است که به تعیین اینکه آیا تأیید کاربر لازم است یا نه کمک میکند.
| دسته خطمشی ایمنی | شرح |
|---|---|
FINANCIAL_TRANSACTIONS |
تأیید کنشهای مربوط به پرداختها، تسویهحساب خردهفروشی، یا کالاهای تحت نظارت را مسدود یا راهاندازی میکند. |
SENSITIVE_DATA_MODIFICATION |
از سوابق بهداشتی، مالی، یا دولتی دربرابر تغییرات غیرمجاز محافظت میکند. |
COMMUNICATION_TOOL |
نماینده را از ارسال خودکار ایمیل، پیام گپ، یا پیشنویس منع میکند. |
ACCOUNT_CREATION |
عامل را از ثبت خودکار حسابهای جدید در وبسایتها محدود میکند. |
DATA_MODIFICATION |
اصلاحات کلی سیستم فایل، همرسانی دادهها، و حذف فضای ذخیرهسازی را تنظیم میکند. |
USER_CONSENT_MANAGEMENT |
برای برنماهای موافقت با کوکی و پیاموارههای حریم خصوصی، نیاز به تسلط کاربر دارد. |
LEGAL_TERMS_AND_AGREEMENTS |
از پذیرش خودکار «شرایط خدمات» یا قراردادهای الزامآور قانونی توسط مدل جلوگیری میکند. |
ملغی کردن ایمنی
میتوانید با گذراندن ملغیسازیها، خطمشیهای انتخابی را ملغی کنید:
Python
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Clean up the local folder by archiving old logs.",
tools=[
{
"type": "computer_use",
"environment": "desktop",
"disabled_safety_policies": [
"data_modification"
]
}
]
)
JavaScript
import { GoogleGenAI } from '@google/genai';
const ai = new GoogleGenAI();
const interaction = await ai.interactions.create({
model: "gemini-3.8-flash",
input: "Clean up the local folder by archiving old logs.",
tools: [
{
type: "computer_use",
environment: "desktop",
disabled_safety_policies: [
"data_modification"
]
}
]
});
جاوا
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.DisabledSafetyPolicy;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model("gemini-3.8-flash")
.input(InteractionsInput.of("Clean up the local folder by archiving old logs."))
.tools(
Arrays.asList(
ComputerUse.builder()
.environment(EnvironmentEnum.DESKTOP)
.disabledSafetyPolicies(
Arrays.asList(DisabledSafetyPolicy.DATA_MODIFICATION))
.build()))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
رفتن
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
_, err = client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput("Clean up the local folder by archiving old logs."),
Tools: []interactions.Tool{
interactions.NewTool(interactions.ComputerUse{
Environment: interactions.EnvironmentEnumDesktop.ToPointer(),
DisabledSafetyPolicies: []interactions.DisabledSafetyPolicy{
interactions.DisabledSafetyPolicyDataModification,
},
}),
},
}),
})
if err != nil {
log.Fatal(err)
}
}
تشخیص تزریق پیامواره
«استفاده از رایانه» برای Gemini 3.5 Flash-Lite یا نسخههای جدیدتر از سازوکار ایمنی پیشرفتهای برای شناسایی حملات تزریق پیامواره پشتیبانی میکند. وقتی فعال باشد، این ویژگی بررسی میکند که آیا نماگرفت اضافهشده حاوی دستورالعملهای خصمانه پنهان (برای مثال، «دستورات قبلی را نادیده بگیر») است یا نه و درصورت شناسایی، اجرای آن را مسدود میکند.
تشخیص تزریق پیامواره ویژگیای است که باید با آن موافقت کنید. پیشفرض false است.
مثالهای زیر نشان میدهد چگونه تشخیص تزریق پیامواره را در پیکربندی ابزار «استفاده از رایانه» فعال کنید:
Python
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Search for flight deals and summarize top results.",
tools=[
{
"type": "computer_use",
"environment": "desktop",
"enable_prompt_injection_detection": True,
}
],
)
JavaScript
import { GoogleGenAI } from '@google/genai';
const ai = new GoogleGenAI();
const interaction = await ai.interactions.create({
model: "gemini-3.8-flash",
input: "Search for flight deals and summarize top results.",
tools: [
{
type: "computer_use",
environment: "desktop",
enablePromptInjectionDetection: true,
}
]
});
جاوا
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model("gemini-3.8-flash")
.input(InteractionsInput.of("Search for flight deals and summarize top results."))
.tools(
Arrays.asList(
ComputerUse.builder()
.environment(EnvironmentEnum.DESKTOP)
.enablePromptInjectionDetection(true)
.build()))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
رفتن
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
_, err = client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
Input: interactions.NewInteractionsInput("Search for flight deals and summarize top results."),
Tools: []interactions.Tool{
interactions.NewTool(interactions.ComputerUse{
Environment: interactions.EnvironmentEnumDesktop.ToPointer(),
EnablePromptInjectionDetection: genai.Ptr(true),
}),
},
}),
})
if err != nil {
log.Fatal(err)
}
}
cURL
curl "https://generativelanguage.googleapis.com/v1beta/interactions?key=${GEMINI_API_KEY}" \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini-3.8-flash",
"input": "Search for flight deals and summarize top results.",
"tools": [
{
"type": "computer_use",
"environment": "desktop",
"enable_prompt_injection_detection": true
}
]
}'
تأیید تصمیم ایمنی
پاسخ ممکن است شامل پارامتر safety_decision در آرگومانهای فراخوانی تابع باشد:
{
"steps": [
{
"type": "function_call",
"name": "click",
"arguments": {
"x": 60,
"y": 100,
"safety_decision": {
"explanation": "Must check check-box",
"decision": "require_confirmation"
}
}
}
]
}
اگر safety_decision require_confirmation است، از کاربر نهایی درخواست کنید. اگر کاربر تأیید کرد، safety_acknowledgement را در function_result تنظیم کنید.
Python
def get_safety_confirmation(safety_decision):
# Prompt user for confirmation
print(f"Safety confirmation required: {safety_decision.get('explanation', '')}")
return "CONTINUE" # Or TERMINATE
# Inside execute_function_calls, check for safety_decision:
if 'safety_decision' in function_call.arguments:
decision = get_safety_confirmation(function_call.arguments['safety_decision'])
if decision == "TERMINATE":
break
# Include safety_acknowledgement inside the action result
action_result["safety_acknowledgement"] = True
روالهای مطلوب ایمنی
«استفاده از رایانه» خطرات امنیتی و عملیاتی منحصربهفردی دارد، زیرا مدلی که ازطرف کاربر عمل میکند ممکن است با محتوای غیرقابلاعتماد در صفحهها مواجه شود یا در اجرای کنشها دچار خطا شود. برای محافظت از دادههای کاربر و سیستمها، روالهای مطلوب زیر را پیادهسازی کنید:
- مشارکت انسانی در حلقه (HITL):
- اجرای تأیید کاربر: وقتی پاسخ ایمنی نشان میدهد
require_confirmation، از کاربر بخواهید تأیید کند. ارائه دستورالعملهای ایمنی سفارشی: دستورالعمل سیستم سفارشی را برای تعریف و اجرای مرزهای ایمنی خود پیادهسازی کنید. برای مثال:
Python
from google import genai client = genai.Client() system_instruction = """ ## **RULE 1: Seek User Confirmation (USER_CONFIRMATION)** This is your first and most important check. If the next required action falls into any of the following categories, you MUST stop immediately, and seek the user's explicit permission. **Procedure for Seeking Confirmation:** * **For Consequential Actions:** Perform all preparatory steps (e.g., navigating, filling out forms, typing a message). You will ask for confirmation **AFTER** all necessary information is entered on the screen, but **BEFORE** you perform the final, irreversible action (e.g., before clicking "Send", "Submit", "Confirm Purchase", "Share"). * **For Prohibited Actions:** If the action is strictly forbidden (e.g., accepting legal terms, solving a CAPTCHA), you must first inform the user about the required action and ask for their confirmation to proceed. **USER_CONFIRMATION Categories:** * **Consent and Agreements:** You are FORBIDDEN from accepting, selecting, or agreeing to any of the following on the user's behalf. You must ask the user to confirm before performing these actions. * Terms of Service * Privacy Policies * Cookie consent banners * End User License Agreements (EULAs) * Any other legally significant contracts or agreements. * **Robot Detection:** You MUST NEVER attempt to solve or bypass the following. You must ask the user to confirm before performing these actions. * CAPTCHAs (of any kind) * Any other anti-robot or human-verification mechanisms, even if you are capable. * **Financial Transactions:** * Completing any purchase. * Managing or moving money (e.g., transfers, payments). * Purchasing regulated goods or participating in gambling. * **Sending Communications:** * Sending emails. * Sending messages on any platform (e.g., social media, chat apps). * Posting content on social media or forums. * **Accessing or Modifying Sensitive Information:** * Health, financial, or government records (e.g., medical history, tax forms, passport status). * Revealing or modifying sensitive personal identifiers (e.g., SSN, bank account number, credit card number). * **User Data Management:** * Accessing, downloading, or saving files from the web. * Sharing or sending files/data to any third party. * Transferring user data between systems. * **Browser Data Usage:** * Accessing or managing Chrome browsing history, bookmarks, autofill data, or saved passwords. * **Security and Identity:** * Logging into any user account. * Any action that involves misrepresentation or impersonation (e.g., creating a fan account, posting as someone else). * **Insurmountable Obstacles:** If you are technically unable to interact with a user interface element or are stuck in a loop you cannot resolve, ask the user to take over. --- ## **RULE 2: Default Behavior (ACTUATE)** If an action does **NOT** fall under the conditions for `USER_CONFIRMATION`, your default behavior is to **Actuate**. **Actuation Means:** You MUST proactively perform all necessary steps to move the user's request forward. Continue to actuate until you either complete the non-consequential task or encounter a condition defined in Rule 1. * **Example 1:** If asked to send money, you will navigate to the payment portal, enter the recipient's details, and enter the amount. You will then **STOP** as per Rule 1 and ask for confirmation before clicking the final "Send" button. * **Example 2:** If asked to post a message, you will navigate to the site, open the post composition window, and write the full message. You will then **STOP** as per Rule 1 and ask for confirmation before clicking the final "Post" button. After the user has confirmed, remember to get the user's latest screen before continuing to perform actions. # Final Response Guidelines: Write final response to the user in the following cases: - User confirmation - When the task is complete or you have enough information to respond to the user """ interaction = client.interactions.create( model="gemini-3.8-flash", system_instruction=system_instruction, input="Prepare a draft but do not send.", tools=[{ "type": "computer_use", "environment": "browser" }] )JavaScript
import { GoogleGenAI } from '@google/genai'; const ai = new GoogleGenAI(); const systemInstruction = ` ## **RULE 1: Seek User Confirmation (USER_CONFIRMATION)** This is your first and most important check. If the next required action falls into any of the following categories, you MUST stop immediately, and seek the user's explicit permission. **Procedure for Seeking Confirmation:** * **For Consequential Actions:** Perform all preparatory steps (e.g., navigating, filling out forms, typing a message). You will ask for confirmation **AFTER** all necessary information is entered on the screen, but **BEFORE** you perform the final, irreversible action (e.g., before clicking "Send", "Submit", "Confirm Purchase", "Share"). * **For Prohibited Actions:** If the action is strictly forbidden (e.g., accepting legal terms, solving a CAPTCHA), you must first inform the user about the required action and ask for their confirmation to proceed. **USER_CONFIRMATION Categories:** * **Consent and Agreements:** You are FORBIDDEN from accepting, selecting, or agreeing to any of the following on the user's behalf. You must ask the user to confirm before performing these actions. * Terms of Service * Privacy Policies * Cookie consent banners * End User License Agreements (EULAs) * Any other legally significant contracts or agreements. * **Robot Detection:** You MUST NEVER attempt to solve or bypass the following. You must ask the user to confirm before performing these actions. * CAPTCHAs (of any kind) * Any other anti-robot or human-verification mechanisms, even if you are capable. * **Financial Transactions:** * Completing any purchase. * Managing or moving money (e.g., transfers, payments). * Purchasing regulated goods or participating in gambling. * **Sending Communications:** * Sending emails. * Sending messages on any platform (e.g., social media, chat apps). * Posting content on social media or forums. * **Accessing or Modifying Sensitive Information:** * Health, financial, or government records (e.g., medical history, tax forms, passport status). * Revealing or modifying sensitive personal identifiers (e.g., SSN, bank account number, credit card number). * **User Data Management:** * Accessing, downloading, or saving files from the web. * Sharing or sending files/data to any third party. * Transferring user data between systems. * **Browser Data Usage:** * Accessing or managing Chrome browsing history, bookmarks, autofill data, or saved passwords. * **Security and Identity:** * Logging into any user account. * Any action that involves misrepresentation or impersonation (e.g., creating a fan account, posting as someone else). * **Insurmountable Obstacles:** If you are technically unable to interact with a user interface element or are stuck in a loop you cannot resolve, ask the user to take over. --- ## **RULE 2: Default Behavior (ACTUATE)** If an action does **NOT** fall under the conditions for \`USER_CONFIRMATION\`, your default behavior is to **Actuate**. **Actuation Means:** You MUST proactively perform all necessary steps to move the user's request forward. Continue to actuate until you either complete the non-consequential task or encounter a condition defined in Rule 1. * **Example 1:** If asked to send money, you will navigate to the payment portal, enter the recipient's details, and enter the amount. You will then **STOP** as per Rule 1 and ask for confirmation before clicking the final "Send" button. * **Example 2:** If asked to post a message, you will navigate to the site, open the post composition window, and write the full message. You will then **STOP** as per Rule 1 and ask for confirmation before clicking the final "Post" button. After the user has confirmed, remember to get the user's latest screen before continuing to perform actions. # Final Response Guidelines: Write final response to the user in the following cases: - User confirmation - When the task is complete or you have enough information to respond to the user `; const interaction = await ai.interactions.create({ model: "gemini-3.8-flash", system_instruction: systemInstruction, input: "Prepare a draft but do not send.", tools: [{ type: "computer_use", environment: "browser" }] });
- اجرای تأیید کاربر: وقتی پاسخ ایمنی نشان میدهد
جاوا
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;
Client client = new Client();
String systemInstruction =
"## **RULE 1: Seek User Confirmation (USER_CONFIRMATION)**\n\n"
+ "This is your first and most important check. If the next required action falls "
+ "into any of the following categories, you MUST stop immediately, and seek the "
+ "user's explicit permission.\n\n"
+ "## **RULE 2: Default Behavior (ACTUATE)**\n\n"
+ "If an action does **NOT** fall under the conditions for `USER_CONFIRMATION`, "
+ "your default behavior is to **Actuate**.";
CreateModelInteraction params =
CreateModelInteraction.builder()
.model("gemini-3.8-flash")
.systemInstruction(systemInstruction)
.input(InteractionsInput.of("Prepare a draft but do not send."))
.tools(
Arrays.asList(
ComputerUse.builder().environment(EnvironmentEnum.BROWSER).build()))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
رفتن
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
systemInstruction := "## **RULE 1: Seek User Confirmation (USER_CONFIRMATION)**\n\n" +
"This is your first and most important check. If the next required action falls " +
"into any of the following categories, you MUST stop immediately, and seek the " +
"user's explicit permission.\n\n" +
"## **RULE 2: Default Behavior (ACTUATE)**\n\n" +
"If an action does **NOT** fall under the conditions for `USER_CONFIRMATION`, " +
"your default behavior is to **Actuate**."
_, err = client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.8-flash"),
SystemInstruction: genai.Ptr(systemInstruction),
Input: interactions.NewInteractionsInput("Prepare a draft but do not send."),
Tools: []interactions.Tool{
interactions.NewTool(interactions.ComputerUse{
Environment: interactions.EnvironmentEnumBrowser.ToPointer(),
}),
},
}),
})
if err != nil {
log.Fatal(err)
}
}
- محیط اجرای امن: نمایندهتان را در محیطی امن و جعبهشنی اجرا کنید تا تأثیر بالقوه آن محدود شود. این میتواند ماشین مجازی (VM) جعبه شنی، ظرف (برای نمونه، Docker)، یا نمایه مرورگر اختصاصی با اجازههای محدود باشد. برای راهنمایی درباره راهاندازی جعبه شنی بااستفاده از Docker، به پیادهسازی مرجع GitHub مراجعه کنید.
- پاکسازی ورودی: همه نوشتارهای تولیدشده توسط کاربر در پیاموارهها را پاکسازی کنید تا خطر دستورالعملهای ناخواسته یا تزریق پیامواره را کاهش دهید. این یک لایه امنیتی مفید است، اما جایگزین محیط اجرای امن نیست.
- نردههای محافظ محتوا: از نردههای محافظ و «میاناهای برنامهسازی کاربردی ایمنی محتوا» برای ارزیابی دروندادهای کاربر، دروندادها و بروندادهای ابزار، و پاسخهای عامل ازنظر مناسب بودن، تزریق پیامواره، و تشخیص گریز از محدودیت استفاده کنید.
- فهرستهای مجاز و فهرستهای مسدود: سازوکارهای فیلتر کردن را برای کنترل اینکه مدل در کجا میتواند پیمایش کند و چه کارهایی میتواند انجام دهد پیادهسازی کنید. فهرست مسدود وبسایتهای ممنوعه نقطه شروع خوبی است، درحالیکه فهرست مجاز محدودتر حتی امنتر است.
- قابلیت مشاهده و ثبت وقایع: گزارشهای دقیق برای اشکالزدایی، حسابرسی، و پاسخ به حادثه را حفظ کنید. کارخواه شما باید پیاموارهها،
نماگرفتها، کنشهای پیشنهادی مدل (
function_call)، پاسخهای ایمنی، و همه کنشهایی را که درنهایت توسط کارخواه اجرا میشوند گزارش کند. - مدیریت محیط: مطمئن شوید محیط «میانای گرافیکی کاربر» یکپارچه باشد. بالاپرهای غیرمنتظره، اعلانها، یا تغییرات در چیدمان میتواند مدل را گیج کند. درصورت امکان، هر تکلیف جدید را از حالت پاک و شناختهشدهای شروع کنید.
نسخههای مدل
میتوانید از «استفاده از رایانه» با مدلهای زیر استفاده کنید:
- Gemini 3.8 Flash (
gemini-3.8-flash): مدل توصیهشده برای استفاده در رایانه، با ویژگیهای تعامل واسط کاربر با دقت بالا و فراخوانی ابزار قابلاعتماد. - Gemini 3.5 Flash-Lite (
gemini-3.5-flash-lite): مدلی با تأخیر کم و مقرونبهصرفه که از استفاده در رایانه پشتیبانی میکند. - پیشنمایش Gemini 3 Flash (
gemini-3-flash-preview): مدل پیشنمایش پشتیبانی از استفاده در رایانه.
قدم بعدی چیست
- در محیط نمایشی Browserbase، «استفاده از رایانه» را آزمایش کنید.
- برای نمونه کد، پیادهسازی مرجع را بررسی کنید.
- با ابزارهای دیگر «میانای برنامهسازی کاربردی Gemini» آشنا شوید: