Использование компьютера

Инструмент "Использование компьютера" позволяет создавать агентов управления браузером, мобильным устройством и компьютером, которые взаимодействуют с задачами и автоматизируют их. Используя скриншоты, модель может "видеть" экран компьютера и "действовать", генерируя определенные действия в интерфейсе, например клики мышью и нажатия клавиш. Как и в случае с вызовом функций, вам потребуется реализовать клиентскую среду выполнения, чтобы получать и выполнять действия, связанные с использованием компьютера.

Список поддерживаемых моделей приведен в разделе Версии моделей. Модели Gemini 3.x поддерживают несколько расширенных функций:

С помощью функции "Использование компьютера" можно создавать агентов, которые:

  • Автоматизируйте повторяющиеся задачи по вводу данных или заполнению форм на сайтах.
  • Автоматическое тестирование веб-приложений и пользовательских потоков
  • проводить исследования на разных сайтах (например, собирать информацию о товарах, ценах и отзывах на сайтах электронной торговли, чтобы принять решение о покупке);

Ниже приведен минимальный пример инициализации клиента и отправки запроса модели с включенным инструментом computer_use для браузера:

Python

from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input="Search for 'Gemini API' on Google.",
    tools=[{"type": "computer_use", "environment": "browser"}]
)

print(interaction)

JavaScript

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

const interaction = await ai.interactions.create({
  model: 'gemini-3.8-flash',
  input: "Search for 'Gemini API' on Google.",
  tools: [{ type: "computer_use", environment: "browser" }]
});

console.log(interaction);

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model("gemini-3.8-flash")
        .input(InteractionsInput.of("Search for 'Gemini API' on Google."))
        .tools(
            Arrays.asList(
                ComputerUse.builder().environment(EnvironmentEnum.BROWSER).build()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction);

Проложить маршрут

package main

import (
    "context"
    "fmt"
    "log"

    "google.golang.org/genai"
    "google.golang.org/genai/interactions/models/interactions"
    "google.golang.org/genai/interactions/models/operations"
)

func main() {
    ctx := context.Background()
    client, err := genai.NewClient(ctx, nil)
    if err != nil {
        log.Fatal(err)
    }

    res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
        Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
            Model: interactions.Model("gemini-3.8-flash"),
            Input: interactions.NewInteractionsInput("Search for 'Gemini API' on Google."),
            Tools: []interactions.Tool{
                interactions.NewTool(interactions.ComputerUse{
                    Environment: interactions.EnvironmentEnumBrowser.ToPointer(),
                }),
            },
        }),
    })
    if err != nil {
        log.Fatal(err)
    }

    fmt.Println(res.Interaction)
}


Как работает функция "Использование компьютера"

Чтобы создать агента с моделью "Использование компьютера", необходимо настроить непрерывный цикл между приложением и API. Вот что будет делать ваш код на каждом этапе:

  1. Как отправить запрос модели
    • Ваше приложение отправляет запрос к API, содержащий инструмент "Использование компьютера", настройки конфигурации (например, целевую среду), запрос пользователя и скриншот текущего экрана.
  2. Получение ответа модели
    • Модель анализирует экран и запрос и возвращает ответ, который включает предложенный function_call, представляющий действие в интерфейсе (например, клик, прокрутку или нажатие клавиши).
    • Для моделей Gemini 3.x в ответе также приводится обоснование intent, объясняющее, почему модель выбрала именно это действие.
    • Ответ также может содержать safety_decision от внутренней системы безопасности, которая классифицирует действие как обычное/разрешенное, require_confirmation (требующее одобрения пользователя) или заблокированное.
  3. Выполните полученное действие
    • Если действие разрешено (или пользователь подтверждает его), ваш клиентский код анализирует function_call, масштабирует нормализованные координаты в соответствии с областью просмотра и выполняет действие в целевой среде с помощью инструментов автоматизации (например, Playwright). Если действие заблокировано, клиент должен остановить выполнение или обработать прерывание.
  4. Захват нового состояния среды
    • После того как действие будет выполнено, приложение сделает новый скриншот и отправит его обратно модели в событии function_result, чтобы запросить следующий шаг.

Затем процесс повторяется с шага 2, и модель запрашивает следующее действие, пока задача не будет выполнена или не будет прервана.

Обзор функции "Работа за компьютером"

Как реализовать функцию "Управление компьютером"

Прежде чем использовать инструмент "Использование компьютера", вам нужно настроить:

  • Безопасная среда выполнения. Запустите агента в изолированной виртуальной машине или контейнере, чтобы отделить его от хост-системы и ограничить его потенциальное влияние. Эталонная реализация включает готовую к использованию песочницу на основе Docker, которую можно использовать в качестве отправной точки.
  • Обработчик действий на стороне клиента. Реализуйте клиентскую логику для выполнения координат, ввода текста и создания скриншотов.

В приведенных ниже примерах в качестве среды выполнения используется веб-браузер, а в качестве клиентского обработчика – Playwright.

0. Как настроить Playwright

Сначала установите необходимые пакеты:

pip install google-genai playwright
playwright install chromium

Затем инициализируйте экземпляр браузера Playwright для выполнения:

from playwright.sync_api import sync_playwright

# 1. Configure screen dimensions for the target environment
SCREEN_WIDTH = 1440
SCREEN_HEIGHT = 900

# 2. Start the Playwright browser
# In production, utilize a sandboxed environment.
playwright = sync_playwright().start()
# Set headless=False to see the actions performed on your screen
browser = playwright.chromium.launch(headless=False)

# 3. Create a context and page with the specified dimensions
context = browser.new_context(
    viewport={"width": SCREEN_WIDTH, "height": SCREEN_HEIGHT}
)
page = context.new_page()

# 4. Navigate to an initial page to start the task
page.goto("https://www.google.com")

# The 'page', 'SCREEN_WIDTH', and 'SCREEN_HEIGHT' variables
# will be used in the steps below.

1. Как отправить запрос модели

Инициализируйте клиентскую библиотеку и настройте инструмент "Использование компьютеров". Обратите внимание, что при отправке запроса не нужно указывать размер экрана. Модель прогнозирует пиксельные координаты, масштабированные по высоте и ширине экрана.

Python

Используйте google-genai Python SDK (версии 2.7.0 или более поздней), чтобы настроить запрос, предназначенный для браузера:

from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model='gemini-3.8-flash',
    input="Find a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th",
    tools=[
        {
            "type": "computer_use",
            "environment": "browser",
            "enable_prompt_injection_detection": True
        }
    ]
)

print(interaction)

JavaScript

Чтобы настроить запрос, предназначенный для браузера, используйте @google/genai Node.js SDK:

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

const interaction = await ai.interactions.create({
  model: 'gemini-3.8-flash',
  input: "Find a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th",
  tools: [
    {
      type: "computer_use",
      environment: "browser",
      enable_prompt_injection_detection: true
    }
  ]
});

console.log(interaction);

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model("gemini-3.8-flash")
        .input(
            InteractionsInput.of(
                "Find a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th"))
        .tools(
            Arrays.asList(
                ComputerUse.builder()
                    .environment(EnvironmentEnum.BROWSER)
                    .enablePromptInjectionDetection(true)
                    .build()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction);

Проложить маршрут

package main

import (
    "context"
    "fmt"
    "log"

    "google.golang.org/genai"
    "google.golang.org/genai/interactions/models/interactions"
    "google.golang.org/genai/interactions/models/operations"
)

func main() {
    ctx := context.Background()
    client, err := genai.NewClient(ctx, nil)
    if err != nil {
        log.Fatal(err)
    }

    res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
        Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
            Model: interactions.Model("gemini-3.8-flash"),
            Input: interactions.NewInteractionsInput("Find a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th"),
            Tools: []interactions.Tool{
                interactions.NewTool(interactions.ComputerUse{
                    Environment:                    interactions.EnvironmentEnumBrowser.ToPointer(),
                    EnablePromptInjectionDetection: genai.Ptr(true),
                }),
            },
        }),
    })
    if err != nil {
        log.Fatal(err)
    }

    fmt.Println(res.Interaction)
}

REST

Чтобы отправить запрос с помощью curl:

curl -X POST \
  "https://generativelanguage.googleapis.com/v1beta/interactions" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.8-flash",
    "input": "Find me a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th. Start by navigating directly to flights.google.com",
    "tools": [
      {
        "type": "computer_use",
        "environment": "browser",
        "enable_prompt_injection_detection": true
      }
    ]
  }'

2. Получение ответа модели

В ответе модели предлагается вызвать функцию, содержащую координаты и специальное обоснование, объясняющее действие:

{
  "steps": [
    {
      "type": "function_call",
      "name": "click",
      "arguments": {
        "x": 450,
        "y": 120,
        "intent": "Click the search box to type the destination."
      }
    }
  ]
}

3. Выполнять полученные действия

Ваше приложение должно проанализировать координаты ответа, масштабировать их из нормализованных координат 1000 x 1000 и выполнить действие:

Python

from typing import Any, List, Tuple
import time

def denormalize_x(x: int, screen_width: int) -> int:
    """Convert normalized x coordinate (0-1000) to actual pixel coordinate."""
    return int(x / 1000 * screen_width)

def denormalize_y(y: int, screen_height: int) -> int:
    """Convert normalized y coordinate (0-1000) to actual pixel coordinate."""
    return int(y / 1000 * screen_height)

def execute_function_calls(interaction, page, screen_width, screen_height):
    results = []
    function_calls = [
        step for step in interaction.steps if step.type == "function_call"
    ]

    for function_call in function_calls:
        action_result = {}
        fname = function_call.name
        args = function_call.arguments
        print(f"  -> Executing: {fname} (Intent: {args.get('intent', 'N/A')})")

        try:
            if fname == "open_app":
                pass # Handled / already open
            elif fname in ("click", "double_click", "triple_click", "middle_click", "right_click", "move", "long_press"):
                actual_x = denormalize_x(args["x"], screen_width)
                actual_y = denormalize_y(args["y"], screen_height)

                if fname == "click":
                    page.mouse.click(actual_x, actual_y)
                elif fname == "double_click":
                    page.mouse.dblclick(actual_x, actual_y)
                elif fname == "right_click":
                    page.mouse.click(actual_x, actual_y, button="right")
                elif fname == "middle_click":
                    page.mouse.click(actual_x, actual_y, button="middle")
                elif fname == "move":
                    page.mouse.move(actual_x, actual_y)
            elif fname == "type":
                actual_x = denormalize_x(args["x"], screen_width) if "x" in args else None
                actual_y = denormalize_y(args["y"], screen_height) if "y" in args else None
                text = args["text"]
                press_enter = args.get("press_enter", False)

                if actual_x is not None and actual_y is not None:
                    page.mouse.click(actual_x, actual_y)
                # Clear field first
                page.keyboard.press("Meta+A")
                page.keyboard.press("Backspace")
                page.keyboard.type(text)
                if press_enter:
                    page.keyboard.press("Enter")
            elif fname == "navigate":
                page.goto(args["url"])
            elif fname == "go_back":
                page.go_back()
            elif fname == "go_forward":
                page.go_forward()
            elif fname == "wait":
                time.sleep(args.get("seconds", 1))
            else:
                print(f"Warning: Custom or unhandled function {fname}")

            page.wait_for_load_state(timeout=5000)
            time.sleep(1)

        except Exception as e:
            print(f"Error executing {fname}: {e}")
            action_result = {"error": str(e)}

        results.append((fname, function_call.id, action_result))

    return results

JavaScript

function denormalizeX(x, screenWidth) {
    // Convert normalized x coordinate (0-1000) to actual pixel coordinate.
    return Math.floor((x / 1000) * screenWidth);
}

function denormalizeY(y, screenHeight) {
    // Convert normalized y coordinate (0-1000) to actual pixel coordinate.
    return Math.floor((y / 1000) * screenHeight);
}

async function executeFunctionCalls(interaction, page, screenWidth, screenHeight) {
    const results = [];
    const functionCalls = interaction.steps.filter(step => step.type === "function_call");

    for (const functionCall of functionCalls) {
        const actionResult = {};
        const fname = functionCall.name;
        const args = functionCall.arguments;
        console.log(`  -> Executing: ${fname} (Intent: ${args.intent || 'N/A'})`);

        try {
            if (fname === "open_app") {
                // Handled / already open
            } else if (["click", "double_click", "triple_click", "middle_click", "right_click", "move", "long_press"].includes(fname)) {
                const actualX = denormalizeX(args.x, screenWidth);
                const actualY = denormalizeY(args.y, screenHeight);

                if (fname === "click") {
                    await page.mouse.click(actualX, actualY);
                } else if (fname === "double_click") {
                    await page.mouse.dblclick(actualX, actualY);
                } else if (fname === "right_click") {
                    await page.mouse.click(actualX, actualY, { button: "right" });
                } else if (fname === "middle_click") {
                    await page.mouse.click(actualX, actualY, { button: "middle" });
                } else if (fname === "move") {
                    await page.mouse.move(actualX, actualY);
                }
            } else if (fname === "type") {
                const actualX = args.x !== undefined ? denormalizeX(args.x, screenWidth) : null;
                const actualY = args.y !== undefined ? denormalizeY(args.y, screenHeight) : null;
                const text = args.text;
                const pressEnter = args.press_enter || false;

                if (actualX !== null && actualY !== null) {
                    await page.mouse.click(actualX, actualY);
                }
                // Clear field first
                await page.keyboard.press("Meta+A");
                await page.keyboard.press("Backspace");
                await page.keyboard.type(text);
                if (pressEnter) {
                    await page.keyboard.press("Enter");
                }
            } else if (fname === "navigate") {
                await page.goto(args.url);
            } else if (fname === "go_back") {
                await page.goBack();
            } else if (fname === "go_forward") {
                await page.goForward();
            } else if (fname === "wait") {
                await new Promise(resolve => setTimeout(resolve, (args.seconds || 1) * 1000));
            } else {
                console.log(`Warning: Custom or unhandled function ${fname}`);
            }

            await page.waitForLoadState('load', { timeout: 5000 }).catch(() => {});
            await new Promise(resolve => setTimeout(resolve, 1000));
        } catch (e) {
            console.log(`Error executing ${fname}: ${e}`);
            actionResult.error = e.message;
        }

        results.push([fname, functionCall.id, actionResult]);
    }

    return results;
}

Java

import com.google.genai.gaos.models.interactions.FunctionCallStep;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.Step;
import java.util.ArrayList;
import java.util.Collections;
import java.util.HashMap;
import java.util.List;
import java.util.Map;

class ActionExecutor {
  int denormalizeX(int x, int screenWidth) {
    return (int) (x / 1000.0 * screenWidth);
  }

  int denormalizeY(int y, int screenHeight) {
    return (int) (y / 1000.0 * screenHeight);
  }

  List<Map<String, Object>> executeFunctionCalls(
      Interaction interaction, int screenWidth, int screenHeight) {
    List<Map<String, Object>> results = new ArrayList<>();

    for (Step step : interaction.steps().orElse(Collections.emptyList())) {
      if (step instanceof FunctionCallStep) {
        FunctionCallStep functionCall = (FunctionCallStep) step;
        String fname = functionCall.name().orElse("");
        Map<String, Object> args = functionCall.arguments().orElse(Collections.emptyMap());
        Map<String, Object> actionResult = new HashMap<>();

        System.out.println(
            "  -> Executing: " + fname + " (Intent: " + args.getOrDefault("intent", "N/A") + ")");

        try {
          if (fname.equals("click")) {
            int actualX = denormalizeX(((Number) args.get("x")).intValue(), screenWidth);
            int actualY = denormalizeY(((Number) args.get("y")).intValue(), screenHeight);
            // Perform mouse click at (actualX, actualY) using your browser automation library
          } else if (fname.equals("type")) {
            String text = (String) args.get("text");
            // Type text into active element using your browser automation library
          } else if (fname.equals("navigate")) {
            String url = (String) args.get("url");
            // Navigate browser to url
          }
        } catch (Exception e) {
          actionResult.put("error", e.getMessage());
        }

        Map<String, Object> entry = new HashMap<>();
        entry.put("name", fname);
        entry.put("callId", functionCall.id().orElse(""));
        entry.put("result", actionResult);
        results.add(entry);
      }
    }
    return results;
  }
}

Проложить маршрут

package main

import (
    "fmt"

    "google.golang.org/genai/interactions/models/interactions"
)

func denormalizeX(x, screenWidth int) int {
    return int(float64(x) / 1000.0 * float64(screenWidth))
}

func denormalizeY(y, screenHeight int) int {
    return int(float64(y) / 1000.0 * float64(screenHeight))
}

func executeFunctionCalls(interaction *interactions.Interaction, screenWidth, screenHeight int) []map[string]any {
    var results []map[string]any

    for _, step := range interaction.Steps {
        if functionCall := step.FunctionCallStep; functionCall != nil {
            fname := functionCall.Name
            args := functionCall.Arguments
            actionResult := map[string]any{}

            intent := args["intent"]
            if intent == nil {
                intent = "N/A"
            }
            fmt.Printf("  -> Executing: %s (Intent: %v)\n", fname, intent)

            switch fname {
            case "click":
                xVal, _ := args["x"].(float64)
                yVal, _ := args["y"].(float64)
                actualX := denormalizeX(int(xVal), screenWidth)
                actualY := denormalizeY(int(yVal), screenHeight)
                _ = actualX
                _ = actualY
                // Perform mouse click at (actualX, actualY) using your browser automation library
            case "type":
                text, _ := args["text"].(string)
                _ = text
                // Type text into active element using your browser automation library
            case "navigate":
                url, _ := args["url"].(string)
                _ = url
                // Navigate browser to url
            }

            results = append(results, map[string]any{
                "name":   fname,
                "callId": functionCall.ID,
                "result": actionResult,
            })
        }
    }
    return results
}

func main() {
    // Example helper usage with an Interaction response
}

4. Зафиксируйте новое состояние среды

После выполнения действий отправьте результат выполнения функции обратно модели, чтобы она могла использовать эту информацию для создания следующего действия. Если было выполнено несколько действий (параллельных вызовов), в следующем ходе пользователя необходимо отправить команду function_result для каждого из них.

Python

import json
import base64

def get_function_responses(page, results):
    screenshot_bytes = page.screenshot(type="png")
    current_url = page.url
    function_responses = []
    for name, call_id, result in results:
        function_responses.append({
            "type": "function_result",
            "name": name,
            "call_id": call_id,
            "result": [
                {
                    "type": "text",
                    "text": json.dumps({"url": current_url, **result})
                },
                {
                    "type": "image",
                    "data": base64.b64encode(screenshot_bytes).decode("utf-8"),
                    "mime_type": "image/png"
                }
            ]
        })
    return function_responses

JavaScript

async function getFunctionResponses(page, results) {
    const screenshotBuffer = await page.screenshot({ type: 'png' });
    const screenshotBase64 = screenshotBuffer.toString('base64');
    const currentUrl = page.url();
    const functionResponses = [];

    for (const [name, callId, result] of results) {
        functionResponses.push({
            type: "function_result",
            name: name,
            call_id: callId,
            result: [
                {
                    type: "text",
                    text: JSON.stringify({ url: currentUrl, ...result })
                },
                {
                    type: "image",
                    data: screenshotBase64,
                    mime_type: "image/png"
                }
            ]
        });
    }
    return functionResponses;
}

Java

import com.google.genai.gaos.models.interactions.FunctionResultStep;
import com.google.genai.gaos.models.interactions.FunctionResultStepResultUnion;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Step;
import com.google.genai.gaos.models.interactions.TextContent;
import java.util.ArrayList;
import java.util.Arrays;
import java.util.Base64;
import java.util.List;
import java.util.Map;

class StateCapturer {
  List<Step> getFunctionResponses(
      byte[] screenshotBytes, String currentUrl, List<Map<String, Object>> results) {
    List<Step> functionResponses = new ArrayList<>();
    String base64Screenshot = Base64.getEncoder().encodeToString(screenshotBytes);

    for (Map<String, Object> entry : results) {
      String name = (String) entry.get("name");
      String callId = (String) entry.get("callId");
      String jsonResult = String.format("{\"url\": \"%s\"}", currentUrl);

      FunctionResultStep responseStep =
          FunctionResultStep.builder()
              .name(name)
              .callId(callId)
              .result(
                  FunctionResultStepResultUnion.of(
                      Arrays.asList(
                          TextContent.builder().text(jsonResult).build(),
                          ImageContent.builder()
                              .data(base64Screenshot)
                              .mimeType(ImageContentMimeType.IMAGE_PNG)
                              .build())))
              .build();
      functionResponses.add(responseStep);
    }
    return functionResponses;
  }
}

Проложить маршрут

package main

import (
    "encoding/base64"
    "fmt"

    "google.golang.org/genai"
    "google.golang.org/genai/interactions/models/interactions"
)

func getFunctionResponses(screenshotBytes []byte, currentURL string, results []map[string]any) []interactions.Step {
    var functionResponses []interactions.Step
    base64Screenshot := base64.StdEncoding.EncodeToString(screenshotBytes)

    for _, entry := range results {
        name, _ := entry["name"].(string)
        callID, _ := entry["callId"].(string)
        jsonResult := fmt.Sprintf(`{"url": "%s"}`, currentURL)

        responseStep := interactions.NewStep(interactions.FunctionResultStep{
            Name:   genai.Ptr(name),
            CallID: callID,
            Result: interactions.NewFunctionResultStepResultUnion([]interactions.FunctionResultSubcontent{
                interactions.NewFunctionResultSubcontent(interactions.TextContent{
                    Text: jsonResult,
                }),
                interactions.NewFunctionResultSubcontent(interactions.ImageContent{
                    Data:     genai.Ptr(base64Screenshot),
                    MimeType: interactions.ImageContentMimeType("image/png").ToPointer(),
                }),
            }),
        })
        functionResponses = append(functionResponses, responseStep)
    }
    return functionResponses
}

func main() {
    // Example helper usage to build FunctionResultStep responses
}

После того как вы определите, как собирать и форматировать данные о состоянии среды, вы можете объединить все эти шаги в непрерывный цикл выполнения.

Как создать цикл агента

Чтобы включить многоэтапные взаимодействия, объедините четыре шага из раздела Как реализовать использование компьютера в один цикл. Этот цикл продолжается, пока задача не будет выполнена.

Не забывайте правильно управлять историей чата, добавляя в нее ответы модели и функции на каждом шаге.

Python

import time
from typing import Any, List, Tuple
from playwright.sync_api import sync_playwright

from google import genai

client = genai.Client()

# Constants for screen dimensions
SCREEN_WIDTH = 1440
SCREEN_HEIGHT = 900

# Setup Playwright
print("Initializing browser...")
playwright = sync_playwright().start()
browser = playwright.chromium.launch(headless=False)
context = browser.new_context(viewport={"width": SCREEN_WIDTH, "height": SCREEN_HEIGHT})
page = context.new_page()

# Define helper functions. Copy/paste from steps 3 and 4
# def denormalize_x(...)
# def denormalize_y(...)
# def execute_function_calls(...)
# def get_function_responses(...)

try:
    # Go to initial page
    page.goto("https://ai.google.dev/gemini-api/docs")

    # Take initial screenshot
    initial_screenshot = page.screenshot(type="png")
    USER_PROMPT = "Go to ai.google.dev/gemini-api/docs and search for pricing."
    print(f"Goal: {USER_PROMPT}")

    # First interaction
    interaction = client.interactions.create(
        model='gemini-3.8-flash',
        input=[
            {"type": "text", "text": USER_PROMPT},
            {"type": "image", "data": base64.b64encode(initial_screenshot).decode("utf-8"), "mime_type": "image/png"}
        ],
        tools=[{
            "type": "computer_use",
            "environment": "browser",
            "enable_prompt_injection_detection": True
        }]
    )

    # Agent Loop
    turn_limit = 5
    for i in range(turn_limit):
        print(f"\n--- Turn {i+1} ---")

        has_function_calls = any(
            step.type == "function_call"
            for step in interaction.steps
        )
        if not has_function_calls:
            text_response = " ".join([
                content_block.text for step in interaction.steps if step.type == "model_output"
                for content_block in step.content if content_block.type == "text"
            ])
            print("Agent finished:", text_response)
            break

        print("Executing actions...")
        results = execute_function_calls(interaction, page, SCREEN_WIDTH, SCREEN_HEIGHT)

        print("Capturing state...")
        function_responses = get_function_responses(page, results)

        # Continue conversation with function responses
        interaction = client.interactions.create(
            model='gemini-3.8-flash',
            previous_interaction_id=interaction.id,
            input=function_responses,
            tools=[{
                "type": "computer_use",
                "environment": "browser",
                "enable_prompt_injection_detection": True
            }]
        )

finally:
    # Cleanup
    print("\nClosing browser...")
    browser.close()
    playwright.stop()

JavaScript

import { chromium } from 'playwright';
import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

// Constants for screen dimensions
const SCREEN_WIDTH = 1440;
const SCREEN_HEIGHT = 900;

console.log("Initializing browser...");
const browser = await chromium.launch({ headless: false });
const context = await browser.newContext({
    viewport: { width: SCREEN_WIDTH, height: SCREEN_HEIGHT }
});
const page = await context.newPage();

// Define helper functions. Copy/paste from steps 3 and 4:
// function denormalizeX(...)
// function denormalizeY(...)
// async function executeFunctionCalls(...)
// async function getFunctionResponses(...)

try {
    // Go to initial page
    await page.goto("https://ai.google.dev/gemini-api/docs");

    // Take initial screenshot
    const initialScreenshotBuffer = await page.screenshot({ type: 'png' });
    const initialScreenshotBase64 = initialScreenshotBuffer.toString('base64');
    const USER_PROMPT = "Go to ai.google.dev/gemini-api/docs and search for pricing.";
    console.log(`Goal: ${USER_PROMPT}`);

    // First interaction
    let interaction = await ai.interactions.create({
        model: 'gemini-3.8-flash',
        input: [
            { type: 'text', text: USER_PROMPT },
            { type: 'image', data: initialScreenshotBase64, mime_type: 'image/png' }
        ],
        tools: [{
            type: 'computer_use',
            environment: 'browser',
            enable_prompt_injection_detection: true
        }]
    });

    // Agent Loop
    const turnLimit = 5;
    for (let i = 0; i < turnLimit; i++) {
        console.log(`\n--- Turn ${i + 1} ---`);

        const hasFunctionCalls = interaction.steps.some(step => step.type === "function_call");
        if (!hasFunctionCalls) {
            const textResponses = [];
            for (const step of interaction.steps) {
                if (step.type === "model_output") {
                    for (const contentBlock of step.content || []) {
                        if (contentBlock.type === "text") {
                            textResponses.push(contentBlock.text);
                        }
                    }
                }
            }
            console.log("Agent finished:", textResponses.join(" "));
            break;
        }

        console.log("Executing actions...");
        const results = await executeFunctionCalls(interaction, page, SCREEN_WIDTH, SCREEN_HEIGHT);

        console.log("Capturing state...");
        const functionResponses = await getFunctionResponses(page, results);

        // Continue conversation with function responses
        interaction = await ai.interactions.create({
            model: 'gemini-3.8-flash',
            previous_interaction_id: interaction.id,
            input: functionResponses,
            tools: [{
                type: 'computer_use',
                environment: 'browser',
                enable_prompt_injection_detection: true
            }]
        });
    }
} finally {
    // Cleanup
    console.log("\nClosing browser...");
    await browser.close();
}

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.Content;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.FunctionCallStep;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.ModelOutputStep;
import com.google.genai.gaos.models.interactions.Step;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.ArrayList;
import java.util.Arrays;
import java.util.Base64;
import java.util.Collections;
import java.util.List;

Client client = new Client();

// Constants for screen dimensions
int screenWidth = 1440;
int screenHeight = 900;

// Capture initial screenshot from browser driver (e.g. Playwright)
byte[] initialScreenshot = new byte[0];
String base64Screenshot = Base64.getEncoder().encodeToString(initialScreenshot);
String userPrompt = "Go to ai.google.dev/gemini-api/docs and search for pricing.";
System.out.println("Goal: " + userPrompt);

ComputerUse computerUseTool =
    ComputerUse.builder()
        .environment(EnvironmentEnum.BROWSER)
        .enablePromptInjectionDetection(true)
        .build();

CreateModelInteraction initialParams =
    CreateModelInteraction.builder()
        .model("gemini-3.8-flash")
        .input(
            InteractionsInput.ofContent(
                Arrays.asList(
                    TextContent.builder().text(userPrompt).build(),
                    ImageContent.builder()
                        .data(base64Screenshot)
                        .mimeType(ImageContentMimeType.IMAGE_PNG)
                        .build())))
        .tools(Arrays.asList(computerUseTool))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(initialParams)).interaction().get();

int turnLimit = 5;
for (int i = 0; i < turnLimit; i++) {
  System.out.println("\n--- Turn " + (i + 1) + " ---");

  boolean hasFunctionCalls =
      interaction.steps().orElse(Collections.emptyList()).stream()
          .anyMatch(step -> step instanceof FunctionCallStep);

  if (!hasFunctionCalls) {
    StringBuilder textResponse = new StringBuilder();
    for (Step step : interaction.steps().orElse(Collections.emptyList())) {
      if (step instanceof ModelOutputStep) {
        for (Content contentBlock :
            ((ModelOutputStep) step).content().orElse(Collections.emptyList())) {
          if (contentBlock instanceof TextContent) {
            textResponse.append(((TextContent) contentBlock).text().orElse("")).append(" ");
          }
        }
      }
    }
    System.out.println("Agent finished: " + textResponse.toString().trim());
    break;
  }

  System.out.println("Executing actions and capturing state...");
  // Execute function calls against browser driver and capture List<Step> functionResponses
  List<Step> functionResponses = new ArrayList<>();

  CreateModelInteraction nextParams =
      CreateModelInteraction.builder()
          .model("gemini-3.8-flash")
          .previousInteractionId(interaction.id().get())
          .input(InteractionsInput.ofStep(functionResponses))
          .tools(Arrays.asList(computerUseTool))
          .build();

  interaction =
      client.interactions.create(CreateInteractionRequestBody.of(nextParams)).interaction().get();
}

Проложить маршрут

package main

import (
    "context"
    "encoding/base64"
    "fmt"
    "log"
    "strings"

    "google.golang.org/genai"
    "google.golang.org/genai/interactions/models/interactions"
    "google.golang.org/genai/interactions/models/operations"
)

func main() {
    ctx := context.Background()
    client, err := genai.NewClient(ctx, nil)
    if err != nil {
        log.Fatal(err)
    }

    // Constants for screen dimensions
    screenWidth := 1440
    screenHeight := 900
    _ = screenWidth
    _ = screenHeight

    // Capture initial screenshot from browser driver (e.g. Playwright)
    initialScreenshot := []byte{}
    base64Screenshot := base64.StdEncoding.EncodeToString(initialScreenshot)
    userPrompt := "Go to ai.google.dev/gemini-api/docs and search for pricing."
    fmt.Println("Goal:", userPrompt)

    computerUseTool := interactions.NewTool(interactions.ComputerUse{
        Environment:                    interactions.EnvironmentEnumBrowser.ToPointer(),
        EnablePromptInjectionDetection: genai.Ptr(true),
    })

    res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
        Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
            Model: interactions.Model("gemini-3.8-flash"),
            Input: interactions.NewInteractionsInput([]interactions.Content{
                interactions.NewContent(interactions.TextContent{Text: userPrompt}),
                interactions.NewContent(interactions.ImageContent{
                    Data:     genai.Ptr(base64Screenshot),
                    MimeType: interactions.ImageContentMimeType("image/png").ToPointer(),
                }),
            }),
            Tools: []interactions.Tool{computerUseTool},
        }),
    })
    if err != nil {
        log.Fatal(err)
    }
    interaction := res.Interaction

    turnLimit := 5
    for i := 0; i < turnLimit; i++ {
        fmt.Printf("\n--- Turn %d ---\n", i+1)

        hasFunctionCalls := false
        for _, step := range interaction.Steps {
            if step.FunctionCallStep != nil {
                hasFunctionCalls = true
                break
            }
        }

        if !hasFunctionCalls {
            var parts []string
            for _, step := range interaction.Steps {
                if outStep := step.ModelOutputStep; outStep != nil {
                    for _, contentBlock := range outStep.Content {
                        if textContent := contentBlock.TextContent; textContent != nil {
                            parts = append(parts, textContent.GetText())
                        }
                    }
                }
            }
            fmt.Println("Agent finished:", strings.TrimSpace(strings.Join(parts, " ")))
            break
        }

        fmt.Println("Executing actions and capturing state...")
        // Execute function calls against browser driver and capture []interactions.Step functionResponses
        var functionResponses []interactions.Step

        nextRes, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
            Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
                Model:                 interactions.Model("gemini-3.8-flash"),
                PreviousInteractionID: interaction.ID,
                Input:                 interactions.NewInteractionsInput(functionResponses),
                Tools:                 []interactions.Tool{computerUseTool},
            }),
        })
        if err != nil {
            log.Fatal(err)
        }
        interaction = nextRes.Interaction
    }
}

Поддерживаемые среды

Модели Gemini 3.x поддерживают три среды, указанные в computer_use конфигурациях:

Среда браузера (ENVIRONMENT_BROWSER)

Доступные действия в инструменте браузера:

Название команды Описание Аргументы (в вызове функции)
click Нажимает левую кнопку мыши в указанных координатах. y: int (0-999)
x: int (0-999)
intent: str
double_click Дважды нажимает на координаты. y: int (0-999)
x: int (0-999)
intent: str
triple_click Трижды нажмите на координаты. y: int (0-999)
x: int (0-999)
intent: str
middle_click Нажатие на среднюю кнопку мыши в координатах. y: int (0-999)
x: int (0-999)
intent: str
right_click Нажимает правой кнопкой мыши на указанных координатах. y: int (0-999)
x: int (0-999)
intent: str
mouse_down Нажимает и удерживает кнопку мыши в указанных координатах. y: int (0-999)
x: int (0-999)
intent: str
mouse_up Отпускает кнопку мыши в указанных координатах. y: int (0-999)
x: int (0-999)
intent: str
move Перемещает курсор в указанное положение. y: int (0-999)
x: int (0-999)
intent: str
type; Вводит текст. text: str
press_enter: bool (необязательно, по умолчанию false)
intent: str
drag_and_drop Перетаскивает объект из начальной координаты в конечную. start_y: int (0-999)
start_x: int (0-999)
end_y: int (0-999)
end_x: int (0-999)
intent: str
wait Приостанавливает выполнение на указанное количество секунд. seconds: int (необязательный, по умолчанию 1)
intent: str
press_key Нажимает указанную клавишу и отпускает ее. key: str
intent: str
key_down Нажимает и удерживает указанную клавишу. key: str
intent: str
key_up Отпускает указанную клавишу. key: str
intent: str
hotkey Нажимает указанное сочетание клавиш. keys: List[str]
intent: str
take_screenshot Возвращает скриншот текущего экрана. intent: стр
Прокрутка Прокручивает экран вверх, вниз, влево или вправо на заданное количество пикселей. y: int (0–999)
x: int (0–999)
direction: str ("up", "down", "left", "right")
magnitude_in_pixels: int (0–999, необязательный параметр, по умолчанию 300)
intent: str
go_back Переход к предыдущей веб-странице в истории браузера. intent: стр
navigate Переход непосредственно на указанный URL. url: str
intent: str
go_forward Переход к следующей веб-странице в истории браузера. intent: стр

Мобильная среда (ENVIRONMENT_MOBILE)

Действия, оптимизированные для среды Android:

Название команды Описание Аргументы (в вызове функции)
open_app Открывает приложение по названию. app_name: str
intent: str
click Нажимает левую кнопку мыши в указанных координатах. y: int (0-999)
x: int (0-999)
intent: str
list_apps Показывает список приложений, установленных на устройстве, с их названиями и названиями пакетов. intent: стр
wait Приостанавливает выполнение на указанное количество секунд. seconds: int (необязательный, по умолчанию 1)
intent: str
go_back Возвращает на предыдущий экран или веб-страницу. intent: стр
type; Вводит текст. text: str
press_enter: bool (необязательно, по умолчанию false)
intent: str
drag_and_drop Перетаскивает объект из начальной координаты в конечную. start_y: int (0-999)
start_x: int (0-999)
end_y: int (0-999)
end_x: int (0-999)
intent: str
long_press Выполняет долгое нажатие на экране в указанных координатах. y: int (0–999)
x: int (0–999)
seconds: int (необязательный параметр, по умолчанию 2)
intent: str
press_key Нажимает указанную клавишу и отпускает ее. key: str
intent: str
take_screenshot Возвращает скриншот текущего экрана. intent: стр

Среда рабочего стола (ENVIRONMENT_DESKTOP)

Команды курсора на уровне ОС для сред рабочего стола:

Название команды Описание Аргументы (в вызове функции)
click Нажимает левую кнопку мыши в указанных координатах. y: int (0-999)
x: int (0-999)
intent: str
double_click Дважды нажимает на координаты. y: int (0-999)
x: int (0-999)
intent: str
triple_click Трижды нажмите на координаты. y: int (0-999)
x: int (0-999)
intent: str
middle_click Нажатие на среднюю кнопку мыши в координатах. y: int (0-999)
x: int (0-999)
intent: str
right_click Нажимает правой кнопкой мыши на указанных координатах. y: int (0-999)
x: int (0-999)
intent: str
mouse_down Нажимает и удерживает кнопку мыши в указанных координатах. y: int (0-999)
x: int (0-999)
intent: str
mouse_up Отпускает кнопку мыши в указанных координатах. y: int (0-999)
x: int (0-999)
intent: str
move Перемещает курсор в указанное положение. y: int (0-999)
x: int (0-999)
intent: str
type; Вводит текст. text: str
press_enter: bool (необязательно, по умолчанию false)
intent: str
drag_and_drop Перетаскивает объект из начальной координаты в конечную. start_y: int (0-999)
start_x: int (0-999)
end_y: int (0-999)
end_x: int (0-999)
intent: str
wait Приостанавливает выполнение на указанное количество секунд. seconds: int (необязательный, по умолчанию 1)
intent: str
press_key Нажимает указанную клавишу и отпускает ее. key: str
intent: str
key_down Нажимает и удерживает указанную клавишу. key: str
intent: str
key_up Отпускает указанную клавишу. key: str
intent: str
hotkey Нажимает указанное сочетание клавиш. keys: List[str]
intent: str
take_screenshot Возвращает скриншот текущего экрана. intent: стр
Прокрутка Прокручивает экран вверх, вниз, влево или вправо на заданное количество пикселей. y: int (0–999)
x: int (0–999)
direction: str ("up", "down", "left", "right")
magnitude_in_pixels: int (0–999, необязательный параметр, по умолчанию 300)
intent: str

Пользовательские функции

Вы можете расширить функциональность модели, добавив пользовательские функции. Например, в сценариях с участием человека в контуре (HITL) можно исключить стандартные предопределенные действия и зарегистрировать настраиваемые действия.

Python

Исключите стандартные предопределенные действия браузера (например, click) и зарегистрируйте специальный инструмент yield_to_user:

from google import genai

client = genai.Client()

yield_to_user_tool = {
    "type": "function",
    "name": "yield_to_user",
    "description": "Yields control back to the user for assistance or verification when an automated action is unsafe or ambiguous.",
    "parameters": {
        "type": "object",
        "properties": {
            "reason": {
                "type": "string",
                "description": "The reason why the agent is yielding control to the human."
            }
        },
        "required": ["reason"]
    }
}

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input="Click the submit button. If you need a second factor authentication code, ask me.",
    tools=[
        {
            "type": "computer_use",
            "environment": "mobile",
            "excluded_predefined_functions": ["click"]
        },
        yield_to_user_tool
    ]
)

JavaScript

Исключите стандартные предопределенные действия браузера (например, click) и зарегистрируйте специальный инструмент yield_to_user:

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

const yieldToUserTool = {
    type: "function",
    name: "yield_to_user",
    description: "Yields control back to the user for assistance or verification when an automated action is unsafe or ambiguous.",
    parameters: {
        type: "object",
        properties: {
            reason: {
                type: "string",
                description: "The reason why the agent is yielding control to the human."
            }
        },
        required: ["reason"]
    }
};

const interaction = await ai.interactions.create({
    model: "gemini-3.8-flash",
    input: "Click the submit button. If you need a second factor authentication code, ask me.",
    tools: [
        {
            type: "computer_use",
            environment: "mobile",
            excluded_predefined_functions: ["click"]
        },
        yieldToUserTool
    ]
});

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Function;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;
import java.util.Collections;
import java.util.HashMap;
import java.util.Map;

Client client = new Client();

Map<String, Object> reasonProp = new HashMap<>();
reasonProp.put("type", "string");
reasonProp.put("description", "The reason why the agent is yielding control to the human.");

Map<String, Object> properties = new HashMap<>();
properties.put("reason", reasonProp);

Map<String, Object> parameters = new HashMap<>();
parameters.put("type", "object");
parameters.put("properties", properties);
parameters.put("required", Collections.singletonList("reason"));

Function yieldToUserTool =
    Function.builder()
        .name("yield_to_user")
        .description(
            "Yields control back to the user for assistance or verification when an automated action is unsafe or ambiguous.")
        .parameters(parameters)
        .build();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model("gemini-3.8-flash")
        .input(
            InteractionsInput.of(
                "Click the submit button. If you need a second factor authentication code, ask me."))
        .tools(
            Arrays.asList(
                ComputerUse.builder()
                    .environment(EnvironmentEnum.MOBILE)
                    .excludedPredefinedFunctions(Arrays.asList("click"))
                    .build(),
                yieldToUserTool))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

Проложить маршрут

package main

import (
    "context"
    "log"

    "google.golang.org/genai"
    "google.golang.org/genai/interactions/models/interactions"
    "google.golang.org/genai/interactions/models/operations"
)

func main() {
    ctx := context.Background()
    client, err := genai.NewClient(ctx, nil)
    if err != nil {
        log.Fatal(err)
    }

    yieldToUserTool := interactions.NewTool(interactions.Function{
        Name:        genai.Ptr("yield_to_user"),
        Description: genai.Ptr("Yields control back to the user for assistance or verification when an automated action is unsafe or ambiguous."),
        Parameters: map[string]any{
            "type": "object",
            "properties": map[string]any{
                "reason": map[string]any{
                    "type":        "string",
                    "description": "The reason why the agent is yielding control to the human.",
                },
            },
            "required": []string{"reason"},
        },
    })

    _, err = client.Interactions.Create(ctx, operations.CreateInteractionRequest{
        Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
            Model: interactions.Model("gemini-3.8-flash"),
            Input: interactions.NewInteractionsInput("Click the submit button. If you need a second factor authentication code, ask me."),
            Tools: []interactions.Tool{
                interactions.NewTool(interactions.ComputerUse{
                    Environment:                 interactions.EnvironmentEnumMobile.ToPointer(),
                    ExcludedPredefinedFunctions: []string{"click"},
                }),
                yieldToUserTool,
            },
        }),
    })
    if err != nil {
        log.Fatal(err)
    }
}

Как управлять уровнями рассуждений

Для агентов, использующих компьютер, можно настроить разные уровни мышления, чтобы сбалансировать качество действий и скорость выполнения. Для стандартных задач автоматизации обычно достаточно более низких уровней.

Безопасность

Как настроить правила безопасности

В моделях Gemini 3.x есть встроенные категории сервисов безопасности, которые помогают определить, требуется ли подтверждение пользователя.

Категория правил безопасности Описание
FINANCIAL_TRANSACTIONS Блокирует или активирует подтверждение для действий, связанных с платежами, оформлением покупки или товарами с ограниченным оборотом.
SENSITIVE_DATA_MODIFICATION Защищает медицинские, финансовые и государственные записи от несанкционированного изменения.
COMMUNICATION_TOOL Запрещает агенту самостоятельно отправлять электронные письма, сообщения в чате или черновики.
ACCOUNT_CREATION Запрещает агенту самостоятельно регистрировать новые аккаунты на сайтах.
DATA_MODIFICATION Регулирует общие изменения файловой системы, передачу данных и удаление хранилища.
USER_CONSENT_MANAGEMENT Требуется перехват управления для баннеров с запросом согласия на использование файлов cookie и уведомлений о конфиденциальности.
LEGAL_TERMS_AND_AGREEMENTS Запрещает модели самостоятельно принимать Условия использования или юридически обязывающие договоры.

Переопределение настроек безопасности

Вы можете переопределить некоторые правила, передав переопределения:

Python

from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input="Clean up the local folder by archiving old logs.",
    tools=[
        {
            "type": "computer_use",
            "environment": "desktop",
            "disabled_safety_policies": [
                "data_modification"
            ]
        }
    ]
)

JavaScript

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

const interaction = await ai.interactions.create({
    model: "gemini-3.8-flash",
    input: "Clean up the local folder by archiving old logs.",
    tools: [
        {
            type: "computer_use",
            environment: "desktop",
            disabled_safety_policies: [
                "data_modification"
            ]
        }
    ]
});

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.DisabledSafetyPolicy;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model("gemini-3.8-flash")
        .input(InteractionsInput.of("Clean up the local folder by archiving old logs."))
        .tools(
            Arrays.asList(
                ComputerUse.builder()
                    .environment(EnvironmentEnum.DESKTOP)
                    .disabledSafetyPolicies(
                        Arrays.asList(DisabledSafetyPolicy.DATA_MODIFICATION))
                    .build()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

Проложить маршрут

package main

import (
    "context"
    "log"

    "google.golang.org/genai"
    "google.golang.org/genai/interactions/models/interactions"
    "google.golang.org/genai/interactions/models/operations"
)

func main() {
    ctx := context.Background()
    client, err := genai.NewClient(ctx, nil)
    if err != nil {
        log.Fatal(err)
    }

    _, err = client.Interactions.Create(ctx, operations.CreateInteractionRequest{
        Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
            Model: interactions.Model("gemini-3.8-flash"),
            Input: interactions.NewInteractionsInput("Clean up the local folder by archiving old logs."),
            Tools: []interactions.Tool{
                interactions.NewTool(interactions.ComputerUse{
                    Environment: interactions.EnvironmentEnumDesktop.ToPointer(),
                    DisabledSafetyPolicies: []interactions.DisabledSafetyPolicy{
                        interactions.DisabledSafetyPolicyDataModification,
                    },
                }),
            },
        }),
    })
    if err != nil {
        log.Fatal(err)
    }
}

Обнаружение внедрения запроса

Использование компьютера для Gemini 3.5 Flash-Lite или более поздней версии поддерживает продвинутый механизм безопасности для обнаружения атак с внедрением запросов. Если эта функция включена, она проверяет, содержит ли скриншот скрытые инструкции, которые могут быть использованы злоумышленниками (например, "Игнорировать предыдущие команды"), и блокирует выполнение таких инструкций.

Обнаружение внедрений запросов – это необязательная функция. Значение по умолчанию – false.

Ниже приведены примеры того, как включить обнаружение инъекций запросов в конфигурации инструмента "Использование компьютера".

Python

from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input="Search for flight deals and summarize top results.",
    tools=[
        {
            "type": "computer_use",
            "environment": "desktop",
            "enable_prompt_injection_detection": True,
        }
    ],
)

JavaScript

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

const interaction = await ai.interactions.create({
    model: "gemini-3.8-flash",
    input: "Search for flight deals and summarize top results.",
    tools: [
        {
            type: "computer_use",
            environment: "desktop",
            enablePromptInjectionDetection: true,
        }
    ]
});

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model("gemini-3.8-flash")
        .input(InteractionsInput.of("Search for flight deals and summarize top results."))
        .tools(
            Arrays.asList(
                ComputerUse.builder()
                    .environment(EnvironmentEnum.DESKTOP)
                    .enablePromptInjectionDetection(true)
                    .build()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

Проложить маршрут

package main

import (
    "context"
    "log"

    "google.golang.org/genai"
    "google.golang.org/genai/interactions/models/interactions"
    "google.golang.org/genai/interactions/models/operations"
)

func main() {
    ctx := context.Background()
    client, err := genai.NewClient(ctx, nil)
    if err != nil {
        log.Fatal(err)
    }

    _, err = client.Interactions.Create(ctx, operations.CreateInteractionRequest{
        Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
            Model: interactions.Model("gemini-3.8-flash"),
            Input: interactions.NewInteractionsInput("Search for flight deals and summarize top results."),
            Tools: []interactions.Tool{
                interactions.NewTool(interactions.ComputerUse{
                    Environment:                    interactions.EnvironmentEnumDesktop.ToPointer(),
                    EnablePromptInjectionDetection: genai.Ptr(true),
                }),
            },
        }),
    })
    if err != nil {
        log.Fatal(err)
    }
}

cURL

curl "https://generativelanguage.googleapis.com/v1beta/interactions?key=${GEMINI_API_KEY}" \
-H 'Content-Type: application/json' \
-d '{
  "model": "gemini-3.8-flash",
  "input": "Search for flight deals and summarize top results.",
  "tools": [
    {
      "type": "computer_use",
      "environment": "desktop",
      "enable_prompt_injection_detection": true
    }
  ]
}'

Подтверждение решения о безопасности

Ответ может содержать параметр safety_decision в аргументах вызова функции:

{
  "steps": [
    {
      "type": "function_call",
      "name": "click",
      "arguments": {
        "x": 60,
        "y": 100,
        "safety_decision": {
          "explanation": "Must check check-box",
          "decision": "require_confirmation"
        }
      }
    }
  ]
}

Если значение safety_decision – require_confirmation, покажите запрос конечному пользователю. Если пользователь подтверждает действие, установите значение safety_acknowledgement в function_result.

Python

def get_safety_confirmation(safety_decision):
    # Prompt user for confirmation
    print(f"Safety confirmation required: {safety_decision.get('explanation', '')}")
    return "CONTINUE" # Or TERMINATE

# Inside execute_function_calls, check for safety_decision:
if 'safety_decision' in function_call.arguments:
    decision = get_safety_confirmation(function_call.arguments['safety_decision'])
    if decision == "TERMINATE":
        break
    # Include safety_acknowledgement inside the action result
    action_result["safety_acknowledgement"] = True

Рекомендации по обеспечению безопасности

Использование компьютера представляет уникальные риски для безопасности и работы, поскольку модель, действующая от имени пользователя, может столкнуться с ненадежным контентом на экранах или допустить ошибки при выполнении действий. Чтобы защитить данные пользователей и системы, следуйте приведенным ниже рекомендациям.

  1. Человек в контуре (HITL):
    • Требовать подтверждения от пользователя. Если ответ системы безопасности содержит значение require_confirmation, запросите у пользователя подтверждение.
    • Задайте собственные инструкции по безопасности. Реализуйте собственную системную инструкцию, чтобы определить и обеспечить соблюдение собственных границ безопасности. Пример:

      Python

      from google import genai
      
      client = genai.Client()
      
      system_instruction = """
      ## **RULE 1: Seek User Confirmation (USER_CONFIRMATION)**
      
      This is your first and most important check. If the next required action falls
      into any of the following categories, you MUST stop immediately, and seek the
      user's explicit permission.
      
      **Procedure for Seeking Confirmation:**
      * **For Consequential Actions:** Perform all preparatory steps (e.g., navigating,
        filling out forms, typing a message). You will ask for confirmation **AFTER**
        all necessary information is entered on the screen, but **BEFORE** you perform
        the final, irreversible action (e.g., before clicking "Send", "Submit",
        "Confirm Purchase", "Share").
      * **For Prohibited Actions:** If the action is strictly forbidden (e.g., accepting
        legal terms, solving a CAPTCHA), you must first inform the user about the
        required action and ask for their confirmation to proceed.
      
      **USER_CONFIRMATION Categories:**
      
      *   **Consent and Agreements:** You are FORBIDDEN from accepting, selecting, or
          agreeing to any of the following on the user's behalf. You must ask the
          user to confirm before performing these actions.
          *   Terms of Service
          *   Privacy Policies
          *   Cookie consent banners
          *   End User License Agreements (EULAs)
          *   Any other legally significant contracts or agreements.
      *   **Robot Detection:** You MUST NEVER attempt to solve or bypass the
          following. You must ask the user to confirm before performing these actions.
          *   CAPTCHAs (of any kind)
          *   Any other anti-robot or human-verification mechanisms, even if you are
              capable.
      *   **Financial Transactions:**
          *   Completing any purchase.
          *   Managing or moving money (e.g., transfers, payments).
          *   Purchasing regulated goods or participating in gambling.
      *   **Sending Communications:**
          *   Sending emails.
          *   Sending messages on any platform (e.g., social media, chat apps).
          *   Posting content on social media or forums.
      *   **Accessing or Modifying Sensitive Information:**
          *   Health, financial, or government records (e.g., medical history, tax
              forms, passport status).
          *   Revealing or modifying sensitive personal identifiers (e.g., SSN, bank
              account number, credit card number).
      *   **User Data Management:**
          *   Accessing, downloading, or saving files from the web.
          *   Sharing or sending files/data to any third party.
          *   Transferring user data between systems.
      *   **Browser Data Usage:**
          *   Accessing or managing Chrome browsing history, bookmarks, autofill data,
              or saved passwords.
      *   **Security and Identity:**
          *   Logging into any user account.
          *   Any action that involves misrepresentation or impersonation (e.g.,
              creating a fan account, posting as someone else).
      *   **Insurmountable Obstacles:** If you are technically unable to interact with
          a user interface element or are stuck in a loop you cannot resolve, ask the
          user to take over.
      ---
      
      ## **RULE 2: Default Behavior (ACTUATE)**
      
      If an action does **NOT** fall under the conditions for `USER_CONFIRMATION`,
      your default behavior is to **Actuate**.
      
      **Actuation Means:**  You MUST proactively perform all necessary steps to move
      the user's request forward. Continue to actuate until you either complete the
      non-consequential task or encounter a condition defined in Rule 1.
      
      *   **Example 1:** If asked to send money, you will navigate to the payment
          portal, enter the recipient's details, and enter the amount. You will then
          **STOP** as per Rule 1 and ask for confirmation before clicking the final
          "Send" button.
      *   **Example 2:** If asked to post a message, you will navigate to the site,
          open the post composition window, and write the full message. You will then
          **STOP** as per Rule 1 and ask for confirmation before clicking the final
          "Post" button.
      
          After the user has confirmed, remember to get the user's latest screen
          before continuing to perform actions.
      
      # Final Response Guidelines:
      Write final response to the user in the following cases:
      - User confirmation
      - When the task is complete or you have enough information to respond to the user
      """
      
      interaction = client.interactions.create(
          model="gemini-3.8-flash",
          system_instruction=system_instruction,
          input="Prepare a draft but do not send.",
          tools=[{
              "type": "computer_use",
              "environment": "browser"
          }]
      )
      

      JavaScript

      import { GoogleGenAI } from '@google/genai';
      
      const ai = new GoogleGenAI();
      
      const systemInstruction = `
      ## **RULE 1: Seek User Confirmation (USER_CONFIRMATION)**
      
      This is your first and most important check. If the next required action falls
      into any of the following categories, you MUST stop immediately, and seek the
      user's explicit permission.
      
      **Procedure for Seeking Confirmation:**
      * **For Consequential Actions:** Perform all preparatory steps (e.g., navigating,
        filling out forms, typing a message). You will ask for confirmation **AFTER**
        all necessary information is entered on the screen, but **BEFORE** you perform
        the final, irreversible action (e.g., before clicking "Send", "Submit",
        "Confirm Purchase", "Share").
      * **For Prohibited Actions:** If the action is strictly forbidden (e.g., accepting
        legal terms, solving a CAPTCHA), you must first inform the user about the
        required action and ask for their confirmation to proceed.
      
      **USER_CONFIRMATION Categories:**
      
      *   **Consent and Agreements:** You are FORBIDDEN from accepting, selecting, or
          agreeing to any of the following on the user's behalf. You must ask the
          user to confirm before performing these actions.
          *   Terms of Service
          *   Privacy Policies
          *   Cookie consent banners
          *   End User License Agreements (EULAs)
          *   Any other legally significant contracts or agreements.
      *   **Robot Detection:** You MUST NEVER attempt to solve or bypass the
          following. You must ask the user to confirm before performing these actions.
          *   CAPTCHAs (of any kind)
          *   Any other anti-robot or human-verification mechanisms, even if you are
              capable.
      *   **Financial Transactions:**
          *   Completing any purchase.
          *   Managing or moving money (e.g., transfers, payments).
          *   Purchasing regulated goods or participating in gambling.
      *   **Sending Communications:**
          *   Sending emails.
          *   Sending messages on any platform (e.g., social media, chat apps).
          *   Posting content on social media or forums.
      *   **Accessing or Modifying Sensitive Information:**
          *   Health, financial, or government records (e.g., medical history, tax
              forms, passport status).
          *   Revealing or modifying sensitive personal identifiers (e.g., SSN, bank
              account number, credit card number).
      *   **User Data Management:**
          *   Accessing, downloading, or saving files from the web.
          *   Sharing or sending files/data to any third party.
          *   Transferring user data between systems.
      *   **Browser Data Usage:**
          *   Accessing or managing Chrome browsing history, bookmarks, autofill data,
              or saved passwords.
      *   **Security and Identity:**
          *   Logging into any user account.
          *   Any action that involves misrepresentation or impersonation (e.g.,
              creating a fan account, posting as someone else).
      *   **Insurmountable Obstacles:** If you are technically unable to interact with
          a user interface element or are stuck in a loop you cannot resolve, ask the
          user to take over.
      ---
      
      ## **RULE 2: Default Behavior (ACTUATE)**
      
      If an action does **NOT** fall under the conditions for \`USER_CONFIRMATION\`,
      your default behavior is to **Actuate**.
      
      **Actuation Means:**  You MUST proactively perform all necessary steps to move
      the user's request forward. Continue to actuate until you either complete the
      non-consequential task or encounter a condition defined in Rule 1.
      
      *   **Example 1:** If asked to send money, you will navigate to the payment
          portal, enter the recipient's details, and enter the amount. You will then
          **STOP** as per Rule 1 and ask for confirmation before clicking the final
          "Send" button.
      *   **Example 2:** If asked to post a message, you will navigate to the site,
          open the post composition window, and write the full message. You will then
          **STOP** as per Rule 1 and ask for confirmation before clicking the final
          "Post" button.
      
          After the user has confirmed, remember to get the user's latest screen
          before continuing to perform actions.
      
      # Final Response Guidelines:
      Write final response to the user in the following cases:
      - User confirmation
      - When the task is complete or you have enough information to respond to the user
      `;
      
      const interaction = await ai.interactions.create({
          model: "gemini-3.8-flash",
          system_instruction: systemInstruction,
          input: "Prepare a draft but do not send.",
          tools: [{
              type: "computer_use",
              environment: "browser"
          }]
      });
      

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.EnvironmentEnum;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

String systemInstruction =
    "## **RULE 1: Seek User Confirmation (USER_CONFIRMATION)**\n\n"
        + "This is your first and most important check. If the next required action falls "
        + "into any of the following categories, you MUST stop immediately, and seek the "
        + "user's explicit permission.\n\n"
        + "## **RULE 2: Default Behavior (ACTUATE)**\n\n"
        + "If an action does **NOT** fall under the conditions for `USER_CONFIRMATION`, "
        + "your default behavior is to **Actuate**.";

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model("gemini-3.8-flash")
        .systemInstruction(systemInstruction)
        .input(InteractionsInput.of("Prepare a draft but do not send."))
        .tools(
            Arrays.asList(
                ComputerUse.builder().environment(EnvironmentEnum.BROWSER).build()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

Проложить маршрут

package main

import (
    "context"
    "log"

    "google.golang.org/genai"
    "google.golang.org/genai/interactions/models/interactions"
    "google.golang.org/genai/interactions/models/operations"
)

func main() {
    ctx := context.Background()
    client, err := genai.NewClient(ctx, nil)
    if err != nil {
        log.Fatal(err)
    }

    systemInstruction := "## **RULE 1: Seek User Confirmation (USER_CONFIRMATION)**\n\n" +
        "This is your first and most important check. If the next required action falls " +
        "into any of the following categories, you MUST stop immediately, and seek the " +
        "user's explicit permission.\n\n" +
        "## **RULE 2: Default Behavior (ACTUATE)**\n\n" +
        "If an action does **NOT** fall under the conditions for `USER_CONFIRMATION`, " +
        "your default behavior is to **Actuate**."

    _, err = client.Interactions.Create(ctx, operations.CreateInteractionRequest{
        Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
            Model:             interactions.Model("gemini-3.8-flash"),
            SystemInstruction: genai.Ptr(systemInstruction),
            Input:             interactions.NewInteractionsInput("Prepare a draft but do not send."),
            Tools: []interactions.Tool{
                interactions.NewTool(interactions.ComputerUse{
                    Environment: interactions.EnvironmentEnumBrowser.ToPointer(),
                }),
            },
        }),
    })
    if err != nil {
        log.Fatal(err)
    }
}
  1. Безопасная среда выполнения. Запускайте агента в безопасной изолированной среде, чтобы ограничить его потенциальное влияние. Это может быть виртуальная машина в изолированной среде, контейнер (например, Docker) или отдельный профиль браузера с ограниченными разрешениями. Инструкции по настройке песочницы с помощью Docker можно найти в справочной реализации GitHub.
  2. Очистка входных данных. Очищайте весь пользовательский текст в запросах, чтобы снизить риск непреднамеренных инструкций или внедрения запросов. Это полезный уровень защиты, но он не заменяет безопасную среду выполнения.
  3. Ограничения на контент. Используйте ограничения и Content Safety API, чтобы оценивать ввод данных пользователем, входные и выходные данные Способов ввода текста и ответы агента на предмет соответствия, внедрения запросов и обнаружения джейлбрейка.
  4. Белые и черные списки. Используйте механизмы фильтрации, чтобы контролировать, куда может переходить модель и что она может делать. Черный список запрещенных сайтов – это хорошая отправная точка, а более строгий белый список обеспечивает ещё большую безопасность.
  5. Наблюдаемость и ведение журналов. Ведите подробные журналы для отладки, аудита и реагирования на инциденты. Ваш клиент должен регистрировать запросы, скриншоты, предложенные моделью действия (function_call), ответы, связанные с безопасностью, и все действия, выполненные клиентом.
  6. Управление средой. Убедитесь, что среда графического интерфейса пользователя согласована. Неожиданные всплывающие окна, уведомления или изменения в макете могут запутать модель. По возможности начинайте каждую новую задачу с известного чистого состояния.

Версии модели

Функция "Использование компьютера" доступна для следующих моделей:

  • Gemini 3.8 Flash (gemini-3.8-flash) – рекомендованная модель для использования на компьютере, обеспечивающая высокую точность взаимодействия с интерфейсом и надежный вызов инструментов.
  • Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) – модель с низкой задержкой и экономичным использованием ресурсов, поддерживающая работу на компьютере.
  • Gemini 3 Flash Preview (gemini-3-flash-preview) – предварительная версия модели, поддерживающая использование компьютера.

Дальнейшие действия