Computernutzung

Mit dem Tool „Computernutzung“ können Sie Browser-, Mobil- und Desktop-Steuerungsagenten erstellen, die mit Aufgaben interagieren und diese automatisieren. Anhand von Screenshots kann das Modell einen Computerbildschirm "sehen" und "agieren", indem es bestimmte UI-Aktionen wie Mausklicks und Tastatureingaben generiert. Ähnlich wie bei Funktionsaufrufen müssen Sie die clientseitige Ausführungsumgebung implementieren, um die Computernutzungsaktionen zu empfangen und auszuführen.

Eine Liste der unterstützten Modelle finden Sie unter Modellversionen. Die Gemini 3.x-Modelle unterstützen mehrere erweiterte Funktionen:

  • Unterstützung mehrerer Umgebungen: Erstellung von Agenten für Browser-, Mobil- und Desktop-Umgebungen.
  • Optimierte Aktionen mit Absichten: Aktionen enthalten ein intent Feld, das die Begründung des Modells für jeden Schritt erläutert.
  • Konfigurierbare Sicherheitsrichtlinien: Feinabstimmung des Sicherheitsverhaltens mit integrierten Richtlinienkategorien und Überschreibungen.
  • Erkennung von Eingabeaufforderungen: optionale Screenshot-Überprüfung zur Erkennung versteckter gegnerischer Anweisungen.

Mit Computer Use können Sie Agenten erstellen, die Folgendes leisten:

  • Wiederholte Dateneingaben oder das Ausfüllen von Formularen auf Websites automatisieren
  • Führen Sie automatisierte Tests von Webanwendungen und Benutzerabläufen durch.
  • Führen Sie Recherchen auf verschiedenen Websites durch (z. B. sammeln Sie Produktinformationen, Preise und Bewertungen von E-Commerce-Websites, um eine Kaufentscheidung zu treffen).

Hier ist ein Minimalbeispiel für die Initialisierung des Clients und das Senden einer Eingabeaufforderung an das Modell mit aktiviertem computer_use-Tool für eine Browserumgebung:

Python

from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input="Search for 'Gemini API' on Google.",
    tools=[{"type": "computer_use", "environment": "browser"}]
)

print(interaction)

JavaScript

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

const interaction = await ai.interactions.create({
  model: 'gemini-3.8-flash',
  input: "Search for 'Gemini API' on Google.",
  tools: [{ type: "computer_use", environment: "browser" }]
});

console.log(interaction);

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model(Model.of("gemini-3.8-flash"))
        .input(InteractionsInput.of("Click the Submit button on the screen."))
        .tools(Arrays.asList(new ComputerUse()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction.outputText().orElse(""));


Wie die Computernutzung funktioniert

Um einen Agenten mit dem Computernutzungsmodell zu erstellen, müssen Sie eine kontinuierliche Schleife zwischen Ihrer Anwendung und der API einrichten. Hier ist, was Ihr Code in jedem Schritt ausführt:

  1. Sende eine Anfrage an das Modell
    • Ihre Anwendung sendet eine API-Anfrage, die das Tool „Computernutzung“, Ihre Konfigurationseinstellungen (wie die Zielumgebung), die Benutzereingabeaufforderung und einen Screenshot des aktuellen Bildschirms enthält.
  2. Empfange die Modellantwort
    • Das Modell analysiert den Bildschirm und die Eingabeaufforderung und gibt eine Antwort zurück, die ein vorgeschlagenes function_call enthält, das eine UI-Aktion darstellt (z. B. einen Klick, Scrollen oder Tastendruck).
    • Bei Gemini 3.x Modellen enthält die Antwort auch eine Begründung intent, die erklärt, warum das Modell diese Aktion gewählt hat.
    • Die Antwort kann auch ein safety_decision von einem internen Sicherheitssystem enthalten, das die Aktion als regulär/zulässig, require_confirmation (Benutzergenehmigung erforderlich) oder blockiert einstuft.
  3. Führe die empfangene Aktion aus
    • Wenn die Aktion erlaubt ist (oder der Benutzer sie bestätigt), analysiert Ihr clientseitiger Code function_call, skaliert die normalisierten Koordinaten, um sie an Ihren Viewport anzupassen, und führt die Aktion in Ihrer Zielumgebung mithilfe von Automatisierungstools (wie z. B. Playwright) aus. Wenn die Aktion blockiert wird, sollte Ihr Client die Ausführung stoppen oder die Unterbrechung behandeln.
  4. Erfasse den neuen Umgebungszustand
    • Nach Abschluss der Aktion erstellt Ihre Anwendung einen neuen Screenshot und sendet diesen in einem function_result an das Modell zurück, um den nächsten Schritt anzufordern.

Dieser Prozess wiederholt sich dann ab Schritt 2, wobei das Modell so lange zur nächsten Aktion aufgefordert wird, bis die Aufgabe abgeschlossen oder beendet ist.

Übersicht zur Computernutzung

Wie man die Computernutzung umsetzt

Bevor Sie mit dem Computernutzungstool arbeiten, müssen Sie Folgendes einrichten:

  • Sichere Ausführungsumgebung: Führen Sie Ihren Agenten in einer Sandbox-VM oder einem Container aus, um ihn von Ihrem Hostsystem zu isolieren und seine potenziellen Auswirkungen zu begrenzen. Die Referenzimplementierung beinhaltet eine sofort einsatzbereite Docker-basierte Sandbox, die Sie als Ausgangspunkt verwenden können.
  • Clientseitiger Aktionshandler: Implementieren Sie clientseitige Logik, um Koordinaten auszuführen, Text einzugeben und Screenshots zu erstellen.

Die folgenden Beispiele verwenden einen Webbrowser als Ausführungsumgebung und Playwright als clientseitigen Handler.

0. Dramatiker einstellen

Installieren Sie zunächst die benötigten Pakete:

pip install google-genai playwright
playwright install chromium

Initialisieren Sie anschließend eine Playwright-Browserinstanz, die für die Ausführung verwendet werden soll:

from playwright.sync_api import sync_playwright

# 1. Configure screen dimensions for the target environment
SCREEN_WIDTH = 1440
SCREEN_HEIGHT = 900

# 2. Start the Playwright browser
# In production, utilize a sandboxed environment.
playwright = sync_playwright().start()
# Set headless=False to see the actions performed on your screen
browser = playwright.chromium.launch(headless=False)

# 3. Create a context and page with the specified dimensions
context = browser.new_context(
    viewport={"width": SCREEN_WIDTH, "height": SCREEN_HEIGHT}
)
page = context.new_page()

# 4. Navigate to an initial page to start the task
page.goto("https://www.google.com")

# The 'page', 'SCREEN_WIDTH', and 'SCREEN_HEIGHT' variables
# will be used in the steps below.

1. Sende eine Anfrage an das Modell

Initialisieren Sie die Clientbibliothek und konfigurieren Sie das Tool zur Computernutzung. Beachten Sie, dass Sie die Anzeigegröße bei einer Anfrage nicht angeben müssen. Das Modell sagt Pixelkoordinaten voraus, die auf die Höhe und Breite des Bildschirms skaliert werden.

Gemini 3.x

Python

Verwenden Sie das google-genai Python SDK (Version 2.7.0 oder höher), um eine Anfrage für die Browserumgebung zu konfigurieren:

from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model='gemini-3.8-flash',
    input="Find a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th",
    tools=[
        {
            "type": "computer_use",
            "environment": "browser",
            "enable_prompt_injection_detection": True
        }
    ]
)

print(interaction)

JavaScript

Verwenden Sie das @google/genai Node.js SDK, um eine Anfrage zu konfigurieren, die auf die Browserumgebung ausgerichtet ist:

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

const interaction = await ai.interactions.create({
  model: 'gemini-3.8-flash',
  input: "Find a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th",
  tools: [
    {
      type: "computer_use",
      environment: "browser",
      enable_prompt_injection_detection: true
    }
  ]
});

console.log(interaction);

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model(Model.of("gemini-3.8-flash"))
        .input(InteractionsInput.of("Click the Submit button on the screen."))
        .tools(Arrays.asList(new ComputerUse()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction.outputText().orElse(""));

REST

So senden Sie eine Anfrage mit curl:

curl -X POST \
  "https://generativelanguage.googleapis.com/v1beta/interactions" \
  -H "x-goog-api-key: $GEMINI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.8-flash",
    "input": "Find me a flight from SF to Hawaii on Jun 30th, coming back on Jul 6th. Start by navigating directly to flights.google.com",
    "tools": [
      {
        "type": "computer_use",
        "environment": "browser",
        "enable_prompt_injection_detection": true
      }
    ]
  }'

Gemini 2.5 (Legacy)

Python

from google import genai

client = genai.Client()

# Specify predefined functions to exclude (optional)
excluded_functions = ["drag_and_drop"]

interaction = client.interactions.create(
    model='gemini-2.5-computer-use-preview-10-2025',
    input="Search for highly rated smart fridges on Google Shopping.",
    tools=[
        {
            "type": "computer_use",
            "environment": "browser",
            "excluded_predefined_functions": excluded_functions
        }
    ]
)

print(interaction)

JavaScript

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

// Specify predefined functions to exclude (optional)
const excludedFunctions = ["drag_and_drop"];

const interaction = await ai.interactions.create({
  model: 'gemini-2.5-computer-use-preview-10-2025',
  input: "Search for highly rated smart fridges on Google Shopping.",
  tools: [
    {
      type: "computer_use",
      environment: "browser",
      excluded_predefined_functions: excludedFunctions
    }
  ]
});

console.log(interaction);

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model(Model.of("gemini-3.8-flash"))
        .input(InteractionsInput.of("Click the Submit button on the screen."))
        .tools(Arrays.asList(new ComputerUse()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction.outputText().orElse(""));

2. Antwort des Modells erhalten

Das Antwortmodell schlägt einen Funktionsaufruf vor. Bei Gemini 3.x-Modellen enthält die Antwort neben Koordinaten auch eine maßgeschneiderte Absicht für das logische Denken. Im Folgenden finden Sie Beispiele für beide Antworten:

Gemini 3.x

{
  "steps": [
    {
      "type": "function_call",
      "name": "click",
      "arguments": {
        "x": 450,
        "y": 120,
        "intent": "Click the search box to type the destination."
      }
    }
  ]
}

Gemini 2.5 (Legacy)

{
  "steps": [
    {
      "type": "model_output",
      "content": [
        {
          "type": "text",
          "text": "I will type the search query into the search bar."
        }
      ]
    },
    {
      "type": "function_call",
      "name": "type_text_at",
      "arguments": {
        "x": 371,
        "y": 470,
        "text": "highly rated smart fridges",
        "press_enter": true
      }
    }
  ]
}

3. Erhaltene Aktionen ausführen

Ihre Anwendung muss die Antwortkoordinaten analysieren, die Aktion ausführen und sie von den normalisierten 1000x1000-Koordinaten skalieren.

Der folgende Code verarbeitet sowohl Legacy-Tool-Befehle (click_at, type_text_at) als auch moderne optimierte Befehle (click, type).

Python

from typing import Any, List, Tuple
import time

def denormalize_x(x: int, screen_width: int) -> int:
    """Convert normalized x coordinate (0-1000) to actual pixel coordinate."""
    return int(x / 1000 * screen_width)

def denormalize_y(y: int, screen_height: int) -> int:
    """Convert normalized y coordinate (0-1000) to actual pixel coordinate."""
    return int(y / 1000 * screen_height)

def execute_function_calls(interaction, page, screen_width, screen_height):
    results = []
    function_calls = [
        step for step in interaction.steps if step.type == "function_call"
    ]

    for function_call in function_calls:
        action_result = {}
        fname = function_call.name
        args = function_call.arguments
        print(f"  -> Executing: {fname} (Intent: {args.get('intent', 'N/A')})")

        try:
            if fname in ("open_web_browser", "open_app"):
                pass # Handled / already open
            elif fname in ("click", "click_at", "double_click", "triple_click", "middle_click", "right_click", "move", "long_press"):
                actual_x = denormalize_x(args["x"], screen_width)
                actual_y = denormalize_y(args["y"], screen_height)

                if fname in ("click", "click_at"):
                    page.mouse.click(actual_x, actual_y)
                elif fname == "double_click":
                    page.mouse.dblclick(actual_x, actual_y)
                elif fname == "right_click":
                    page.mouse.click(actual_x, actual_y, button="right")
                elif fname == "middle_click":
                    page.mouse.click(actual_x, actual_y, button="middle")
                elif fname == "move":
                    page.mouse.move(actual_x, actual_y)
            elif fname in ("type", "type_text_at"):
                actual_x = denormalize_x(args["x"], screen_width) if "x" in args else None
                actual_y = denormalize_y(args["y"], screen_height) if "y" in args else None
                text = args["text"]
                press_enter = args.get("press_enter", False)

                if actual_x is not None and actual_y is not None:
                    page.mouse.click(actual_x, actual_y)
                # Clear field first
                page.keyboard.press("Meta+A")
                page.keyboard.press("Backspace")
                page.keyboard.type(text)
                if press_enter:
                    page.keyboard.press("Enter")
            elif fname == "navigate":
                page.goto(args["url"])
            elif fname == "go_back":
                page.go_back()
            elif fname == "go_forward":
                page.go_forward()
            elif fname == "wait":
                time.sleep(args.get("seconds", 1))
            else:
                print(f"Warning: Custom or unhandled function {fname}")

            page.wait_for_load_state(timeout=5000)
            time.sleep(1)

        except Exception as e:
            print(f"Error executing {fname}: {e}")
            action_result = {"error": str(e)}

        results.append((fname, function_call.id, action_result))

    return results

JavaScript

function denormalizeX(x, screenWidth) {
    // Convert normalized x coordinate (0-1000) to actual pixel coordinate.
    return Math.floor((x / 1000) * screenWidth);
}

function denormalizeY(y, screenHeight) {
    // Convert normalized y coordinate (0-1000) to actual pixel coordinate.
    return Math.floor((y / 1000) * screenHeight);
}

async function executeFunctionCalls(interaction, page, screenWidth, screenHeight) {
    const results = [];
    const functionCalls = interaction.steps.filter(step => step.type === "function_call");

    for (const functionCall of functionCalls) {
        const actionResult = {};
        const fname = functionCall.name;
        const args = functionCall.arguments;
        console.log(`  -> Executing: ${fname} (Intent: ${args.intent || 'N/A'})`);

        try {
            if (fname === "open_web_browser" || fname === "open_app") {
                // Handled / already open
            } else if (["click", "click_at", "double_click", "triple_click", "middle_click", "right_click", "move", "long_press"].includes(fname)) {
                const actualX = denormalizeX(args.x, screenWidth);
                const actualY = denormalizeY(args.y, screenHeight);

                if (fname === "click" || fname === "click_at") {
                    await page.mouse.click(actualX, actualY);
                } else if (fname === "double_click") {
                    await page.mouse.dblclick(actualX, actualY);
                } else if (fname === "right_click") {
                    await page.mouse.click(actualX, actualY, { button: "right" });
                } else if (fname === "middle_click") {
                    await page.mouse.click(actualX, actualY, { button: "middle" });
                } else if (fname === "move") {
                    await page.mouse.move(actualX, actualY);
                }
            } else if (fname === "type" || fname === "type_text_at") {
                const actualX = args.x !== undefined ? denormalizeX(args.x, screenWidth) : null;
                const actualY = args.y !== undefined ? denormalizeY(args.y, screenHeight) : null;
                const text = args.text;
                const pressEnter = args.press_enter || false;

                if (actualX !== null && actualY !== null) {
                    await page.mouse.click(actualX, actualY);
                }
                // Clear field first
                await page.keyboard.press("Meta+A");
                await page.keyboard.press("Backspace");
                await page.keyboard.type(text);
                if (pressEnter) {
                    await page.keyboard.press("Enter");
                }
            } else if (fname === "navigate") {
                await page.goto(args.url);
            } else if (fname === "go_back") {
                await page.goBack();
            } else if (fname === "go_forward") {
                await page.goForward();
            } else if (fname === "wait") {
                await new Promise(resolve => setTimeout(resolve, (args.seconds || 1) * 1000));
            } else {
                console.log(`Warning: Custom or unhandled function ${fname}`);
            }

            await page.waitForLoadState('load', { timeout: 5000 }).catch(() => {});
            await new Promise(resolve => setTimeout(resolve, 1000));
        } catch (e) {
            console.log(`Error executing ${fname}: ${e}`);
            actionResult.error = e.message;
        }

        results.push([fname, functionCall.id, actionResult]);
    }

    return results;
}

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model(Model.of("gemini-3.8-flash"))
        .input(InteractionsInput.of("Click the Submit button on the screen."))
        .tools(Arrays.asList(new ComputerUse()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction.outputText().orElse(""));

4. Neuen Umgebungsstatus erfassen

Senden Sie nach der Ausführung der Aktionen das Ergebnis der Funktionsausführung zurück an das Modell, damit es diese Informationen zum Generieren der nächsten Aktion verwenden kann. Wenn mehrere Aktionen (parallele Aufrufe) ausgeführt wurden, müssen Sie im nächsten Nutzerzug für jede Aktion ein function_result senden.

Python

import json
import base64

def get_function_responses(page, results):
    screenshot_bytes = page.screenshot(type="png")
    current_url = page.url
    function_responses = []
    for name, call_id, result in results:
        function_responses.append({
            "type": "function_result",
            "name": name,
            "call_id": call_id,
            "result": [
                {
                    "type": "text",
                    "text": json.dumps({"url": current_url, **result})
                },
                {
                    "type": "image",
                    "data": base64.b64encode(screenshot_bytes).decode("utf-8"),
                    "mime_type": "image/png"
                }
            ]
        })
    return function_responses

JavaScript

async function getFunctionResponses(page, results) {
    const screenshotBuffer = await page.screenshot({ type: 'png' });
    const screenshotBase64 = screenshotBuffer.toString('base64');
    const currentUrl = page.url();
    const functionResponses = [];

    for (const [name, callId, result] of results) {
        functionResponses.push({
            type: "function_result",
            name: name,
            call_id: callId,
            result: [
                {
                    type: "text",
                    text: JSON.stringify({ url: currentUrl, ...result })
                },
                {
                    type: "image",
                    data: screenshotBase64,
                    mime_type: "image/png"
                }
            ]
        });
    }
    return functionResponses;
}

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model(Model.of("gemini-3.8-flash"))
        .input(InteractionsInput.of("Click the Submit button on the screen."))
        .tools(Arrays.asList(new ComputerUse()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction.outputText().orElse(""));

Nachdem Sie festgelegt haben, wie der Umgebungsstatus erfasst und formatiert werden soll, können Sie alle diese Schritte in einem kontinuierlichen Ausführungszyklus kombinieren.

Agent-Schleife erstellen

Um Interaktionen mit mehreren Schritten zu ermöglichen, kombinieren Sie die vier Schritte aus dem Abschnitt Computer Use implementieren in einem einzigen Loop. In dieser Schleife werden so lange Aktionen angefordert und die Ergebnisse an das Modell zurückgegeben, bis die Aufgabe abgeschlossen ist.

Denken Sie daran, den Unterhaltungsverlauf richtig zu verwalten, indem Sie die Modellantworten und Ihre Funktionsantworten in jedem Schritt an den Verlauf anhängen.

Python

import time
from typing import Any, List, Tuple
from playwright.sync_api import sync_playwright

from google import genai

client = genai.Client()

# Constants for screen dimensions
SCREEN_WIDTH = 1440
SCREEN_HEIGHT = 900

# Setup Playwright
print("Initializing browser...")
playwright = sync_playwright().start()
browser = playwright.chromium.launch(headless=False)
context = browser.new_context(viewport={"width": SCREEN_WIDTH, "height": SCREEN_HEIGHT})
page = context.new_page()

# Define helper functions. Copy/paste from steps 3 and 4
# def denormalize_x(...)
# def denormalize_y(...)
# def execute_function_calls(...)
# def get_function_responses(...)

try:
    # Go to initial page
    page.goto("https://ai.google.dev/gemini-api/docs")

    # Take initial screenshot
    initial_screenshot = page.screenshot(type="png")
    USER_PROMPT = "Go to ai.google.dev/gemini-api/docs and search for pricing."
    print(f"Goal: {USER_PROMPT}")

    # First interaction
    interaction = client.interactions.create(
        model='gemini-3.8-flash',
        input=[
            {"type": "text", "text": USER_PROMPT},
            {"type": "image", "data": base64.b64encode(initial_screenshot).decode("utf-8"), "mime_type": "image/png"}
        ],
        tools=[{
            "type": "computer_use",
            "environment": "browser",
            "enable_prompt_injection_detection": True
        }]
    )

    # Agent Loop
    turn_limit = 5
    for i in range(turn_limit):
        print(f"\n--- Turn {i+1} ---")

        has_function_calls = any(
            step.type == "function_call"
            for step in interaction.steps
        )
        if not has_function_calls:
            text_response = " ".join([
                content_block.text for step in interaction.steps if step.type == "model_output"
                for content_block in step.content if content_block.type == "text"
            ])
            print("Agent finished:", text_response)
            break

        print("Executing actions...")
        results = execute_function_calls(interaction, page, SCREEN_WIDTH, SCREEN_HEIGHT)

        print("Capturing state...")
        function_responses = get_function_responses(page, results)

        # Continue conversation with function responses
        interaction = client.interactions.create(
            model='gemini-3.8-flash',
            previous_interaction_id=interaction.id,
            input=function_responses,
            tools=[{
                "type": "computer_use",
                "environment": "browser",
                "enable_prompt_injection_detection": True
            }]
        )

finally:
    # Cleanup
    print("\nClosing browser...")
    browser.close()
    playwright.stop()

JavaScript

import { chromium } from 'playwright';
import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

// Constants for screen dimensions
const SCREEN_WIDTH = 1440;
const SCREEN_HEIGHT = 900;

console.log("Initializing browser...");
const browser = await chromium.launch({ headless: false });
const context = await browser.newContext({
    viewport: { width: SCREEN_WIDTH, height: SCREEN_HEIGHT }
});
const page = await context.newPage();

// Define helper functions. Copy/paste from steps 3 and 4:
// function denormalizeX(...)
// function denormalizeY(...)
// async function executeFunctionCalls(...)
// async function getFunctionResponses(...)

try {
    // Go to initial page
    await page.goto("https://ai.google.dev/gemini-api/docs");

    // Take initial screenshot
    const initialScreenshotBuffer = await page.screenshot({ type: 'png' });
    const initialScreenshotBase64 = initialScreenshotBuffer.toString('base64');
    const USER_PROMPT = "Go to ai.google.dev/gemini-api/docs and search for pricing.";
    console.log(`Goal: ${USER_PROMPT}`);

    // First interaction
    let interaction = await ai.interactions.create({
        model: 'gemini-3.8-flash',
        input: [
            { type: 'text', text: USER_PROMPT },
            { type: 'image', data: initialScreenshotBase64, mime_type: 'image/png' }
        ],
        tools: [{
            type: 'computer_use',
            environment: 'browser',
            enable_prompt_injection_detection: true
        }]
    });

    // Agent Loop
    const turnLimit = 5;
    for (let i = 0; i < turnLimit; i++) {
        console.log(`\n--- Turn ${i + 1} ---`);

        const hasFunctionCalls = interaction.steps.some(step => step.type === "function_call");
        if (!hasFunctionCalls) {
            const textResponses = [];
            for (const step of interaction.steps) {
                if (step.type === "model_output") {
                    for (const contentBlock of step.content || []) {
                        if (contentBlock.type === "text") {
                            textResponses.push(contentBlock.text);
                        }
                    }
                }
            }
            console.log("Agent finished:", textResponses.join(" "));
            break;
        }

        console.log("Executing actions...");
        const results = await executeFunctionCalls(interaction, page, SCREEN_WIDTH, SCREEN_HEIGHT);

        console.log("Capturing state...");
        const functionResponses = await getFunctionResponses(page, results);

        // Continue conversation with function responses
        interaction = await ai.interactions.create({
            model: 'gemini-3.8-flash',
            previous_interaction_id: interaction.id,
            input: functionResponses,
            tools: [{
                type: 'computer_use',
                environment: 'browser',
                enable_prompt_injection_detection: true
            }]
        });
    }
} finally {
    // Cleanup
    console.log("\nClosing browser...");
    await browser.close();
}

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model(Model.of("gemini-3.8-flash"))
        .input(InteractionsInput.of("Click the Submit button on the screen."))
        .tools(Arrays.asList(new ComputerUse()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction.outputText().orElse(""));

Unterstützte Umgebungen (Gemini 3.x)

Gemini 3.x-Modelle unterstützen drei Umgebungen, die in den computer_use-Konfigurationen angegeben sind:

Browserumgebung (ENVIRONMENT_BROWSER)

Verfügbare Aktionen im Browsertool:

Befehlsname Beschreibung Argumente (im Funktionsaufruf)
click Linksklicks an der Koordinate. y: int (0–999)
x: int (0–999)
intent: str
double_click Doppelklicks an der Koordinate. y: int (0–999)
x: int (0–999)
intent: str
triple_click Dreifachklicks an der Koordinate. y: int (0–999)
x: int (0–999)
intent: str
middle_click Mit der mittleren Maustaste auf die Koordinate klicken. y: int (0–999)
x: int (0–999)
intent: str
right_click Rechtsklicks an der Koordinate. y: int (0–999)
x: int (0–999)
intent: str
mouse_down Drückt die Maustaste an der Koordinate und hält sie gedrückt. y: int (0–999)
x: int (0–999)
intent: str
mouse_up Lässt die Maustaste an der Koordinate los. y: int (0–999)
x: int (0–999)
intent: str
move Bewegt den Cursor an die angegebene Position. y: int (0–999)
x: int (0–999)
intent: str
type Text eingeben text: str
press_enter: bool (optional, Standardwert: false)
intent: str
drag_and_drop Zieht ein Element von der Startkoordinate zur Endkoordinate. start_y: int (0–999)
start_x: int (0–999)
end_y: int (0–999)
end_x: int (0–999)
intent: str
wait Hält die Ausführung für eine bestimmte Anzahl von Sekunden an. seconds: int (optional, Standardwert: 1)
intent: str
press_key Drückt die angegebene Taste und lässt sie wieder los. key: str
intent: str
key_down Drückt und hält die angegebene Taste. key: str
intent: str
key_up Gibt den angegebenen Schlüssel frei. key: str
intent: str
Tastenkürzel Drückt die angegebene Tastenkombination. keys: List[str]
intent: str
take_screenshot Gibt einen Screenshot des aktuellen Bildschirms zurück. intent: str
scroll Scrollt an einer Koordinate um eine bestimmte Anzahl von Pixeln nach oben, unten, links oder rechts. y: int (0–999)
x: int (0–999)
direction: str ("up", "down", "left", "right")
magnitude_in_pixels: int (0–999, optional, Standardwert 300)
intent: str
go_back Navigiert zurück zur vorherigen Webseite im Browserverlauf. intent: str
navigate Navigiert direkt zu einer angegebenen URL. url: str
intent: str
go_forward Navigiert vorwärts zur nächsten Webseite im Browserverlauf. intent: str

Mobile Umgebung (ENVIRONMENT_MOBILE)

Android-optimierte Umgebungsvorgänge:

Befehlsname Beschreibung Argumente (im Funktionsaufruf)
open_app Öffnet eine Anwendung anhand ihres Namens. app_name: str
intent: str
click Linksklicks an der Koordinate. y: int (0–999)
x: int (0–999)
intent: str
list_apps Listet die auf dem Gerät verfügbaren Anwendungen auf und gibt ihre Namen und Paketnamen zurück. intent: str
wait Hält die Ausführung für eine bestimmte Anzahl von Sekunden an. seconds: int (optional, Standardwert: 1)
intent: str
go_back Navigiert zurück zum vorherigen Bildschirm oder zur vorherigen Webseite. intent: str
type Text eingeben text: str
press_enter: bool (optional, Standardwert: false)
intent: str
drag_and_drop Zieht ein Element von der Startkoordinate zur Endkoordinate. start_y: int (0–999)
start_x: int (0–999)
end_y: int (0–999)
end_x: int (0–999)
intent: str
long_press Führt einen langen Druck auf eine Koordinate auf dem Bildschirm aus. y: int (0–999)
x: int (0–999)
seconds: int (optional, Standardwert: 2)
intent: str
press_key Drückt die angegebene Taste und lässt sie wieder los. key: str
intent: str
take_screenshot Gibt einen Screenshot des aktuellen Bildschirms zurück. intent: str

Desktopumgebung (ENVIRONMENT_DESKTOP)

Betriebssystemebene – Cursorbefehle für Desktopumgebungen:

Befehlsname Beschreibung Argumente (im Funktionsaufruf)
click Linksklicks an der Koordinate. y: int (0–999)
x: int (0–999)
intent: str
double_click Doppelklicks an der Koordinate. y: int (0–999)
x: int (0–999)
intent: str
triple_click Dreifachklicks an der Koordinate. y: int (0–999)
x: int (0–999)
intent: str
middle_click Mit der mittleren Maustaste auf die Koordinate klicken. y: int (0–999)
x: int (0–999)
intent: str
right_click Rechtsklicks an der Koordinate. y: int (0–999)
x: int (0–999)
intent: str
mouse_down Drückt die Maustaste an der Koordinate und hält sie gedrückt. y: int (0–999)
x: int (0–999)
intent: str
mouse_up Lässt die Maustaste an der Koordinate los. y: int (0–999)
x: int (0–999)
intent: str
move Bewegt den Cursor an die angegebene Position. y: int (0–999)
x: int (0–999)
intent: str
type Text eingeben text: str
press_enter: bool (optional, Standardwert: false)
intent: str
drag_and_drop Zieht ein Element von der Startkoordinate zur Endkoordinate. start_y: int (0–999)
start_x: int (0–999)
end_y: int (0–999)
end_x: int (0–999)
intent: str
wait Hält die Ausführung für eine bestimmte Anzahl von Sekunden an. seconds: int (optional, Standardwert: 1)
intent: str
press_key Drückt die angegebene Taste und lässt sie wieder los. key: str
intent: str
key_down Drückt und hält die angegebene Taste. key: str
intent: str
key_up Gibt den angegebenen Schlüssel frei. key: str
intent: str
Tastenkürzel Drückt die angegebene Tastenkombination. keys: List[str]
intent: str
take_screenshot Gibt einen Screenshot des aktuellen Bildschirms zurück. intent: str
scroll Scrollt an einer Koordinate um eine bestimmte Anzahl von Pixeln nach oben, unten, links oder rechts. y: int (0–999)
x: int (0–999)
direction: str ("up", "down", "left", "right")
magnitude_in_pixels: int (0–999, optional, Standardwert 300)
intent: str

Unterstützte Legacy-UI-Aktionen (Gemini 2.5)

Für Legacy-Modelle (gemini-2.5-computer-use-preview-10-2025) werden die folgenden Aktionen unterstützt:

Befehlsname Beschreibung Argumente (im Funktionsaufruf) Beispiel für Funktionsaufruf
open_web_browser Öffnet den Webbrowser. {"name": "open_web_browser", "arguments": {}}
wait_5_seconds Die Ausführung wird für 5 Sekunden unterbrochen. {"name": "wait_5_seconds", "arguments": {}}
go_back Navigiert zur vorherigen Seite im Verlauf. {"name": "go_back", "arguments": {}}
go_forward Navigiert zur nächsten Seite im Verlauf. {"name": "go_forward", "arguments": {}}
search Navigiert zur Standardsuchmaschine. {"name": "search", "arguments": {}}
navigate Leitet den Browser direkt zur angegebenen URL weiter. url: str {"name": "navigate", "arguments": {"url": "https://www.wikipedia.org"}}
click_at Klicks an einer bestimmten Koordinate. y: int (0–999), x: int (0–999) {"name": "click_at", "arguments": {"y": 300, "x": 500}}
hover_at Führt den Mauszeiger an eine bestimmte Koordinate. y: int (0–999), x: int (0–999) {"name": "hover_at", "arguments": {"y": 150, "x": 250}}
type_text_at Gibt an einer Koordinate Text aus. y: int (0-999), x: int (0-999), text: str, press_enter: bool (Optional, Standardwert True), clear_before_typing: bool (Optional, Standardwert True) {"name": "type_text_at", "arguments": {"y": 250, "x": 400, "text": "search", "press_enter": false}}
key_combination Drücken Sie Tasten oder Tastenkombinationen. keys: str {"name": "key_combination", "arguments": {"keys": "Control+A"}}
scroll_document Scrollt die gesamte Webseite durch. direction: str {"name": "scroll_document", "arguments": {"direction": "down"}}
scroll_at Scrollt an den Koordinaten (x,y). y: int, x: int, direction: str, magnitude: int (Optional, Standardwert 800) {"name": "scroll_at", "arguments": {"y": 500, "x": 500, "direction": "down"}}
drag_and_drop Ziehen zwischen zwei Koordinaten. y: int, x: int, destination_y: int, destination_x: int {"name": "drag_and_drop", "arguments": {"y": 100, "destination_y": 500, "destination_x": 500, "x": 100}}

Benutzerdefinierte Funktionen

Sie können die Funktionalität des Modells erweitern, indem Sie benutzerdefinierte Funktionen einbinden. Beispielsweise können Sie in Human-in-the-Loop (HITL)-Szenarien vordefinierte Standardaktionen ausschließen und benutzerdefinierte Aktionen registrieren.

Gemini 3.x Custom Tooling

Python

Standardmäßige vordefinierte Browseraktionen (wie z. B. click) ausschließen und ein benutzerdefiniertes yield_to_user-Tool registrieren:

from google import genai

client = genai.Client()

yield_to_user_tool = {
    "type": "function",
    "name": "yield_to_user",
    "description": "Yields control back to the user for assistance or verification when an automated action is unsafe or ambiguous.",
    "parameters": {
        "type": "object",
        "properties": {
            "reason": {
                "type": "string",
                "description": "The reason why the agent is yielding control to the human."
            }
        },
        "required": ["reason"]
    }
}

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input="Click the submit button. If you need a second factor authentication code, ask me.",
    tools=[
        {
            "type": "computer_use",
            "environment": "mobile",
            "excluded_predefined_functions": ["click"]
        },
        yield_to_user_tool
    ]
)

JavaScript

Schließen Sie standardmäßige vordefinierte Browseraktionen wie click aus und registrieren Sie ein benutzerdefiniertes yield_to_user-Tool:

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

const yieldToUserTool = {
    type: "function",
    name: "yield_to_user",
    description: "Yields control back to the user for assistance or verification when an automated action is unsafe or ambiguous.",
    parameters: {
        type: "object",
        properties: {
            reason: {
                type: "string",
                description: "The reason why the agent is yielding control to the human."
            }
        },
        required: ["reason"]
    }
};

const interaction = await ai.interactions.create({
    model: "gemini-3.8-flash",
    input: "Click the submit button. If you need a second factor authentication code, ask me.",
    tools: [
        {
            type: "computer_use",
            environment: "mobile",
            excluded_predefined_functions: ["click"]
        },
        yieldToUserTool
    ]
});

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model(Model.of("gemini-3.8-flash"))
        .input(InteractionsInput.of("Click the Submit button on the screen."))
        .tools(Arrays.asList(new ComputerUse()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction.outputText().orElse(""));

Gemini 2.5 (Legacy) Kundenspezifische Werkzeuge

Python

from google import genai

client = genai.Client()

# Define custom tools here
custom_functions = [...]  # Describe parameters as function declarations

excluded_functions = [
    "open_web_browser",
    "wait_5_seconds",
    "go_back",
    "go_forward",
    "search",
    "navigate",
    "hover_at",
    "scroll_document",
    "key_combination",
    "drag_and_drop",
]

interaction = client.interactions.create(
    model='gemini-2.5-computer-use-preview-10-2025',
    input="Open Chrome, then long-press at 200,400.",
    tools=[
        {
            "type": "computer_use",
            "environment": "browser",
            "excluded_predefined_functions": excluded_functions
        },
        *custom_functions
    ]
)

print(interaction)

JavaScript

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

// Define custom tools here
const customFunctions = [...]; // Describe parameters as function declarations

const excludedFunctions = [
    "open_web_browser",
    "wait_5_seconds",
    "go_back",
    "go_forward",
    "search",
    "navigate",
    "hover_at",
    "scroll_document",
    "key_combination",
    "drag_and_drop",
];

const interaction = await ai.interactions.create({
    model: 'gemini-2.5-computer-use-preview-10-2025',
    input: "Open Chrome, then long-press at 200,400.",
    tools: [
        {
            type: "computer_use",
            environment: "browser",
            excluded_predefined_functions: excludedFunctions
        },
        ...customFunctions
    ]
});

console.log(interaction);

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model(Model.of("gemini-3.8-flash"))
        .input(InteractionsInput.of("Click the Submit button on the screen."))
        .tools(Arrays.asList(new ComputerUse()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction.outputText().orElse(""));

Steuerung der Denkebenen (Zwillinge 3.x)

Bei Agenten, die Computer einsetzen, können Sie verschiedene Denkebenen konfigurieren, um ein Gleichgewicht zwischen Aktionsqualität und Ausführungsgeschwindigkeit herzustellen. Bei Standardautomatisierungsaufgaben wird im Allgemeinen ein gutes Gleichgewicht zwischen niedrigerem Denkvermögen und anderen Faktoren erzielt.

Sicherheit

Konfiguration von Sicherheitsrichtlinien (Gemini 3.x)

Die Gemini 3.x-Modelle beinhalten integrierte Sicherheitsdienstkategorien, die automatisch ermitteln, ob eine Benutzerbestätigung erforderlich ist.

Kategorie der Sicherheitsrichtlinien Beschreibung
FINANCIAL_TRANSACTIONS Blockiert oder löst eine Bestätigung für Aktionen im Zusammenhang mit Zahlungen, Kassenvorgängen im Einzelhandel oder regulierten Waren aus.
SENSITIVE_DATA_MODIFICATION Schützt Gesundheits-, Finanz- oder Regierungsdaten vor unbefugter Änderung.
COMMUNICATION_TOOL Verhindert, dass der Agent selbstständig E-Mails, Chatnachrichten oder Entwürfe versendet.
ACCOUNT_CREATION Verhindert, dass der Agent selbstständig neue Konten auf Webseiten registriert.
DATA_MODIFICATION Regelt allgemeine Dateisystemänderungen, Datenaustausch und Speicherlöschung.
USER_CONSENT_MANAGEMENT Erfordert die Zustimmung des Nutzers für Cookie-Einwilligungsbanner und Datenschutzhinweise.
LEGAL_TERMS_AND_AGREEMENTS Verhindert, dass das Modell selbstständig Nutzungsbedingungen oder rechtsverbindliche Verträge akzeptiert.

Sicherheitsüberbrückungen

Sie können einzelne Richtlinien durch Überschreibungen außer Kraft setzen:

Python

from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-3.8-flash",
    input="Clean up the local folder by archiving old logs.",
    tools=[
        {
            "type": "computer_use",
            "environment": "desktop",
            "disabled_safety_policies": [
                "data_modification"
            ]
        }
    ]
)

JavaScript

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

const interaction = await ai.interactions.create({
    model: "gemini-3.8-flash",
    input: "Clean up the local folder by archiving old logs.",
    tools: [
        {
            type: "computer_use",
            environment: "desktop",
            disabled_safety_policies: [
                "data_modification"
            ]
        }
    ]
});

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model(Model.of("gemini-3.8-flash"))
        .input(InteractionsInput.of("Click the Submit button on the screen."))
        .tools(Arrays.asList(new ComputerUse()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction.outputText().orElse(""));

Schnelle Injektionserkennung (Gemini 3.x)

Computer Use for Gemini 3.5 Flash oder später unterstützt einen fortschrittlichen Sicherheitsmechanismus zur Erkennung von Prompt-Injection-Angriffen. Wenn diese Funktion aktiviert ist, prüft sie, ob ein beigefügter Screenshot versteckte Angriffsanweisungen enthält (z. B. „Vorherige Befehle ignorieren“) und blockiert die Ausführung, wenn sie erkannt wird.

Die Erkennung von Fehlinjektionen ist eine optionale Funktion. Der Standardwert ist false.

Die folgenden Beispiele veranschaulichen, wie Sie die Erkennung von Prompt-Injection in der Konfiguration Ihres Computer Use-Tools aktivieren:

Python

from google import genai

client = genai.Client()

interaction = client.interactions.create(
    model="gemini-3.5-flash",
    input="Search for flight deals and summarize top results.",
    tools=[
        {
            "type": "computer_use",
            "environment": "desktop",
            "enable_prompt_injection_detection": True,
        }
    ],
)

JavaScript

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI();

const interaction = await ai.interactions.create({
    model: "gemini-3.5-flash",
    input: "Search for flight deals and summarize top results.",
    tools: [
        {
            type: "computer_use",
            environment: "desktop",
            enablePromptInjectionDetection: true,
        }
    ]
});

Java

import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.ComputerUse;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.Arrays;

Client client = new Client();

CreateModelInteraction params =
    CreateModelInteraction.builder()
        .model(Model.of("gemini-3.8-flash"))
        .input(InteractionsInput.of("Click the Submit button on the screen."))
        .tools(Arrays.asList(new ComputerUse()))
        .build();

Interaction interaction =
    client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();

System.out.println(interaction.outputText().orElse(""));

cURL

curl "https://generativelanguage.googleapis.com/v1beta/interactions?key=${GEMINI_API_KEY}" \
-H 'Content-Type: application/json' \
-d '{
  "model": "gemini-3.5-flash",
  "input": "Search for flight deals and summarize top results.",
  "tools": [
    {
      "type": "computer_use",
      "environment": "desktop",
      "enable_prompt_injection_detection": true
    }
  ]
}'

Sicherheitsentscheidung bestätigen

Die Antwort kann einen safety_decision-Parameter in den Funktionsaufrufargumenten enthalten:

{
  "steps": [
    {
      "type": "function_call",
      "name": "click_at",
      "arguments": {
        "x": 60,
        "y": 100,
        "safety_decision": {
          "explanation": "Must check check-box",
          "decision": "require_confirmation"
        }
      }
    }
  ]
}

Wenn safety_decision gleich require_confirmation ist, zeige dem Endbenutzer eine entsprechende Meldung an. Wenn der Benutzer bestätigt, setze safety_acknowledgement in function_result.

Python

def get_safety_confirmation(safety_decision):
    # Prompt user for confirmation
    print(f"Safety confirmation required: {safety_decision.get('explanation', '')}")
    return "CONTINUE" # Or TERMINATE

# Inside execute_function_calls, check for safety_decision:
if 'safety_decision' in function_call.arguments:
    decision = get_safety_confirmation(function_call.arguments['safety_decision'])
    if decision == "TERMINATE":
        break
    # Include safety_acknowledgement inside the action result
    action_result["safety_acknowledgement"] = True

Best Practices für die Sicherheit

Die Computernutzung birgt besondere Sicherheits- und Betriebsrisiken, da ein Modell, das im Auftrag eines Benutzers agiert, auf nicht vertrauenswürdige Inhalte auf dem Bildschirm stoßen oder Fehler bei der Ausführung von Aktionen machen könnte. Um Benutzerdaten und Systeme zu schützen, sollten Sie folgende bewährte Verfahren umsetzen:

  1. Human-in-the-Loop (HITL):
    • Benutzerbestätigung erzwingen: Wenn die Sicherheitsreaktion require_confirmation anzeigt (oder eine ältere Sicherheitsentscheidung dies erfordert), fordern Sie den Benutzer zur Genehmigung auf.
    • Benutzerdefinierte Sicherheitsanweisungen bereitstellen: Implementieren Sie eine benutzerdefinierte Systemanweisung, um Ihre eigenen Sicherheitsgrenzen zu definieren und durchzusetzen. Beispiel:

      Python

      from google import genai
      
      client = genai.Client()
      
      system_instruction = """
      ## **RULE 1: Seek User Confirmation (USER_CONFIRMATION)**
      
      This is your first and most important check. If the next required action falls
      into any of the following categories, you MUST stop immediately, and seek the
      user's explicit permission.
      
      **Procedure for Seeking Confirmation:**
      * **For Consequential Actions:** Perform all preparatory steps (e.g., navigating,
        filling out forms, typing a message). You will ask for confirmation **AFTER**
        all necessary information is entered on the screen, but **BEFORE** you perform
        the final, irreversible action (e.g., before clicking "Send", "Submit",
        "Confirm Purchase", "Share").
      * **For Prohibited Actions:** If the action is strictly forbidden (e.g., accepting
        legal terms, solving a CAPTCHA), you must first inform the user about the
        required action and ask for their confirmation to proceed.
      
      **USER_CONFIRMATION Categories:**
      
      *   **Consent and Agreements:** You are FORBIDDEN from accepting, selecting, or
          agreeing to any of the following on the user's behalf. You must ask the
          user to confirm before performing these actions.
          *   Terms of Service
          *   Privacy Policies
          *   Cookie consent banners
          *   End User License Agreements (EULAs)
          *   Any other legally significant contracts or agreements.
      *   **Robot Detection:** You MUST NEVER attempt to solve or bypass the
          following. You must ask the user to confirm before performing these actions.
          *   CAPTCHAs (of any kind)
          *   Any other anti-robot or human-verification mechanisms, even if you are
              capable.
      *   **Financial Transactions:**
          *   Completing any purchase.
          *   Managing or moving money (e.g., transfers, payments).
          *   Purchasing regulated goods or participating in gambling.
      *   **Sending Communications:**
          *   Sending emails.
          *   Sending messages on any platform (e.g., social media, chat apps).
          *   Posting content on social media or forums.
      *   **Accessing or Modifying Sensitive Information:**
          *   Health, financial, or government records (e.g., medical history, tax
              forms, passport status).
          *   Revealing or modifying sensitive personal identifiers (e.g., SSN, bank
              account number, credit card number).
      *   **User Data Management:**
          *   Accessing, downloading, or saving files from the web.
          *   Sharing or sending files/data to any third party.
          *   Transferring user data between systems.
      *   **Browser Data Usage:**
          *   Accessing or managing Chrome browsing history, bookmarks, autofill data,
              or saved passwords.
      *   **Security and Identity:**
          *   Logging into any user account.
          *   Any action that involves misrepresentation or impersonation (e.g.,
              creating a fan account, posting as someone else).
      *   **Insurmountable Obstacles:** If you are technically unable to interact with
          a user interface element or are stuck in a loop you cannot resolve, ask the
          user to take over.
      ---
      
      ## **RULE 2: Default Behavior (ACTUATE)**
      
      If an action does **NOT** fall under the conditions for `USER_CONFIRMATION`,
      your default behavior is to **Actuate**.
      
      **Actuation Means:**  You MUST proactively perform all necessary steps to move
      the user's request forward. Continue to actuate until you either complete the
      non-consequential task or encounter a condition defined in Rule 1.
      
      *   **Example 1:** If asked to send money, you will navigate to the payment
          portal, enter the recipient's details, and enter the amount. You will then
          **STOP** as per Rule 1 and ask for confirmation before clicking the final
          "Send" button.
      *   **Example 2:** If asked to post a message, you will navigate to the site,
          open the post composition window, and write the full message. You will then
          **STOP** as per Rule 1 and ask for confirmation before clicking the final
          "Post" button.
      
          After the user has confirmed, remember to get the user's latest screen
          before continuing to perform actions.
      
      # Final Response Guidelines:
      Write final response to the user in the following cases:
      - User confirmation
      - When the task is complete or you have enough information to respond to the user
      """
      
      interaction = client.interactions.create(
          model="gemini-3.8-flash",
          system_instruction=system_instruction,
          input="Prepare a draft but do not send.",
          tools=[{
              "type": "computer_use",
              "environment": "browser"
          }]
      )
      

      JavaScript

      import { GoogleGenAI } from '@google/genai';
      
      const ai = new GoogleGenAI();
      
      const systemInstruction = `
      ## **RULE 1: Seek User Confirmation (USER_CONFIRMATION)**
      
      This is your first and most important check. If the next required action falls
      into any of the following categories, you MUST stop immediately, and seek the
      user's explicit permission.
      
      **Procedure for Seeking Confirmation:**
      * **For Consequential Actions:** Perform all preparatory steps (e.g., navigating,
        filling out forms, typing a message). You will ask for confirmation **AFTER**
        all necessary information is entered on the screen, but **BEFORE** you perform
        the final, irreversible action (e.g., before clicking "Send", "Submit",
        "Confirm Purchase", "Share").
      * **For Prohibited Actions:** If the action is strictly forbidden (e.g., accepting
        legal terms, solving a CAPTCHA), you must first inform the user about the
        required action and ask for their confirmation to proceed.
      
      **USER_CONFIRMATION Categories:**
      
      *   **Consent and Agreements:** You are FORBIDDEN from accepting, selecting, or
          agreeing to any of the following on the user's behalf. You must ask the
          user to confirm before performing these actions.
          *   Terms of Service
          *   Privacy Policies
          *   Cookie consent banners
          *   End User License Agreements (EULAs)
          *   Any other legally significant contracts or agreements.
      *   **Robot Detection:** You MUST NEVER attempt to solve or bypass the
          following. You must ask the user to confirm before performing these actions.
          *   CAPTCHAs (of any kind)
          *   Any other anti-robot or human-verification mechanisms, even if you are
              capable.
      *   **Financial Transactions:**
          *   Completing any purchase.
          *   Managing or moving money (e.g., transfers, payments).
          *   Purchasing regulated goods or participating in gambling.
      *   **Sending Communications:**
          *   Sending emails.
          *   Sending messages on any platform (e.g., social media, chat apps).
          *   Posting content on social media or forums.
      *   **Accessing or Modifying Sensitive Information:**
          *   Health, financial, or government records (e.g., medical history, tax
              forms, passport status).
          *   Revealing or modifying sensitive personal identifiers (e.g., SSN, bank
              account number, credit card number).
      *   **User Data Management:**
          *   Accessing, downloading, or saving files from the web.
          *   Sharing or sending files/data to any third party.
          *   Transferring user data between systems.
      *   **Browser Data Usage:**
          *   Accessing or managing Chrome browsing history, bookmarks, autofill data,
              or saved passwords.
      *   **Security and Identity:**
          *   Logging into any user account.
          *   Any action that involves misrepresentation or impersonation (e.g.,
              creating a fan account, posting as someone else).
      *   **Insurmountable Obstacles:** If you are technically unable to interact with
          a user interface element or are stuck in a loop you cannot resolve, ask the
          user to take over.
      ---
      
      ## **RULE 2: Default Behavior (ACTUATE)**
      
      If an action does **NOT** fall under the conditions for \`USER_CONFIRMATION\`,
      your default behavior is to **Actuate**.
      
      **Actuation Means:**  You MUST proactively perform all necessary steps to move
      the user's request forward. Continue to actuate until you either complete the
      non-consequential task or encounter a condition defined in Rule 1.
      
      *   **Example 1:** If asked to send money, you will navigate to the payment
          portal, enter the recipient's details, and enter the amount. You will then
          **STOP** as per Rule 1 and ask for confirmation before clicking the final
          "Send" button.
      *   **Example 2:** If asked to post a message, you will navigate to the site,
          open the post composition window, and write the full message. You will then
          **STOP** as per Rule 1 and ask for confirmation before clicking the final
          "Post" button.
      
          After the user has confirmed, remember to get the user's latest screen
          before continuing to perform actions.
      
      # Final Response Guidelines:
      Write final response to the user in the following cases:
      - User confirmation
      - When the task is complete or you have enough information to respond to the user
      `;
      
      const interaction = await ai.interactions.create({
          model: "gemini-3.8-flash",
          system_instruction: systemInstruction,
          input: "Prepare a draft but do not send.",
          tools: [{
              type: "computer_use",
              environment: "browser"
          }]
      });
      

Java

java import com.google.genai.Client; import com.google.genai.gaos.models.interactions.ComputerUse; import com.google.genai.gaos.models.interactions.CreateModelInteraction; import com.google.genai.gaos.models.interactions.Interaction; import com.google.genai.gaos.models.interactions.InteractionsInput; import com.google.genai.gaos.models.interactions.Model; import com.google.genai.gaos.models.operations.CreateInteractionRequestBody; import java.util.Arrays; Client client = new Client(); CreateModelInteraction params = CreateModelInteraction.builder() .model(Model.of("gemini-3.8-flash")) .input(InteractionsInput.of("Click the Submit button on the screen.")) .tools(Arrays.asList(new ComputerUse())) .build(); Interaction interaction = client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get(); System.out.println(interaction.outputText().orElse(""));
  1. Sichere Ausführungsumgebung: Führen Sie Ihren Agenten in einer sicheren, isolierten Umgebung aus, um seine potenziellen Auswirkungen zu begrenzen. Dies kann eine isolierte virtuelle Maschine (VM), ein Container (z.B. Docker) oder ein dediziertes Browserprofil mit eingeschränkten Berechtigungen sein. Siehe die GitHub-Referenzimplementierung für eine Anleitung zur Einrichtung der Sandbox mit Docker.
  2. Eingabebereinigung:Bereinigen Sie alle von Nutzern generierten Texte in Prompts, um das Risiko unbeabsichtigter Anweisungen oder Prompt-Injection zu minimieren. Dies ist eine hilfreiche Sicherheitsebene, aber kein Ersatz für eine sichere Ausführungsumgebung.
  3. Inhaltsschutzmaßnahmen: Nutzen Sie Schutzmaßnahmen und APIs zur Inhaltssicherheit, um Benutzereingaben, Tool-Ein- und -Ausgaben sowie die Antworten des Agenten auf Angemessenheit, Prompt-Injection und Jailbreak-Erkennung zu überprüfen.
  4. Zulassungslisten und Sperrlisten: Implementieren Sie Filtermechanismen, um zu steuern, wohin das Modell navigieren und was es tun kann. Eine Sperrliste verbotener Websites ist ein guter Ausgangspunkt, während eine restriktivere Zulassungsliste noch mehr Sicherheit bietet.
  5. Beobachtbarkeit und Protokollierung:Detaillierte Logs für das Debugging, die Prüfung und die Incident Response führen. Ihr Client sollte Eingabeaufforderungen, Screenshots, vom Modell vorgeschlagene Aktionen (function_call), Sicherheitsreaktionen und alle letztendlich vom Client ausgeführten Aktionen protokollieren.
  6. Umgebungsverwaltung: Sicherstellen, dass die GUI-Umgebung konsistent ist. Unerwartete Pop-ups, Benachrichtigungen oder Layoutänderungen können das Modell verwirren. Beginnen Sie nach Möglichkeit jede neue Aufgabe mit einem bekannten, sauberen Zustand.

Modellversionen

Sie können die Computernutzung mit den folgenden Modellen verwenden:

  • Gemini 3.8 Flash (gemini-3.8-flash): Das empfohlene Modell für die Computernutzung mit hochpräziser UI-Interaktion und zuverlässigem Tool-Aufruf.
  • Gemini 3.7 Flash (gemini-3.7-flash): Das bisherige stabile Modell für die Computernutzung mit optimierten Aktionen mit Intents, Unterstützung für Browser-, Mobilgeräte- und Desktopumgebungen, konfigurierbaren Sicherheitsrichtlinien und Erkennung von Prompt-Injection.
  • Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite): Ein Modell mit niedriger Latenz und geringem Kostenaufwand, das die Nutzung am Computer unterstützt.
  • Gemini 3.5 Flash (gemini-3.5-flash): Das bisherige stabile Modell, das die Nutzung auf Computern unterstützt.
  • Gemini 3 Flash (Vorabversion) (gemini-3-flash-preview): Vorabversion des Modells, das die Nutzung von Computern unterstützt.
  • Gemini 2.5 (Legacy-Vorabversion) (gemini-2.5-computer-use-preview-10-2025): Legacy-Vorabversion, die für die browserbasierte Computernutzung optimiert ist.

Nächste Schritte