Модели Gemini Robotics ER могут планировать задачи и анализировать пространство, определяя, какие действия необходимо предпринять и какие объекты переместить для достижения цели. На этой странице показан пример управления операцией захвата и перемещения с помощью пользовательского API робота для организации задачи по размещению предмета в чаше. В этом примере используется стандартная модель Gemini ER 2; пример потоковой обработки см. в руководстве по потоковой обработке Gemini ER 2 .
Полный рабочий код можно найти в руководстве по робототехнике .
Использование пользовательского API робота
Этот пример демонстрирует организацию задач с помощью пользовательского API робота. В нем представлен фиктивный API, разработанный для операции захвата и перемещения. Задача состоит в том, чтобы поднять синий блок и поместить его в оранжевую миску:

В этом примере используется следующий фиктивный API робота:
Python
def move(x, y, high):
print(f"Mock Robot: Moving to coordinates: {x}, {y}, {'high above table' if high else 'down at table level'}")
def setGripperState(opened):
print(f"Mock Robot: {'Opening gripper' if opened else 'Closing gripper'}")
robot_origin_y = 300
robot_origin_x = 500
move_function = {
"type": "function",
"name": "move",
"description": "Moves the arm to the given coordinates.",
"parameters": {
"type": "object",
"properties": {
"x": {"type": "integer", "description": "X coordinate relative to the origin"},
"y": {"type": "integer", "description": "Y coordinate relative to the origin"},
"high": {"type": "boolean", "description": "Set to True to lift the robot arm above the scene for avoiding obstacles. Set to False to place the gripper on the surface."}
},
"required": ["x", "y", "high"]
}
}
set_gripper_state_function = {
"type": "function",
"name": "setGripperState",
"description": "Opens or closes the robot's gripper.",
"parameters": {
"type": "object",
"properties": {
"opened": {"type": "boolean", "description": "True opens the gripper, False closes the gripper."}
},
"required": ["opened"]
}
}
В следующем примере модель отправляет запрос и изображение вместе с определениями инструментов. Затем запускается агентный цикл: после каждого ответа модели выполняются все запрошенные вызовы функций ( move , setGripperState ), результаты возвращаются модели с использованием previous_interaction_id , и цикл повторяется до тех пор, пока модель не перестанет вызывать функции или не будет достигнут лимит шагов.
Python
prompt = (
"You are a robotic arm with six degrees-of-freedom. "
f"The origin point for calculating the moves is at normalized point y={robot_origin_y}, x={robot_origin_x}. "
"Use this as the new (0,0) for calculating moves, allowing x and y to be negative.\n\n"
"Find the blue block and the orange bowl. Calculate their coordinates relative to the origin.\n"
"Perform a pick and place operation where you pick up the blue block and place it into the orange bowl. "
"Call the appropriate sequence of functions to complete this operation."
)
# 1. Initial Interaction
interaction = client.interactions.create(
model=MODEL_ID,
input=[{"type": "user_input", "content": [
{"type": "image", "data": img_b64, "mime_type": "image/png"},
{"type": "text", "text": prompt}
]}],
tools=[move_function, set_gripper_state_function],
generation_config={"thinking_level": "low"}
)
print("\n--- Executing Orchestrated Plan ---")
max_steps = 15 # Safety limit to prevent infinite loops
step_count = 0
# 2. The Agentic Loop
while step_count < max_steps:
step_count += 1
# Check if the model wants to call any functions
tool_calls = [step for step in interaction.steps if step.type == "function_call"]
if not tool_calls:
# If no tools were called, the model is finished with the sequence
print("Sequence complete.")
if interaction.output_text:
print(f"Model Summary: {interaction.output_text}")
break
function_results = []
for step in tool_calls:
function_name = step.name
arguments = step.arguments
# Execute the mock function
if function_name == "move":
move(**arguments)
elif function_name == "setGripperState":
setGripperState(**arguments)
else:
print(f"Unknown function: {function_name}")
# 3. Create a result object to tell the model the function succeeded
function_results.append({
"type": "function_result",
"name": step.name,
"call_id": step.id,
"result": [{"type": "text", "text": '{"status": "success"}'}]
})
# 4. Send the results back to the model, passing previous_interaction_id
# so it remembers the conversation history and generates the NEXT step
interaction = client.interactions.create(
model=MODEL_ID,
previous_interaction_id=interaction.id,
tools=[move_function, set_gripper_state_function],
input=function_results
)
Ниже показан возможный результат работы модели на основе запроса и фиктивного API робота. Результат включает в себя вывод вызовов функций робота, которые модель последовательно объединила.
--- Executing Orchestrated Plan ---
Mock Robot: Opening gripper
Mock Robot: Moving to coordinates: 160, 440, high above table
Mock Robot: Moving to coordinates: 160, 440, down at table level
Mock Robot: Closing gripper
Mock Robot: Moving to coordinates: 160, 440, high above table
Mock Robot: Moving to coordinates: -250, 60, high above table
Mock Robot: Moving to coordinates: -250, 60, down at table level
Mock Robot: Opening gripper
Mock Robot: Moving to coordinates: -250, 60, high above table
Sequence complete.
Model Summary: I have completed the task of picking up the blue block and placing it into the orange bowl.
Что дальше?
- Робототехника с потоковой передачей данных — потоковая передача данных в реальном времени с вызовом функций (только для Gemini Robotics ER 2).
- Анализ видеоматериалов — отслеживание хода выполнения задачи по видео (только для ER 2).
- Пространственное мышление — примеры указания, отслеживания и построения ограничивающих рамок.