Gemini 3.5 Flash (gemini-3.5-flash) is a deprecated earlier-generation Flash
model in the Gemini 3 series. Requests to gemini-3.5-flash are automatically
routed to Gemini 3.6 Flash. For our
newest Flash model, see Gemini 3.8 Flash.
This guide is preserved for historical reference and documents the baseline Gemini 3.x API changes and parameter recommendations.
Model overview
| Model | Model ID | Status |
|---|---|---|
| Gemini 3.5 Flash | gemini-3.5-flash |
Deprecated (automatically routed to gemini-3.6-flash). |
Gemini 3.5 Flash supported the 1M token context window, 65k max output tokens, thinking, and the same set of tools and platform features as Gemini 3 Flash, including Computer Use (Preview).
For complete specs, see the models overview.
Quickstart
All examples in this guide use the Interactions API. The GenerateContent API is also supported; the same configuration options and recommendations apply.
Python
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.5-flash",
input="Explain how parallel agentic execution works in three sentences."
)
print(interaction.output_text)
JavaScript
import { GoogleGenAI } from "@google/genai";
const client = new GoogleGenAI({});
async function main() {
const interaction = await client.interactions.create({
model: "gemini-3.5-flash",
input: "Explain how parallel agentic execution works in three sentences.",
});
console.log(interaction.output_text);
}
main();
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.5-flash"))
.input(
InteractionsInput.of(
"Explain how parallel agentic execution works in three sentences."))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println(interaction.outputText().orElse(""));
Go
package main
import (
"context"
"fmt"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.5-flash"),
Input: interactions.NewInteractionsInput("Explain how parallel agentic execution works in three sentences."),
}),
})
if err != nil {
log.Fatal(err)
}
if res.Interaction.OutputText != nil {
fmt.Println(*res.Interaction.OutputText)
}
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini-3.5-flash",
"input": "Explain how parallel agentic execution works in three sentences."
}'
What's new
- Sustained performance: Optimized for agentic and coding tasks at scale.
- Agentic execution: Sub-agent deployment, problem solving, and rapid agentic loops at scale.
- Coding: Iterative coding cycles, rapid exploration, and prototyping to test alternate paths and dynamically explore solutions.
- Long horizon: Multi-step workflows and tool use at scale.
- Thought preservation: The model maintains intermediate reasoning across multi-turn conversations automatically. No API changes needed.
- New default effort level: Default thinking effort changed from
hightomedium. See New default effort level for details. - Improved
lowthinking:lowis significantly improved for code and agentic tasks that require fewer steps, offering strong quality at lower latency and cost.
Choosing the right Flash model
Since gemini-3.5-flash is deprecated and automatically routed to
gemini-3.6-flash, we recommend migrating to one of the following active
models:
- Gemini 3.8 Flash: Our latest and most intelligent Flash model for complex coding and agentic workflows. See the Gemini 3.8 Flash guide.
- Gemini 3.6 Flash: Direct replacement for
gemini-3.5-flashwith reduced token usage and lower output pricing. See the Gemini 3.6 Flash and 3.5 Flash-Lite guide. - Gemini 3.5 Flash-Lite / Gemini 3.1 Flash-Lite: For low-cost, high-volume tasks that don't require full Flash reasoning depth, use Gemini 3.5 Flash-Lite or Gemini 3.1 Flash-Lite.
Behavioral changes
New default effort level: medium
The default thinking effort is now medium, changed from high in Gemini 3
Flash Preview. medium yields very good results across a wide range of tasks
while being faster and more cost-efficient. For complex problems, high
encourages the model to think more deeply.
| Effort level | When to use |
|---|---|
minimal |
Optimized for response speed. Chat-like use cases, quick factual answers, simpler tool calls. |
low |
Code and agentic tasks that require lower latency and fewer steps. Also works well for analysis and writing tasks that require some thinking. |
medium (default) |
Best quality for most tasks. Recommended for complex code and agentic use cases. |
high |
Maximizes the model's ability to think and use tools. Best for complex reasoning, hard math, and the most difficult code or agent tasks. Allows extended thoughts and function calls. |
To override the default, set thinking_level in your config:
Python
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="gemini-3.5-flash",
input="Prove that the square root of 2 is irrational.",
generation_config={"thinking_level": "high"},
)
print(interaction.output_text)
JavaScript
import { GoogleGenAI } from "@google/genai";
const client = new GoogleGenAI({});
async function main() {
const interaction = await client.interactions.create({
model: "gemini-3.5-flash",
input: "Prove that the square root of 2 is irrational.",
generationConfig: { thinkingLevel: "high" },
});
console.log(interaction.output_text);
}
main();
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.GenerationConfig;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.ThinkingLevel;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.5-flash"))
.input(InteractionsInput.of("Prove that the square root of 2 is irrational."))
.generationConfig(GenerationConfig.builder().thinkingLevel(ThinkingLevel.HIGH).build())
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println(interaction.outputText().orElse(""));
Go
package main
import (
"context"
"fmt"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.5-flash"),
Input: interactions.NewInteractionsInput("Prove that the square root of 2 is irrational."),
GenerationConfig: &interactions.GenerationConfig{
ThinkingLevel: interactions.ThinkingLevelHigh.ToPointer(),
},
}),
})
if err != nil {
log.Fatal(err)
}
if res.Interaction.OutputText != nil {
fmt.Println(*res.Interaction.OutputText)
}
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini-3.5-flash",
"input": "Prove that the square root of 2 is irrational.",
"generation_config": {"thinking_level": "high"}
}'
The following table shows which thinking levels are supported per model:
| Thinking Level | Gemini 3.5 Flash | Gemini 3.1 Pro | Gemini 3.1 Flash-Lite | Gemini 3 Flash | Description |
|---|---|---|---|---|---|
minimal |
Supported | Not supported | Supported (Default) | Supported | Matches the "no thinking" setting for most queries. Note, minimal does not guarantee that thinking is off, the model may reason very minimally for complex tasks. |
low |
Supported | Supported | Supported | Supported | Minimizes latency and cost. |
medium |
Supported (Default) | Supported | Supported | Supported | Balanced thinking for most tasks. |
high |
Supported (Dynamic) | Supported (Default, Dynamic) | Supported (Dynamic) | Supported (Default, Dynamic) | Maximizes reasoning depth. |
Thought preservation
The model maintains intermediate reasoning across multi-turn conversations automatically. When present in the conversation history, reasoning context carries forward, which improves performance on complex multi-step tasks like iterative debugging and code refactoring. No API changes needed:
- Interactions API: Thoughts are already preserved automatically. No change in behavior.
- GenerateContent API: Beginning with Gemini 3.5 Flash, the model uses
reasoning context from all previous turns when thought signatures are
present in the conversation history. To enable this, pass the full,
unmodified conversation history (including
thought signatures) in
contents. The SDKs handle this automatically.
Parameter updates and best practices in Gemini 3.x
The following apply to all Gemini 3.x models, including Gemini 3.5 Flash.
temperature,top_p,top_k: we strongly recommend not changing the default values. Gemini 3's reasoning capabilities are optimized for the default settings.- Use
thinking_levelinstead ofthinking_budget. - Function calling response matching:
id,name, and response count must match the preceding calls. - Multimodal function responses: include multimodal content inside the function response, not outside it.
- Inline instructions in function responses: append to the function response text, not as separate parts.
- Reduce unnecessary tool calls: Use lower thinking levels or experiment with system instructions to reduce tool calls in agentic workflows.
See the sections below for how to update your code.
Sampling parameters (no longer recommended)
temperature, top_p, and top_k are no longer recommended for all Gemini
3.x models. Gemini 3's reasoning capabilities are optimized for the default
settings. Remove these parameters from all requests.
# ⚠️ Remove these parameters (not recommended)
generation_config = {
"temperature": 0.7,
"top_p": 0.9,
"top_k": 40,
}
To ensure determinism, we recommend defining a system instruction with explicit rules for your specific use case.
thinking_budget (no longer recommended)
The raw numeric thinking_budget parameter is no longer recommended across all
Gemini 3.x models. Use the thinking_level string enum instead.
# ⚠️ Before (not recommended)
generation_config = {
"thinking": {"thinking_budget": 7500},
}
# ✅ After
generation_config = {
"thinking": {"thinking_level": "medium"},
}
Available values: minimal, low, medium (default), and high.
Function calling: strict response matching
The Interactions API already errors on mismatched function responses. The
GenerateContent API does not yet error, but mismatched responses cause the model
to return empty responses with finish_reason: STOP in most cases. Always
follow these conventions:
| Requirement | Details |
|---|---|
Include id |
Every FunctionResponse must include the id from the corresponding FunctionCall |
Match name |
The name in the response must match the name in the call |
| Match counts | Return exactly one FunctionResponse for each FunctionCall received |
Python
# ✅ Include matching call_id and name in the function_result
final_interaction = client.interactions.create(
model="gemini-3.5-flash",
previous_interaction_id=interaction.id,
tools=[my_tool],
input=[{
"type": "function_result",
"name": fc_step.name,
"call_id": fc_step.id,
"result": [{"type": "text", "text": json.dumps(result)}],
}],
)
JavaScript
// ✅ Include matching call_id and name in the function_result
const finalInteraction = await client.interactions.create({
model: "gemini-3.5-flash",
previousInteractionId: interaction.id,
tools: [myTool],
input: [{
type: "function_result",
name: fcStep.name,
call_id: fcStep.id,
result: [{ type: "text", text: JSON.stringify(result) }],
}],
});
Java
import java.util.Arrays;
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Function;
import com.google.genai.gaos.models.interactions.FunctionResultStep;
import com.google.genai.gaos.models.interactions.FunctionResultStepResultUnion;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.List;
Client client = new Client();
// Assumes interactionId, callId, functionName, myTool, and resultJson from previous step
String interactionId = "interaction-id-123";
String callId = "call-id-123";
String functionName = "get_weather";
Function myTool = Function.builder().name(functionName).build();
String resultJson = "{\"temperature\": \"72F\"}";
// ✅ Include matching callId and name in the FunctionResultStep
FunctionResultStep functionResult =
FunctionResultStep.builder()
.name(functionName)
.callId(callId)
.result(
FunctionResultStepResultUnion.of(
Arrays.asList(TextContent.builder().text(resultJson).build())))
.build();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.5-flash"))
.previousInteractionId(interactionId)
.tools(Arrays.asList(myTool))
.input(InteractionsInput.ofStep(Arrays.asList(functionResult)))
.build();
Interaction finalInteraction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
Go
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
// Assumes interactionID, callID, functionName, myTool, and resultJSON from previous step
interactionID := "interaction-id-123"
callID := "call-id-123"
functionName := "get_weather"
myTool := interactions.NewTool(interactions.Function{Name: genai.Ptr(functionName)})
resultJSON := `{"temperature": "72F"}`
// ✅ Include matching CallID and Name in the FunctionResultStep
functionResult := interactions.NewStep(interactions.FunctionResultStep{
Name: genai.Ptr(functionName),
CallID: callID,
Result: interactions.NewFunctionResultStepResultUnion([]interactions.FunctionResultSubcontent{
interactions.NewFunctionResultSubcontent(interactions.TextContent{Text: resultJSON}),
}),
})
finalRes, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.5-flash"),
PreviousInteractionID: genai.Ptr(interactionID),
Tools: []interactions.Tool{myTool},
Input: interactions.NewInteractionsInput([]interactions.Step{functionResult}),
}),
})
if err != nil {
log.Fatal(err)
}
_ = finalRes
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini-3.5-flash",
"previous_interaction_id": "<INTERACTION_ID>",
"tools": [...],
"input": [{
"type": "function_result",
"name": "my_function",
"call_id": "<CALL_ID>",
"result": [{"type": "text", "text": "..."}]
}]
}'
Multimodal function responses
We often see clients provide images outside function response. This can lead to unexpected model behavior (e.g. thought leakage) and result in lower quality outputs. Follow the recommendation in Multimodal Function Responses API docs instead and include multimodal content in the function response parts that you send to the model. The model can process this multimodal content in its next turn to produce a more informed response.
Python
# ✅ Include multimodal content in the function response
final_interaction = client.interactions.create(
model="gemini-3.5-flash",
previous_interaction_id=interaction.id,
input=[
{
"type": "function_result",
"name": tool_call.name,
"call_id": tool_call.id,
"result": [
{"type": "text", "text": "instrument.jpg"},
{
"type": "image",
"mime_type": "image/jpeg",
"data": base64_image_data,
},
],
}
],
)
JavaScript
// ✅ Include multimodal content in the function response
const finalInteraction = await client.interactions.create({
model: "gemini-3.5-flash",
previousInteractionId: interaction.id,
input: [{
type: "function_result",
name: toolCall.name,
call_id: toolCall.id,
result: [
{ type: "text", text: "instrument.jpg" },
{
type: "image",
mime_type: "image/jpeg",
data: base64ImageData,
},
],
}],
});
Java
import java.util.Arrays;
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.FunctionResultStep;
import com.google.genai.gaos.models.interactions.FunctionResultStepResultUnion;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.List;
Client client = new Client();
// Assumes interactionId, callId, functionName, and base64ImageData from previous step
String interactionId = "interaction-id-123";
String callId = "call-id-123";
String functionName = "get_instrument_image";
String base64ImageData = "iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk+A8AAQUBAScY42YAAAAASUVORK5CYII=";
// ✅ Include multimodal content in the function response
FunctionResultStep functionResult =
FunctionResultStep.builder()
.name(functionName)
.callId(callId)
.result(
FunctionResultStepResultUnion.of(
Arrays.asList(
TextContent.builder().text("instrument.jpg").build(),
ImageContent.builder()
.mimeType(ImageContentMimeType.IMAGE_JPEG)
.data(base64ImageData)
.build())))
.build();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.5-flash"))
.previousInteractionId(interactionId)
.input(InteractionsInput.ofStep(Arrays.asList(functionResult)))
.build();
Interaction finalInteraction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
Go
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
// Assumes interactionID, callID, functionName, and base64ImageData from previous step
interactionID := "interaction-id-123"
callID := "call-id-123"
functionName := "get_instrument_image"
base64ImageData := "iVBORw0KGgoAAAANSUhEUgAAAAEAAAABCAQAAAC1HAwCAAAAC0lEQVR42mNk+A8AAQUBAScY42YAAAAASUVORK5CYII="
// ✅ Include multimodal content in the function response
functionResult := interactions.NewStep(interactions.FunctionResultStep{
Name: genai.Ptr(functionName),
CallID: callID,
Result: interactions.NewFunctionResultStepResultUnion([]interactions.FunctionResultSubcontent{
interactions.NewFunctionResultSubcontent(interactions.TextContent{
Text: "instrument.jpg",
}),
interactions.NewFunctionResultSubcontent(interactions.ImageContent{
MimeType: interactions.ImageContentMimeType("image/jpeg").ToPointer(),
Data: genai.Ptr(base64ImageData),
}),
}),
})
finalRes, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.5-flash"),
PreviousInteractionID: genai.Ptr(interactionID),
Input: interactions.NewInteractionsInput([]interactions.Step{functionResult}),
}),
})
if err != nil {
log.Fatal(err)
}
_ = finalRes
}
Inline instructions in function responses
We often see clients provide additional instructions along with function responses
as subsequent Parts. This can lead to unexpected model behavior (e.g.
thought leakage) and result in lower quality outputs. Instead, append any extra
instructions to the end of the function response text separated by two newlines.
Python
# ✅ Append inline instructions to the end of the function response separated by two newlines
result_text = f"{json.dumps(result)}\n\n<your inline instructions>"
final_interaction = client.interactions.create(
model="gemini-3.5-flash",
previous_interaction_id=interaction.id,
tools=[my_tool],
input=[{
"type": "function_result",
"name": fc_step.name,
"call_id": fc_step.id,
"result": [{"type": "text", "text": result_text}],
}],
)
JavaScript
// ✅ Append inline instructions to the end of the function response separated by two newlines
const resultText = `${JSON.stringify(result)}\n\n<your inline instructions>`;
const finalInteraction = await client.interactions.create({
model: "gemini-3.5-flash",
previousInteractionId: interaction.id,
tools: [myTool],
input: [{
type: "function_result",
name: fcStep.name,
call_id: fcStep.id,
result: [{ type: "text", text: resultText }],
}],
});
Java
import java.util.Arrays;
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Function;
import com.google.genai.gaos.models.interactions.FunctionResultStep;
import com.google.genai.gaos.models.interactions.FunctionResultStepResultUnion;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.util.List;
Client client = new Client();
// Assumes interactionId, callId, functionName, myTool, and resultJson from previous step
String interactionId = "interaction-id-123";
String callId = "call-id-123";
String functionName = "get_weather";
Function myTool = Function.builder().name(functionName).build();
String resultJson = "{\"temperature\": \"72F\"}";
// ✅ Append inline instructions to the end of the function response separated by two newlines
String resultText = resultJson + "\n\n<your inline instructions>";
FunctionResultStep functionResult =
FunctionResultStep.builder()
.name(functionName)
.callId(callId)
.result(
FunctionResultStepResultUnion.of(
Arrays.asList(TextContent.builder().text(resultText).build())))
.build();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.5-flash"))
.previousInteractionId(interactionId)
.tools(Arrays.asList(myTool))
.input(InteractionsInput.ofStep(Arrays.asList(functionResult)))
.build();
Interaction finalInteraction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
Go
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
// Assumes interactionID, callID, functionName, myTool, and resultJSON from previous step
interactionID := "interaction-id-123"
callID := "call-id-123"
functionName := "get_weather"
myTool := interactions.NewTool(interactions.Function{Name: genai.Ptr(functionName)})
resultJSON := `{"temperature": "72F"}`
// ✅ Append inline instructions to the end of the function response separated by two newlines
resultText := resultJSON + "\n\n<your inline instructions>"
functionResult := interactions.NewStep(interactions.FunctionResultStep{
Name: genai.Ptr(functionName),
CallID: callID,
Result: interactions.NewFunctionResultStepResultUnion([]interactions.FunctionResultSubcontent{
interactions.NewFunctionResultSubcontent(interactions.TextContent{Text: resultText}),
}),
})
finalRes, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("gemini-3.5-flash"),
PreviousInteractionID: genai.Ptr(interactionID),
Tools: []interactions.Tool{myTool},
Input: interactions.NewInteractionsInput([]interactions.Step{functionResult}),
}),
})
if err != nil {
log.Fatal(err)
}
_ = finalRes
}
Reducing unnecessary tool calls
If you experience an overuse of tool calls, two techniques help minimize them:
Start by reducing the thinking level (
medium,low, orminimal): Higher thinking levels encourage the model to use more tools to explore and verify, so lowering the level can reduce tool calls.Add a system instruction: If overuse persists after adjusting the thinking level, consider a prompt that restricts tool usage. For example:
You have a limited action budget of <n> tool calls. Use them efficiently.
Gemini 3 family features
Gemini 3.5 Flash inherits all Gemini 3 family capabilities, including Computer Use. Features introduced in Gemini 3 that carry forward:
- Thinking: Encrypted reasoning context preserved across API calls. Automatic in the Interactions API; implicit in GenerateContent.
- Structured outputs with tools: Combine JSON mode with built-in tools (Search, URL context, code execution, function calling).
- Multimodal function responses: Return images, audio, and other media in function call results.
- Code execution with images: Execute code that processes and generates images.
- Combined tool use: Use built-in tools and custom function calling in the same request.
- Media resolution:
Fine-grained control over token allocation for image, video, and PDF inputs.
Gemini 3 models support per-content-item resolution settings (
low,medium,high,ultra_high) for mixed-fidelity prompts. - Thought signatures: Encrypted representations of the model's internal reasoning. Required for multi-turn function calling in stateless mode; managed automatically by the Interactions API and the official SDKs.
Prompting best practices
Gemini 3.x models are reasoning models, which changes how you should prompt.
- Precise instructions: Be concise. Gemini 3.x responds best to direct, clear instructions. Verbose or complex prompt engineering techniques designed for older models may cause the model to over-analyze.
- Output verbosity: By default, Gemini 3 and 3.1 is less verbose and prefers direct, efficient answers. If your use case requires a conversational tone, steer the model explicitly in your prompt (for example, "Explain this as a friendly, talkative assistant").
- Context management: When working with large datasets (such as entire books, codebases, or long videos), place your specific instructions or questions at the end of the prompt, after the data context. Anchor the model's reasoning by starting your question with a phrase like, "Based on the preceding information...".
Learn more about prompt design strategies in the prompt engineering guide.
Limitations
- Image segmentation is not supported in Gemini 3.x. For segmentation workloads, continue using Gemini 2.5 Flash with thinking off.
FAQ
What is the knowledge cutoff for Gemini 3.5 Flash? Gemini 3.5 Flash has a knowledge cutoff of January 2025. For more recent information, use the Search Grounding tool.
What are the context window limits? Gemini 3.5 Flash supports a 1 million token input context window and up to 65k output tokens.
Will my old
thinking_budgetcode still work? Yes,thinking_budgetis still supported for backward compatibility, but we recommend migrating tothinking_levelfor more predictable performance. Don't use both in the same request.Does Gemini 3.5 Flash support the Batch API? Yes. See the Batch API guide for details.
Is Context Caching supported? Yes, Context Caching is supported.
Which tools are supported? Gemini 3.5 Flash supports Google Search, Grounding with Google Maps, File Search, Code Execution, URL Context, and standard Function Calling, including combined tool use, and Computer Use.
Other models
What's next
- Learn more about prompt design strategies in the prompt engineering guide.
- Get started with the Gemini 3 Cookbook
- Learn about Gemini API optimization and inference