تعالج نماذج Gemini ونماذج الذكاء الاصطناعي التوليدي الأخرى الإدخالات والمخرجات بدقة على مستوى وحدة تُعرف باسم الرمز المميز.
في نماذج Gemini، يعادل الرمز المميز الواحد 4 أحرف تقريبًا. ويعادل 100 رمز مميز من 60 إلى 80 كلمة باللغة الإنجليزية تقريبًا.
لمحة عن الرموز المميّزة
يمكن أن تكون الرموز المميّزة أحرفًا مفردة، مثل z، أو كلمات كاملة، مثل cat. ويتم تقسيم الكلمات الطويلة إلى عدة رموز مميّزة. تُعرف مجموعة جميع الرموز المميّزة التي يستخدمها النموذج باسم المفردات، وتُعرف عملية تقسيم النص إلى رموز مميّزة باسم الترميز.
عند تفعيل الفوترة، يتم تحديد تكلفة طلب إلى Gemini API جزئيًا من خلال عدد الرموز المميّزة للإدخال والإخراج، لذا قد يكون من المفيد معرفة كيفية عدّ الرموز المميّزة.
عدّ الرموز المميّزة
يتم ترميز جميع الإدخالات والمخرجات من Gemini API، بما في ذلك النصوص وملفات الصور والوسائط الأخرى غير النصية.
يمكنك عدّ الرموز المميّزة بالطرق التالية:
استدعاء
count_tokensباستخدام إدخال الطلب. تعرض هذه الدالة إجمالي عدد الرموز المميّزة في الإدخال فقط. يمكنك إجراء هذا الاستدعاء قبل إرسال الإدخال للتحقق من حجم طلباتك.استخدام الحقل
usageفي استجابة التفاعل. تعرض هذه الدالة عدد الرموز المميّزة للإدخال (total_input_tokens) والإخراج (total_output_tokens) والتفكير (total_thought_tokens) والمحتوى المخزّن مؤقتًا (total_cached_tokens) واستخدام الأدوات (total_tool_use_tokens) والإجمالي (total_tokens).
عدّ الرموز المميّزة النصية
Python
# This will only work for SDK newer than 2.0.0
from google import genai
client = genai.Client()
prompt = "The quick brown fox jumps over the lazy dog."
# Count tokens before sending
total_tokens = client.models.count_tokens(
model="gemini-3.8-flash",
contents=prompt
)
print("total_tokens:", total_tokens.total_tokens)
# Get usage from interaction
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=prompt
)
print(interaction.usage)
JavaScript
// This will only work for SDK newer than 2.0.0
import { GoogleGenAI } from '@google/genai';
const client = new GoogleGenAI({});
const prompt = "The quick brown fox jumps over the lazy dog.";
// Count tokens before sending
const countResponse = await client.models.countTokens({
model: "gemini-3.8-flash",
contents: prompt,
});
console.log(countResponse.totalTokens);
// Get usage from interaction
const interaction = await client.interactions.create({
model: "gemini-3.8-flash",
input: prompt,
});
console.log(interaction.usage);
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import com.google.genai.types.CountTokensResponse;
Client client = new Client();
String prompt = "The quick brown fox jumps over the lazy dog.";
// Count tokens before sending
CountTokensResponse countResponse =
client.models.countTokens("gemini-3.8-flash", prompt, null);
System.out.println("total_tokens: " + countResponse.totalTokens().orElse(0));
// Get usage from interaction
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.8-flash"))
.input(InteractionsInput.of(prompt))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println(interaction.usage().orElse(null));
REST
# Specifies the API revision to avoid breaking changes when they become default
curl -X POST "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.8-flash:countTokens" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"contents": [{"parts": [{"text": "The quick brown fox."}]}]}'
عدّ الرموز المميّزة للمحادثة المترابطة
يمكنك عدّ الرموز المميّزة في سجلّ المحادثات باستخدام previous_interaction_id:
Python
# This will only work for SDK newer than 2.0.0
# First interaction
interaction1 = client.interactions.create(
model="gemini-3.8-flash",
input="Hi, my name is Bob"
)
# Second interaction continues the conversation
interaction2 = client.interactions.create(
model="gemini-3.8-flash",
input="What's my name?",
previous_interaction_id=interaction1.id
)
# Usage includes tokens from both turns
print(f"Input tokens: {interaction2.usage.total_input_tokens}")
print(f"Output tokens: {interaction2.usage.total_output_tokens}")
print(f"Total tokens: {interaction2.usage.total_tokens}")
JavaScript
// This will only work for SDK newer than 2.0.0
// First interaction
const interaction1 = await client.interactions.create({
model: "gemini-3.8-flash",
input: "Hi, my name is Bob"
});
// Second interaction continues the conversation
const interaction2 = await client.interactions.create({
model: "gemini-3.8-flash",
input: "What's my name?",
previous_interaction_id: interaction1.id
});
console.log(`Input tokens: ${interaction2.usage.total_input_tokens}`);
console.log(`Output tokens: ${interaction2.usage.total_output_tokens}`);
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.Usage;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
Client client = new Client();
// First interaction
CreateModelInteraction params1 =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.8-flash"))
.input(InteractionsInput.of("Hi, my name is Bob"))
.build();
Interaction interaction1 =
client.interactions.create(CreateInteractionRequestBody.of(params1)).interaction().get();
// Second interaction continues the conversation
CreateModelInteraction params2 =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.8-flash"))
.input(InteractionsInput.of("What's my name?"))
.previousInteractionId(interaction1.id().orElse(""))
.build();
Interaction interaction2 =
client.interactions.create(CreateInteractionRequestBody.of(params2)).interaction().get();
// Usage includes tokens from both turns
if (interaction2.usage().isPresent()) {
Usage usage = interaction2.usage().get();
System.out.println("Input tokens: " + usage.totalInputTokens().orElse(0));
System.out.println("Output tokens: " + usage.totalOutputTokens().orElse(0));
System.out.println("Total tokens: " + usage.totalTokens().orElse(0));
}
عدّ الرموز المميّزة المتعددة الوسائط
يتم ترميز جميع الإدخالات إلى Gemini API، بما في ذلك الصور والفيديوهات والتسجيلات الصوتية. في ما يلي نقاط أساسية حول الترميز:
- الصور: يتم احتساب الصور التي يبلغ حجمها 384×384 بكسل أو أقل في كلا البُعدَين على أنّها 258 رمزًا مميزًا. ويتم تقسيم الصور الأكبر حجمًا إلى مربّعات بحجم 768×768 بكسل، ويتم احتساب كل مربّع على أنّه 258 رمزًا مميزًا.
- الفيديوهات: يتم احتساب 263 رمزًا مميزًا في الثانية (ينطبق ذلك على المعالجة الثابتة). بالنسبة إلى المعالجة المستندة إلى الذكاء الاصطناعي الوكيل، يختلف استخدام الرموز المميّزة. يمكنك الاطّلاع على استخدام الرموز المميّزة للفيديوهات حسب وضع المعالجة.
- التسجيلات الصوتية: يتم احتساب 32 رمزًا مميزًا في الثانية
الرموز المميّزة للصور
Python
# This will only work for SDK newer than 2.0.0
uploaded_file = client.files.upload(file="path/to/image.jpg")
# Count tokens for image + text
total_tokens = client.models.count_tokens(
model="gemini-3.8-flash",
contents=["Tell me about this image", uploaded_file]
)
print(f"Total tokens: {total_tokens}")
# Generate with image
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": "Tell me about this image"},
{"type": "image", "uri": uploaded_file.uri, "mime_type": uploaded_file.mime_type}
]
)
print(interaction.usage)
JavaScript
// This will only work for SDK newer than 2.0.0
const uploadedFile = await client.files.upload({
file: "path/to/image.jpg",
config: { mimeType: "image/jpeg" }
});
// Count tokens
const countResponse = await client.models.countTokens({
model: "gemini-3.8-flash",
contents: [
{ text: "Tell me about this image" },
{ fileData: { fileUri: uploadedFile.uri, mimeType: uploadedFile.mimeType } }
]
});
console.log(countResponse.totalTokens);
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import com.google.genai.types.Content;
import com.google.genai.types.CountTokensResponse;
import com.google.genai.types.File;
import com.google.genai.types.Part;
import com.google.genai.types.UploadFileConfig;
import java.util.Arrays;
Client client = new Client();
File uploadedFile =
client.files.upload(
new java.io.File("path/to/image.jpg"),
UploadFileConfig.builder().mimeType("image/jpeg").build());
// Count tokens for image + text
CountTokensResponse countResponse =
client.models.countTokens(
"gemini-3.8-flash",
Arrays.asList(
Content.fromParts(
Part.fromText("Tell me about this image"),
Part.fromUri(
uploadedFile.uri().orElse(""), uploadedFile.mimeType().orElse("image/jpeg")))),
null);
System.out.println("Total tokens: " + countResponse.totalTokens().orElse(0));
// Generate with image
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("gemini-3.8-flash"))
.input(
InteractionsInput.ofContent(
Arrays.asList(
TextContent.builder().text("Tell me about this image").build(),
ImageContent.builder()
.uri(uploadedFile.uri().orElse(""))
.mimeType(
ImageContentMimeType.of(uploadedFile.mimeType().orElse("image/jpeg")))
.build())))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
System.out.println(interaction.usage().orElse(null));
مثال على البيانات المضمّنة:
Python
# This will only work for SDK newer than 2.0.0
import base64
with open('image.jpg', 'rb') as f:
image_bytes = f.read()
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": "Describe this image"},
{
"type": "image",
"data": base64.b64encode(image_bytes).decode('utf-8'),
"mime_type": "image/jpeg"
}
]
)
print(interaction.usage)
الرموز المميّزة للفيديوهات
Python
# This will only work for SDK newer than 2.0.0
import time
video_file = client.files.upload(file="path/to/video.mp4")
while not video_file.state or video_file.state.name != "ACTIVE":
print("Processing video...")
time.sleep(5)
video_file = client.files.get(name=video_file.name)
# A 60-second video is approximately 100 * 60 = 6,000 tokens
total_tokens = client.models.count_tokens(
model="gemini-3.8-flash",
contents=["Summarize this video", video_file]
)
print(f"Total tokens: {total_tokens}")
# Generate with video
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": "Summarize this video"},
{"type": "video", "uri": video_file.uri, "mime_type": video_file.mime_type}
]
)
print(interaction.usage)
استخدام الرموز المميّزة للفيديوهات حسب وضع المعالجة
يعتمد استخدام الرموز المميّزة للفيديوهات على وضع المعالجة:
| وضع المعالجة | احتساب الرموز المميّزة | معدّل الاستخدام |
|---|---|---|
| ثابت (تلقائي) | 100 رمز مميز في الثانية تقريبًا بشكلٍ تلقائي (دقة منخفضة) أو 300 رمز مميز في الثانية تقريبًا (دقة عالية) يتم أخذ عيّنات من جميع اللقطات بمعدّل لقطة واحدة في الثانية. | يمكن توقّعه، ويتناسب مع مدة الفيديو. |
| مستند إلى الذكاء الاصطناعي الوكيل | يختلف حسب مدى تعقيد المحتوى. لا يحمِّل النموذج سوى النص و/أو اللقطات و/أو التسجيل الصوتي اللازم للإجابة عن الطلب. | يتم استخدام رموز مميّزة أقل بنسبة تصل إلى% 88 للمحتوى الطويل. |
باستخدام المعالجة المستندة إلى الذكاء الاصطناعي الوكيل، قد تستخدم محاضرة مدتها ساعة واحدة حوالي 108 ألف رمز مميز، بينما قد تستخدم حوالي 1.08 مليون رمز مميز في الوضع الثابت، وذلك حسب الطلب والمحتوى.
لمعرفة استخدام الرموز المميّزة الفعلي لطلب معيّن، يمكنك فحص interaction.usage. يتم الإبلاغ عن الرموز المميّزة للفيديوهات المستندة إلى الذكاء الاصطناعي الوكيل في الحقول التالية:
- الطلب الأولي (مرجع الفيديو + طلب المستخدم):
total_input_tokens - التفكير في التنقّل:
total_thought_tokens - النص واللقطات والتسجيل الصوتي الذي يتم تحميله عند الطلب:
total_tool_use_tokens - الإجابة النهائية:
total_output_tokens
الرموز المميّزة للتسجيلات الصوتية
Python
# This will only work for SDK newer than 2.0.0
audio_file = client.files.upload(file="path/to/audio.mp3")
# A 60-second audio clip is approximately 32 * 60 = 1,920 tokens
total_tokens = client.models.count_tokens(
model="gemini-3.8-flash",
contents=["Transcribe this audio", audio_file]
)
print(f"Total tokens: {total_tokens}")
# Generate with audio
interaction = client.interactions.create(
model="gemini-3.8-flash",
input=[
{"type": "text", "text": "Transcribe this audio"},
{"type": "audio", "uri": audio_file.uri, "mime_type": audio_file.mime_type}
]
)
print(interaction.usage)
عدّ الرموز المميّزة لتعليمات النظام
يتم احتساب تعليمات النظام كجزء من الرموز المميّزة للإدخال:
Python
# This will only work for SDK newer than 2.0.0
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="Hello!",
system_instruction="You are a helpful assistant who speaks like a pirate."
)
# system_instruction tokens included in total_input_tokens
print(f"Input tokens: {interaction.usage.total_input_tokens}")
عدّ الرموز المميّزة للأدوات
يتم أيضًا احتساب الأدوات (الدوال وتنفيذ التعليمات البرمجية و"بحث Google"):
Python
# This will only work for SDK newer than 2.0.0
tools = [
{
"type": "function",
"name": "get_weather",
"description": "Get current weather",
"parameters": {
"type": "object",
"properties": {
"location": {"type": "string"}
}
}
}
]
interaction = client.interactions.create(
model="gemini-3.8-flash",
input="What's the weather in Tokyo?",
tools=tools
)
print(f"Input tokens: {interaction.usage.total_input_tokens}")
print(f"Tool use tokens: {interaction.usage.total_tool_use_tokens}")
قدرة الاستيعاب
لكل نموذج من نماذج Gemini عدد أقصى من الرموز المميّزة التي يمكنه معالجتها. وتحدّد قدرة الاستيعاب الحدّ المجمّع للرموز المميّزة للإدخال والإخراج.
الحصول على حجم قدرة الاستيعاب آليًا
Python
# This will only work for SDK newer than 2.0.0
model_info = client.models.get(model="gemini-3.8-flash")
print(f"Input token limit: {model_info.input_token_limit}")
print(f"Output token limit: {model_info.output_token_limit}")
JavaScript
// This will only work for SDK newer than 2.0.0
const modelInfo = await client.models.get({ model: "gemini-3.8-flash" });
console.log(`Input token limit: ${modelInfo.inputTokenLimit}`);
console.log(`Output token limit: ${modelInfo.outputTokenLimit}`);
Java
import com.google.genai.Client;
import com.google.genai.types.Model;
Client client = new Client();
Model modelInfo = client.models.get("gemini-3.8-flash", null);
System.out.println("Input token limit: " + modelInfo.inputTokenLimit().orElse(0));
System.out.println("Output token limit: " + modelInfo.outputTokenLimit().orElse(0));
يمكنك الاطّلاع على أحجام قدرة الاستيعاب في صفحة النماذج.
الخطوات التالية
- إنشاء النص: أساسيات الإنشاء
- التخزين المؤقت: خفض التكاليف باستخدام التخزين المؤقت
- التسعير: التعرّف على التكاليف