O Lyria 3.5 é a família de modelos de geração de música do Google, disponível pela API Gemini. Com o Lyria 3.5, é possível gerar áudio estéreo de alta qualidade em 44, 1 kHz com base em comandos de texto ou imagens. Esses modelos oferecem coerência estrutural, incluindo vocais, letras sincronizadas e arranjos instrumentais completos.
A família Lyria inclui os seguintes modelos:
| Modelo | ID do modelo | Ideal para | Duração | Saída |
|---|---|---|---|---|
| Lyria 3 Clip | lyria-3-clip-preview |
Clipes curtos, loops, prévias | 30 segundos | MP3 |
| Lyria 3.5 | lyria-3.5 |
Músicas completas com versos, refrões e pontes | Alguns minutos (controláveis usando o comando) | MP3 |
Os dois modelos podem ser usados com a nova API Interactions, que aceita entradas multimodais (texto e imagens) e produz áudio estéreo de alta fidelidade de 44,1 kHz.
Gerar um videoclipe
O modelo Lyria 3 Clip sempre gera um clipe de 30 segundos. Para gerar um
clipe, chame o método interactions.create com um comando de texto. A resposta sempre inclui a letra e a estrutura da música geradas, além do áudio no esquema steps.
Python
import base64
from google import genai
client = genai.Client()
interaction = client.interactions.create(
model="lyria-3-clip-preview",
input="A short instrumental acoustic guitar piece.",
)
generated_audio = interaction.output_audio
if generated_audio:
with open("music.mp3", "wb") as f:
f.write(base64.b64decode(generated_audio.data))
lyrics = interaction.output_text
if lyrics:
print(f"Lyrics:\n{lyrics}")
JavaScript
import { GoogleGenAI } from '@google/genai';
import * as fs from 'fs';
const client = new GoogleGenAI({});
const interaction = await client.interactions.create({
model: 'lyria-3-clip-preview',
input: 'A short instrumental acoustic guitar piece.',
});
const generatedAudio = interaction.output_audio;
if (generatedAudio) {
fs.writeFileSync('music.mp3', Buffer.from(generatedAudio.data, 'base64'));
}
const lyrics = interaction.output_text;
if (lyrics) {
console.log(`Lyrics:\n${lyrics}`);
}
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.Base64;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("lyria-3-clip-preview"))
.input(InteractionsInput.of("A short instrumental acoustic guitar piece."))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
if (interaction.outputAudio().isPresent() && interaction.outputAudio().get().data().isPresent()) {
byte[] audioBytes = Base64.getDecoder().decode(interaction.outputAudio().get().data().get());
Files.write(Paths.get("music.mp3"), audioBytes);
}
interaction.outputText().ifPresent(lyrics -> System.out.println("Lyrics:\n" + lyrics));
Go
package main
import (
"context"
"encoding/base64"
"fmt"
"log"
"os"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("lyria-3-clip-preview"),
Input: interactions.NewInteractionsInput("A short instrumental acoustic guitar piece."),
}),
})
if err != nil {
log.Fatal(err)
}
if res.Interaction.OutputAudio != nil && res.Interaction.OutputAudio.Data != nil {
audioBytes, err := base64.StdEncoding.DecodeString(*res.Interaction.OutputAudio.Data)
if err != nil {
log.Fatal(err)
}
if err := os.WriteFile("music.mp3", audioBytes, 0644); err != nil {
log.Fatal(err)
}
}
if res.Interaction.OutputText != nil {
fmt.Printf("Lyrics:\n%s\n", *res.Interaction.OutputText)
}
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "Content-Type: application/json" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-d '{
"model": "lyria-3-clip-preview",
"input": "A short instrumental acoustic guitar piece."
}'
É possível recuperar os dados de música gerados usando a propriedade interaction.output_audio, que retorna o último bloco de áudio gerado. Também é possível recuperar
a letra e a estrutura da música usando a propriedade interaction.output_text. Para detalhes sobre propriedades de conveniência, consulte a
Visão geral das interações.
Gerar uma música completa
Use o modelo lyria-3.5 para gerar músicas completas que duram alguns minutos. O modelo Pro entende a estrutura musical e pode criar composições com versos, refrões e pontes distintos. Você pode influenciar a
duração especificando-a no comando (por exemplo, "crie uma música de 2 minutos") ou usando
carimbos de data/hora para definir a estrutura.
Python
interaction = client.interactions.create(
model="lyria-3.5",
input="An epic cinematic orchestral piece about a journey home. Starts with a solo piano intro, builds through sweeping strings, and climaxes with a massive wall of sound.",
)
JavaScript
const interaction = await client.interactions.create({
model: 'lyria-3.5',
input: 'A beautiful piano melody.',
});
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("lyria-3.5"))
.input(
InteractionsInput.of(
"An epic cinematic orchestral piece about a journey home. Starts with a solo piano intro, builds through sweeping strings, and climaxes with a massive wall of sound."))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
Go
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("lyria-3.5"),
Input: interactions.NewInteractionsInput(
"An epic cinematic orchestral piece about a journey home. Starts with a solo piano intro, builds through sweeping strings, and climaxes with a massive wall of sound.",
),
}),
})
if err != nil {
log.Fatal(err)
}
_ = res
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "Content-Type: application/json" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-d '{
"model": "lyria-3.5",
"input": "A beautiful piano melody."
}'
Selecionar formato de saída
Por padrão, os modelos do Lyria 3.5 geram áudio no formato MP3. Para o Lyria 3.5, também é possível solicitar a saída no formato WAV definindo o response_format.
Python
interaction = client.interactions.create(
model="lyria-3.5",
input="A beautiful piano melody.",
response_format={"type": "audio"},
)
JavaScript
const interaction = await client.interactions.create({
model: 'lyria-3.5',
input: 'A beautiful piano melody.',
response_format: {
type: 'audio',
},
});
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.AudioResponseFormat;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.CreateModelInteractionResponseFormat;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.ResponseFormat;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("lyria-3.5"))
.input(InteractionsInput.of("A beautiful piano melody."))
.responseFormat(
CreateModelInteractionResponseFormat.of(
ResponseFormat.of(AudioResponseFormat.builder().build())))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
Go
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("lyria-3.5"),
Input: interactions.NewInteractionsInput("A beautiful piano melody."),
ResponseFormat: genai.Ptr(interactions.NewCreateModelInteractionResponseFormat(
interactions.NewResponseFormat(interactions.AudioResponseFormat{}),
)),
}),
})
if err != nil {
log.Fatal(err)
}
_ = res
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "lyria-3.5",
"input": "A beautiful piano melody.",
"response_format": {
"type": "audio"
}
}'
Analise a resposta
A resposta da Lyria 3.5 contém vários blocos de conteúdo no esquema steps.
As interações retornam uma sequência de etapas, em que as etapas model_output contêm o
conteúdo gerado.
Os blocos de conteúdo de texto contêm a letra gerada ou uma descrição JSON da estrutura da música.
Os blocos de conteúdo do tipo audio contêm os dados de áudio codificados em base64.
Python
lyrics = []
audio_data = None
generated_audio = interaction.output_audio
if generated_audio:
with open("output.mp3", "wb") as f:
f.write(base64.b64decode(generated_audio.data))
lyrics = interaction.output_text
if lyrics:
print(f"Lyrics:\n{lyrics}")
JavaScript
const lyrics = [];
let audioData = null;
const generatedAudio = interaction.output_audio;
if (generatedAudio) {
fs.writeFileSync("output.mp3", Buffer.from(generatedAudio.data, 'base64'));
}
const lyrics = interaction.output_text;
if (lyrics) {
console.log("Lyrics:\n" + lyrics);
}
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.Base64;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("lyria-3.5"))
.input(InteractionsInput.of("A song about a starry night."))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
if (interaction.outputAudio().isPresent() && interaction.outputAudio().get().data().isPresent()) {
byte[] audioBytes = Base64.getDecoder().decode(interaction.outputAudio().get().data().get());
Files.write(Paths.get("output.mp3"), audioBytes);
}
if (interaction.outputText().isPresent()) {
System.out.println("Lyrics:\n" + interaction.outputText().get());
}
Go
package main
import (
"context"
"encoding/base64"
"fmt"
"log"
"os"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("lyria-3.5"),
Input: interactions.NewInteractionsInput("A song about a starry night."),
}),
})
if err != nil {
log.Fatal(err)
}
if res.Interaction.OutputAudio != nil && res.Interaction.OutputAudio.Data != nil {
audioBytes, err := base64.StdEncoding.DecodeString(*res.Interaction.OutputAudio.Data)
if err != nil {
log.Fatal(err)
}
if err := os.WriteFile("output.mp3", audioBytes, 0644); err != nil {
log.Fatal(err)
}
}
if res.Interaction.OutputText != nil {
fmt.Printf("Lyrics:\n%s\n", *res.Interaction.OutputText)
}
}
REST
# The output from the REST API is a JSON object containing base64 encoded data.
# You can extract the text or the audio data using a tool like jq.
# To extract the audio and save it to a file:
curl ... | jq -r '.steps[] | select(.type=="model_output") | .content[] | select(.type=="audio") | .data' | base64 -d > output.mp3
Letras e músicas intercaladas
Como a saída do Lyria 3.5 é complexa, contendo etapas e blocos separados para letras geradas (texto) e a música em si (áudio), as propriedades de conveniência oferecem um atalho rápido e recomendado.
No entanto, se você quiser controle programático total sobre a linha do tempo bruta de etapas
retornadas pelo servidor (como registrar blocos de conteúdo individuais à medida que são
recebidos), itere manualmente em steps:
Python
lyrics = []
audio_data = None
for step in interaction.steps:
if step.type == "model_output":
for content_block in step.content:
if content_block.type == "audio":
audio_data = base64.b64decode(content_block.data)
elif content_block.type == "text":
lyrics.append(content_block.text)
if lyrics:
print("Lyrics:\n" + "\n".join(lyrics))
if audio_data:
with open("output.mp3", "wb") as f:
f.write(audio_data)
JavaScript
const lyrics = [];
let audioData = null;
for (const step of interaction.steps) {
if (step.type === 'model_output') {
for (const contentBlock of step.content) {
if (contentBlock.type === 'audio') {
audioData = Buffer.from(contentBlock.data, 'base64');
} else if (contentBlock.type === 'text') {
lyrics.push(contentBlock.text);
}
}
}
}
if (lyrics.length) {
console.log("Lyrics:\n" + lyrics.join("\n"));
}
if (audioData) {
fs.writeFileSync("output.mp3", audioData);
}
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.AudioContent;
import com.google.genai.gaos.models.interactions.Content;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.ModelOutputStep;
import com.google.genai.gaos.models.interactions.Step;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.ArrayList;
import java.util.Base64;
import java.util.List;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("lyria-3.5"))
.input(InteractionsInput.of("A song about a starry night."))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
List<String> lyrics = new ArrayList<>();
byte[] audioData = null;
if (interaction.steps().isPresent()) {
for (Step step : interaction.steps().get()) {
if (step instanceof ModelOutputStep) {
ModelOutputStep outputStep = (ModelOutputStep) step;
if (outputStep.content().isPresent()) {
for (Content contentBlock : outputStep.content().get()) {
if (contentBlock instanceof AudioContent) {
AudioContent audioBlock = (AudioContent) contentBlock;
if (audioBlock.data().isPresent()) {
audioData = Base64.getDecoder().decode(audioBlock.data().get());
}
} else if (contentBlock instanceof TextContent) {
TextContent textBlock = (TextContent) contentBlock;
textBlock.text().ifPresent(lyrics::add);
}
}
}
}
}
}
if (!lyrics.isEmpty()) {
System.out.println("Lyrics:\n" + String.join("\n", lyrics));
}
if (audioData != null) {
Files.write(Paths.get("output.mp3"), audioData);
}
Go
package main
import (
"context"
"encoding/base64"
"fmt"
"log"
"os"
"strings"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("lyria-3.5"),
Input: interactions.NewInteractionsInput("A song about a starry night."),
}),
})
if err != nil {
log.Fatal(err)
}
var lyrics []string
var audioData []byte
for _, step := range res.Interaction.Steps {
if step.ModelOutputStep != nil {
for _, contentBlock := range step.ModelOutputStep.Content {
if contentBlock.AudioContent != nil && contentBlock.AudioContent.Data != nil {
decoded, err := base64.StdEncoding.DecodeString(*contentBlock.AudioContent.Data)
if err != nil {
log.Fatal(err)
}
audioData = decoded
} else if contentBlock.TextContent != nil {
lyrics = append(lyrics, contentBlock.TextContent.Text)
}
}
}
}
if len(lyrics) > 0 {
fmt.Printf("Lyrics:\n%s\n", strings.Join(lyrics, "\n"))
}
if audioData != nil {
if err := os.WriteFile("output.mp3", audioData, 0644); err != nil {
log.Fatal(err)
}
}
}
Gerar música com base em imagens
O Lyria 3.5 aceita entradas multimodais. Você pode fornecer até 10 imagens com seu comando de texto na lista input, e o modelo vai compor músicas inspiradas no conteúdo visual.
Python
import base64
with open("desert_sunset.jpg", "rb") as f:
image_bytes = f.read()
image_b64 = base64.b64encode(image_bytes).decode("utf-8")
response = client.interactions.create(
model="lyria-3.5",
input=[
{
"type": "text",
"text": "An atmospheric ambient track inspired by the mood and colors in this image.",
},
{
"type": "image",
"mime_type": "image/jpeg",
"data": image_b64,
},
],
)
JavaScript
import * as fs from "fs";
const imageBytes = fs.readFileSync("desert_sunset.jpg").toString("base64");
const interaction = await client.interactions.create({
model: "lyria-3.5",
input: [
{
type: "text",
text: "An atmospheric ambient track inspired by the mood and colors in this image.",
},
{
type: "image",
mime_type: "image/jpeg",
data: imageBytes,
},
],
});
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.Content;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.ImageContent;
import com.google.genai.gaos.models.interactions.ImageContentMimeType;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.interactions.TextContent;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
import java.nio.file.Files;
import java.nio.file.Paths;
import java.util.Arrays;
import java.util.Base64;
import java.util.List;
Client client = new Client();
byte[] imageBytes = Files.readAllBytes(Paths.get("desert_sunset.jpg"));
String imageB64 = Base64.getEncoder().encodeToString(imageBytes);
Content textContent =
TextContent.builder()
.text("An atmospheric ambient track inspired by the mood and colors in this image.")
.build();
Content imageContent =
ImageContent.builder()
.mimeType(ImageContentMimeType.IMAGE_JPEG)
.data(imageB64)
.build();
List<Content> contents = Arrays.asList(textContent, imageContent);
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("lyria-3.5"))
.input(InteractionsInput.ofContent(contents))
.build();
Interaction response =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
Go
package main
import (
"context"
"encoding/base64"
"log"
"os"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
imageBytes, err := os.ReadFile("desert_sunset.jpg")
if err != nil {
log.Fatal(err)
}
imageB64 := base64.StdEncoding.EncodeToString(imageBytes)
contents := []interactions.Content{
interactions.NewContent(interactions.TextContent{
Text: "An atmospheric ambient track inspired by the mood and colors in this image.",
}),
interactions.NewContent(interactions.ImageContent{
MimeType: interactions.ImageContentMimeTypeImageJpeg.ToPointer(),
Data: genai.Ptr(imageB64),
}),
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("lyria-3.5"),
Input: interactions.NewInteractionsInput(contents),
}),
})
if err != nil {
log.Fatal(err)
}
_ = res
}
REST
# Pass base64 encoded image data directly:
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "lyria-3.5",
"input": [
{"type": "text", "text": "An atmospheric ambient track inspired by the mood and colors in this image."},
{"type": "image", "mime_type": "image/jpeg", "data": "/9j/4AAQSkZJRgABAQEASABIAAD/2wBDAP//////////////////////////////////////////////////////////////////////////////////////wgALCAABAAEBAREA/8QAFBABAAAAAAAAAAAAAAAAAAAAAP/aAAgBAQABPxA="}
]
}'
Fornecer letras personalizadas
Você pode escrever suas próprias letras e incluí-las no comando. Use tags de seção
como [Verse], [Chorus] e [Bridge] para ajudar o modelo a entender a
estrutura da música:
Python
prompt = """
Create a dreamy indie pop song with the following lyrics:
[Verse 1]
Walking through the neon glow,
city lights reflect below,
every shadow tells a story,
every corner, fading glory.
[Chorus]
We are the echoes in the night,
burning brighter than the light,
hold on tight, don't let me go,
we are the echoes down below.
[Verse 2]
Footsteps lost on empty streets,
rhythms sync to heartbeats,
whispers carried by the breeze,
dancing through the autumn leaves.
"""
interaction = client.interactions.create(
model="lyria-3.5",
input=prompt,
)
JavaScript
const prompt = `
Create a dreamy indie pop song with the following lyrics:
[Verse 1]
Walking through the neon glow,
city lights reflect below,
every shadow tells a story,
every corner, fading glory.
[Chorus]
We are the echoes in the night,
burning brighter than the light,
hold on tight, don't let me go,
we are the echoes down below.
[Verse 2]
Footsteps lost on empty streets,
rhythms sync to heartbeats,
whispers carried by the breeze,
dancing through the autumn leaves.
`;
const interaction = await client.interactions.create({
model: 'lyria-3.5',
input: prompt,
});
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
Client client = new Client();
String prompt =
"Create a dreamy indie pop song with the following lyrics:\n\n"
+ "[Verse 1]\n"
+ "Walking through the neon glow,\n"
+ "city lights reflect below,\n"
+ "every shadow tells a story,\n"
+ "every corner, fading glory.\n\n"
+ "[Chorus]\n"
+ "We are the echoes in the night,\n"
+ "burning brighter than the light,\n"
+ "hold on tight, don't let me go,\n"
+ "we are the echoes down below.\n\n"
+ "[Verse 2]\n"
+ "Footsteps lost on empty streets,\n"
+ "rhythms sync to heartbeats,\n"
+ "whispers carried by the breeze,\n"
+ "dancing through the autumn leaves.";
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("lyria-3.5"))
.input(InteractionsInput.of(prompt))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
Go
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
prompt := "Create a dreamy indie pop song with the following lyrics:\n\n" +
"[Verse 1]\n" +
"Walking through the neon glow,\n" +
"city lights reflect below,\n" +
"every shadow tells a story,\n" +
"every corner, fading glory.\n\n" +
"[Chorus]\n" +
"We are the echoes in the night,\n" +
"burning brighter than the light,\n" +
"hold on tight, don't let me go,\n" +
"we are the echoes down below.\n\n" +
"[Verse 2]\n" +
"Footsteps lost on empty streets,\n" +
"rhythms sync to heartbeats,\n" +
"whispers carried by the breeze,\n" +
"dancing through the autumn leaves."
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("lyria-3.5"),
Input: interactions.NewInteractionsInput(prompt),
}),
})
if err != nil {
log.Fatal(err)
}
_ = res
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "lyria-3.5",
"input": "Create a dreamy indie pop song with the following lyrics: ..."
}'
Controlar a marcação de tempo e a estrutura
É possível especificar exatamente o que acontece em momentos específicos da música usando carimbos de data/hora. Isso é útil para controlar quando os instrumentos entram, quando as letras são entregues e como a música progride:
Python
prompt = """
[0:00 - 0:10] Intro: Begin with a soft lo-fi beat and muffled
vinyl crackle.
[0:10 - 0:30] Verse 1: Add a warm Fender Rhodes piano melody
and gentle vocals singing about a rainy morning.
[0:30 - 0:50] Chorus: Full band with upbeat drums and soaring
synth leads. The lyrics are hopeful and uplifting.
[0:50 - 1:00] Outro: Fade out with the piano melody alone.
"""
interaction = client.interactions.create(
model="lyria-3.5",
input=prompt,
)
JavaScript
const prompt = `
[0:00 - 0:10] Intro: Begin with a soft lo-fi beat and muffled
vinyl crackle.
[0:10 - 0:30] Verse 1: Add a warm Fender Rhodes piano melody
and gentle vocals singing about a rainy morning.
[0:30 - 0:50] Chorus: Full band with upbeat drums and soaring
synth leads. The lyrics are hopeful and uplifting.
[0:50 - 1:00] Outro: Fade out with the piano melody alone.
`;
const interaction = await client.interactions.create({
model: 'lyria-3.5',
input: prompt,
});
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
Client client = new Client();
String prompt =
"[0:00 - 0:10] Intro: Begin with a soft lo-fi beat and muffled vinyl crackle.\n"
+ "[0:10 - 0:30] Verse 1: Add a warm Fender Rhodes piano melody and gentle vocals singing about a rainy morning.\n"
+ "[0:30 - 0:50] Chorus: Full band with upbeat drums and soaring synth leads. The lyrics are hopeful and uplifting.\n"
+ "[0:50 - 1:00] Outro: Fade out with the piano melody alone.";
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("lyria-3.5"))
.input(InteractionsInput.of(prompt))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
Go
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
prompt := "[0:00 - 0:10] Intro: Begin with a soft lo-fi beat and muffled vinyl crackle.\n" +
"[0:10 - 0:30] Verse 1: Add a warm Fender Rhodes piano melody and gentle vocals singing about a rainy morning.\n" +
"[0:30 - 0:50] Chorus: Full band with upbeat drums and soaring synth leads. The lyrics are hopeful and uplifting.\n" +
"[0:50 - 1:00] Outro: Fade out with the piano melody alone."
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("lyria-3.5"),
Input: interactions.NewInteractionsInput(prompt),
}),
})
if err != nil {
log.Fatal(err)
}
_ = res
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "lyria-3.5",
"input": "[0:00 - 0:10] Intro: ..."
}'
Gerar músicas instrumentais
Para música de fundo, trilhas sonoras de jogos ou qualquer caso de uso em que não sejam necessários vocais, peça ao modelo para produzir músicas apenas instrumentais:
Python
interaction = client.interactions.create(
model="lyria-3-clip-preview",
input="A bright chiptune melody in C Major, retro 8-bit video game style. Instrumental only, no vocals.",
)
JavaScript
const interaction = await client.interactions.create({
model: 'lyria-3-clip-preview',
input: 'A bright chiptune melody in C Major, retro 8-bit video game style. Instrumental only, no vocals.',
});
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("lyria-3-clip-preview"))
.input(
InteractionsInput.of(
"A bright chiptune melody in C Major, retro 8-bit video game style. Instrumental only, no vocals."))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
Go
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("lyria-3-clip-preview"),
Input: interactions.NewInteractionsInput(
"A bright chiptune melody in C Major, retro 8-bit video game style. Instrumental only, no vocals.",
),
}),
})
if err != nil {
log.Fatal(err)
}
_ = res
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "lyria-3-clip-preview",
"input": "A bright chiptune melody in C Major, retro 8-bit video game style. Instrumental only, no vocals."
}'
Gerar músicas em diferentes idiomas
O Lyria 3.5 gera letras no idioma do seu comando. Para gerar uma música com letras em francês, escreva o comando nesse idioma. O modelo adapta o estilo vocal e a pronúncia para corresponder ao idioma.
Python
interaction = client.interactions.create(
model="lyria-3.5",
input="Crée une chanson pop romantique en français sur un coucher de soleil à Paris. Utilise du piano et de la guitare acoustique.",
)
JavaScript
const interaction = await client.interactions.create({
model: 'lyria-3.5',
input: 'Crée une chanson pop romantique en français sur un coucher de soleil à Paris. Utilise du piano et de la guitare acoustique.',
});
Java
import com.google.genai.Client;
import com.google.genai.gaos.models.interactions.CreateModelInteraction;
import com.google.genai.gaos.models.interactions.Interaction;
import com.google.genai.gaos.models.interactions.InteractionsInput;
import com.google.genai.gaos.models.interactions.Model;
import com.google.genai.gaos.models.operations.CreateInteractionRequestBody;
Client client = new Client();
CreateModelInteraction params =
CreateModelInteraction.builder()
.model(Model.of("lyria-3.5"))
.input(
InteractionsInput.of(
"Crée une chanson pop romantique en français sur un coucher de soleil à Paris. Utilise du piano et de la guitare acoustique."))
.build();
Interaction interaction =
client.interactions.create(CreateInteractionRequestBody.of(params)).interaction().get();
Go
package main
import (
"context"
"log"
"google.golang.org/genai"
"google.golang.org/genai/interactions/models/interactions"
"google.golang.org/genai/interactions/models/operations"
)
func main() {
ctx := context.Background()
client, err := genai.NewClient(ctx, nil)
if err != nil {
log.Fatal(err)
}
res, err := client.Interactions.Create(ctx, operations.CreateInteractionRequest{
Body: operations.NewCreateInteractionRequestBody(interactions.CreateModelInteraction{
Model: interactions.Model("lyria-3.5"),
Input: interactions.NewInteractionsInput(
"Crée une chanson pop romantique en français sur un coucher de soleil à Paris. Utilise du piano et de la guitare acoustique.",
),
}),
})
if err != nil {
log.Fatal(err)
}
_ = res
}
REST
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "lyria-3.5",
"input": "Crée une chanson pop romantique en français sur un coucher de soleil à Paris. Utilise du piano et de la guitare acoustique."
}'
Inteligência do modelo
O Lyria 3.5 analisa seu processo de comando em que o modelo usa a estrutura musical (introdução, verso, refrão, ponte etc.) com base no seu comando. Isso acontece antes da geração do áudio e garante coerência estrutural e musicalidade.
Guia para a criação de comandos
Para saber como criar comandos eficazes para gêneros musicais, instrumentos, estrutura de músicas, letras personalizadas e estilos de interpretação vocal, consulte o guia de comandos do Lyria.
Práticas recomendadas
- Itere primeiro com o Clipe. Use o modelo
lyria-3-clip-previewmais rápido para testar comandos antes de gerar um conteúdo completo comlyria-3.5. - Faça uma descrição específica. Comandos vagos produzem resultados genéricos. Mencione instrumentos, BPM, tom, humor e estrutura para ter o melhor resultado.
- Use o mesmo idioma. Use o comando no idioma em que você quer a letra.
- Use tags de seção. As tags
[Verse],[Chorus]e[Bridge]oferecem ao modelo uma estrutura clara para seguir. - Separe a letra das instruções. Ao fornecer letras personalizadas, separe-as claramente das instruções de direção musical.
Limitações
- Segurança: todos os comandos são verificados por filtros de segurança. Os comandos que acionam os filtros são bloqueados. Isso inclui comandos que pedem vozes de artistas específicos ou a geração de letras protegidas por direitos autorais.
- Marca-d'água: todos os áudios gerados incluem uma marca-d'água digital do SynthID para identificação. Essa marca-d'água é imperceptível ao ouvido humano e não afeta a experiência de audição.
- Edição multiturno: a geração de música é um processo de turno único. A edição iterativa ou o refinamento de um clipe gerado com vários comandos não são compatíveis com a versão atual do Lyria 3.5.
- Duração: o modelo de clipe sempre gera clipes de 30 segundos. O modelo Pro gera músicas que duram alguns minutos. A duração exata pode ser influenciada pelo comando.
- Determinismo: os resultados podem variar entre chamadas, mesmo com o mesmo comando.
A seguir
- Confira os preços dos modelos do Lyria 3.5.
- Teste a geração de músicas em streaming e em tempo real com o Lyria RealTime.
- Gere conversas com vários locutores usando os modelos de TTS.
- Saiba como gerar imagens ou vídeos.
- Saiba como o Gemini pode entender arquivos de áudio.
- Converse em tempo real com o Gemini usando a API Live.