جمینی میتواند ورودی صوتی را تجزیه و تحلیل کرده و پاسخهای متنی تولید کند.
پایتون
from google import genai
import base64
client = genai.Client()
uploaded_file = client.files.upload(file="path/to/sample.mp3")
interaction = client.interactions.create(
model="gemini-3.7-flash",
input=[
{"type": "text", "text": "Describe this audio clip"},
{
"type": "audio",
"uri": uploaded_file.uri,
"mime_type": uploaded_file.mime_type
}
]
)
print(interaction.output_text)
جاوا اسکریپت
import { GoogleGenAI } from "@google/genai";
const client = new GoogleGenAI({});
const uploadedFile = await client.files.upload({
file: "path/to/sample.mp3",
config: { mime_type: "audio/mp3" }
});
const interaction = await client.interactions.create({
model: "gemini-3.7-flash",
input: [
{type: "text", text: "Describe this audio clip"},
{
type: "audio",
uri: uploadedFile.uri,
mime_type: uploadedFile.mimeType
}
]
});
console.log(interaction.output_text);
استراحت
# First upload the file, then use the URI:
curl -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H 'Content-Type: application/json' \
-d '{
"model": "gemini-3.7-flash",
"input": [
{"type": "text", "text": "Describe this audio clip"},
{
"type": "audio",
"uri": "YOUR_FILE_URI",
"mime_type": "audio/mp3"
}
]
}'
نمای کلی
Gemini میتواند ورودی صوتی را تجزیه و تحلیل و درک کند و پاسخهای متنی تولید کند و موارد استفادهای مانند موارد زیر را امکانپذیر سازد:
- توصیف، خلاصه کردن یا پاسخ به سوالات مربوط به محتوای صوتی
- رونویسی و ترجمه (گفتار به متن)
- شناسایی گوینده (شناسایی گویندگان مختلف)
- تشخیص احساسات در گفتار و موسیقی
- تجزیه و تحلیل بخشهای خاص با مهرهای زمانی
برای تعاملات صوتی و تصویری در لحظه، به Live API مراجعه کنید. برای مدلهای اختصاصی تبدیل گفتار به متن با پشتیبانی از رونویسی در لحظه، از Google Cloud Speech-to-Text API استفاده کنید.
تبدیل گفتار به متن
این مثال نحوه رونویسی، ترجمه و خلاصهسازی گفتار با استفاده از مهرهای زمانی، تشخیص خاطرات گوینده و تشخیص احساسات را با استفاده از خروجیهای ساختاریافته نشان میدهد.
پایتون
from google import genai
client = genai.Client()
YOUTUBE_URL = "https://www.youtube.com/watch?v=ku-N-eS1lgM"
prompt = """
Process the audio file and generate a detailed transcription.
Requirements:
1. Identify distinct speakers (e.g., Speaker 1, Speaker 2).
2. Provide accurate timestamps for each segment (Format: MM:SS).
3. Detect the primary language of each segment.
4. If not English, provide the English translation.
5. Identify the primary emotion: Happy, Sad, Angry, or Neutral.
6. Provide a brief summary at the beginning.
"""
response_schema = {
"type": "object",
"properties": {
"summary": {"type": "string"},
"segments": {
"type": "array",
"items": {
"type": "object",
"properties": {
"speaker": {"type": "string"},
"timestamp": {"type": "string"},
"content": {"type": "string"},
"language": {"type": "string"},
"emotion": {
"type": "string",
"enum": ["happy", "sad", "angry", "neutral"]