The latest Gemini models, like Gemini 3.5 Flash, are available to use with Firebase AI Logic! Learn more.

Gemini 2.0 Flash and Flash-Lite models were shut down on June 1, 2026. To avoid service disruption, update to a newer model like gemini-3.1-flash-lite. Learn more.

All Imagen models will shut down on June 24, 2026. Learn about migrating your apps to use Nano Banana.

Google uses AI technology to translate content into your preferred language. AI translations can contain errors.

Gemini API を使用して音声ファイルを分析する

Gemini モデルに、インライン（base64 エンコード）または URL で提供された音声ファイルを分析させることができます。Firebase AI Logic を使用すると、アプリから直接このリクエストを行うことができます。

この機能を使用すると、次のようなことができます。

音声コンテンツの説明、要約、質問への回答
音声コンテンツの文字起こし
タイムスタンプを使用して音声の特定のセグメントを分析する

コードサンプルに移動ストリーミングレスポンスのコードに移動

音声の操作に関するその他のオプションについては、他のガイドをご覧ください
構造化出力の生成マルチターンチャット双方向ストリーミング

始める前に

Gemini API プロバイダをクリックして、このページでプロバイダ固有のコンテンツとコードを表示します。

まだ行っていない場合は、スタートガイドに沿って、記載されている手順（ Firebase プロジェクトの設定、アプリと Firebase の連携、SDK の追加、選択した Gemini API プロバイダのバックエンドサービスの初期化、 GenerativeModel インスタンスの作成）を完了します。

プロンプトのテストと反復処理には、 Google AI Studioを使用することをおすすめします。

サンプル音声ファイルが必要ですか？

MIME タイプが audio/mp3 の一般公開されているこちらのファイルを使用できます (ファイルの表示またはダウンロード)。 https://storage.googleapis.com/cloud-samples-data/generative-ai/audio/pixel.mp3

音声ファイル（base64 エンコード）からテキストを生成する

このサンプルを試す前に、このガイドの始める前にのセクションを完了して、プロジェクトとアプリを設定してください。
このセクションでは、選択した Gemini API プロバイダのボタンをクリックして、このページにプロバイダ固有のコンテンツを表示します。

Gemini モデルにテキストを生成させるには、テキストと音声でプロンプトを表示し、入力ファイルの mimeType とファイル自体を指定します。入力ファイルの要件と推奨事項については、このページの後半をご覧ください。

Swift

`generateContent()` を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力からテキストを生成できます。generateContent()


import FirebaseAILogic

// Initialize the Gemini Developer API backend service.
let ai = FirebaseAI.firebaseAI(backend: .googleAI())

// Create a `GenerativeModel` instance with a model that supports your use case.
let model = ai.generativeModel(modelName: "gemini-3.5-flash")


// Provide the audio as `Data`
guard let audioData = try? Data(contentsOf: audioURL) else {
    print("Error loading audio data.")
    return // Or handle the error appropriately
}

// Specify the appropriate audio MIME type
let audio = InlineDataPart(data: audioData, mimeType: "audio/mpeg")


// Provide a text prompt to include with the audio
let prompt = "Transcribe what's said in this audio recording."

// To generate text output, call `generateContent` with the audio and text prompt
let response = try await model.generateContent(audio, prompt)

// Print the generated text, handling the case where it might be nil
print(response.text ?? "No text in response.")

Kotlin

`generateContent()` を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力からテキストを生成できます。generateContent()

^{Kotlin の場合、この SDK のメソッドは suspend 関数であり、
Coroutine スコープから呼び出す必要があります。}


// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a model that supports your use case.
val model = Firebase.ai(backend = GenerativeBackend.googleAI())
                        .generativeModel("gemini-3.5-flash")


val contentResolver = applicationContext.contentResolver

val inputStream = contentResolver.openInputStream(audioUri)

if (inputStream != null) {  // Check if the audio loaded successfully
    inputStream.use { stream ->
        val bytes = stream.readBytes()

        // Provide a prompt that includes the audio specified above and text
        val prompt = content {
            inlineData(bytes, "audio/mpeg")  // Specify the appropriate audio MIME type
            text("Transcribe what's said in this audio recording.")
        }

        // To generate text output, call `generateContent` with the prompt
        val response = model.generateContent(prompt)

        // Log the generated text, handling the case where it might be null
        Log.d(TAG, response.text?: "")
    }
} else {
    Log.e(TAG, "Error getting input stream for audio.")
    // Handle the error appropriately
}

Java

`generateContent()` を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力からテキストを生成できます。generateContent()

^{Java の場合、この SDK のメソッドは
ListenableFuture を返します。}


// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a model that supports your use case.
GenerativeModel ai = FirebaseAI.getInstance(GenerativeBackend.googleAI())
        .generativeModel("gemini-3.5-flash");

// Use the GenerativeModelFutures Java compatibility layer which offers
// support for ListenableFuture and Publisher APIs
GenerativeModelFutures model = GenerativeModelFutures.from(ai);


ContentResolver resolver = getApplicationContext().getContentResolver();

try (InputStream stream = resolver.openInputStream(audioUri)) {
    File audioFile = new File(new URI(audioUri.toString()));
    int audioSize = (int) audioFile.length();
    byte audioBytes = new byte[audioSize];
    if (stream != null) {
        stream.read(audioBytes, 0, audioBytes.length);
        stream.close();

        // Provide a prompt that includes the audio specified above and text
        Content prompt = new Content.Builder()
              .addInlineData(audioBytes, "audio/mpeg")  // Specify the appropriate audio MIME type
              .addText("Transcribe what's said in this audio recording.")
              .build();

        // To generate text output, call `generateContent` with the prompt
        ListenableFuture<GenerateContentResponse> response = model.generateContent(prompt);
        Futures.addCallback(response, new FutureCallback<GenerateContentResponse>() {
            @Override
            public void onSuccess(GenerateContentResponse result) {
                String text = result.getText();
                Log.d(TAG, (text == null) ? "" : text);
            }
            @Override
            public void onFailure(Throwable t) {
                Log.e(TAG, "Failed to generate a response", t);
            }
        }, executor);
    } else {
        Log.e(TAG, "Error getting input stream for file.");
        // Handle the error appropriately
    }
} catch (IOException e) {
    Log.e(TAG, "Failed to read the audio file", e);
} catch (URISyntaxException e) {
    Log.e(TAG, "Invalid audio file", e);
}

Web

`generateContent()` を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力からテキストを生成できます。generateContent()


import { initializeApp } from "firebase/app";
import { getAI, getGenerativeModel, GoogleAIBackend } from "firebase/ai";

// TODO(developer) Replace the following with your app's Firebase configuration
// See: https://firebase.google.com/docs/web/learn-more#config-object
const firebaseConfig = {
  // ...
};

// Initialize FirebaseApp
const firebaseApp = initializeApp(firebaseConfig);

// Initialize the Gemini Developer API backend service.
const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });

// Create a `GenerativeModel` instance with a model that supports your use case.
const model = getGenerativeModel(ai, { model: "gemini-3.5-flash" });


// Converts a File object to a Part object.
async function fileToGenerativePart(file) {
  const base64EncodedDataPromise = new Promise((resolve) => {
    const reader = new FileReader();
    reader.onloadend = () => resolve(reader.result.split(','));
    reader.readAsDataURL(file);
  });
  return {
    inlineData: { data: await base64EncodedDataPromise, mimeType: file.type },
  };
}

async function run() {
  // Provide a text prompt to include with the audio
  const prompt = "Transcribe what's said in this audio recording.";

  // Prepare audio for input
  const fileInputEl = document.querySelector("input[type=file]");
  const audioPart = await fileToGenerativePart(fileInputEl.files);

  // To generate text output, call `generateContent` with the text and audio
  const result = await model.generateContent([prompt, audioPart]);

  // Log the generated text, handling the case where it might be undefined
  console.log(result.response.text() ?? "No text in response.");
}

run();

Dart

generateContent() を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力からテキストを生成できます。


import 'package:firebase_ai/firebase_ai.dart';
import 'package:firebase_core/firebase_core.dart';
import 'firebase_options.dart';

// Initialize FirebaseApp
await Firebase.initializeApp(
  options: DefaultFirebaseOptions.currentPlatform,
);

// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a model that supports your use case.
final model =
      FirebaseAI.googleAI().generativeModel(model: 'gemini-3.5-flash');


// Provide a text prompt to include with the audio
final prompt = TextPart("Transcribe what's said in this audio recording.");

// Prepare audio for input
final audio = await File('audio0.mp3').readAsBytes();

// Provide the audio as `Data` with the appropriate audio MIME type
final audioPart = InlineDataPart('audio/mpeg', audio);

// To generate text output, call `generateContent` with the text and audio
final response = await model.generateContent([
  Content.multi([prompt,audioPart])
]);

// Print the generated text
print(response.text);

Unity

GenerateContentAsync() を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力からテキストを生成できます。


using Firebase;
using Firebase.AI;

// Initialize the Gemini Developer API backend service.
var ai = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI());

// Create a `GenerativeModel` instance with a model that supports your use case.
var model = ai.GetGenerativeModel(modelName: "gemini-3.5-flash");


// Provide a text prompt to include with the audio
var prompt = ModelContent.Text("Transcribe what's said in this audio recording.");

// Provide the audio as `data` with the appropriate audio MIME type
var audio = ModelContent.InlineData("audio/mpeg",
      System.IO.File.ReadAllBytes(System.IO.Path.Combine(
        UnityEngine.Application.streamingAssetsPath, "audio0.mp3")));

// To generate text output, call `GenerateContentAsync` with the text and audio
var response = await model.GenerateContentAsync(new [] { prompt, audio });

// Print the generated text
UnityEngine.Debug.Log(response.Text ?? "No text in response.");

ユースケースとアプリに適したモデルを選択する方法をご覧ください。

レスポンスをストリーミングする

モデル生成の結果全体を待つのではなく、ストリーミングを使用して部分的な結果を処理することで、インタラクションを高速化できます。レスポンスをストリーミングするには、generateContentStream を呼び出します。

例を見る: 音声ファイルから生成されたテキストをストリーミングする

Swift

generateContentStream() を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力から生成されたテキストをストリーミングできます。


import FirebaseAILogic

// Initialize the Gemini Developer API backend service.
let ai = FirebaseAI.firebaseAI(backend: .googleAI())

// Create a `GenerativeModel` instance with a model that supports your use case.
let model = ai.generativeModel(modelName: "gemini-3.5-flash")


// Provide the audio as `Data`
guard let audioData = try? Data(contentsOf: audioURL) else {
    print("Error loading audio data.")
    return // Or handle the error appropriately
}

// Specify the appropriate audio MIME type
let audio = InlineDataPart(data: audioData, mimeType: "audio/mpeg")


// Provide a text prompt to include with the audio
let prompt = "Transcribe what's said in this audio recording."

// To stream generated text output, call `generateContentStream` with the audio and text prompt
let contentStream = try model.generateContentStream(audio, prompt)

// Print the generated text, handling the case where it might be nil
for try await chunk in contentStream {
    if let text = chunk.text {
        print(text)
    }
}

Kotlin

generateContentStream() を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力から生成されたテキストをストリーミングできます。

^{Kotlin の場合、この SDK のメソッドは suspend 関数であり、
Coroutine スコープから呼び出す必要があります。}


// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a model that supports your use case.
val model = Firebase.ai(backend = GenerativeBackend.googleAI())
                        .generativeModel("gemini-3.5-flash")


val contentResolver = applicationContext.contentResolver

val inputStream = contentResolver.openInputStream(audioUri)

if (inputStream != null) {  // Check if the audio loaded successfully
    inputStream.use { stream ->
        val bytes = stream.readBytes()

        // Provide a prompt that includes the audio specified above and text
        val prompt = content {
            inlineData(bytes, "audio/mpeg")  // Specify the appropriate audio MIME type
            text("Transcribe what's said in this audio recording.")
        }

        // To stream generated text output, call `generateContentStream` with the prompt
        var fullResponse = ""
        model.generateContentStream(prompt).collect { chunk ->
            // Log the generated text, handling the case where it might be null
            Log.d(TAG, chunk.text?: "")
            fullResponse += chunk.text?: ""
        }
    }
} else {
    Log.e(TAG, "Error getting input stream for audio.")
    // Handle the error appropriately
}

Java

generateContentStream() を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力から生成されたテキストをストリーミングできます。

^{Java の場合、この SDK のストリーミングメソッドは Reactive Streams library から
Publisher 型を返します。}


// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a model that supports your use case.
GenerativeModel ai = FirebaseAI.getInstance(GenerativeBackend.googleAI())
        .generativeModel("gemini-3.5-flash");

// Use the GenerativeModelFutures Java compatibility layer which offers
// support for ListenableFuture and Publisher APIs
GenerativeModelFutures model = GenerativeModelFutures.from(ai);


ContentResolver resolver = getApplicationContext().getContentResolver();

try (InputStream stream = resolver.openInputStream(audioUri)) {
    File audioFile = new File(new URI(audioUri.toString()));
    int audioSize = (int) audioFile.length();
    byte audioBytes = new byte[audioSize];
    if (stream != null) {
        stream.read(audioBytes, 0, audioBytes.length);
        stream.close();

        // Provide a prompt that includes the audio specified above and text
        Content prompt = new Content.Builder()
              .addInlineData(audioBytes, "audio/mpeg")  // Specify the appropriate audio MIME type
              .addText("Transcribe what's said in this audio recording.")
              .build();

        // To stream generated text output, call `generateContentStream` with the prompt
        Publisher<GenerateContentResponse> streamingResponse =
                model.generateContentStream(prompt);

        StringBuilder fullResponse = new StringBuilder();
        streamingResponse.subscribe(new Subscriber<GenerateContentResponse>() {
            @Override
            public void onNext(GenerateContentResponse generateContentResponse) {
                String chunk = generateContentResponse.getText();
                String text = (chunk == null) ? "" : chunk;
                Log.d(TAG, text);
                fullResponse.append(text);
            }
            @Override
            public void onComplete() {
                Log.d(TAG, fullResponse.toString());
            }
            @Override
            public void onError(Throwable t) {
                Log.e(TAG, "Failed to generate a response", t);
            }
            @Override
            public void onSubscribe(Subscription s) {
            }
         });
    } else {
        Log.e(TAG, "Error getting input stream for file.");
        // Handle the error appropriately
    }
} catch (IOException e) {
    Log.e(TAG, "Failed to read the audio file", e);
} catch (URISyntaxException e) {
    Log.e(TAG, "Invalid audio file", e);
}

Web

generateContentStream() を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力から生成されたテキストをストリーミングできます。


import { initializeApp } from "firebase/app";
import { getAI, getGenerativeModel, GoogleAIBackend } from "firebase/ai";

// TODO(developer) Replace the following with your app's Firebase configuration
// See: https://firebase.google.com/docs/web/learn-more#config-object
const firebaseConfig = {
  // ...
};

// Initialize FirebaseApp
const firebaseApp = initializeApp(firebaseConfig);

// Initialize the Gemini Developer API backend service.
const ai = getAI(firebaseApp, { backend: new GoogleAIBackend() });

// Create a `GenerativeModel` instance with a model that supports your use case.
const model = getGenerativeModel(ai, { model: "gemini-3.5-flash" });


// Converts a File object to a Part object.
async function fileToGenerativePart(file) {
  const base64EncodedDataPromise = new Promise((resolve) => {
    const reader = new FileReader();
    reader.onloadend = () => resolve(reader.result.split(','));
    reader.readAsDataURL(file);
  });
  return {
    inlineData: { data: await base64EncodedDataPromise, mimeType: file.type },
  };
}

async function run() {
  // Provide a text prompt to include with the audio
  const prompt = "Transcribe what's said in this audio recording.";

  // Prepare audio for input
  const fileInputEl = document.querySelector("input[type=file]");
  const audioPart = await fileToGenerativePart(fileInputEl.files);

  // To stream generated text output, call `generateContentStream` with the text and audio
  const result = await model.generateContentStream([prompt, audioPart]);

  // Log the generated text
  for await (const chunk of result.stream) {
    const chunkText = chunk.text();
    console.log(chunkText);
  }
}

run();

Dart

`generateContentStream()` を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力から生成されたテキストをストリーミングできます。generateContentStream()


import 'package:firebase_ai/firebase_ai.dart';
import 'package:firebase_core/firebase_core.dart';
import 'firebase_options.dart';

// Initialize FirebaseApp
await Firebase.initializeApp(
  options: DefaultFirebaseOptions.currentPlatform,
);

// Initialize the Gemini Developer API backend service.
// Create a `GenerativeModel` instance with a model that supports your use case.
final model =
      FirebaseAI.googleAI().generativeModel(model: 'gemini-3.5-flash');


// Provide a text prompt to include with the audio
final prompt = TextPart("Transcribe what's said in this audio recording.");

// Prepare audio for input
final audio = await File('audio0.mp3').readAsBytes();

// Provide the audio as `Data` with the appropriate audio MIME type
final audioPart = InlineDataPart('audio/mpeg', audio);

// To stream generated text output, call `generateContentStream` with the text and audio
final response = await model.generateContentStream([
  Content.multi([prompt, audioPart])
]);

// Print the generated text
await for (final chunk in response) {
  print(chunk.text);
}

Unity

GenerateContentStreamAsync() を呼び出して、テキストと単一の音声ファイルのマルチモーダル入力から生成されたテキストをストリーミングできます。


using Firebase;
using Firebase.AI;

// Initialize the Gemini Developer API backend service.
var ai = FirebaseAI.GetInstance(FirebaseAI.Backend.GoogleAI());

// Create a `GenerativeModel` instance with a model that supports your use case.
var model = ai.GetGenerativeModel(modelName: "gemini-3.5-flash");


// Provide a text prompt to include with the audio
var prompt = ModelContent.Text("Transcribe what's said in this audio recording.");

// Provide the audio as `data` with the appropriate audio MIME type
var audio = ModelContent.InlineData("audio/mpeg",
      System.IO.File.ReadAllBytes(System.IO.Path.Combine(
        UnityEngine.Application.streamingAssetsPath, "audio0.mp3")));

// To stream generated text output, call `GenerateContentStreamAsync` with the text and audio
var responseStream = model.GenerateContentStreamAsync(new [] { prompt, audio });

// Print the generated text
await foreach (var response in responseStream) {
  if (!string.IsNullOrWhiteSpace(response.Text)) {
    UnityEngine.Debug.Log(response.Text);
  }
}

ユースケースとアプリに適したモデルを選択する方法をご覧ください。

入力音声ファイルの要件と推奨事項

インラインデータとして提供されるファイルは転送中に base64 にエンコードされるため、リクエストのサイズが大きくなります。リクエストが大きすぎると、HTTP 413 エラーが発生します。

次の詳細については、サポートされている入力ファイルと要件のページをご覧ください。

リクエストでファイルを提供するさまざまなオプション（インライン、ファイルの URL または URI を使用）
音声ファイルの要件とおすすめの方法

サポートされている音声 MIME タイプ

Gemini マルチモーダルモデルは、次の音声 MIME タイプをサポートしています:

AAC - audio/aac
FLAC - audio/flac
MP3 - audio/mp3
MPA - audio/m4a
MPEG - audio/mpeg
MPGA - audio/mpga
MP4 - audio/mp4
OPUS - audio/opus
PCM - audio/pcm
WAV - audio/wav
WEBM - audio/webm

リクエストあたりの上限

リクエストあたりの最大ファイル数: 1 つの音声ファイル

Google アシスタントの機能

長いプロンプトをモデルに送信する前に、トークンをカウントする方法を確認する。
を設定して、マルチモーダルリクエストに大きなファイルを含めることができるようにし、プロンプトでファイルを提供するソリューションをより管理しやすくする。Cloud Storage for Firebaseファイルには、画像、PDF、動画、音声を含めることができます。
本番環境の準備について検討する（本番環境チェックリストを参照）。
- できるだけ早くFirebase App Checkを設定して、未承認のクライアントによるGemini APIの不正使用を防ぐ。
- Integrate Firebase Remote Config を統合して、新しいバージョンのアプリをリリースせずに、アプリ内の値（モデル名など）を更新する。

他の機能を試す

マルチターン会話（チャット）を構築する。
テキストのみのプロンプトからテキストを生成する。
構造化出力（JSON など）を生成するテキストとマルチモーダルプロンプトの両方から。
テキストとマルチモーダルプロンプトの両方から画像を生成して編集する。
入出力（音声を含む）を Gemini Live API を使用してストリーミングする。
ツール（関数呼び出しやGoogle 検索によるグラウンディング）を使用して、Geminiモデルをアプリの他の部分や外部システム、情報に接続する。

コンテンツ生成を制御する方法を確認する

プロンプト設計（ベストプラクティス、戦略、プロンプトの例など）を理解する。
モデルパラメータを構成する（最大出力トークン数、繰り返し出力トークンの確率など）。
安全性設定を使用して、有害と見なされる可能性のあるレスポンスを取得する可能性を調整する。

Google AI Studio を使用して、プロンプトとモデル構成を試したり、生成されたコードスニペットを取得したりすることもできます Google AI Studio。

サポートされているモデルの詳細

さまざまなユースケースで利用できるモデルとその割り当てと料金について説明します。

フィードバックを送信する Firebase AI Logicの使用に関する

Gemini API を使用して音声ファイルを分析する コレクションでコンテンツを整理 必要に応じて、コンテンツの保存と分類を行います。

始める前に

音声ファイル（base64 エンコード）からテキストを生成する

Swift

Kotlin

Java

Web

Dart

Unity

レスポンスをストリーミングする

例を見る: 音声ファイルから生成されたテキストをストリーミングする

Swift

Kotlin

Java

Web

Dart

Unity

入力音声ファイルの要件と推奨事項

サポートされている音声 MIME タイプ

リクエストあたりの上限

Google アシスタントの機能

他の機能を試す

コンテンツ生成を制御する方法を確認する

サポートされているモデルの詳細

Gemini API を使用して音声ファイルを分析する