Pocket TTS

Pocket TTS 音声翻訳

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

登録 5000文字の制限を設けました

SSML タグでテキストを囲み、正確な制御を行う:

<speak><prosody rate="slow">Slow speech</prosody></speak>

選択したモデルが理解するタグ - クリックしてテキストにドラッグします:

このモデルは単純テキストを読み込み、インラインタグは無視されます。タグベースの感情を表現するには、Orpheus や Bark のような表現モデルに切り替えてください。

カスタム発音を定義 (単語=発音):

-12 +12
0.5x 2.0x
ピパー、VITS、MeloTTS をフリーで使用
生成したオーディオがここに表示されます。モデルを選択し、テキストを入力して、生成をクリックします。
オーディオを作成しましたName
0:00
音声をダウンロード ダウンロード リンクは24時間で失効します
無料階級:個人用。 商用ライセンス $5/月から
これを自分の声にしよう 30秒で声をクローン
TTS.aiが気に入りましたか?友達に教えてあげましょう!

情報 Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

適合する: Lightweight deployment, CPU-only environments, quick voice cloning

すべてブラウズ Pocket TTS 声

概要

開発者
Kyutai
ライセンス
MIT
動物
free
スピード
fast
声のクローン
はい
言語
English, French
最大文字数
1000

Pocket TTS 声

Alba

English
自由 Female

Azelma

English
自由 Female

Cosette

English
自由 Female

Eponine

English
自由 Female

Fantine

English
自由 Female

Fantine (French)

French
自由 Female

Javert

English
自由 Male

Jean

English
自由 Male

Jean (French)

French
自由 Male

Marius

English
自由 Male

Pocket TTS よくある質問

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Yes. Pocket TTS is MIT-licensed and in the free tier. It supports English and French.
← すべての声