Pocket TTS

Pocket TTS 음성 인식

A compact 100M-parameter CPU model from Kyutai (makers of Moshi) with single-sample voice cloning.

가입하기 5,000자 한도

정확한 제어를 위해 SSML 태그로 텍스트를 래핑하십시오.

<speak><prosody rate="slow">Slow speech</prosody></speak>

선택한 모델이 이해하는 태그 — 텍스트에 드래그하려면 클릭하세요:

이 모델은 일반 텍스트를 읽기 때문에 인라인 태그는 무시됩니다. 태그 기반 감정을 위해서는 Orpheus 또는 Bark과 같은 표현 모델로 전환하십시오.

사용자 지정 발음 정의 (단어 = 발음):

-12 +12
0.5x 2.0x
파이퍼, VITS, MeloTTS와 무료
생성된 오디오가 여기에 나타납니다. 모델을 선택하고 텍스트를 입력한 다음 생성 을 클릭합니다.
오디오가 성공적으로 생성되었습니다
0:00
오디오 다운로드 .srt 파일 다운로드 링크는 24시간 이내에 만료됩니다.
무료 계층: 개인용. 상업용 라이센스 최저 $5/mo
이것을 당신의 목소리로 만들어라 30초만에 목소리 복제
TTS.ai가 마음에 드시나요? 친구들에게 알려주세요!

정보 Pocket TTS

Pocket TTS comes from Kyutai, the lab behind the Moshi speech model, and is built around a transformer paired with the Mimi codec. At just 100M parameters it runs efficiently on CPU, yet it still supports zero-shot voice cloning from a single audio sample — an unusual feature at this size. It covers English and French and handles up to 1,000 characters per request at fast (~2s) speeds. The small footprint and ~1GB VRAM make it a natural fit for edge deployment and low-resource or CPU-only environments where quick voice cloning is needed.

최적화된 용도: Lightweight deployment, CPU-only environments, quick voice cloning

모두 찾아보기 Pocket TTS 목소리

한눈에

개발자
Kyutai
라이선스
MIT
free
속도
fast
음성 복제
언어
English, French
최대 문자수
1000

Pocket TTS 목소리

Alba

English
자유 Female

Azelma

English
자유 Female

Cosette

English
자유 Female

Eponine

English
자유 Female

Fantine

English
자유 Female

Fantine (French)

French
자유 Female

Javert

English
자유 Male

Jean

English
자유 Male

Jean (French)

French
자유 Male

Marius

English
자유 Male

Pocket TTS TTS — 자주 묻는 질문

Yes. Pocket TTS does zero-shot voice cloning from a single reference sample (about 3 seconds), which is notable for a model this small.

Yes. At 100M parameters it runs efficiently on CPU and needs only about 1GB VRAM if a GPU is used, making it well suited to edge and low-resource deployment.

Yes. Pocket TTS is MIT-licensed and in the free tier. It supports English and French.
← 모든 음성