Resona:基于 Chatterbox TTS 的 AI 语音生成与克隆平台
Resona - AI Voice Generation
Resona 是一个 AI 语音生成与克隆平台,面向需要制作漫画解说视频的 YouTube 创作者,支持从脚本生成一致旁白、克隆自定义音色并集中管理音频。平台基于开源 Chatterbox TTS 模型,自托管于 Modal 并搭配 FastAPI 推理服务,前端采用 Next.js、React、TypeScript、tRPC、Prisma 等技术栈,支持私有音频存储、组织工作区和按用量计费。
This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend
What I Built
I built Resona, an AI voice generation and cloning platform for a friend who wanted to create manhwa and manga recap videos on YouTube.
Recording narration for every video was time-consuming, so I built Resona to make it easier to generate consistent voiceovers from scripts, create custom voices, and manage generated audio in one place.
Demo
Live Demo: https://resonapro.vercel.app/
Brief Walkthrough: https://youtu.be/Bt2X_5rsFO0
Quick Demo: https://youtu.be/h0urRp9gXrU
Code
GitHub: https://github.com/verkiya/Resona
How I Built It
Resona is built around Chatterbox TTS, an open-source text-to-speech model that I self-host on Modal with a FastAPI inference service.
The platform uses Next.js, React, TypeScript, tRPC, Prisma, PostgreSQL, Clerk, AWS S3, Polar, Sentry, WaveSurfer.js, and RecordRTC.
It supports text-to-speech generation, custom voice cloning, private audio storage, organization workspaces, and usage-based billing.
Why Does Open Innovation Matter?
Using open-source Chatterbox TTS let me own the voice-generation layer instead of relying entirely on a closed API.
I could self-host the model, control the inference infrastructure, and build the rest of the platform around it. It also keeps the model layer flexible, so it can be changed or improved independently from the application.
Prize Categories
Best Use of ElevenLabs
I used ElevenLabs to generate the narration for the Resona demo video.
来源:Google AI:DEV 作者专属(RSS) · dev.to