The Full Flow — From Zero to AI Voice
Using a community AI voice in Echo takes five steps. Each one is covered in detail below, but here is the overview so you know what to expect before you start.
- ●Step 1 — Find a voice model on AIVoices.gg. You can also browse Voice-Models.com, Hugging Face, or the AI Hub community
- ●Step 2 — Download the .pth file (and .index file if available). This is the raw model format used by the RVC community
- ●Step 3 — Convert the .pth file to .onnx format using the free browser-based converter at voicechanger.live/tools/model-converter. This takes a few seconds and never uploads your file
- ●Step 4 — Import the .onnx file into Echo. Open Echo → Voice Library → Import → select your .onnx file
- ●Step 5 — Select the voice, click Start, and talk. Your voice is transformed in real-time through any app that uses your microphone
Understanding the File Types
RVC voice models involve three file types. You will encounter all three during the process, so knowing what each one does saves confusion.
- ●.pth (PyTorch Model) — The standard format used by the RVC community. This is what you download from Hugging Face. A .pth file is typically 50–100 MB and contains the trained neural network that defines the target voice. You cannot use .pth files directly in Echo — they need to be converted to .onnx first
- ●.index (FAISS Index) — An optional companion file that improves voice quality by referencing real acoustic samples from the original training audio. For real-time voice changing in Echo, the .index file is not required — Echo's ONNX pipeline handles conversion without it. However, if you use other RVC tools for offline conversion (song covers, voiceovers), the .index file makes a noticeable quality difference. Always download it when available
- ●.onnx (Open Neural Network Exchange) — The optimized format that Echo uses. Converting .pth to .onnx produces an identical voice but in a format that runs natively without Python or PyTorch — meaning faster inference, lower memory, and no heavy dependencies. The conversion is free and takes seconds
Step 1 — Where to Find Voice Models
The RVC community has produced thousands of voice models covering characters, celebrities, original voices, and custom creations. Quality varies enormously — a model trained on clean audio sounds convincing, while a poorly trained one sounds robotic. Here are the main sources.
- ●AIVoices.gg — Start here to browse community AI voice models by character or voice type
- ●Voice-Models.com — Another dedicated directory for discovering downloadable RVC models
- ●Hugging Face — Many creators host RVC models in their own repositories. Search for "[character name] RVC" or "RVC v2 model". Repositories often include both .pth and .index files
- ●AI Hub Discord — A community option for recommendations, niche models, and model requests. Join at discord.gg/aihub
- ●Reddit (r/RVC) — Model recommendations, quality reviews, and community-vetted suggestions. Good for discovering which models are considered the best for specific characters
Step 2 — How to Pick a Good Model
Not all models are worth downloading. The difference between a well-trained and poorly-trained model is dramatic. Here is how to spot the good ones.
- ●Listen to audio previews — On Hugging Face, many models have audio demos. A 5-second preview tells you more than any description. If it sounds good in the preview, it will sound good in Echo
- ●Check the training data duration — Models trained on 15–40 minutes of clean audio generally sound best. Under 5 minutes tends to sound thin or generic
- ●Look for v2 models — RVC v2 models use 768-dimensional feature vectors and sound more natural than v1 (256-dimensional). Always prefer v2
- ●Read the description — Good creators document their parameters. A model described as "trained on 25 min clean vocal, 300 epochs, RMVPE, 40kHz" is more trustworthy than one with no description
- ●Check download counts and ratings — On Hugging Face, popular models with high ratings have been community-validated. On Discord, ask in the model-sharing channels for recommendations
- ●Check if an .index file is included — Models with an .index file generally produce better quality, especially for offline use. It is a good quality signal even if you do not use the file in Echo
Step 3 — Convert .pth to .onnx
Echo uses .onnx format for voice models — it is faster, lighter, and does not require Python. After downloading a .pth model from the community, you need to convert it once. This is free and takes seconds.
- ●Go to voicechanger.live/tools/model-converter in your browser
- ●Select your downloaded .pth file. The file never leaves your computer — conversion happens entirely in your browser
- ●Choose the model version (v1 or v2). If you are unsure, try v2 first — most modern models are v2
- ●Click Convert and download the resulting .onnx file
- ●The output is identical in voice quality to the original .pth — only the format changes
Step 4 — Import into Echo
With your .onnx file ready, importing into Echo takes about 10 seconds.
- ●Open Echo and go to the Voice Library (the voices tab in the sidebar)
- ●Click the Import button at the top of the library
- ●Select your .onnx file from wherever you saved it
- ●The voice appears in your library immediately — no restart needed
- ●There is no limit on the number of models you can import. Import as many as you want
Step 5 — Start Using Your Voice
Once imported, select the voice from your library and click Start. Echo captures your microphone input, runs it through the AI model in real-time, and outputs the converted voice through a virtual audio cable. Any application that uses a microphone — Discord, games, OBS, Zoom — can use the transformed voice by selecting the virtual cable as its input device.
- ●Pitch offset — If the voice sounds too high or too low, adjust the pitch offset in Echo. For male-to-female conversions, try +12. For female-to-male, try -12. Fine-tune from there based on what sounds natural
- ●GPU acceleration — Enable GPU mode for the lowest latency. Echo supports NVIDIA (CUDA), AMD/Intel (DirectML), and Apple Silicon (CoreML)
- ●Virtual audio cable — Make sure VB-Cable (Windows) or BlackHole (macOS) is installed and selected as the output device in Echo. Then select it as the microphone in Discord, OBS, or your game
About the .index File
You will see .index files alongside many .pth downloads. Here is exactly when they matter and when you can ignore them.
- ●For real-time voice changing in Echo — The .index file is NOT required. Echo's ONNX pipeline runs the full voice conversion without needing the FAISS index. Your voice will sound great without it
- ●For offline voice conversion (song covers, voiceovers) — The .index file IS useful. Tools like w-okada and Applio use the index during inference to retrieve real acoustic details from the training data, which reduces artifacts and improves fidelity on longer audio
- ●For model quality assessment — A model that ships with an .index file is usually higher quality. It means the creator took the extra step to generate the index, which correlates with better overall training practices
- ●Storage — .index files are typically 10–50 MB. Download them when available and keep them with your .pth files in case you want to use offline tools later
Troubleshooting
If a voice does not sound right after importing, the problem is almost always one of these. Most can be fixed without finding a new model.
- ●Voice sounds robotic or metallic — The model is likely over-trained. Try a different model from the same character, or look for one with fewer training epochs
- ●Pitch sounds wrong — Adjust the pitch offset in Echo. Male-to-female typically needs +12, female-to-male needs -12. Experiment within that range
- ●Voice sounds nothing like the target — You may have a v1 model in a v2 pipeline or vice versa. Re-convert with the correct version selected in the model converter
- ●Output sounds muffled — The model was probably trained on low-quality audio. Find a different model trained on cleaner source material
- ●Crackling or popping — This is usually an audio pipeline issue, not a model problem. Increase the block size in Echo settings, or check that your GPU is not overloaded
- ●Conversion sounds generic — For offline tools, this usually means the .index file is missing. For Echo, try a different model — some models simply convert better than others
Want to Create Your Own Voice Model?
If you cannot find the voice you want in community repositories, you can train a custom RVC model from scratch. This is a more advanced process that requires a GPU and some patience, but the tools are free and the community is extremely helpful. You need 10–30 minutes of clean, isolated vocal audio from the target voice. The recommended training tool is Applio (applio.org) — it has the best interface and most active development. For users without a local GPU, Google Colab notebooks provide free cloud GPU access for training. We have a complete step-by-step training guide that covers dataset preparation with UVR5, training parameters, f0 extraction methods, evaluating checkpoints, and troubleshooting common problems.
Ethical Use of Voice Models
RVC voice models are powerful technology that should be used responsibly. Do not create models of real people without their knowledge or consent. Do not use voice models to impersonate someone for deception, fraud, or harassment. If you create AI-generated voice content, label it clearly. Several jurisdictions have enacted legislation targeting AI voice impersonation — using a cloned voice to mislead people is increasingly a criminal offense. The RVC community thrives on mutual respect between creators, model trainers, and users.