guides15 min read · Updated 2026-05-31

How to Find, Convert & Import RVC Voice Models

Everything you need to go from zero to a working AI voice in Echo — where to find models, how to pick good ones, converting formats, importing, and troubleshooting.

The Full Flow — From Zero to AI Voice

Using a community AI voice in Echo takes five steps. Each one is covered in detail below, but here is the overview so you know what to expect before you start.

  • ●Step 1 — Find a voice model on AIVoices.gg. You can also browse Voice-Models.com, Hugging Face, or the AI Hub community
  • ●Step 2 — Download the .pth file (and .index file if available). This is the raw model format used by the RVC community
  • ●Step 3 — Convert the .pth file to .onnx format using the free browser-based converter at voicechanger.live/tools/model-converter. This takes a few seconds and never uploads your file
  • ●Step 4 — Import the .onnx file into Echo. Open Echo → Voice Library → Import → select your .onnx file
  • ●Step 5 — Select the voice, click Start, and talk. Your voice is transformed in real-time through any app that uses your microphone

Understanding the File Types

RVC voice models involve three file types. You will encounter all three during the process, so knowing what each one does saves confusion.

  • ●.pth (PyTorch Model) — The standard format used by the RVC community. This is what you download from Hugging Face. A .pth file is typically 50–100 MB and contains the trained neural network that defines the target voice. You cannot use .pth files directly in Echo — they need to be converted to .onnx first
  • ●.index (FAISS Index) — An optional companion file that improves voice quality by referencing real acoustic samples from the original training audio. For real-time voice changing in Echo, the .index file is not required — Echo's ONNX pipeline handles conversion without it. However, if you use other RVC tools for offline conversion (song covers, voiceovers), the .index file makes a noticeable quality difference. Always download it when available
  • ●.onnx (Open Neural Network Exchange) — The optimized format that Echo uses. Converting .pth to .onnx produces an identical voice but in a format that runs natively without Python or PyTorch — meaning faster inference, lower memory, and no heavy dependencies. The conversion is free and takes seconds

Step 1 — Where to Find Voice Models

The RVC community has produced thousands of voice models covering characters, celebrities, original voices, and custom creations. Quality varies enormously — a model trained on clean audio sounds convincing, while a poorly trained one sounds robotic. Here are the main sources.

  • ●AIVoices.gg — Start here to browse community AI voice models by character or voice type
  • ●Voice-Models.com — Another dedicated directory for discovering downloadable RVC models
  • ●Hugging Face — Many creators host RVC models in their own repositories. Search for "[character name] RVC" or "RVC v2 model". Repositories often include both .pth and .index files
  • ●AI Hub Discord — A community option for recommendations, niche models, and model requests. Join at discord.gg/aihub
  • ●Reddit (r/RVC) — Model recommendations, quality reviews, and community-vetted suggestions. Good for discovering which models are considered the best for specific characters

Step 2 — How to Pick a Good Model

Not all models are worth downloading. The difference between a well-trained and poorly-trained model is dramatic. Here is how to spot the good ones.

  • ●Listen to audio previews — On Hugging Face, many models have audio demos. A 5-second preview tells you more than any description. If it sounds good in the preview, it will sound good in Echo
  • ●Check the training data duration — Models trained on 15–40 minutes of clean audio generally sound best. Under 5 minutes tends to sound thin or generic
  • ●Look for v2 models — RVC v2 models use 768-dimensional feature vectors and sound more natural than v1 (256-dimensional). Always prefer v2
  • ●Read the description — Good creators document their parameters. A model described as "trained on 25 min clean vocal, 300 epochs, RMVPE, 40kHz" is more trustworthy than one with no description
  • ●Check download counts and ratings — On Hugging Face, popular models with high ratings have been community-validated. On Discord, ask in the model-sharing channels for recommendations
  • ●Check if an .index file is included — Models with an .index file generally produce better quality, especially for offline use. It is a good quality signal even if you do not use the file in Echo

Step 3 — Convert .pth to .onnx

Echo uses .onnx format for voice models — it is faster, lighter, and does not require Python. After downloading a .pth model from the community, you need to convert it once. This is free and takes seconds.

  • ●Go to voicechanger.live/tools/model-converter in your browser
  • ●Select your downloaded .pth file. The file never leaves your computer — conversion happens entirely in your browser
  • ●Choose the model version (v1 or v2). If you are unsure, try v2 first — most modern models are v2
  • ●Click Convert and download the resulting .onnx file
  • ●The output is identical in voice quality to the original .pth — only the format changes

Step 4 — Import into Echo

With your .onnx file ready, importing into Echo takes about 10 seconds.

  • ●Open Echo and go to the Voice Library (the voices tab in the sidebar)
  • ●Click the Import button at the top of the library
  • ●Select your .onnx file from wherever you saved it
  • ●The voice appears in your library immediately — no restart needed
  • ●There is no limit on the number of models you can import. Import as many as you want

Step 5 — Start Using Your Voice

Once imported, select the voice from your library and click Start. Echo captures your microphone input, runs it through the AI model in real-time, and outputs the converted voice through a virtual audio cable. Any application that uses a microphone — Discord, games, OBS, Zoom — can use the transformed voice by selecting the virtual cable as its input device.

  • ●Pitch offset — If the voice sounds too high or too low, adjust the pitch offset in Echo. For male-to-female conversions, try +12. For female-to-male, try -12. Fine-tune from there based on what sounds natural
  • ●GPU acceleration — Enable GPU mode for the lowest latency. Echo supports NVIDIA (CUDA), AMD/Intel (DirectML), and Apple Silicon (CoreML)
  • ●Virtual audio cable — Make sure VB-Cable (Windows) or BlackHole (macOS) is installed and selected as the output device in Echo. Then select it as the microphone in Discord, OBS, or your game

About the .index File

You will see .index files alongside many .pth downloads. Here is exactly when they matter and when you can ignore them.

  • ●For real-time voice changing in Echo — The .index file is NOT required. Echo's ONNX pipeline runs the full voice conversion without needing the FAISS index. Your voice will sound great without it
  • ●For offline voice conversion (song covers, voiceovers) — The .index file IS useful. Tools like w-okada and Applio use the index during inference to retrieve real acoustic details from the training data, which reduces artifacts and improves fidelity on longer audio
  • ●For model quality assessment — A model that ships with an .index file is usually higher quality. It means the creator took the extra step to generate the index, which correlates with better overall training practices
  • ●Storage — .index files are typically 10–50 MB. Download them when available and keep them with your .pth files in case you want to use offline tools later

Troubleshooting

If a voice does not sound right after importing, the problem is almost always one of these. Most can be fixed without finding a new model.

  • ●Voice sounds robotic or metallic — The model is likely over-trained. Try a different model from the same character, or look for one with fewer training epochs
  • ●Pitch sounds wrong — Adjust the pitch offset in Echo. Male-to-female typically needs +12, female-to-male needs -12. Experiment within that range
  • ●Voice sounds nothing like the target — You may have a v1 model in a v2 pipeline or vice versa. Re-convert with the correct version selected in the model converter
  • ●Output sounds muffled — The model was probably trained on low-quality audio. Find a different model trained on cleaner source material
  • ●Crackling or popping — This is usually an audio pipeline issue, not a model problem. Increase the block size in Echo settings, or check that your GPU is not overloaded
  • ●Conversion sounds generic — For offline tools, this usually means the .index file is missing. For Echo, try a different model — some models simply convert better than others

Want to Create Your Own Voice Model?

If you cannot find the voice you want in community repositories, you can train a custom RVC model from scratch. This is a more advanced process that requires a GPU and some patience, but the tools are free and the community is extremely helpful. You need 10–30 minutes of clean, isolated vocal audio from the target voice. The recommended training tool is Applio (applio.org) — it has the best interface and most active development. For users without a local GPU, Google Colab notebooks provide free cloud GPU access for training. We have a complete step-by-step training guide that covers dataset preparation with UVR5, training parameters, f0 extraction methods, evaluating checkpoints, and troubleshooting common problems.

Ethical Use of Voice Models

RVC voice models are powerful technology that should be used responsibly. Do not create models of real people without their knowledge or consent. Do not use voice models to impersonate someone for deception, fraud, or harassment. If you create AI-generated voice content, label it clearly. Several jurisdictions have enacted legislation targeting AI voice impersonation — using a cloned voice to mislead people is increasingly a criminal offense. The RVC community thrives on mutual respect between creators, model trainers, and users.

FAQ

What file format does Echo use for voice models?
Echo uses .onnx format. Community models are distributed as .pth files, so you need to convert them first using the free browser-based converter at voicechanger.live/tools/model-converter. The conversion takes seconds, runs entirely in your browser, and produces identical voice quality.
How many voice models can I import?
There is no limit. Import as many models as your disk space allows. They are stored locally and you can switch between them instantly during a conversation.
Are RVC voice models free?
The vast majority are free. The community shares models openly on Hugging Face and Discord servers. Thousands of high-quality character, celebrity, and original voice models are available at no cost.
Do I need the .index file?
Not for real-time voice changing in Echo. Echo’s ONNX pipeline works without it. The .index file is useful for offline conversion tools (w-okada, Applio) where it improves quality on longer audio like song covers. Download it when available in case you want to use offline tools later.
What pitch offset should I use?
For male-to-female conversions, start at +12 and adjust. For female-to-male, start at -12. Same-gender conversions usually need 0 or small adjustments. Fine-tune by ear until the voice sounds natural.
What is the difference between v1 and v2 models?
RVC v2 models use higher-dimensional feature vectors (768 vs 256) and sound more natural. Always prefer v2 models. When converting in the model converter, make sure to select the correct version — v1 and v2 are not interchangeable.
Can I train my own voice model?
Yes. You need 10–30 minutes of clean vocal audio and a GPU. The recommended tool is Applio (applio.org). See our complete training guide for the full walkthrough.
Are voice models legal?
Creating and using voice models is legal in most jurisdictions. Using a model to impersonate a real person for fraud or harassment is illegal and increasingly prosecuted. Use voice models responsibly.

Ready to try it?

Download Echo and experience AI-powered voice conversion for yourself.

Download Echo