RVC troubleshooting guide
Start by finding where the problem is
Most RVC problems get worse when you start changing random settings. The better move is to split the voice chain into three parts: your raw microphone input, the AI voice conversion output, and the final app that receives the voice.
This is the same basic troubleshooting pattern used by real-time RVC tools like w-okada Voice Changer: first prove whether the microphone audio is clean, then prove whether the converted voice sounds right, then check Discord, OBS, VRChat, Roblox, or whatever app is receiving the virtual microphone.
In Echo, start with Monitor / Hear Myself through headphones. If your unconverted microphone already sounds noisy, clipped, echoey, or too quiet, fix that before touching the RVC model. If the Echo monitor sounds good but Discord sounds bad, the problem is probably Discord input, virtual cable routing, or extra noise processing in the target app.
If the RVC voice sounds robotic or metallic
Robotic RVC output usually comes from one of four places: weak input audio, a weak model, pitch mismatch, or settings that do not give the model enough context. Applio's inference guide points to the same root causes: improve input quality, check model training, clean the audio, verify dataset quality, and experiment with advanced settings.
First, test a different voice model. If one model sounds metallic and another sounds natural with the same microphone, the model is the problem. It may be undertrained, overtrained, trained on noisy audio, or simply not a good match for your speaking range.
Next, raise Extra Context one step. Extra Context gives the model more surrounding audio so syllables connect smoothly across chunks. In Echo, the default is 4096. If the voice is stable but metallic or chopped, try 8192, then 16384. If it becomes laggy or unstable, move one step back.
If the voice identity is strong but artifacts get worse when Index Rate is high, lower Index Rate. Index retrieval can improve target-speaker character, but too much blending can also pull in weird artifacts when the model, index file, and input voice do not match cleanly.
If the voice crackles, pops, or drops out
Crackling is usually a performance problem, not a "bad voice" problem. The computer is being asked to process audio faster than it can reliably handle, so the stream breaks into pops, gaps, or crunchy audio.
Raise Chunk Size one step. Echo exposes 2048, 3072, 4096, 6144, 8192, 10240, and 16384. Lower Chunk Size means less delay, but more GPU or CPU pressure. Higher Chunk Size is more stable, but adds delay. Echo's balanced default is 6144, and the conservative presets move higher for older hardware or CPU mode.
Close GPU-heavy apps while testing. Games, screen recording, video rendering, browser video, and AI tools can compete with RVC inference. If the voice only crackles while a game is open, you may need a more conservative preset while playing.
If you are on CPU mode, start conservative. CPU mode can work, but it has much less headroom than GPU acceleration. Use a higher Chunk Size first, then lower it only after the audio is stable.
If the voice has too much delay
Delay comes from several small buffers stacking together: Chunk Size, inference time, Crossfade, Extra Context, virtual audio routing, and the receiving app. The trap is lowering everything at once until the voice crackles.
Start from a stable preset. Then lower Chunk Size one step and test for a full minute of normal talking. If it stays clean, lower one more step. If crackles appear, go back up. This gives you the lowest stable setting for your hardware instead of a random low-latency guess.
Keep Crossfade and Extra Context reasonable. Crossfade smooths the boundary between chunks, and Extra Context helps the model keep continuity. Lowering both can reduce processing work, but if you go too low the voice may sound chopped, metallic, or unstable.
Test Echo directly before blaming Discord or OBS. If Echo's monitor feels responsive but the voice is delayed in another app, check that app's input device, noise suppression, monitoring, and audio buffer settings.
If the pitch or gender conversion sounds wrong
RVC changes voice identity, but pitch still matters. If your natural voice is much deeper or higher than the target model, the output can sound fake even when the model is good.
Use Pitch Shift in semitones. For a deep voice going into a higher target voice, move pitch upward gradually. For a high voice going into a lower target voice, move pitch downward gradually. Echo onboarding uses +8 to +12 as a common starting range for deep-to-higher voices and -4 to -10 for high-to-lower voices, but you should tune by ear.
Use RMVPE as the default pitch extractor. The RVC ecosystem generally treats RMVPE as the modern reliable choice, and Echo defaults to RMVPE. CREPE is still worth testing if one specific model behaves strangely, especially with singing-style material or unusual pitch tracking.
Do not overcorrect. If +12 sounds cartoonish, try +6 or +8. If -10 sounds muddy, try -4 or -6. The goal is not to force a perfect octave jump; it is to put your performance into a range the target model can handle naturally.
If background noise follows the AI voice
RVC does not magically know which parts of your microphone are "you" and which parts are your fan, keyboard, room echo, or speakers. If the input noise is loud enough, the model may convert it into weird breathy or buzzy artifacts.
Fix the room first: use headphones, move speakers away from the mic, reduce fan noise, lower keyboard noise, and avoid talking in a very reflective room. Echo can help with Neural Denoise, Highpass Filter, Silence Threshold, and Echo Cancellation, but these are cleanup tools, not a replacement for a clean signal.
Be careful with double noise suppression. If Echo, Discord, OBS, and a headset app all process the same voice, the RVC output can get thin, gated, or watery. Use only the filters you actually need and test one change at a time.
If a custom model is noisy no matter what microphone you use, the noise may be baked into the model. Applio's dataset guide is very clear on this point: clean, consistent, low-noise source audio matters. Background music, reverb, clicks, coughs, and room noise can survive training unless the dataset is cleaned before training.
If the model or index file is the problem
Model file problems are boring, but they cause a lot of fake "quality" issues. Applio uses `.pth` and `.index` files for many workflows. Echo uses `.onnx` models for live use, so a `.pth` model normally needs to be converted before you use it live in Echo.
Keep filenames simple. Applio's troubleshooting docs call out spaces, special characters, and non-standard path characters as common sources of errors. For model work, boring names are good names: short, plain English, no symbols, no emoji, no nested mystery folders.
Treat the index file as optional power, not magic. Index Rate requires a matching index file. If the base model sounds bad, the index will not save it. If the base model sounds good, a small amount of index blending can sometimes add character, but too much can add artifacts.
If you trained a model yourself and every live app sounds bad, go back to the dataset. Use vocal isolation tools like UVR5 to remove music, reverb, backing vocals, and noise before retraining. A clean model is easier to tune than a messy model with heroic settings.
Platform checks for Discord, OBS, VRChat, and games
When Echo sounds good in your headphones but bad elsewhere, stop tuning RVC and check routing. The target app should use the virtual microphone or virtual cable that receives Echo's processed output, not your physical microphone.
Discord: select the Echo virtual mic or virtual cable as Input Device. Then test with Discord's mic test. If the voice gets chopped up, try disabling extra processing like automatic gain control, echo cancellation, or noise suppression one at a time.
OBS and Streamlabs: add the virtual microphone as the mic source. Avoid monitoring the same signal twice, because double monitoring creates delay and can cause feedback. If you need filters in OBS, start with none, then add only what the stream actually needs.
VRChat, Roblox, Fortnite, Valorant, and other games: pick the same virtual microphone in the game voice settings. If a game has its own voice activation threshold, lower or raise it after RVC is working, because converted voices can trigger thresholds differently than your raw microphone.
A quick fix order that actually works
Use this order when you are stuck: check raw mic, switch to a known-good voice model, reset Pitch Shift near 0, use RMVPE, use Echo's balanced preset, raise Chunk Size if crackling appears, raise Extra Context if the voice is metallic, lower Index Rate if artifacts appear, then test routing in Discord or OBS.
Only change one setting at a time. Speak normally for at least a few sentences after each change. RVC problems can sound similar, so the fastest path is boring and methodical: one symptom, one setting, one test.
If nothing helps, collect a short description before asking for support: your GPU/CPU, model format, whether the issue happens in Echo monitor or only in another app, Chunk Size, Extra Context, Crossfade, Pitch Shift, Pitch Extractor, and whether you are using an index file. That information turns "it sounds bad" into something people can actually debug.