What is RVC Training?
RVC (Retrieval-based Voice Conversion) training teaches a neural network to reproduce a specific voice. You provide clean audio samples of a target voice, and the model learns that voice's characteristics — timbre, resonance, formant structure, breathiness — so it can convert any input speech into that voice in real time. The trained model does not memorize the training audio. It learns the voice's identity, then applies that identity to new speech while preserving the original speaker's words, timing, emotion, and intonation.
What You Need Before Training
Training requires three things: clean audio data, a training environment, and patience. The quality of your dataset is the single biggest factor in model quality — a short, clean dataset will always beat a long, noisy one.
- ●10-30 minutes of clean, isolated vocal audio — no background music, no noise, no reverb, no echo. 40+ minutes of diverse audio can improve naturalness further, but stay under 60 minutes to avoid diminishing returns
- ●Audio must be a single speaker with reasonably consistent recording conditions
- ●Lossless format preferred: WAV or FLAC, 44.1kHz or 48kHz sample rate, mono channel. Avoid low-bitrate MP3s — compression artifacts get baked into the model
- ●A GPU with at least 6 GB VRAM (NVIDIA recommended for CUDA support). 8 GB+ VRAM is ideal for comfortable batch sizes
- ●Training software: Applio (applio.org — recommended, most active development, best UI), Mangio-RVC, or the original RVC WebUI
- ●No GPU? Use a cloud service: Google Colab (colab.research.google.com — free tier with T4 GPU), Kaggle, Paperspace, or RunPod
Step 1 — Prepare Your Dataset
Dataset preparation is where most people fail. A model trained on noisy, reverberant, or inconsistent audio will sound noisy, reverberant, and inconsistent. The community-standard cleaning pipeline uses Ultimate Vocal Remover (UVR5 — github.com/Anjok07/ultimatevocalremovergui), a free tool that separates vocals from instrumentals and removes reverb using AI.
- ●Source your audio: The best sources are isolated vocal tracks (use a vocal remover to extract vocals from songs), podcast recordings, audiobook narration, interview clips, or direct microphone recordings in a quiet room
- ●Cleaning pipeline: Remove instrumentals (UVR5 with Kim Vocal 2 or MDX-NET model) → Remove reverb and echo (UVR5 VR-DeEchoDeReverb) → Extract main vocals → Remove residual noise. Do not over-process — each pass slightly degrades quality
- ●Ensure single speaker only: Remove any segments with overlapping voices, background chatter, or other speakers
- ●Include variety: Different pitches, emotions, speaking speeds, and vocal registers help the model generalize. Monotone datasets produce monotone models
- ●Trim silence and non-speech: Remove long pauses, coughs, and non-vocal sounds. Keep natural breathing — removing all breaths can make the model sound robotic
- ●Do NOT cut words in the middle: Keep phrases intact to preserve natural intonation patterns
- ●File paths: Ensure your dataset folder path contains no spaces or special characters — this is a common cause of training failures across all RVC tools
Step 2 — Choose Your Training Tool
Three major open-source tools handle RVC training. Each takes the same dataset and produces the same .pth model format — the difference is in the user experience.
- ●Applio — The most popular and actively developed RVC training tool. Clean UI, automated dataset preprocessing, TensorBoard integration for monitoring loss curves, and support for all major f0 extraction methods. Recommended for most users
- ●Mangio-RVC / RVC-WebUI — The community fork of the original RVC project. Runs through a Gradio web interface in your browser. More manual configuration but well-documented. Good alternative if Applio does not work on your system
- ●Google Colab notebooks — Browser-based training using free cloud GPUs. No local GPU required. Search for "RVC training Colab" for community notebooks. Best for users without a dedicated GPU
Step 3 — Configure Training Parameters
Understanding these parameters helps you get better results. Start with the recommended defaults and adjust based on what you hear.
- ●Epochs (200-500): An epoch is one full pass through your training data. Start with 200-300 and evaluate. More epochs can improve quality, but over-training causes robotic artifacts. There is no magic number — your ears are the judge
- ●Batch Size (depends on VRAM): 6 GB VRAM → batch size 4-6. 8 GB VRAM → batch size 6-8. 12 GB+ VRAM → batch size 8-12. Higher batch sizes train faster but use more memory. If you get CUDA out-of-memory errors, reduce the batch size
- ●Sample Rate: 40kHz is the community standard and works well for most voices. 48kHz captures slightly more high-frequency detail but requires more compute. Use 40kHz unless you have a specific reason to go higher
- ●Save Frequency: Save checkpoints every 10-50 epochs so you can test intermediate versions. This is critical for finding the sweet spot before over-training
- ●RVC Version: Always use v2 for new models. v2 uses 768-dimensional feature vectors (vs 256 in v1), producing more accurate and natural-sounding conversions
F0 Extraction Methods Explained
The f0 (fundamental frequency) extraction method determines how pitch is detected from your training audio. This choice significantly affects model quality.
- ●RMVPE (Recommended) — Robust Multiscale Waveform-based Pitch Estimation. The current community standard. Fast, handles noisy audio well, and produces reliable pitch tracking across all voice types. Use this unless you have a specific reason not to
- ●Crepe — Higher pitch accuracy than RMVPE, especially for singing voices and clean recordings. Slower to process. Best for high-quality studio recordings where maximum fidelity matters
- ●FCPE — Fast, modern algorithm with high pitch precision. Good alternative to RMVPE when speed matters. Slightly less robust with noisy audio
- ●Hybrid (RMVPE + FCPE) — Some training tools offer combined methods. Can improve results in edge cases but adds processing time. Try single methods first
Step 4 — Train and Monitor
Launch training and monitor progress using TensorBoard (integrated in Applio) or by testing saved checkpoints. Training typically takes 1-4 hours on a modern GPU depending on dataset size and epoch count. Watch the loss curve — it should decrease steadily and then flatten. When it flattens, the model has learned what it can from the data. Continuing beyond this point risks over-training. Save checkpoints regularly and test them by running inference on a sample audio file. Listen for naturalness, clarity, and whether the output actually sounds like the target voice. The best checkpoint is rarely the last one — most experienced trainers find the sweet spot somewhere between epoch 200 and 400.
Step 5 — Generate the Index File
After training completes, generate the FAISS index file. This step is often overlooked but significantly improves output quality. The index file contains a searchable database of feature vectors from your training audio. During inference, the model uses this index to retrieve real acoustic patterns from the target voice, reducing "timbre leakage" where the original speaker's voice bleeds through. In Applio, index generation is a one-click step after training. Always include the .index file alongside the .pth model when sharing or importing.
Step 6 — Use Your Model
Your trained model produces two key files: the .pth model file (50-100 MB) and the .index file. To use them in real time, import them into a voice changer application. Echo Live accepts .onnx format models — convert your .pth file using the free converter at voicechanger.live/tools/model-converter (browser-based, no upload). For testing during training, most training tools include a built-in inference tab where you can test the .pth model directly without conversion.
Troubleshooting Common Problems
If your model does not sound right, the cause is almost always one of these issues:
- ●Model sounds robotic or metallic — Over-training. Roll back to an earlier checkpoint (try epoch 150-250). If all checkpoints sound robotic, your dataset likely has reverb or noise that needs cleaning
- ●Model sounds like the original speaker, not the target — The dataset is too short or too homogeneous. Add more variety (different pitches, emotions) and ensure at least 15 minutes of clean audio
- ●Model sounds muffled or low quality — The source audio was low-bitrate or over-compressed. Always use lossless WAV or FLAC. Also check that you are not over-processing with too many UVR5 passes
- ●Training crashes with CUDA out of memory — Reduce batch size. If already at minimum, reduce sample rate or close other GPU-intensive applications
- ●Training loss is not decreasing — Check that your dataset is not full of silence or non-speech audio. Verify the file format is correct (WAV/FLAC, mono, correct sample rate)
- ●Model works well for some voices but not others — RVC conversion quality depends on pitch range similarity between the input and target voice. Male-to-male and female-to-female conversions generally produce the best results