explainers10 min read · Updated 2026-05-28

What is RVC?

Retrieval-based Voice Conversion is the open-source AI technology that powers the most realistic real-time voice changers. Here is how it works — from the neural network architecture to the community ecosystem.

RVC in Plain Language

RVC stands for Retrieval-based Voice Conversion. It is an open-source AI system that transforms one person's voice into another's — in real time — while preserving everything about how you speak: your words, your timing, your emotion, and your intonation. Only the voice identity changes. The result sounds like a genuinely different person said your exact words, not a filtered or pitch-shifted version of your own voice.

How RVC Works — The Four-Stage Pipeline

When you speak through an RVC model, your audio passes through four distinct processing stages before reaching the output. Understanding these stages explains why RVC sounds dramatically more realistic than traditional voice effects.

  • ●Stage 1 — Content Feature Extraction: A neural network called HuBERT (or its variant ContentVec) analyzes your raw audio and extracts speaker-invariant content features — essentially what you are saying and how you are saying it, stripped of your voice identity. RVC v2 uses 768-dimensional feature vectors (up from 256 in v1), capturing more nuanced speech detail.
  • ●Stage 2 — Pitch Extraction: A separate model called RMVPE (Robust Multiscale Waveform-based Pitch Estimation) extracts the fundamental frequency (F0) of your speech — your pitch contour, melody, and prosody. This is what preserves your natural intonation in the converted output.
  • ●Stage 3 — Index Retrieval (FAISS): This is the "retrieval" in Retrieval-based Voice Conversion. During training, features from the target voice are stored in a vector database. During inference, FAISS (Facebook AI Similarity Search) finds the closest matching voice characteristics for each segment of your speech. This reduces "timbre leakage" — where the original voice bleeds through — by grounding the conversion in real samples from the target voice.
  • ●Stage 4 — Neural Synthesis (VITS / HiFi-GAN Vocoder): A generative model based on VITS (Variational Inference with adversarial learning for end-to-end Text-to-Speech) combines your extracted content features, pitch data, and retrieved voice characteristics to synthesize a completely new audio waveform. The output is not a modified version of your voice — it is entirely new audio generated by the neural network.

Why RVC Sounds More Realistic Than Pitch Shifting

Traditional voice changers apply audio effects to your existing voice — pitch shifting, formant shifting, EQ filtering. This is like applying an Instagram filter to a photograph: you can change the appearance, but it is obviously still the same image underneath. The original voice characteristics always bleed through, especially in sustained speech. RVC works fundamentally differently. It deconstructs your speech into content and pitch, discards your voice identity entirely, and reconstructs the audio from scratch using a trained neural network. The output is a new voice, not a filtered version of the old one. This is why RVC voices are often indistinguishable from real human speech — because they are synthesized by a model that learned how humans actually sound.

RVC Model Files Explained

When people share or download RVC voice models, three file types are involved. Understanding what each does helps you get the best results.

  • ●.pth (PyTorch Model) — The core voice model containing the trained neural network weights. This is the primary file required for voice conversion. A typical .pth model is 50-100 MB and captures the target voice's complete vocal characteristics: timbre, resonance, breathiness, and formant structure.
  • ●.index (FAISS Index) — An optional but recommended companion file that contains a searchable database of feature vectors from the training audio. During inference, the model uses this index to retrieve real acoustic patterns from the training data, significantly improving output quality and reducing artifacts. Without the .index file, conversion still works but may sound more generic.
  • ●.onnx (Open Neural Network Exchange) — An optimized, portable model format. Converting a .pth model to .onnx enables faster inference, lower memory usage, and compatibility with non-Python runtimes. Applications like Echo Live use .onnx models because they can be loaded natively without requiring a Python interpreter or PyTorch installation.

RVC v1 vs v2

RVC v2 is the current standard. The most significant improvement over v1 is the feature extraction dimensionality: v1 used 256-dimensional feature vectors from HuBERT, while v2 uses 768-dimensional vectors. This higher dimensionality captures more detailed voice characteristics, producing more natural and accurate conversions — especially for cross-gender voice changing, where the difference is most audible. RVC v2 also benefits from improved training stability and reduced inference latency. The trade-off is slightly higher compute requirements. Importantly, v1 and v2 models are not interchangeable — a model trained on v1 cannot be used in a v2 pipeline without retraining. For new projects, always use v2.

RVC vs Other Voice AI Technologies

RVC is not the only voice AI technology, but it occupies a specific niche that makes it ideal for real-time voice changing.

  • ●RVC vs Pitch Shifting — Pitch shifting modifies frequency. RVC generates entirely new audio. Pitch shifting sounds robotic; RVC sounds human. There is no comparison in quality.
  • ●TTS and RVC — TTS generates speech from text; RVC changes the voice in audio. Echo Live combines both: Kokoro or Supertonic generates high-quality speech, then RVC applies any character voice with a compatible model. Generate locally, export a WAV, or save the clip to the Soundboard. Custom model imports require Pro.
  • ●RVC vs SVC (Singing Voice Conversion) — SVC is RVC's cousin, optimized for singing. It handles sustained notes, vibrato, and musical phrasing better than standard RVC. Use SVC for AI song covers; use RVC for real-time voice changing.
  • ●RVC vs GPT-SoVITS / VALL-E — These newer architectures offer zero-shot voice cloning: convert your voice to a target with just a few seconds of reference audio, no training required. The trade-off is higher compute cost, less consistent results, and they are not yet practical for low-latency real-time use. RVC remains the most battle-tested option for real-time conversion.

Real-Time Performance

RVC's architecture is efficient enough to run in real time on consumer hardware. A single inference pass typically takes 10-30ms on a modern GPU (NVIDIA GTX 1650 or better), making it fast enough for live conversation, gaming callouts, and streaming. GPU acceleration via CUDA (NVIDIA), DirectML (AMD/Intel on Windows), or CoreML (Apple Silicon) is strongly recommended for real-time use. CPU-only inference is possible but adds significant latency that makes live conversation impractical for most users.

Where to Find Voice Models

The RVC community has produced thousands of free voice models. Start with AIVoices.gg to find community voices. Voice-Models.com is another dedicated directory, while Hugging Face and the AI Hub community offer additional creator uploads and recommendations. Models exist for original voices, characters, celebrities, and custom creations. Quality varies — models trained on clean, isolated vocal data with 10-30 minutes of audio produce the best results.

The Open-Source Ecosystem

RVC was originally developed by the RVC-Project team on GitHub and released as open-source software. This openness is what makes the ecosystem possible: anyone can train a model, anyone can build an application on top of RVC, and anyone can contribute improvements. Major projects built on RVC include Applio (the leading training tool), w-okada Voice Changer (a multi-model voice conversion server), and Echo Live (a native desktop voice changer focused on real-time RVC with integrated DSP effects and soundboard). The technology is free and permissive — there are no licensing fees or usage restrictions on the core RVC engine.

Getting Started with RVC

The fastest way to use RVC is through a desktop application that handles the pipeline for you. Echo Live (voicechanger.live) packages the entire RVC inference engine into a native installer — no Python, no command line, no manual model management. For users who want to train custom voice models, Applio is the most popular open-source trainer. For technical users who want maximum control across multiple model architectures (including Beatrice and so-vits-svc alongside RVC), w-okada Voice Changer provides a configurable server-based setup.

FAQ

What does RVC stand for?
RVC stands for Retrieval-based Voice Conversion. The "retrieval" refers to the FAISS index lookup that matches your speech features against stored target voice characteristics during inference.
Is RVC free to use?
Yes. RVC is open-source software. The core technology, training tools, and thousands of community voice models are all free. Applications built on RVC (like Echo Live) may have their own pricing — Echo Live is free to use with an optional Pro upgrade for unlimited voices.
Do I need a GPU for RVC?
For real-time voice changing, a dedicated GPU (NVIDIA GTX 1650 or better, or Apple Silicon Mac) is strongly recommended. CPU-only inference works for batch processing but adds too much latency for live conversation.
Can I train my own RVC voice model?
Yes. You need 10-30 minutes of clean, isolated vocal audio from the target voice and a GPU with at least 6 GB VRAM. Applio is the most popular free training tool. Training typically takes 1-4 hours.
What is a .pth file?
A .pth file is a PyTorch model file containing the trained neural network weights of an RVC voice model. It is the core file needed for voice conversion. A typical model is 50-100 MB.
What is the .index file for?
The .index file is a FAISS search index containing feature vectors from the training audio. It improves conversion quality by letting the model retrieve real acoustic patterns from the target voice during inference. It is optional but recommended.
What is the difference between .pth and .onnx?
.pth is the native PyTorch format used for training and Python-based inference. .onnx is an optimized portable format that runs without Python — it is faster, uses less memory, and works with native applications like Echo Live. The voice quality is identical.
Is RVC the same as voice cloning?
RVC is a form of voice cloning applied in real time. It converts your live speech into a target voice. This is different from TTS-based voice cloning, which generates speech from text. RVC preserves your natural speech patterns, emotion, and timing — only the voice identity changes.

Ready to try it?

Download Echo and experience AI-powered voice conversion for yourself.

Download Echo