Javascript must be enabled to continue!
Wav2wav: Wave-to-Wave Voice Conversion
View through CrossRef
Voice conversion is the task of changing the speaker characteristics of input speech while preserving its linguistic content. It can be used in various areas, such as entertainment, medicine, and education. The quality of the converted speech is crucial for voice conversion algorithms to be useful in these various applications. Deep learning-based voice conversion algorithms, which have been showing promising results recently, generally consist of three modules: a feature extractor, feature converter, and vocoder. The feature extractor accepts the waveform as the input and extracts speech feature vectors for further processing. These speech feature vectors are later synthesized back into waveforms by the vocoder. The feature converter module performs the actual voice conversion; therefore, many previous studies separately focused on improving this module. These works combined the separately trained vocoder to synthesize the final waveform. Since the feature converter and the vocoder are trained independently, the output of the converter may not be compatible with the input of the vocoder, which causes performance degradation. Furthermore, most voice conversion algorithms utilize mel-spectrogram-based speech feature vectors without modification. These feature vectors have performed well in a variety of speech-processing areas but could be further optimized for voice conversion tasks. To address these problems, we propose a novel wave-to-wave (wav2wav) voice conversion method that integrates the feature extractor, the feature converter, and the vocoder into a single module and trains the system in an end-to-end manner. We evaluated the efficiency of the proposed method using the VCC2018 dataset.
Title: Wav2wav: Wave-to-Wave Voice Conversion
Description:
Voice conversion is the task of changing the speaker characteristics of input speech while preserving its linguistic content.
It can be used in various areas, such as entertainment, medicine, and education.
The quality of the converted speech is crucial for voice conversion algorithms to be useful in these various applications.
Deep learning-based voice conversion algorithms, which have been showing promising results recently, generally consist of three modules: a feature extractor, feature converter, and vocoder.
The feature extractor accepts the waveform as the input and extracts speech feature vectors for further processing.
These speech feature vectors are later synthesized back into waveforms by the vocoder.
The feature converter module performs the actual voice conversion; therefore, many previous studies separately focused on improving this module.
These works combined the separately trained vocoder to synthesize the final waveform.
Since the feature converter and the vocoder are trained independently, the output of the converter may not be compatible with the input of the vocoder, which causes performance degradation.
Furthermore, most voice conversion algorithms utilize mel-spectrogram-based speech feature vectors without modification.
These feature vectors have performed well in a variety of speech-processing areas but could be further optimized for voice conversion tasks.
To address these problems, we propose a novel wave-to-wave (wav2wav) voice conversion method that integrates the feature extractor, the feature converter, and the vocoder into a single module and trains the system in an end-to-end manner.
We evaluated the efficiency of the proposed method using the VCC2018 dataset.
Related Results
Makna Voice Over dalam Pemberitaan Feature di Televisi
Makna Voice Over dalam Pemberitaan Feature di Televisi
Abstract. Voice Over or what is known as VO is being discussed a lot, not only about the profession, but also from the industry side and the various voice over techniques used. Due...
Determinants of employee-prohibitive voice behavior : mediating role of psychological safety
Determinants of employee-prohibitive voice behavior : mediating role of psychological safety
The study of the determinants of employee-prohibitive voice behavior, a relatively new area of research, has recently gained significant attention. Psychological safety has emerged...
Speech, communication, and neuroimaging in Parkinson's disease : characterisation and intervention outcomes
Speech, communication, and neuroimaging in Parkinson's disease : characterisation and intervention outcomes
<p dir="ltr">Most individuals with Parkinson's disease (PD) experience changes in speech, voice or communication. Speech changes often manifest as hypokinetic dysarthria, a m...
Speech, communication, and neuroimaging in Parkinson's disease : characterisation and intervention outcomes
Speech, communication, and neuroimaging in Parkinson's disease : characterisation and intervention outcomes
<p dir="ltr">Most individuals with Parkinson's disease (PD) experience changes in speech, voice or communication. Speech changes often manifest as hypokinetic dysarthria, a m...
Speech, communication, and neuroimaging in Parkinson's disease : Characterisation and intervention outcomes
Speech, communication, and neuroimaging in Parkinson's disease : Characterisation and intervention outcomes
<p dir="ltr">Most individuals with Parkinson's disease (PD) experience changes in speech, voice or communication. Speech changes often manifest as hypokinetic dysarthria, a m...
Brain mechanism of unfamiliar and familiar voice processing: an activation likelihood estimation meta-analysis
Brain mechanism of unfamiliar and familiar voice processing: an activation likelihood estimation meta-analysis
Interpersonal communication through vocal information is very important for human society. During verbal interactions, our vocal cord vibrations convey important information regard...
How to speak and vocal hygiene
How to speak and vocal hygiene
An abnormal tongue shape, pitch difference or voice quality can lead to difficulty communicating effectively. Common among teachers are voice issues, which can be uncomfortable and...
VOICE COMMERCE (V-COMMERCE): EXPLORING THE INTEGRATION OF VOICE ASSISTANTS IN ONLINE SHOPPING
VOICE COMMERCE (V-COMMERCE): EXPLORING THE INTEGRATION OF VOICE ASSISTANTS IN ONLINE SHOPPING
Voice commerce, also known as v-commerce, refers to the process of purchasing goods and services using voice commands or virtual assistants, typically through devices equipped with...

