Audio Demo · Electrolaryngeal Voice Conversion · Interspeech 2026

A Preclinical Study of Electrolaryngeal Voice Conversion for a Novel Nasal Electrolarynx: Feature Choice and Data Augmentation

Abstract

Electrolaryngeal (EL) speech often exhibits reduced intelligibility and naturalness due to imperfect excitation and fixed pitch. In this preclinical study, we present the first systematic investigation of electrolaryngeal voice conversion (ELVC) for speech produced by a newly invented nasal electrolarynx (NEL), extending previous work that mainly targets cervical EL (CEL) speech. We examine the choice of features under NEL’s distinctive acoustics by comparing Mel-spectrograms and WavLM features as input to a seq2seq ELVC system. To address NEL data scarcity, we propose an LLE-VC-based augmentation strategy to synthesize paired sNEL–sNL speech for VC pretraining. Experiments reveal device-dependent trends: WavLM features benefit CEL-oriented conversion and augmentation, whereas Mel-spectrogram inputs better preserve NEL-specific spectral characteristics and yield better intelligibility-related metrics and listener preferences for NEL-to-NL conversion.

Architecture and Experimental Settings

Overview of CEL and NEL speech production

Fig. 1. Overview of CEL and NEL speech production. The leftmost panel shows the use of a traditional cervical EL (CEL) device, the rightmost panel shows the use of a nasal EL (NEL) device, and the two middle panels illustrate the excitation pathways of the two EL devices.

Training process of VTN-VC and ETN-VC

Fig. 2. The training process of VTN-VC and ETN-VC models.

VTN-VC
Three stages: decoder pretraining and encoder pretraining on a large-scale NL corpus (COSPRO), followed by VC training on the paired EL–NL dataset.
ETN-VC
Same two pretraining stages, plus an additional VC pretraining stage on the synthetic sNEL–sNL pairs combined with the NEL–NL pairs; the final VC training stage then uses the NEL–NL pairs only.
Data Augmentation
sNL speech is generated by F5-TTS from the 10,000-sentence TWnews corpus; sNEL speech is obtained by converting sNL with LLE-VC (WavLM-Large layer 6 for neighbor search), reconstructing either WavLM features or the corresponding Mel-spectrograms.

EL Dataset

  • CEL–NL: 320 parallel utterances (TMHINT sentences)
  • NEL–NL: 320 parallel utterances (TMHINT sentences)
  • All recorded by the same healthy speaker; NL and CEL in a studio, NEL in a hospital ward
  • 240 / 40 / 40 pairs for training / development / evaluation

Pretraining Data

  • COSPRO: 45k utterances (44 h) for decoder and encoder pretraining
  • TWnews: 10k sNEL–sNL synthetic pairs for VC pretraining (1k / 5k / 10k subsets compared)
  • Features: 80-dim Mel-spectrogram (16 kHz) or 1024-dim WavLM-Large layer-6 features; HiFi-GAN vocoder trained per feature type

Evaluation Metrics

Spectral Distortion

  • MCDMel-cepstral distortion ↓

Intelligibility

  • CERCharacter error rate (Whisper-Large ASR) ↓
  • SERSyllable error rate (Mandarin syllable recognizer trained on MATBN) ↓
Error Rate = (S + D + I) / N × 100%

F0 / Pitch

  • F0 RMSERoot mean square error of F0 ↓
  • F0 CORRCorrelation of F0 contours ↑

Duration

  • DDURAverage absolute duration difference ↓

Neural Assessment

  • UTMOSMOS prediction ↑
  • MOSA-Net+Quality and intelligibility scores ↑
  • SpeechBERTScoreSemantic fidelity ↑

Subjective

  • A/B Test30 native Mandarin listeners select the more intelligible sample or "no preference."

Experimental Results

Table I. VTN-VC on CEL-to-NL and NEL-to-NL

Table I. Performance comparison of VTN-VC models on CEL/NEL-to-NL tasks. Metrics of unprocessed raw speech are included for reference. (↑ indicates higher is better, ↓ indicates lower is better.) CER and SER are reported as mean ± 95% confidence interval.

ModelFeatureEL TypeMCD ↓f0RMSE ↓f0CORR ↑DDUR ↓CER (%) ↓SER (%) ↓
✗✗CEL10.7934.10-0.080.8875.3 ± 6.1569.3 ± 5.2
NEL13.7548.810.022.6890.8 ± 2.4189.5 ± 6.3
VTN-VCMel-spectrogramCEL5.8834.460.270.1173.3 ± 6.3364.5 ± 5.6
NEL5.6634.620.250.2572.0 ± 6.6458.8 ± 6.4
WavLMCEL5.9733.570.300.2069.0 ± 6.6557.5 ± 6.3
NEL6.1233.320.310.2971.5 ± 7.6862.0 ± 5.6

Although NEL speech yields worse metric values and higher ASR errors than CEL speech, the overall conversion difficulty is comparable. Feature preferences are opposite across devices: Mel-spectrograms work better for NEL-to-NL, whereas WavLM features are more effective for CEL-to-NL.

Zero-shot VC baselines
SystemCEL SER (%) ↓NEL SER (%) ↓
SeedVC71.884.8
MKL-VC85.099.5
Vevo84.0101.5

SER of several zero-/few-shot VC systems on CEL and NEL speech. These results suggest that general-purpose zero-shot VC is insufficient for CEL/NEL speech, motivating the need to tailor and adapt VC models to their unique acoustic characteristics.

Table II. ETN-VC with data augmentation

Table II. ETN-VC trained on the TMHINT training set with 1k/5k/10k TWnews utterances. CER and SER are reported as mean ± 95% confidence interval.

ModelFeatureVC Pretraining DataMCD ↓f0RMSE ↓f0CORR ↑DDUR ↓CER (%) ↓SER (%) ↓
TWnewsTMHINT
ETN-VCMel-spectrogram1K✓5.6433.970.270.1466.5 ± 7.7258.8 ± 5.1
5K✓5.6333.230.300.1666.3 ± 7.6259.3 ± 5.4
10K✓5.5634.110.280.1463.0 ± 7.3753.8 ± 4.8
WavLM1K✓6.0133.390.310.2771.3 ± 7.0161.8 ± 5.6
5K✓6.0734.180.290.2868.8 ± 6.7558.0 ± 5.8
10K✓6.0934.000.290.2765.8 ± 7.3954.3 ± 5.6

ETN-VC (with data augmentation) outperforms VTN-VC regardless of the feature type, and the performance improves with increasing augmentation data. Compared with VTN-VC using Mel-spectrograms (Table I), the best ETN-VC model with Mel-spectrograms reduces the CER from 72.0% to 63.0% and the SER from 58.8% to 53.8%.

Table III. Neural assessment scores

Table III. SpeechBERTScore, UTMOS, and MOSA-Net+ scores for NEL, NL, and converted speech (mean ± 95% confidence interval).

ModelSpeechBERTScore ↑UTMOS ↑MOSA-Net+
Quality ↑Intelligibility ↑
NEL Speech0.523 ± 0.0061.299 ± 0.0052.153 ± 0.0620.642 ± 0.022
VTN-VC Mel0.674 ± 0.0082.469 ± 0.0713.511 ± 0.0760.969 ± 0.004
VTN-VC WavLM0.678 ± 0.0082.944 ± 0.0923.783 ± 0.0930.980 ± 0.004
ETN-VC Mel0.675 ± 0.0092.666 ± 0.0843.584 ± 0.0790.973 ± 0.004
ETN-VC WavLM0.674 ± 0.0072.998 ± 0.0943.788 ± 0.0730.981 ± 0.003
NL Speech1.000 ± 0.0003.261 ± 0.0694.510 ± 0.0760.996 ± 0.001

WavLM-based VC models achieve higher non-intrusive neural assessment scores (UTMOS and MOSA-Net+), suggesting smoother prosody and more natural signal flow. In contrast, Mel-spectrogram-based models perform better on intelligibility-related metrics (CER/SER) and in A/B tests, highlighting the trade-off between naturalness and intelligibility. SpeechBERTScore differences are small across systems, indicating good semantic preservation after conversion. Given the medical context of NEL speech, intelligibility is the focus of this study.

A/B test results

Fig. 4. A/B test results on intelligibility. The bars represent the percentage of subjects' voting for each system. Three sets of A/B tests were conducted with 30 participants (native Mandarin speakers, 18 males and 12 females, aged 20–50, self-reported normal hearing). Mel-spectrogram-based VTN-VC and ETN-VC both outperform their WavLM-based counterparts, and ETN-VC with Mel-spectrograms outperforms VTN-VC with Mel-spectrograms, confirming the benefit of data augmentation for nasal ELVC.

Audio Samples

CEL Speech
NEL Speech
VTN-VC CEL (WavLM)
VTN-VC NEL (Mel)
ETN-VC CEL (WavLM)
ETN-VC NEL (Mel)
NL Speech

CEL conversion uses WavLM features; NEL conversion uses Mel-spectrograms.

Sample 1 他捐了很多衣物給災區 tā juān le hěn duō yī wù gěi zāi qū
CEL Speech
NEL Speech
VTN-VC CEL
VTN-VC NEL
ETN-VC CEL
ETN-VC NEL
NL Speech
Sample 2 電視報導那裡發生地震 diàn shì bào dǎo nà lǐ fā shēng dì zhèn
CEL Speech
NEL Speech
VTN-VC CEL
VTN-VC NEL
ETN-VC CEL
ETN-VC NEL
NL Speech
Sample 3 我把不用的傢俱送人了 wǒ bǎ bù yòng de jiā jù sòng rén le
CEL Speech
NEL Speech
VTN-VC CEL
VTN-VC NEL
ETN-VC CEL
ETN-VC NEL
NL Speech
Sample 4 昨天他向我借了三百塊 zuó tiān tā xiàng wǒ jiè le sān bǎi kuài
CEL Speech
NEL Speech
VTN-VC CEL
VTN-VC NEL
ETN-VC CEL
ETN-VC NEL
NL Speech

Spectrogram Analysis

Mel-spectrograms of NL, CEL, NEL, sNEL, and ETN-VC-converted speech

Fig. 3. Mel-spectrograms of NL, CEL, NEL, sNEL, and ETN-VC-converted speech. NEL speech is longer, while sNEL and converted speech follow the NL duration due to frame-wise conversion and seq2seq alignment, respectively. Mel-spectrograms synthesized directly by LLE-VC preserve the mid-high-frequency resonance of NEL speech much better, whereas sNEL speech generated from reconstructed WavLM features fails to preserve this resonance and introduces low-frequency noise.

Extended Mel-spectrogram comparison

Fig. 3 (extended). Extended version of Fig. 3 with two additional panels: sNEL synthesized via the "Mel-spectrogram → waveform → Mel-spectrogram" route, which exhibits a similar but milder loss of the NEL resonance, and the output of VTN-VC with Mel-spectrograms for comparison with ETN-VC.

Conclusion

This paper presents a preclinical, systematic study of ELVC for a newly invented NEL device, extending previous ELVC work that mainly targets CEL speech. To mitigate NEL data scarcity, we propose an LLE-VC-based data augmentation strategy that synthesizes paired sNEL–sNL speech for VC pretraining, yielding consistent gains in NEL-to-NL conversion. Our experiments further show that feature choice is device-dependent: WavLM features benefit CEL-to-NL ELVC, whereas Mel-spectrogram inputs better preserve NEL-specific spectral characteristics and lead to improved intelligibility-related metrics and listener preferences for NEL-to-NL conversion. Because this preclinical study used the speech of a single healthy speaker, validation with real patient NEL speech remains an important next step. Future work will collect patient NEL data for formal perceptual evaluation, including naturalness MOS, and will examine broader SSL features, such as layer sweeps and alternative encoders.