PronounSE: SFX Synthesizer from Language-Independent Vocal Mimicry

Riki Takizawa1, 4, Shigeyuki Hirai2, Asako Kanezaki3, Hitoshi Suda4
1Graduate School of Kyoto Sangyo University, Japan
2Kyoto Sangyo University, Japan
3Institute of Science Tokyo, Japan
4National Institute of Advanved Industrial Science and Technology, Japan

Abstract

Overview image

This study proposes PronounSE, a sound effect synthesis method that utilizes only non-verbal vocal expressions mimicking sound effects, without requiring prior knowledge or experience in sound production techniques. Focusing on human articulatory capabilities, PronounSE learns the correspondence between variations in vocal mimicry and actual sound effects, enabling the generation of sound effects that reflect the nuanced characteristics of vocal input. In this study, we constructed a new dataset centered on explosion sounds and trained PronounSE accordingly, successfully synthesizing explosion sounds that reflect the nuances of vocal mimicry. Furthermore, we designed novel evaluation metrics and conducted comparative experiments with existing methods.

Synthesis Sample

Input Vocal Mimicry: windy nuance with the explosion

Synthesized Audio: nuance-reflective synthesized sound

Comparative Evaluation with Other methods

In this study, we introduced new subjective evaluation metrics focusing on the similarity between the reference and synthesized sounds, as well as the extent to which vocal mimicry nuances are reflected in the synthesized outputs. In addition, we conducted conventional subjective evaluations on the naturalness and sound quality of the synthesized sounds. For comparison, we evaluated two Audio-to-Audio methods: T-Foley [1] and Stable Audio 2.0 [2].

Using an evaluation dataset consisting of vocal mimicry samples and reference sounds from six speakers, each method was used to synthesize audio. An example of the reference sound, the six vocal imitations, and the corresponding synthesized sounds is shown below.

Reference

Vocal Mimicry and Synthesized Sounds

Speaker Vocal Mimicry PronounSE T-Foley Stable Audio 2.0
m-01
m-02
m-03
m-04
m-05
f-01

Objective Fidelity Evaluation with FAD [3] and Cosine Similarity
  Objective evaluation results on FAD between the reference and synthesized sound sets, and the average cosine similarity between the embedding vectors of corresponding sound pairs.

Speaker ID FAD by PANNs ↓ Cosine similarity ± SD ↑
PronounSE T-Foley Stable Audio 2.0 PronounSE T-Foley Stable Audio 2.0
m-01 17.85 54.02 21.35 0.85 ± 0.08 0.73 ± 0.07 0.83 ± 0.06
m-02 17.05 54.47 20.94 0.86 ± 0.06 0.74 ± 0.07 0.84 ± 0.06
m-03 17.28 48.85 20.95 0.86 ± 0.07 0.75 ± 0.07 0.84 ± 0.06
m-04 16.65 48.52 21.46 0.85 ± 0.07 0.75 ± 0.08 0.84 ± 0.06
m-05 19.29 47.89 19.75 0.86 ± 0.07 0.75 ± 0.08 0.85 ± 0.07
f-01 19.28 56.94 22.01 0.85 ± 0.07 0.72 ± 0.07 0.85 ± 0.06
Whole 14.09 55.93 21.15 0.85 ± 0.07 0.74 ± 0.07 0.84 ± 0.06

Subjective Fidelity Evaluation (Fidelity MOS ± 95% CI)
  5-point MOS resuts on the fidelity of synthesized sounds to reference sounds.

Speaker ID PronounSE T-Foley Stable Audio 2.0
m-012.57 ± 0.121.54 ± 0.092.67 ± 0.13
m-022.23 ± 0.111.29 ± 0.062.66 ± 0.12
m-032.67 ± 0.121.44 ± 0.072.71 ± 0.12
m-042.56 ± 0.121.47 ± 0.082.68 ± 0.12
m-052.31 ± 0.141.57 ± 0.092.89 ± 0.12
f-012.22 ± 0.121.31 ± 0.072.52 ± 0.12
Whole2.43 ± 0.051.44 ± 0.032.69 ± 0.05

Subjective Nuance Reflection Evaluation (Attack MOS ± 95% CI, Release MOS ± 95% CI, Overall MOS ± 95% CI)
  5-point MOS results on the nuance reflection of vocal mimicry in the synthesized sounds, evaluated from three perspectives: the attack phase, the release phase, and the overall impression.

Speaker ID Attack (Onset) Release (Decay) Overall Impression
PronounSET-FoleyStable Audio 2.0 PronounSET-FoleyStable Audio 2.0 PronounSET-FoleyStable Audio 2.0
m-01 3.13 ± 0.122.95 ± 0.113.08 ± 0.12 3.03 ± 0.122.87 ± 0.112.98 ± 0.12 3.04 ± 0.122.80 ± 0.112.94 ± 0.12
m-02 2.87 ± 0.132.65 ± 0.122.93 ± 0.12 3.05 ± 0.122.68 ± 0.122.93 ± 0.12 2.92 ± 0.122.45 ± 0.122.85 ± 0.11
m-03 2.92 ± 0.122.55 ± 0.122.90 ± 0.12 2.86 ± 0.122.58 ± 0.122.80 ± 0.11 2.81 ± 0.112.50 ± 0.112.76 ± 0.11
m-04 3.03 ± 0.122.76 ± 0.123.07 ± 0.12 3.12 ± 0.112.81 ± 0.122.93 ± 0.11 3.03 ± 0.112.71 ± 0.122.95 ± 0.12
m-05 2.74 ± 0.132.64 ± 0.122.78 ± 0.11 2.83 ± 0.122.63 ± 0.112.75 ± 0.11 2.69 ± 0.122.61 ± 0.112.63 ± 0.10
f-01 2.85 ± 0.122.84 ± 0.122.86 ± 0.13 2.92 ± 0.122.76 ± 0.122.71 ± 0.12 2.81 ± 0.112.79 ± 0.112.68 ± 0.12
Whole 2.92 ± 0.052.73 ± 0.052.94 ± 0.05 2.97 ± 0.052.72 ± 0.052.85 ± 0.05 2.88 ± 0.052.64 ± 0.052.80 ± 0.05

Subjective Naturalness and Sound Quality Evaluation (Naturallness MOS ± 95% CI, Quality MOS ± 95% CI)
  5-point MOS results on the naturalness and audio quality of the synthesized sounds.

Speaker ID Naturalness (Ground Truth: 3.70 ± 0.11) Sound Quality (Ground Truth: 3.70 ± 0.10)
PronounSET-FoleyStable Audio 2.0 PronounSET-FoleyStable Audio 2.0
m-01 3.50 ± 0.112.86 ± 0.123.29 ± 0.11 3.36 ± 0.112.73 ± 0.123.36 ± 0.10
m-02 3.09 ± 0.122.56 ± 0.113.42 ± 0.10 3.16 ± 0.112.47 ± 0.113.45 ± 0.10
m-03 3.46 ± 0.112.82 ± 0.123.48 ± 0.10 3.32 ± 0.102.55 ± 0.113.52 ± 0.10
m-04 3.55 ± 0.112.95 ± 0.123.57 ± 0.10 3.39 ± 0.102.83 ± 0.123.48 ± 0.10
m-05 2.93 ± 0.122.69 ± 0.113.72 ± 0.10 2.95 ± 0.112.49 ± 0.113.57 ± 0.10
f-01 3.14 ± 0.122.76 ± 0.123.53 ± 0.11 2.99 ± 0.112.69 ± 0.113.45 ± 0.10
Whole 3.28 ± 0.052.77 ± 0.053.50 ± 0.04 3.19 ± 0.042.62 ± 0.053.47 ± 0.04

Links

Paper PronounSE implementation and trained model Evaluation Dataset

Citation

If you use this work, please cite our paper:

   @article{hoge,
        title   = {hogehoge},
        author  = {Riki Takizawa, Shigeyuki Hirai, Asako Kanezaki, Hitoshi Suda},
        journal = {hoge},
        volume  = {hoge},
        number  = {hoge},
        pages   = {hoge},
        year    = {hoge},
        doi     = {hoge}
    }

Reference

[1]: Y. Chung, J.Lee, and J. Nam. "T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis". In: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE.
[2]: Stability AI. "Stable Audio 2.0". https://stability.ai/stable-audio.
[3]: ilgour, K., Zuluaga, M., Roblek, D. and Sharifi, M.: Fr´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms, Proceedings of the 20th Annual Conference of the International Speech Communication Association (INTERSPEECH), pp. 2350–2354