This study proposes PronounSE, a sound effect synthesis method that utilizes only non-verbal vocal expressions mimicking sound effects, without requiring prior knowledge or experience in sound production techniques. Focusing on human articulatory capabilities, PronounSE learns the correspondence between variations in vocal mimicry and actual sound effects, enabling the generation of sound effects that reflect the nuanced characteristics of vocal input. In this study, we constructed a new dataset centered on explosion sounds and trained PronounSE accordingly, successfully synthesizing explosion sounds that reflect the nuances of vocal mimicry. Furthermore, we designed novel evaluation metrics and conducted comparative experiments with existing methods.
Input Vocal Mimicry: windy nuance with the explosion
Synthesized Audio: nuance-reflective synthesized sound
In this study, we introduced new subjective evaluation metrics focusing on the similarity between the reference and synthesized sounds, as well as the extent to which vocal mimicry nuances are reflected in the synthesized outputs. In addition, we conducted conventional subjective evaluations on the naturalness and sound quality of the synthesized sounds. For comparison, we evaluated two Audio-to-Audio methods: T-Foley [1] and Stable Audio 2.0 [2].
Using an evaluation dataset consisting of vocal mimicry samples and reference sounds from six speakers, each method was used to synthesize audio. An example of the reference sound, the six vocal imitations, and the corresponding synthesized sounds is shown below.
Reference
Vocal Mimicry and Synthesized Sounds
| Speaker | Vocal Mimicry | PronounSE | T-Foley | Stable Audio 2.0 |
|---|---|---|---|---|
| m-01 | ||||
| m-02 | ||||
| m-03 | ||||
| m-04 | ||||
| m-05 | ||||
| f-01 |
Objective Fidelity Evaluation with FAD [3] and Cosine Similarity
Objective evaluation results on FAD between the reference and synthesized sound sets, and the average cosine similarity between the embedding vectors of corresponding sound pairs.
| Speaker ID | FAD by PANNs ↓ | Cosine similarity ± SD ↑ | ||||
|---|---|---|---|---|---|---|
| PronounSE | T-Foley | Stable Audio 2.0 | PronounSE | T-Foley | Stable Audio 2.0 | |
| m-01 | 17.85 | 54.02 | 21.35 | 0.85 ± 0.08 | 0.73 ± 0.07 | 0.83 ± 0.06 |
| m-02 | 17.05 | 54.47 | 20.94 | 0.86 ± 0.06 | 0.74 ± 0.07 | 0.84 ± 0.06 |
| m-03 | 17.28 | 48.85 | 20.95 | 0.86 ± 0.07 | 0.75 ± 0.07 | 0.84 ± 0.06 |
| m-04 | 16.65 | 48.52 | 21.46 | 0.85 ± 0.07 | 0.75 ± 0.08 | 0.84 ± 0.06 |
| m-05 | 19.29 | 47.89 | 19.75 | 0.86 ± 0.07 | 0.75 ± 0.08 | 0.85 ± 0.07 |
| f-01 | 19.28 | 56.94 | 22.01 | 0.85 ± 0.07 | 0.72 ± 0.07 | 0.85 ± 0.06 |
| Whole | 14.09 | 55.93 | 21.15 | 0.85 ± 0.07 | 0.74 ± 0.07 | 0.84 ± 0.06 |
Subjective Fidelity Evaluation (Fidelity MOS ± 95% CI)
5-point MOS resuts on the fidelity of synthesized sounds to reference sounds.
| Speaker ID | PronounSE | T-Foley | Stable Audio 2.0 |
|---|---|---|---|
| m-01 | 2.57 ± 0.12 | 1.54 ± 0.09 | 2.67 ± 0.13 |
| m-02 | 2.23 ± 0.11 | 1.29 ± 0.06 | 2.66 ± 0.12 |
| m-03 | 2.67 ± 0.12 | 1.44 ± 0.07 | 2.71 ± 0.12 |
| m-04 | 2.56 ± 0.12 | 1.47 ± 0.08 | 2.68 ± 0.12 |
| m-05 | 2.31 ± 0.14 | 1.57 ± 0.09 | 2.89 ± 0.12 |
| f-01 | 2.22 ± 0.12 | 1.31 ± 0.07 | 2.52 ± 0.12 |
| Whole | 2.43 ± 0.05 | 1.44 ± 0.03 | 2.69 ± 0.05 |
Subjective Nuance Reflection Evaluation (Attack MOS ± 95% CI, Release MOS ± 95% CI, Overall MOS ± 95% CI)
5-point MOS results on the nuance reflection of vocal mimicry in the synthesized sounds, evaluated from three perspectives: the attack phase, the release phase, and the overall impression.
| Speaker ID | Attack (Onset) | Release (Decay) | Overall Impression | ||||||
|---|---|---|---|---|---|---|---|---|---|
| PronounSE | T-Foley | Stable Audio 2.0 | PronounSE | T-Foley | Stable Audio 2.0 | PronounSE | T-Foley | Stable Audio 2.0 | |
| m-01 | 3.13 ± 0.12 | 2.95 ± 0.11 | 3.08 ± 0.12 | 3.03 ± 0.12 | 2.87 ± 0.11 | 2.98 ± 0.12 | 3.04 ± 0.12 | 2.80 ± 0.11 | 2.94 ± 0.12 |
| m-02 | 2.87 ± 0.13 | 2.65 ± 0.12 | 2.93 ± 0.12 | 3.05 ± 0.12 | 2.68 ± 0.12 | 2.93 ± 0.12 | 2.92 ± 0.12 | 2.45 ± 0.12 | 2.85 ± 0.11 |
| m-03 | 2.92 ± 0.12 | 2.55 ± 0.12 | 2.90 ± 0.12 | 2.86 ± 0.12 | 2.58 ± 0.12 | 2.80 ± 0.11 | 2.81 ± 0.11 | 2.50 ± 0.11 | 2.76 ± 0.11 |
| m-04 | 3.03 ± 0.12 | 2.76 ± 0.12 | 3.07 ± 0.12 | 3.12 ± 0.11 | 2.81 ± 0.12 | 2.93 ± 0.11 | 3.03 ± 0.11 | 2.71 ± 0.12 | 2.95 ± 0.12 |
| m-05 | 2.74 ± 0.13 | 2.64 ± 0.12 | 2.78 ± 0.11 | 2.83 ± 0.12 | 2.63 ± 0.11 | 2.75 ± 0.11 | 2.69 ± 0.12 | 2.61 ± 0.11 | 2.63 ± 0.10 |
| f-01 | 2.85 ± 0.12 | 2.84 ± 0.12 | 2.86 ± 0.13 | 2.92 ± 0.12 | 2.76 ± 0.12 | 2.71 ± 0.12 | 2.81 ± 0.11 | 2.79 ± 0.11 | 2.68 ± 0.12 |
| Whole | 2.92 ± 0.05 | 2.73 ± 0.05 | 2.94 ± 0.05 | 2.97 ± 0.05 | 2.72 ± 0.05 | 2.85 ± 0.05 | 2.88 ± 0.05 | 2.64 ± 0.05 | 2.80 ± 0.05 |
Subjective Naturalness and Sound Quality Evaluation (Naturallness MOS ± 95% CI, Quality MOS ± 95% CI)
5-point MOS results on the naturalness and audio quality of the synthesized sounds.
| Speaker ID | Naturalness (Ground Truth: 3.70 ± 0.11) | Sound Quality (Ground Truth: 3.70 ± 0.10) | ||||
|---|---|---|---|---|---|---|
| PronounSE | T-Foley | Stable Audio 2.0 | PronounSE | T-Foley | Stable Audio 2.0 | |
| m-01 | 3.50 ± 0.11 | 2.86 ± 0.12 | 3.29 ± 0.11 | 3.36 ± 0.11 | 2.73 ± 0.12 | 3.36 ± 0.10 |
| m-02 | 3.09 ± 0.12 | 2.56 ± 0.11 | 3.42 ± 0.10 | 3.16 ± 0.11 | 2.47 ± 0.11 | 3.45 ± 0.10 |
| m-03 | 3.46 ± 0.11 | 2.82 ± 0.12 | 3.48 ± 0.10 | 3.32 ± 0.10 | 2.55 ± 0.11 | 3.52 ± 0.10 |
| m-04 | 3.55 ± 0.11 | 2.95 ± 0.12 | 3.57 ± 0.10 | 3.39 ± 0.10 | 2.83 ± 0.12 | 3.48 ± 0.10 |
| m-05 | 2.93 ± 0.12 | 2.69 ± 0.11 | 3.72 ± 0.10 | 2.95 ± 0.11 | 2.49 ± 0.11 | 3.57 ± 0.10 |
| f-01 | 3.14 ± 0.12 | 2.76 ± 0.12 | 3.53 ± 0.11 | 2.99 ± 0.11 | 2.69 ± 0.11 | 3.45 ± 0.10 |
| Whole | 3.28 ± 0.05 | 2.77 ± 0.05 | 3.50 ± 0.04 | 3.19 ± 0.04 | 2.62 ± 0.05 | 3.47 ± 0.04 |
If you use this work, please cite our paper:
@article{hoge,
title = {hogehoge},
author = {Riki Takizawa, Shigeyuki Hirai, Asako Kanezaki, Hitoshi Suda},
journal = {hoge},
volume = {hoge},
number = {hoge},
pages = {hoge},
year = {hoge},
doi = {hoge}
}
[1]: Y. Chung, J.Lee, and J. Nam. "T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis". In: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE.
[2]: Stability AI. "Stable Audio 2.0". https://stability.ai/stable-audio.
[3]: ilgour, K., Zuluaga, M., Roblek, D. and Sharifi, M.: Fr´echet Audio Distance: A Reference-Free Metric for Evaluating Music Enhancement Algorithms, Proceedings of the 20th Annual Conference of the International Speech Communication Association (INTERSPEECH), pp. 2350–2354