STE-GAN: Speech-to-Electromyography Signal Conversion using Generative Adversarial Networks
by ,
Abstract:
With Speech-to-Electromyography Generative Adversarial Network (STE-GAN), we propose a model which can synthesize Electromyography (EMG) signals from acoustic speech. We condition the generator network on representations of the spoken content obtained from a voice conversion model. Given these representations, the generator outputs an EMG signal corresponding to the articulated content of the acoustic speech in the setting of a specific EMG recording session. In comparison to previous work, STE-GAN directly generates EMG signals from acoustic speech. As it uses more speaker-independent content representations as input, it can synthesize EMG signals from speech of speakers who were unseen during training.
Reference:
STE-GAN: Speech-to-Electromyography Signal Conversion using Generative Adversarial Networks (Kevin Scheck, Tanja Schultz), In Proc. INTERSPEECH 2023, 2023.
Bibtex Entry:
@INPROCEEDINGS{Scheck2023STEGAN,
  author={Kevin Scheck and Tanja Schultz},
  title={{STE-GAN: Speech-to-Electromyography Signal Conversion using Generative Adversarial Networks}},
  year=2023,
  booktitle={Proc. INTERSPEECH 2023},
  pages={1174--1178},
  doi={10.21437/Interspeech.2023-174},
  url={https://www.csl.uni-bremen.de/cms/images/documents/publications/ScheckSchultz-Interspeech23.pdf},
  abstract={With Speech-to-Electromyography Generative Adversarial Network (STE-GAN), we propose a model which can synthesize Electromyography (EMG) signals from acoustic speech. We condition the generator network on representations of the spoken content obtained from a voice conversion model. Given these representations, the generator outputs an EMG signal corresponding to the articulated content of the acoustic speech in the setting of a specific EMG recording session. In comparison to previous work, STE-GAN directly generates EMG signals from acoustic speech. As it uses more speaker-independent content representations as input, it can synthesize EMG signals from speech of speakers who were unseen during training.}
}