EDBT 2026 Demo / reviewers in the wild / expert
Chulayuth Asawaroengchai
dblp:348/5227
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Speech recognition and synthesis · 67% Language models and text generation · 33% |
Topics — the 2 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis
speech language model |
0.8 | 1 | 2024 | Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM · ICLR 2024 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
spoken question answering |
0.8 | 1 | 2024 | Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
speech encoder · 0.8spectrogram modeling · 0.8large language model adaptation · 0.8cross-modal chain-of-thought · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Translatotron 3: Speech to Speech Translation with Monolingual DataabstractThis paper presents Translatotron 3, a novel approach to unsupervised direct speech-to-speech translation from monolingual speech-text datasets by combining masked autoencoder, unsupervised embedding mapping, and back-translation. Experimental results in speech-to-speech translation tasks between Spanish and English show that Translatotron 3 outperforms a baseline cascade system, reporting 18.14 BLEU points improvement on the synthesized Unpaired-Conversational dataset. In contrast to supervised approaches that necessitate real paired data, or specialized modeling to replicate para-/non-linguistic information such as pauses, speaking rates, and speaker identity, Translatotron 3 showcases its capability to retain it. Eliya Nachmani, Alon Levkovitch, Yifan Ding 0004, Chulayuth Asawaroengchai, Heiga Zen, Michelle Tadmor Ramanovich |
ICASSP | 4 |
| 2024 | Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLMabstractWe present Spectron, a novel approach to adapting pre-trained large language models (LLMs) to perform spoken question answering (QA) and speech continuation. By endowing the LLM with a pre-trained speech encoder, our model becomes able to take speech inputs and generate speech outputs. The entire system is trained end-to-end and operates directly on spectrograms, simplifying our architecture. Key to our approach is a training objective that jointly supervises speech recognition, text continuation, and speech synthesis using only paired speech-text pairs, enabling a `cross-modal' chain-of-thought within a single decoding pass. Our method surpasses existing spoken language models in speaker preservation and semantic coherence. Furthermore, the proposed model improves upon direct initialization in retaining the knowledge of the original LLM as demonstrated through spoken QA datasets. We release our audio samples and spoken QA dataset via our website. Eliya Nachmani, Alon Levkovitch, Roy Hirsch, Julian Salazar, Chulayuth Asawaroengchai, Soroosh Mariooryad, Ehud Rivlin, R. J. Skerry-Ryan, Michelle Tadmor Ramanovich |
ICLR | 5 |