EDBT 2026 Demo / reviewers in the wild / expert
Heting Gao
dblp:241/5964
· DBLP profile ↗
15ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0002-2857-3842ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards unsupervised speech recognition without pronunciation models
Junrui Ni, Liming Wang 0003, Yang Zhang 0001, Kaizhi Qian, Heting Gao, Mark Hasegawa-Johnson, James R. Glass, Chang Dong Yoo |
Speech Commun. | 5 |
| 2025 | SX-Stitch: An Efficient VMS-UNet Based Framework for Intraoperative Scoliosis X-Ray Image StitchingabstractIn scoliosis surgery, the limited field of view of the C-arm Xray machine restricts the surgeons’ holistic analysis of spinal structures. This paper presents an end-to-end efficient and robust intraoperative X-ray image stitching method for scoliosis surgery, named SX-Stitch. The method is divided into two stages: segmentation and stitching. In the segmentation stage, we propose a medical image segmentation model named Vision Mamba of Spine-UNet (VMS-UNet), which utilizes the state space Mamba to capture long-distance contextual information while maintaining linear computational complexity. Meanwhile, the proposed model incorporates the SimAM attention mechanism, significantly improving the segmentation performance. In the stitching stage, we simplify the alignment process between images to minimize a registration energy function. The total energy function is then optimized to order unordered images and a hybrid energy function is introduced to optimize the best seam, effectively eliminating parallax artifacts. On the clinical dataset, Sx-Stitch demonstrates superiority over SOTA schemes both qualitatively and quantitatively. Heting Gao, Mingde He, Jinqian Liang, Jason Gu |
ICASSP | 2 |
| 2025 | SA-CIM: A 28nm 16Mb RRAM-based Sparsity-Aware Compute-In-Memory Macro for Edge AI Algorithm ProcessingabstractCompute-in-memory (CIM) for edge devices is usually constrained by on-chip resources, including on-chip memory and physical chip size, which hinders the deployment of more complex neural networks. By leveraging the sparsity of neural networks, the overall memory requirements and energy consumption can be reduced. However, Existing sparsity-aware architectures cannot achieve high energy efficiency due to off-chip sparsity control. This work proposes:1) Hybrid sparsity regulation strategy. The sparsity encoding and alignment circuit is designed and implemented, realizing on-chip sparsity detecting and encoding. 2) Sparsity-aware compute-in-memory (CIM) array based on RRAMs. The in-situ deployment of unstructured sparsity is implemented inside the CIM array, and the CIM array and sparsity are tightly coupled by sparsity read/write. This work demonstrates the design and evaluation of SA-CIM: a sparsity-aware CIM macro with 16Mb RRAM with fine-grained sparsity detecting and encoding capacity, achieving energy efficiency of 22.7TOP/W@8b/8b. Hao Ding 0011, Zongwei Wang 0001, Jinshan Li, Shigeng Zhao, Heting Gao, Junbo Ao, Ling Liang 0003, Yimao Cai, Ru Huang 0001 |
ISCAS | 5 |
| 2025 | VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech InteractionabstractRecent Multimodal Large Language Models (MLLMs) have typically focused on integrating visual and textual modalities, with less emphasis placed on the role of speech in enhancing interaction. However, speech plays a crucial role in multimodal dialogue systems, and implementing high-performance in both vision and speech tasks remains a challenge due to the fundamental modality differences. In this paper, we propose a carefully designed multi-stage training methodology that progressively trains LLM to understand both visual and speech information, ultimately enabling fluent vision and speech interaction. Our approach not only preserves strong vision-language capacity, but also enables efficient speech-to-speech dialogue capabilities without separate ASR and TTS modules, significantly accelerating multimodal end-to-end response speed. By comparing against state-of-the-art counterparts across benchmarks for image, video, and speech, we demonstrate that our omni model is equipped with both strong visual and speech capabilities, making omni understanding and interaction. Chaoyou Fu, Haojia Lin, Yifan Zhang 0004, Yunhang Shen, Haoyu Cao 0001, Zuwei Long, Heting Gao, Ke Li 0015, Xiawu Zheng, Rongrong Ji, Xing Sun 0001, Caifeng Shan, Ran He 0001 |
NeurIPS | 9 |
| 2025 | VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language ModelabstractWith the growing requirement for natural human-computer interaction, speech-based systems receive increasing attention as speech is one of the most common forms of daily communication. However, the existing speech models still experience high latency when generating the first audio token during streaming, which poses a significant bottleneck for deployment. To address this issue, we propose VITA-Audio, an end-to-end large speech model with fast audio-text token generation. Specifically, we introduce a lightweight Multiple Cross-modal Token Prediction (MCTP) module that efficiently generates multiple audio tokens within a single model forward pass, which not only accelerates the inference but also significantly reduces the latency for generating the first audio in streaming scenarios. In addition, a four-stage progressive training strategy is explored to achieve model acceleration with minimal loss of speech quality. To our knowledge, VITA-Audio is the first multi-modal large language model capable of generating audio output during the first forward pass, enabling real-time conversational capabilities with minimal latency. VITA-Audio is fully reproducible and is trained on open-source data only. Experimental results demonstrate that our model achieves an inference speedup of 3~5x at the 7B parameter scale, but also significantly outperforms open-source models of similar model size on multiple benchmarks for automatic speech recognition (ASR), text-to-speech (TTS), and spoken question answering (SQA) tasks. Zuwei Long, Yunhang Shen, Chaoyou Fu, Heting Gao, Lijiang Li, Peixian Chen, Mengdan Zhang, Jian Li 0062, Jinlong Peng, Haoyu Cao 0001, Ke Li 0015, Rongrong Ji, Xing Sun 0001 |
NeurIPS | 4 |
| 2025 | SyncDiff: Diffusion-Based Talking Head Synthesis with Bottlenecked Temporal Visual Prior for Improved SynchronizationabstractTalking head synthesis, also known as speech-to-lip synthesis, reconstructs the facial motions that align with the given audio tracks. The synthesized videos are evaluated on mainly two aspects, lip-speech synchronization and image fidelity. Recent studies demonstrate that GAN-based and diffusion-based models achieve state-of-the-art (SOTA) performance on this task, with diffusion-based models achieving superior image fidelity but experiencing lower synchronization compared to their GAN-based counterparts. To this end, we propose SYNcDIFF, a simple yet effective approach to improve diffusion-based models using a temporal pose frame with information bottleneck and facial-informative audio features extracted from AVHuBERT, as conditioning input into the diffusion process. We evaluate SYNcDIFF on two canonical talking head datasets, LRS2 and LRS3 for direct comparison with other SOTA models. Experiments on LRS2/LRS3 datasets show that SYNcDIFF achieves a synchronization score 27.7%/62.3% relatively higher than previous diffusion-based methods, while preserving their high-fidelity characteristics. Xulin Fan, Heting Gao, Ziyi Chen 0005, Peng Chang 0002, Mark Hasegawa-Johnson |
WACV | 2 |
| 2024 | G2PU: Grapheme-To-Phoneme Transducer with Speech UnitsabstractMost phoneme transcripts are generated using forced alignment: typically a grapheme-to-phoneme transducer (G2P) is applied to text sequences to generate candidate phoneme transcripts, which are then time-aligned to the waveform using an acoustic model. This paper demonstrates, for the first time, simultaneous optimization of the G2P, the acoustic model, and the acoustic alignment to a corpus. To this end, we propose G2PU, a joint CTC-attention model consisting of an encoder-decoder G2P network and an encoder-CTC unit-to-phoneme (U2P) network, where the units are extracted from speech. We demonstrate that the G2P and U2P, operating in parallel, produce lower phone error rates than those of state-of-the-art open-source G2P and forced alignment systems. Furthermore, although the G2P and U2P are trained using parallel speech and text, their synergy can be generalized to text-only test corpora if we also train a grapheme-to-unit (G2U) network that generates speech units from text in the absence of parallel speech. Our G2PU model is trained using phoneme transcripts generated by a teacher G2P tool. Our experiments on Chinese and Japanese show that G2PU reduces phoneme error rate by 7% to 29% relative compared to its teacher. Finally, we include case studies to provide insights into the system’s workings. Heting Gao, Mark Hasegawa-Johnson, Chang Dong Yoo |
ICASSP | 1 |
| 2024 | Speech Self-Supervised Learning Using Diffusion Model Synthetic DataabstractWhile self-supervised learning (SSL) in speech has greatly reduced the reliance of speech processing systems on annotated corpora, the success of SSL still hinges on the availability of a large-scale unannotated corpus, which is still often impractical for many low-resource languages or under privacy concerns. Some existing work seeks to alleviate the problem by data augmentation, but most works are confined to introducing perturbations to real speech and do not introduce new variations in speech prosody, speakers, and speech content, which are important for SSL. Motivated by the recent finding that diffusion models have superior capabilities for modeling data distributions, we propose DiffS4L, a pretraining scheme that augments the limited unannotated data with synthetic data with different levels of variations, generated by a diffusion model trained on the limited unannotated data. Finally, an SSL model is pre-trained on the real and the synthetic speech. Our experiments show that DiffS4L can significantly improve the performance of SSL models, such as reducing the WER of the HuBERT pretrained model by 6.26 percentage points in the English ASR task. Notably, we find that the synthetic speech with all levels of variations, i.e. new prosody, new speakers, and even new content (despite the new content being mostly babble), accounts for significant performance improvement. The code is available at github.com/Hertin/DiffS4L. Heting Gao, Kaizhi Qian, Junrui Ni, Chuang Gan 0001, Mark Hasegawa-Johnson, Shiyu Chang, Yang Zhang 0001 |
ICML | 1 |
| 2023 | Mitigating the Exposure Bias in Sentence-Level Grapheme-to-Phoneme (G2P) Transduction
Eunseop Yoon, Hee Suk Yoon, Dhananjaya Gowda, SooHwan Eom, Daehyeok Kim, John B. Harvill, Heting Gao, Mark Hasegawa-Johnson, Chanwoo Kim 0001, Chang Dong Yoo |
INTERSPEECH | 7 |
| 2022 | ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersabstractSelf-supervised learning in speech involves training a speech representation network on a large-scale unannotated speech corpus, and then applying the learned representations to downstream tasks. Since the majority of the downstream tasks of SSL learning in speech largely focus on the content information in speech, the most desirable speech representations should be able to disentangle unwanted variations, such as speaker variations, from the content. However, disentangling speakers is very challenging, because removing the speaker information could easily result in a loss of content as well, and the damage of the latter usually far outweighs the benefit of the former. In this paper, we propose a new SSL method that can achieve speaker disentanglement without severe loss of content. Our approach is adapted from the HuBERT framework, and incorporates disentangling mechanisms to regularize both the teacher labels and the learned representations. We evaluate the benefit of speaker disentanglement on a set of content-related downstream tasks, and observe a consistent and notable performance advantage of our speaker-disentangled representations. Kaizhi Qian, Yang Zhang 0001, Heting Gao, Junrui Ni, Cheng-I Lai, David D. Cox, Mark Hasegawa-Johnson, Shiyu Chang |
ICML | 3 |
| 2022 | WavPrompt: Towards Few-Shot Spoken Language Understanding with Frozen Language Models
Heting Gao, Junrui Ni, Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Mark Hasegawa-Johnson |
INTERSPEECH | 1 |
| 2022 | Unsupervised Text-to-Speech Synthesis by Unsupervised Automatic Speech Recognition
Junrui Ni, Liming Wang 0003, Heting Gao, Kaizhi Qian, Yang Zhang 0001, Shiyu Chang, Mark Hasegawa-Johnson |
INTERSPEECH | 3 |
| 2022 | Seamless equal accuracy ratio for inclusive CTC speech recognitionabstractConcerns have been raised regarding performance disparity in automatic speech recognition (ASR) systems as they provide unequal transcription accuracy for different user groups defined by different attributes that include gender, dialect, and race. In this paper, we propose “equal accuracy ratio”, a novel inclusiveness measure for ASR systems that can be seamlessly integrated into the standard connectionist temporal classification (CTC) training pipeline of an end-to-end neural speech recognizer to increase the recognizer’s inclusiveness. We also create a novel multi-dialect benchmark dataset to study the inclusiveness of ASR, by combining data from existing corpora in seven dialects of English (African American, General American, Latino English, British English, Indian English, Afrikaaner English, and Xhosa English). Experiments on this multi-dialect corpus show that using the equal accuracy ratio as a regularization term along with CTC loss, succeeds in lowering the accuracy gap between user groups and reduces the recognition error rate compared with a non-regularized baseline. Experiments on additional speech corpora that have different user groups also confirm our findings. Heting Gao, Sunghun Kang, Rusty Mina, Dias Issa, John B. Harvill, Leda Sari, Mark Hasegawa-Johnson, Chang Dong Yoo |
Speech Commun. | 1 |
| 2021 | Zero-Shot Cross-Lingual Phonetic Recognition with External Language Embedding
Heting Gao, Junrui Ni, Yang Zhang 0001, Kaizhi Qian, Shiyu Chang, Mark Hasegawa-Johnson |
Interspeech | 1 |
| 2019 | Hierarchical multi-armed bandits for discovering hidden populationsabstractThis paper proposes a novel algorithm to discover hidden individuals in a social network. The problem is increasingly important for social scientists as the populations (e.g., individuals with mental illness) that they study converse online. Since these populations do not use the category (e.g., mental illness) to self-describe, directly querying with text is non-trivial. To by-pass the limitations of network and query re-writing frameworks, we focus on identifying hidden populations through attributed search. We propose a hierarchical Multi-Arm Bandit (DT-TMP) sampler that uses a decision tree coupled with reinforcement learning to query the combinatorial attributed search space by exploring and expanding along high yielding decision-tree branches. A comprehensive set of experiments over a suite of twelve sampling tasks on three online web platforms, and three offline entity datasets reveals that DT-TMP outperforms all baseline samplers by upto a margin of 54% on Twitter and 48% on RateMDs. An extensive ablation study confirms DT-TMP's superior performance under different sampling scenarios. Suhansanu Kumar, Heting Gao, Changyu Wang, Kevin Chen-Chuan Chang, Hari Sundaram |
ASONAM | 2 |