VLDB 2026 Research / reviewers in the wild / expert
Tianchi Sun
dblp:261/6281
· DBLP profile ↗
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-4589-7769ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimization of modular multi-speaker distant conversational speech recognition
Qinwen Hu, Tianchi Sun, Xiaobin Rong |
Comput. Speech Lang. | 2 |
| 2025 | Leveraging Self-Supervised Learning Based Speaker Diarization for MISP 2025 AVSD Challenge
Zeyan Song, Tianchi Sun, Ronghui Hu, Kai Chen 0029 |
INTERSPEECH | 2 |
| 2025 | A Lightweight Hybrid Dual Channel Speech Enhancement System under Low-SNR Conditions
Xiaobin Rong, Tianchi Sun |
INTERSPEECH | 4 |
| 2024 | GTCRN: A Speech Enhancement Model Requiring Ultralow Computational ResourcesabstractWhile modern deep learning-based models have significantly outperformed traditional methods in the area of speech enhancement, they often necessitate a lot of parameters and extensive computational power, making them impractical to be deployed on edge devices in real-world applications. In this paper, we introduce Grouped Temporal Convolutional Recurrent Network (GTCRN), which incorporates grouped strategies to efficiently simplify a competitive model, DPCRN. Additionally, it leverages subband feature extraction modules and temporal recurrent attention modules to enhance its performance. Remarkably, the resulting model demands ultralow computational resources, featuring only 23.7 K parameters and 39.6 MMACs per second. Experimental results show that our proposed model not only surpasses RNNoise, a typical lightweight model with similar computational burden, but also achieves competitive performance when compared to recent baseline models with significantly higher computational resources requirements. Xiaobin Rong, Tianchi Sun, Changbao Zhu |
ICASSP | 2 |
| 2024 | A Lightweight Hybrid Multi-Channel Speech Extraction System with Directional Voice Activity DetectionabstractAlthough deep learning (DL) based end-to-end models have shown outstanding performance in multi-channel speech extraction, their practical applications on edge devices are restricted due to their high computational complexity. In this paper, we propose a hybrid system that can more effectively integrate the generalized sidelobe canceller (GSC) and a lightweight post-filtering model under the assistance of spatial speaker activity information provided by a directional voice activity detection (DVAD) module. In addition to guiding the update of the adaptive blocking matrix (ABM) and the adaptive interference canceller (AIC) used in GSC to alleviate the distortion of the desired speech, DVAD is also utilized as an auxiliary input to the post-filtering model to enhance its capability of interference suppression. The experimental results demonstrate that, with much lower computational costs, our method can achieve comparable performance with a current state-of-the-art end-to-end model on simulated data and generalize even better on real-world data. Tianchi Sun, Changbao Zhu |
ICASSP | 1 |
| 2023 | Convolutional Recurrent MetriCGAN With Spectral Dimension Compression For Full-Band Speech EnhancementabstractMetricGAN and its variations have been proven to be an effective wide-band speech enhancement model. In this paper, we expand it to full-band enhancement by combining our recently proposed learnable spectral dimension compression mapping strategy. The encoder-decoder structure with a time-frequency convolutional recurrent network is utilized as the generator. The proposed model is submitted to the ICASSP Signal Processing Grand Challenge: DNS-5 Challenge (2023). Without using the enrollment speech, it obtains a final score of 0.548 on Track-1 and 0.559 on Track-2. Zhongshu Hou, Qinwen Hu, Tianchi Sun, Changbao Zhu |
ICASSP | 3 |
| 2023 | A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing AidsabstractThis paper summarizes a hybrid multi-channel speech enhancement system for the ICASSP Signal Processing Grand Challenge: Clarity Challenge (Speech Enhancement for Hearing Aids) 2023. The system consists of a rule-based dereverberation module, a multi-channel enhancement module, and a post-processing module. Without using the head rotation information and the enrollment speech, the system can reach an average hearing aid speech perception index (HASPI) score of 0.696 and hearing aid speech quality index (HASQI) score of 0.320 on the official development set. The corresponding scores are 0.729 and 0.316 respectively on the Eval1 set for the challenge ranking. Zhongshu Hou, Wanyu Yang, Tianchi Sun, Xiaobin Rong, Dahan Wang, Kai Chen 0029 |
ICASSP | 5 |
| 2022 | SAR Imaging Based on Deep Unfolded Network With Approximated ObservationabstractCompressed sensing (CS) based synthetic aperture radar (SAR) imaging methods are showing superior potential in imaging performance over classical matched filtering based methods. However, the CS-based methods require much more computational cost to solve the iterative optimization composed of large-scale matrix operators. To hold the improvement of imaging performance and reduce the computational cost, in this paper, we propose a novel SAR imaging method by Deep Unfolded Network (DUN) of Iterative Shrinkage Threshold Algorithm (ISTA) with the approximated observation of Range-Doppler Algorithm (RDA) operator. The proposed method takes the radar echoes as the input to learn the imaging procedure. Firstly, the approximated observation is utilized in SAR imaging model to reduce the size of the DUN. Moreover, we use ISTA as an example to introduce how to establish DUN with approximated observation, in which the detailed structure to handle the complex-valued radar echoes is also designed. Finally, the auto-encoder is utilized to calculate the difference of the echoes rather than the imaging results so that we can train the proposed network by unsupervised learning. The experiments of both point targets, surface targets, and real scenes show that the proposed imaging method is superior in terms of imaging performance and computing efficiency. Tianchi Sun, Ying Luo 0001, Qun Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |