VLDB 2026 Research / reviewers in the wild / expert
Weiguang Chen
dblp:136/7858
· DBLP profile ↗
15ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0002-7945-3197ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Injecting Visual Features into Whisper for Parameter-Efficient Noise-Robust Audio-Visual Speech RecognitionabstractAudio-visual speech recognition (AVSR) aims to enhance the robustness of an automatic speech recognition (ASR) systems by incorporating visual information from lip movements, especially in challenging noisy environments. Nevertheless, most current approaches either involve training from scratch or fully finetuning a pre-trained model, both of which incur significant computational costs and are often impractical for large-scale speech foundation models. This gap highlights the need for more efficient methods to leverage visual and acoustic information in AVSR tasks. To address this challenge, we propose AVWhisper, a parameter-efficient model that integrates visual and acoustic representations by injecting visual features from the AV-HuBERT encoder into the pre-trained Whisper model. Our approach leverages the existing attention mechanisms in Whisper to facilitate cross-modal interaction and integrates auxiliary visual information through lightweight adapters based on Low-Rank Adaptation (LoRA) and prompt-based techniques. Furthermore, a two-phase training strategy is adopted to effectively handle cross-domain differences and visual information injection problems respectively. Extensive experiments on the LRS3-TED dataset demonstrate that AVWhisper consistently outperforms state-of-the-art methods across various noise conditions, offering a more efficient and scalable solution for audio-visual speech recognition. Yue Heng Yeo, Xiao Fu 0001, Weiguang Chen, Wei Xi 0003, Jizhong Zhao |
ICASSP | 5 |
| 2025 | UniArray: Unified Spectral-Spatial Modeling for Array-Geometry-Agnostic Speech SeparationabstractArray-geometry-agnostic speech separation (AGA-SS) aims to develop an effective separation method regardless of the microphone array geometry. Conventional methods rely on permutation-free operations, such as summation or attention mechanisms, to capture spatial information. However, these approaches often incur high computational costs or disrupt the effective use of spatial information during intra- and inter-channel interactions, leading to suboptimal performance. To address these issues, we propose UniArray, a novel approach that abandons the conventional interleaving manner. UniArray consists of three key components: a virtual microphone estimation (VME) module, a feature extraction and fusion module, and a hierarchical dual-path separator. The VME ensures robust performance across arrays with varying channel numbers. The feature extraction and fusion module leverages a spectral feature extraction module and a spatial dictionary learning (SDL) module to extract and fuse frequency-bin-level features, allowing the separator to focus on using the fused features. The hierarchical dual-path separator models feature dependencies along the time and frequency axes while maintaining computational efficiency. Experimental results show that UniArray outperforms state-of-the-art methods in SI-SDRi, WB-PESQ, NB-PESQ, and STOI across both seen and unseen array geometries. Weiguang Chen, Jielong Yang, Chng Eng Siong, Xionghu Zhong |
IEEE Signal Process. Lett. | 1 |
| 2024 | Enhancing Low-Latency Speaker Diarization with Spatial Dictionary LearningabstractThis study proposes a low-latency online speaker diarization framework. Specifically, we design a spatial dictionary learning module shared across different frequency bands, enabling spatial feature learning at each frequency bin. This contributes to reducing the latency constraints of the online diarization system. Additionally, a magnitude-weighted fusion is devised to integrate spectral features. Consequently, the system can extract discriminative speaker embeddings by simultaneously considering spectral and spatial features. Experimental results on the Alimeeting dataset demonstrate a significant improvement in diarization error rates across various latencies, with a relative improvement of 45.80% compared to single-channel online diarization. Moreover, our method surpasses offline direction-of-arrival-based diarization and achieves comparable performance to the second-ranked offline system of the Alimeeting challenge. Weiguang Chen, Xionghu Zhong, Chng Eng Siong |
ICASSP | 1 |
| 2024 | GAMMA: Graspability-Aware Mobile MAnipulation Policy Learning based on Online Grasping Pose FusionabstractMobile manipulation constitutes a fundamental task for robotic assistants and garners significant attention within the robotics community. A critical challenge inherent in mobile manipulation is the effective observation of the target while approaching it for grasping. In this work, we propose a graspability-aware mobile manipulation approach powered by an online grasping pose fusion framework that enables a temporally consistent grasping observation. Specifically, the predicted grasping poses are online organized to eliminate the redundant, outlier grasping poses, which can be encoded as a grasping pose observation state for reinforcement learning. Moreover, on-the-fly fusing the grasping poses enables a direct assessment of graspability, encompassing both the quantity and quality of grasping poses. This assessment can subsequently serve as an observe-to-grasp reward, motivating the agent to prioritize actions that yield detailed observations while approaching the target object for grasping. Through extensive experiments conducted on the Habitat and Isaac Gym simulators, we find that our method attains a good balance between observation and manipulation, yielding high performance under various grasping metrics. Furthermore, we discover that the incorporation of temporal information from grasping poses aids in mitigating the sim-to-real gap, leading to robust performance in challenging real-world experiments. Project page: https://pku-epic.github.io/GAMMA/ Jiazhao Zhang, Nandiraju Gireesh, Jilong Wang 0011, Xiaomeng Fang, Chaoyi Xu, Weiguang Chen, Liu Dai, He Wang 0010 |
ICRA | 6 |
| 2024 | ISFP: Iterative Simultaneous Optical Flow Estimation and Interest Points Extraction NetworkabstractIn this paper, we propose ISFP, an end-to-end deep learning method for robust feature point extraction and optical flow estimation simultaneously. We observe that the structure of the deep network for feature point extraction bears similarities with the feature extraction network used in optical flow. Leveraging this observation, we merge these two tasks into a unified approach, enabling the extraction of feature points and their corresponding optical flow displacements in an end-to-end manner. ISFP serves as a comprehensive deep learning solution for feature point tracking in monocular visual SLAM systems, seamlessly replacing the front-end tasks of feature point extraction and inter-frame tracking. Our network utilizes a pre-trained interest point generation network to automatically generate interest point labels on an annotated optical flow dataset, thus creating a new dataset with ground truth interest points for training. Extensive experiments demonstrate the effectiveness and practical applicability of our approach in real-world scenarios. Compared with RAFT, our method is trained on a single FlyingChairs data set. It improves by 0.3 on the Sintel final data set, maintains comparable accuracy in clean, and the inference speed is 1.5 times that of RAFT. Weiguang Chen, Jichao Jiao, Ben Ding |
IJCNN | 1 |
| 2022 | Spectro-Temporal SubNet for Real-Time Monaural Speech Denoising and Dereverberation
Feifei Xiong, Weiguang Chen, Pengyu Wang 0010, Jinwei Feng |
INTERSPEECH | 2 |
| 2021 | CNN-DMA: A Predictable and Scalable Direct Memory Access Engine for Convolutional Neural Network with Sliding-window FilteringabstractMemory bandwidth utilization has become the key performance bottleneck for state-of-the-art variants of neural network kernels. Current structures such as depth-wise, point-wise and atrous convolutions have already introduced diverse and discontinuous memory access patterns, which impact efficient activation supply due to more frequent cache misses and consequently high-penalty DRAM pre-charging. To handle this, GPU achieves efficient parallelization with sophisticated optimization of CUDA program to reduce memory footprints, which demands high engineering efforts. In this work, we in contrast propose a programmable direct memory access engine for convolutional neural networks (CNN-DMA) supporting a fast supply of activation for independent and scalable computing units. The CNN-DMA favours a predictable activation streaming approach which completely avoids penalties by bus contention, cache misses and less carefully designed low-level programs. Furthermore, we enhance the baseline DMA with the capability of out-of-order data supply to filter out unique sliding-windows to boost the performance of the computing infrastructure. Experiments on state-of-the-art neural networks show that CNN-DMA achieves optimal DRAM access efficiency for point-wise convolution layers, while reduces 30% to 70% rounds of computation with sliding-window filtering. Zheng Wang 0027, Chao Chen 0022, Yongkui Yang, Weiguang Chen, Wenxuan Chen, Weiyu Guo, Zhibin Yu 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2021 | Overlapped Speech Detection Based on Spectral and Spatial Feature Fusion
Weiguang Chen, Van Tung Pham, Chng Eng Siong, Xionghu Zhong |
Interspeech | 1 |
| 2021 | Cramér-Rao Lower Bound for DOA Estimation with an Array of Directional Microphones in Reverberant Environments
Weiguang Chen, Xionghu Zhong |
Interspeech | 1 |
| 2021 | Real-Time Multi-Channel Speech Enhancement Based on Neural Network Masking with Attention Model
Weilong Huang, Weiguang Chen, Jinwei Feng |
Interspeech | 3 |
| 2021 | Multiple Acoustic Source Localization in Microphone Array NetworksabstractThe problem of multiple acoustic source localization using observations from a microphone array network is investigated in this article. Multiple source signals are assumed to be window-disjoint-orthogonal (WDO) on the time-frequency (TF) domain and time delay of arrival (TDOA) measurements are extracted at each TF bin. A Bayesian network model is then proposed to jointly assign the measurements to different sources and estimate the acoustic source locations. Considering that the WDO assumption is usually violated under reverberant and noisy environments, we construct a relational network by coding the distance information between the distributed microphone arrays such that adjacent arrays have higher probabilities of observing the same acoustic source, which is able to mitigate the miss detection issues in adverse environments. A Laplace approximate variational inference method is introduced to estimate the hidden variables in the proposed Bayesian network model. Both simulations and real data experiments are performed. The results show that our proposed method is able to achieve better source localization accuracy than existing methods. Jielong Yang, Xionghu Zhong, Weiguang Chen, Wenwu Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | BBS: Micro-Architecture Benchmarking Blockchain Systems through Machine Learning and Fuzzy SetabstractDue to the decentralization, irreversibility, and traceability, blockchain has attracted significant attention and has been deployed in many critical industries such as banking and logistics. However, the micro-architecture characteristics of blockchain programs still remain unclear. What's worse, the large number of micro-architecture events make understanding the characteristics extremely difficult. We even lack a systematic approach to identify the important events to focus on. In this paper, we propose a novel benchmarking methodology dubbed BBS to characterize blockchain programs at micro-architecture level. The key is to leverage fuzzy set theory to identify important micro-architecture events after the significance of them is quantified by a machine learning based approach. The important events for single programs are employed to characterize the programs while the common important events for multiple programs form an importance vector which is used to measure the similarity between benchmarks. We leverage BBS to characterize seven and six benchmarks from Blockbench and Caliper, respectively. The results show that BBS can reveal interesting findings. Moreover, by leveraging the importance characterization results, we improve that the transaction throughput of Smallbank from Fabric by 70% while reduce the transaction latency by 55%. In addition, we find that three of seven and two of six benchmarks from Blockbench and Caliper are redundant, respectively. Chao Chen 0022, Zihao Su, Weiguang Chen, Tao Li 0006, Zhibin Yu 0001 |
HPCA | 4 |
| 2020 | Directional and Explainable Serendipity RecommendationabstractSerendipity recommendation has attracted more and more attention in recent years; it is committed to providing recommendations which could not only cater to users’ demands but also broaden their horizons. However, existing approaches usually measure user-item relevance with a scalar instead of a vector, ignoring user preference direction, which increases the risk of unrelated recommendations. In addition, reasonable explanations increase users’ trust and acceptance, but there is no work to provide explanations for serendipitous recommendations. To address these limitations, we propose a Directional and Explainable Serendipity Recommendation method named DESR. Specifically, we extract users’ long-term preferences with an unsupervised method based on GMM (Gaussian Mixture Model) and capture their short-term demands with the capsule network at first. Then, we propose the serendipity vector to combine long-term preferences with short-term demands and generate directionally serendipitous recommendations with it. Finally, a back-routing scheme is exploited to offer explanations. Extensive experiments on real-world datasets show that DESR could effectively improve the serendipity and explainability, and give impetus to the diversity, compared with existing serendipity-based methods. Xueqi Li 0002, Weiguang Chen, Jie Wu 0001, Guojun Wang 0001, Kenli Li 0001 |
WWW | 3 |
| 2019 | HAES: A New Hybrid Approach for Movie Recommendation with Elastic SerendipityabstractRecommendation systems provide good guidance for users to find their favorite movies from an overwhelming amount of options. However, most systems excessively pursue the recommendation accuracy and give rise to over-specialization, which triggers the emergence of serendipity. Hence, serendipity recommendation has received more attention in recent years, facing three key challenges: subjectivity in the definition, the lack of data, and users' floating demands for serendipity. To address these challenges, we introduce a new model called HAES, a H ybrid A pproach for movie recommendation with E lastic S erendipity, to recommend serendipitous movies. Specifically, we (1) propose a more objective definition of serendipity, \em content difference and \em genre accuracy, according to the analysis on a real dataset, (2) propose a new algorithm named JohnsonMax to mitigate the data sparsity and build weak ties beneficial to finding serendipitous movies, and (3) define a novel concept of elasticity in the recommendation, to adjust the level of serendipity flexibly and reach a trade-off between accuracy and serendipity. Extensive experiments on real-world datasets show that HAES enhances the serendipity of recommendations while preserving recommendation quality, compared to several widely used methods. Xueqi Li 0002, Weiguang Chen, Jie Wu 0001, Guojun Wang 0001 |
CIKM | 3 |
| 2013 | Accelerating sparse matrix-vector multiplication on GPUs using bit-representation-optimized schemesabstractThe sparse matrix-vector (SpMV) multiplication routine is an important building block used in many iterative algorithms for solving scientific and engineering problems. One of the main challenges of SpMV is its memory-boundedness. Although compression has been proposed previously to improve SpMV performance on CPUs, its use has not been demonstrated on the GPU because of the serial nature of many compression and decompression schemes. In this paper, we introduce a family of bit-representation-optimized (BRO) compression schemes for representing sparse matrices on GPUs. The proposed schemes, BRO-ELL, BRO-COO, and BRO-HYB, perform compression on index data and help to speed up SpMV on GPUs through reduction of memory traffic. Furthermore, we formulate a BRO-aware matrix reordering scheme as a data clustering problem and use it to increase compression ratios. With the proposed schemes, experiments show that average speedups of 1.5x compared to ELLPACK and HYB can be achieved for SpMV on GPUs. Wai Teng Tang, Wen Jun Tan, Rajarshi Ray 0001, Yi Wen Wong, Weiguang Chen, Shyh-Hao Kuo, Rick Siow Mong Goh, Stephen John Turner, Weng-Fai Wong |
SC | 5 |