Guanjun Li

dblp:17/224 · DBLP profile ↗
← Back
19ranked-venue papers
9as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A novel dynamic graph attention aggregation network for multivariate time series classification
Haoyu Gui, Xianghong Tang, Guanjun Li, Chaobin Wang, Jianguang Lu
Pattern Recognit.3
2026 MSSG: Multi-Scale Speaker Graph Network for Active Speaker Detection
abstract
The active speaker detection task is to determine whether a person is speaking or not across a series of video frames. Existing methods heavily rely on facial information within the annotated face bounding boxes for cross-modal learning with audio. This leads to a substantial decline in detection performance when facial cues are unclear, such as in cases of face occlusion or low-resolution facial appearances. In this paper, we extend the perception scale using only face bounding box annotations to model both facial and gestural cues, addressing the over-reliance on facial cues in active speaker detection. We propose a novel graph neural network that models inter-speaker interactions and integrates various cues from individual speakers. The final detection results are obtained through a binary graph node classification task. Our method achieves state-of-the-art performance on the AVA-ActiveSpeaker dataset (mAP: 95.6%) and the ASW dataset (mAP: 99.4%), with a model size only 21% that of the second-best method. Additionally, when facial cues are of poor quality, our method demonstrates a significant performance advantage over existing approaches. The code and model weights will be available athttps://github.com/sdqdlgj/MSSG.
Guanjun Li, Jiangyan Yi, Zhengqi Wen, Ruibo Fu, Yuwang Wang, Jianhua Tao 0001
IEEE Trans. Multim.1
2025 ImViD: Immersive Volumetric Videos for Enhanced VR Engagement
abstract
User engagement is greatly enhanced by fully immersive multi-modal experiences that combine visual and auditory stimuli. Consequently, the next frontier in VR/AR technologies lies in immersive volumetric videos with complete scene capture, large 6-DoF interaction space, multimodal feedback, and high resolution & frame-rate contents. To stimulate the reconstruction of immersive volumetric videos, we introduce ImViD, a multi-view, multi-modal dataset featuring complete space-oriented data capture and various indoor/outdoor scenarios. Our capture rig supports multi-view video-audio capture while on the move, a capability absent in existing datasets, significantly enhancing the completeness, flexibility, and efficiency of data capture.The captured multi-view videos (with synchronized audios) are in 5K resolution at 60FPS, lasting from 1-5 minutes, and include rich foreground-background elements, and complex dynamics. We benchmark existing methods using our dataset and establish a base pipeline for constructing immersive volumetric videos from multi-view audiovisual inputs for 6-DoF multi-modal immersive VR experiences. The benchmark and the reconstruction and interaction results demonstrate the effectiveness of our dataset and baseline method, which we believe will stimulate future research on immersive volumetric video production. Project Page: https://yzxqh.github.io/ImViD/
Zhengxian Yang, Shi Pan, Shengqi Wang, Guanjun Li, Zhengqi Wen, Borong Lin, Jianhua Tao 0001
CVPR6
2025 DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
abstract
In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models.
Ruibo Fu, Zhengqi Wen, Tao Wang 0074, Chunyu Qiang, Jianhua Tao 0001, Chenxing Li, Shuchen Shi, Yuankun Xie, Xuefei Liu, Guanjun Li
ICASSP15
2025 Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
abstract
Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model further enhances detection performance. However, most of the previously proposed fusion methods require fine-tuning the pretrained models, resulting in excessively long training times and hindering model iteration when facing new speech synthesis technology. To address this issue, this paper proposes a feature fusion method based on the Mixture of Experts, which extracts and integrates features relevant to fake audio detection from layer features, guided by a gating network based on the last layer feature, while freezing the pretrained model. Experiments conducted on the ASVspoof2019 and ASVspoof2021 datasets demonstrate that the proposed method achieves competitive performance compared to those requiring fine-tuning.
Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Yuankun Xie, Shuchen Shi, Chenxing Li, Xuefei Liu, Guanjun Li
ICASSP13
2025 MTPareto: A MultiModal Targeted Pareto Framework for Fake News Detection
abstract
Multimodal fake news detection is essential for maintaining the authenticity of Internet multimedia information. Significant differences in form and content of multimodal information lead to intensified optimization conflicts, hindering effective model training as well as reducing the effectiveness of existing fusion methods for bimodal. To address this problem, we propose the MTPareto framework to optimize multimodal fusion, using a Targeted Pareto(TPareto) optimization algorithm for fusion-level-specific objective learning with a certain focus. Based on the designed hierarchical fusion network, the algorithm defines three fusion levels with corresponding losses and implements all-modal-oriented Pareto gradient integration for each. This approach accomplishes superior multimodal fusion by utilizing the information obtained from intermediate fusion to provide positive effects to the entire process. Experiment results on FakeSV and FVC datasets show that the proposed framework outperforms baselines and the TPareto optimization algorithm achieves 2.40% and 1.89% accuracy improvement respectively.
Kaiying Yan, Moyang Liu, Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Xuefei Liu, Guanjun Li
ICASSP8
2025 Multi-scale feature fusion network with temporal dynamic graphs for small-sample FW-UAV fault diagnosis
Guanjun Li, Haoyu Gui, Jianguang Lu, Xianghong Tang, Xiaoyu Gao
Knowl. Based Syst.1
2024 An Output Voltage Tracking Control Method with Current Constraint Capability for Disturbed Lc-Type Three-Phase Inverters
abstract
In this article, a novel output voltage tracking control scheme with current constraint capability is designed for disturbed LC-type three-phase inverters based on sliding-mode control, nonlinear mapping and finite-time disturbance observers. In this control scheme, the current constraints are firstly converted into state constraints for the error state equations. Utilizing nonlinear mapping functions, the state constraint problems are further transformed into the boundedness problems of sliding-mode variables. The sliding-mode controllers which are designed on this basis, ensure the boundedness of the sliding-mode variables, thereby preventing the currents from violating the constraints. By employing finite-time disturbance observers to estimate both mismatched and matched disturbances, and incorporating disturbance compensations into the sliding-mode surfaces and controllers, the system's vulnerability to disturbances is mitigated. With the above design, this control scheme can achieve fast output voltage tracking and current constraints in the presence of mismatched and matched disturbances in the system. The efficacy of the proposed method has been confirmed through simulations.
Qinhao Tang, Saijin Huang, Xiangyu Wang 0003, Xianghui He, Guanjun Li
INDIN5
2024 CATodyNet: Cross-attention temporal dynamic graph neural network for multivariate time series classification
Haoyu Gui, Guanjun Li, Xianghong Tang, Jianguang Lu
Knowl. Based Syst.2
2023 GCC-Speaker: Target Speaker Localization with Optimal Speaker-Dependent Weighting in Multi-Speaker Scenarios
abstract
Existing noise-robust and reverberant-robust localization algorithms fail to localize the target speaker when interfering speakers are present. In this paper, we address the problem of localizing only the target speaker in multi-speaker scenarios and propose a target speaker localization algorithm, called GCC-speaker. Specifically, we modify the weighting of the generalized cross-correlation with phase transform (GCC-PHAT) algorithm and propose an optimal speaker-dependent weighting based on a novel localization-related loss function and data-driven training. The speaker-dependent weighting is responsible for guiding the GCC algorithm to obtain the optimal target speaker localization results. As for the loss function, we constrain the estimated GCC angular spectrum and the estimated direction of arrival (DOA) to be close to their ground truth values, respectively. The experimental results show the superiority of GCC-speaker compared to the existing target speaker localization algorithms for different signal-to-interference ratios, reverberation times and array geometries.
Guanjun Li, Wei Xue 0002, Jiangyan Yi, Jianhua Tao 0001
ICASSP1
2023 Hierarchical graph attention network for temporal knowledge graph reasoning
Pengpeng Shao, Guanjun Li, Dawei Zhang 0001, Jianhua Tao 0001
Neurocomputing3
2021 Deep neural network-based generalized sidelobe canceller for dual-channel far-field speech recognition
Guanjun Li, Shan Liang 0001, Shuai Nie 0001, Zhanlei Yang
Neural Networks1
2021 Exploiting the directional coherence function for multichannel source extraction
Shan Liang 0007, Guanjun Li, Shuai Nie 0001, Zhanlei Yang, Jianhua Tao 0001
Speech Commun.2
2020 Deep Neural Network-Based Generalized Sidelobe Canceller for Robust Multi-Channel Speech Recognition
Guanjun Li, Zhanlei Yang, Longshuai Xiao
INTERSPEECH1
2020 Microphone Array Post-Filter for Target Speech Enhancement Without a Prior Information of Point Interferers
Guanjun Li, Zhanlei Yang, Longshuai Xiao
INTERSPEECH1
2019 Adaptive Dereverberation Using Multi-channel Linear Prediction with Deficient Length Filter
abstract
In almost all adaptive dereverberation algorithms based on the multi-channel linear prediction (MCLP) model, it is assumed that the filter length can cover the reverberation time. However, in many practical situations, a deficient length filter, whose length is less than the reverberation time, is employed in consideration of computational cost. A deficient length filter fails to fully model the late reverberation, resulting in degraded performance. In this paper, we present a new MCLP-based adaptive dereverberation algorithm to improve the dereverberation performance when using a deficient length filter. We introduce a gain and use the filter coefficients estimated from the previous frame to track the MCLP modeling errors of the current frame. The gain and the filter coeffi-cients are jointly optimized and solved by using an alternating minimization technique. Experimental results show the superiority of the proposed algorithm. The shorter the filter length is, the more advantageous the proposed algorithm is.
Guanjun Li
ICASSP1
2019 Direction-Aware Speaker Beam for Multi-Channel Speaker Extraction
Guanjun Li, Shan Liang 0007, Shuai Nie 0001, Meng Yu 0003, Lianwu Chen, Shouye Peng, Changliang Li
INTERSPEECH1
2017 Adaptive fuzzy prescribed performance controller design for a class of uncertain fractional-order nonlinear systems with external disturbances
Heng Liu 0003, Shenggang Li, Jinde Cao, Guanjun Li, Ahmed Alsaedi, Fuad E. Alsaadi
Neurocomputing4
2009 The Dahlquist Constant Approach to Stability Analysis of the Static Neural Networks
Guanjun Li
ISNN (1)1