EDBT 2026 Demo / reviewers in the wild / expert
Yidi Li 0001
dblp:254/8169-1
· DBLP profile ↗
23ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0002-5236-7010ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Document Layout Analysis Through Frequency Decomposition Convolution and Multi-scale Adaptive Linear Attention
Chaoyi Guo, Yidi Li 0001 |
ICIC (10) | 6 |
| 2026 | XS-YOLO: A Lightweight NMS-Free Framework for Small Object Detection in Remote Sensing Imagery
Chaoyi Guo, Zichu Zhang, Muhan Guo, Yidi Li 0001 |
ICIC (18) | 5 |
| 2026 | ASNet: An adaptive scene-aware network for RGB-thermal urban scene semantic segmentation
Zhenxue Chen, Xuewen Rong, Chengyun Liu, Lili Song, Yidi Li 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2026 | Global-local distillation network-based audio-visual speaker tracking with incomplete modalities
Yidi Li 0001, Zhenhuan Xu, Weiwei Wan, Hong Liu 0008 |
Pattern Recognit. | 1 |
| 2025 | Multi-Stage Multimodal Distillation for Audio-Visual Speaker TrackingabstractSpeaker tracking plays a crucial role in various human-robot interaction applications. Recently, leveraging multimodal information, such as audio and visual signals, has become an important strategy for enhancing the robustness of the tracking system. However, current methods face challenges in effectively exploring the complementarity between audio and visual modalities. To this end, we propose an Audio-Visual Tracker based on Multi-Stage Multimodal Distillation (MSMD-AVT), which utilizes an audio-visual knowledge distillation framework to facilitate audio-visual information fusion over multiple stages progressively. MSMD-AVT is constructed based on an audio-visual teacher-student model incorporating three distinct distillation losses. During the feature extraction stage, the feature alignment distillation is designed to ensure that the feature representations from the student network remain consistent with the teacher encoding feature. Moreover, during the feature fusion stage, the fusion guidance distillation is proposed, using deep teacher features to guide the multimodal fusion process in the student network, optimizing the complementary benefits of audio-visual fusion. Finally, the logits distillation is applied during the position estimation stage to help the student model better capture localization features through knowledge transfer and output alignment. Additionally, we present a multimodal fusion module based on a bidirectional cross-attention mechanism in the student network, dynamically adjusting the effectiveness of different modal features for the tracking task by extracting complementary audio-visual contextual information. Extensive experimental results on the widely used AV16.3 dataset indicate that MSMD-AVT significantly outperforms existing state-of-the-art methods in terms of accuracy and robustness. Our code is publicly available at https://github.com/moyitech/MSMD-AVT. Yidi Li 0001, Wenkai Zhao, Zhenhuan Xu, Bin Ren 0005, Nicu Sebe |
ICASSP | 1 |
| 2025 | Vision-Guided Acoustic Localization with Decoupled Inference for Moving Speakers
Yidi Li 0001, Kairan Zhang, Chenxu Yang, Chongwei Yan, Rongshan Gao, Mingliang Dou |
ICIC (19) | 1 |
| 2025 | DSGC-Net: A Dual-Stream Graph Convolutional Network for Crowd Counting via Feature Correlation Mining
Jinqiao Wei, Xionghui Zhao, Yidi Li 0001, Shaoyi Du, Bin Ren 0005, Nicu Sebe |
PRCV (17) | 4 |
| 2025 | GLMKD: Joint global and local mutual knowledge distillation for weakly supervised lesion segmentation in histopathology images
Hangbei Cheng, Xueyu Liu, Jun Zhang 0095, Xiaorong Dong, Xuetao Ma 0001, Xing Chen 0017, Guanghui Yue 0001, Yidi Li 0001, Yongfei Wu |
Expert Syst. Appl. | 10 |
| 2025 | PVAFN: Point-Voxel Attention Fusion Network with Multi-Pooling Enhancing for 3D Object DetectionabstractThe integration of point and voxel representations is becoming more common in Light Detection and Ranging (LiDAR)-based 3D object detection. However, existing fusion strategies suffer from ineffective semantic alignment and contextual information loss, while relying solely on point features within regions of interest leads to geometric detail degradation and limited local–global feature integration. To tackle these challenges, we propose the Point-Voxel Attention Fusion Network (PVAFN), a novel two-stage 3D object detector that introduces a point-voxel attention fusion module based on dual-gated cross-modal interaction and a multi-pooling strategy based on density-space awareness. During the feature extraction and fusion stage, a dual-gated hierarchical attention mechanism is proposed to dynamically fuse three heterogeneous modalities—keypoint-based geometric details, voxel-wise local regularity, and Bird’s-Eye-View (BEV)-level global semantics—through learnable gating functions. In the refinement stage, a density-spatial-aware multi-pooling enhancement module is designed to synergize density-aware cluster pooling and multi-scale spatial-aware pyramid pooling, efficiently capturing key geometric details and fine-grained shape structures. This design enhances the integration of local and global features while enabling adaptive multi-scale context modeling and spatially sensitive feature aggregation. Extensive experiments on the KITTI and Waymo benchmark datasets demonstrate that PVAFN achieves promising detection accuracy in 3D mean Average Precision. • PVAFN: A novel network for 3D object detection, fusing keypoint, voxel, and BEV features dynamically. • Stage-I: Dual-gated hierarchical attention fusion for bidirectional point-voxel feature calibration. • Stage-II: Density-spatial-aware multi-pooling module enhances local and global geometric perception. • SOTA Performance: PVAFN achieves highest AP on KITTI and Waymo benchmarks for autonomous driving. Yidi Li 0001, Bin Ren 0005, Wenhao Li 0002, Hong Liu 0008, Nicu Sebe |
Expert Syst. Appl. | 1 |
| 2025 | 3CNet: Cross-modal cooperative correction network for RGB-T semantic segmentation
Zhenxue Chen, Xuewen Rong, Chengyun Liu, Lili Song, Yidi Li 0001 |
Image Vis. Comput. | 6 |
| 2025 | Multi-instance curriculum learning for histopathology image classification with bias reduction
Zihao Mi, Xueyu Liu, Guanghui Yue 0001, Junhong Yue, Mingqiang Wei, Yidi Li 0001, Yongfei Wu |
Medical Image Anal. | 7 |
| 2025 | Dual Attention Guidance Network for Self-Supervised Monocular Depth EstimationabstractSelf-supervised monocular depth estimation shows great promise since only a single camera is required. However, most existing methods fail to model the geometric structure of objects, leading to poor performance in object boundary depth estimation. To overcome these shortcomings, a dual attention guidance network (DAG-Net), containing two complementary modules termed depth-guided attention module (DAM) and semantic-guided multi-modal attention module (SAM), is proposed in this paper. The DAM utilizes depth features to guide semantic features through multi-head attention. When semantic features are well learned, they guide depth features to learn useful geometric representations through backpropagation. Besides, the SAM is proposed to incorporate multi-modal data from depth estimation and semantic segmentation predictions at different scales. To eliminate the mutual interference between DAM and SAM, we also propose a two-stage training strategy to adjust the convergence direction during the training process. The effectiveness of our proposed DAG-Net is qualitatively and quantitatively verified by various experiments on KITTI, Cityscapes, and Make3D datasets, showing outstanding performance compared with the state-of-the-art methods. Hong Liu 0008, Guoliang Hua, Hao Tang 0005, Yidi Li 0001, Weibo Huang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | STNet: Deep Audio-Visual Fusion Network for Robust Speaker TrackingabstractAudio-visual speaker tracking aims to determine the location of human targets in a scene using signals captured by a multi-sensor platform, whose accuracy and robustness can be improved by multi-modal fusion methods. Recently, several fusion methods have been proposed to model the correlation in multiple modalities. However, for the speaker tracking problem, the cross-modal interaction between audio and visual signals hasn't been well exploited. To this end, we present a novel Speaker Tracking Network (STNet) with a deep audio-visual fusion model in this work. We design a visual-guided acoustic measurement method to fuse heterogeneous cues in a unified localization space, which employs visual observations via a camera model to construct the enhanced acoustic map. For feature fusion, a cross-modal attention module is adopted to jointly model multi-modal contexts and interactions. The correlated information between audio and visual features is further interacted in the fusion model. Moreover, the STNet-based tracker is applied to multi-speaker cases by a quality-aware module, which evaluates the reliability of multi-modal observations to achieve robust tracking in complex scenarios. Experiments on the AV16.3 and CAV3D datasets show that the proposed STNet-based tracker outperforms uni-modal methods and state-of-the-art audio-visual speaker trackers. Yidi Li 0001, Hong Liu 0008, Bing Yang 0004 |
IEEE Trans. Multim. | 1 |
| 2024 | Multi-instance Curriculum Learning for Histopathology Image Classifications with Hard Negative Mining and Positive AugmentationabstractMulti-instance learning (MIL) exhibits advanced and surpassed capabilities in understanding and recognizing complex patterns within gigapixel histopathological images. However, the currents MIL methods for the analysis of the histopathological image still give rise to two main concerns. On one hand, vanilla MIL methods intuitively focus on identifying key instances (easy-to-classify) without considering hard-to-classify instances, which is biased and prone to produce false positive instance. On the other hand, since the positive tissue occupies only a small fraction of histopathological images, it is commonly suffer from class imbalance between positive and negative instances, causing the MIL model to overly focus on the majority class. In light of these issues of bias learning, we propose a multi-instance curriculum learning method that collaboratively incorporates hard negative instance mining and positive instance augmentation to improve model’s classification performance. Specifically, we first initialize the MIL model using easy-to-classify instances, then we mine the hard negative instances (hard-to-classify) and augment the positive instances via the diffusion model. Finally, the MIL model is retrained with memory rehearsal method by combining the mined negative instances and augmented positive instances. Technically, the diffusion model is first designed to generate lesion instances, which optimally augment diverse features to reflect the realistic positive samples with post screening scenario. Extensive experimental results show that the proposed method alleviates model bias in MIL and yields improvements over the state-of-the-art methods on both public datasets and private dataset. Zihao Mi, Xueyu Liu, Guangze Shi 0001, Yidi Li 0001, Yongfei Wu |
BIBM | 5 |
| 2024 | AttA-NET: Attention Aggregation Network for Audio-Visual Emotion RecognitionabstractIn video-based emotion recognition, effective multi-modal fusion techniques are essential to leverage the complementary relationship between audio and visual modalities. Recent attention-based fusion methods are widely leveraged for capturing modal-shared properties. However, they often ignore the modal-specific properties of audio and visual modalities and the unalignment of model-shared emotional semantic features. In this paper, an Attention Aggregation Network (AttA-NET) is proposed to address these challenges. An attention aggregation module is proposed to get modal-shared properties effectively. This module comprises similarity-aware enhancement blocks and a contrastive loss that facilitates aligning audio and visual semantic features. Moreover, an auxiliary uni-modal classifier is introduced to obtain modal-specific properties, in which intra-modal discriminative features are fully extracted. Under joint optimization of uni-modal and multi-modal classification loss, modal-specific information can be infused. Extensive experiments on RAVDESS and PKU-ER datasets validate the superiority of AttA-NET. The code is available at: https://github.com/NariFan2002/AttA-NET. Ruijia Fan, Hong Liu 0008, Yidi Li 0001, Peini Guo, Ti Wang |
ICASSP | 3 |
| 2024 | Adaptive Fourier Decomposition Based Signal Extraction on Weak Electromagnetic FieldabstractShaft-rate electromagnetic (EM) field is a critical feature in the detection of ships and underwater vehicles. However, the signal-to-noise ratio of the shaft-rate EM field is greatly reduced due to the presence of the static EM field, whose main energy is concentrated in the low-frequency section. In order to realize the effective detection of the weak shaft-rate EM field signals under a low signal-to-noise ratio, we propose a signal extraction method based on the Adaptive Fourier Decomposition (AFD) algorithm. At the decomposition stage, we utilize the Nevanlinna factorization and the maximal selection principle in each step, and then iteratively obtain various single components from low-frequency to high-frequency. At the extraction stage, the low-frequency information of the signal is effectively reconstructed by summing the first few components, leading to the retrieval of the shaft-rate EM field signal through residual operations. The experiment results on both synthesized and measured data show that the proposed algorithm converges faster with higher fidelity compared to the existing state-of-the-art method. Zhenhuan Xu, Yongfei Wu, Liming Zhang 0002, Yidi Li 0001 |
ICASSP | 4 |
| 2024 | MVSSC: Meta-reinforcement learning based visual indoor navigation using multi-view semantic spatial contextabstractIn Visual Indoor Navigation (VIN), Deep Reinforcement Learning (DRL) is commonly used by agents to achieve end-to-end mapping from vision to action when navigating toward a target based on observation. However, current DRL-based work suffers from two challenges: partial observability resulting from using solely a single first-person view and poor generalization in the case of unknown scenes and unknown objects. In light of these issues, this paper introduces the integration of multi-view as an expansion of observability and meta-learning as a primary generalization technique into the DRL framework and presents the meta-reinforcement learning method that leverages Multi-View Semantic Spatial Context (MVSSC). Specifically, aiming to explore the informative multi-view context for better efficiency in searching for and navigating to the target, we model the objects’ relationship from two aspects, multi-view semantic context (MVSEC) and multi-view spatial context (MVSPC). MVSEC enables agents to encode prior semantic relationships adaptively via a multi-view modulated graph. Meanwhile, MVSPC enhances the spatial representation of target-related objects’ correlation through similarity grids of multi-view. After adaptively fusing the multi-view and context information under the meta-reinforcement learning framework, our method can encourage efficient target search and robust navigation with stronger generalization performance to unknown scenes and unknown objects. Extensive experimental results on the AI2-THOR simulator demonstrate that our method outperforms the current state-of-the-art approaches. Wanruo Zhang, Hong Liu 0008, Jianbing Wu, Yidi Li 0001 |
Pattern Recognit. Lett. | 4 |
| 2024 | Feature Completion Transformer for Occluded Person Re-IdentificationabstractOccluded person re-identification is a challenging problem due to the destruction of occluders in different camera views. Most existing paradigms focus on visible human body parts through some external models to reduce noise interference. However, the feature misalignment problem caused by discarded occlusions negatively affects the performance of the network. Different from most previous works that discard the occluded regions, we present Feature Completion Transformer (FCFormer) that reduces noise interference and complements missing features in occluded parts. Specifically, Occlusion Instance Augmentation is proposed to simulate real and diverse occlusion situations on the holistic image, which enlarges the occlusion samples in the training set and forms aligned occluded-holistic pairs. To reduce the interference of noise, a two-stream architecture is proposed to learn pairwise discriminative features from aligned image pairs, while obtaining self-aligned occluded-holistic feature level sample-label pairs without additional auxiliary models. To complement the features of occluded regions, a Feature Completion Decoder is designed to aggregate possible information from self-generated occluded features in a self-supervised manner. Further, in order to correlate the completion features with identity information, Feature Completion Consistency loss is introduced to enforce the distribution of the generated completion features to be consistent with the real holistic feature distribution. In addition, we propose the Cross Hard Triplet loss to further bridge the gap between completion features and extracting features under the same ID. Extensive experiments over five challenging datasets demonstrate that the proposed FCFormer achieves superior performance and outperforms the state-of-theart methods by significant margins on Occluded-Duke dataset. Mengyuan Liu 0001, Hong Liu 0008, Wenhao Li 0002, Miaoju Ban, Tianyu Guo 0001, Yidi Li 0001 |
IEEE Trans. Multim. | 7 |
| 2023 | Boosting Person Re-Identification with Viewpoint Contrastive Learning and Adversarial TrainingabstractPerson re-identification (ReID) aims at retrieving a person of interest across multiple cameras. Despite significant progress in person ReID, viewpoint variation remains an obstacle to extracting discriminative features for retrieval. To address this problem, we propose a Viewpoint-Robust Network (VRN) based on contrastive learning and adversarial training to boost person ReID. Specifically, a View-point Confusion (VC) module is proposed to conceal viewpoint information to extract viewpoint-agnostic features. We employ viewpoint contrastive learning to discriminate viewpoints, and then conversely ignore the viewpoint information by adversarial training. Besides, an ID Prototype (IDP) module further enhances the network by introducing a confidence-weighted IDP as a viewpoint-robust ID representation and conducting contrastive metric learning with an IDP triplet loss. Extensive experiments demonstrate the proposed method achieves state-of-the-art performance on widely used datasets Market1501 and MSMT17. Visualization of retrieval results illustrates the effectiveness and robustness of the proposed method. Xingyue Shi, Hong Liu 0008, Wei Shi 0009, Zihui Zhou, Yidi Li 0001 |
ICASSP | 5 |
| 2023 | Self-Supervised 3D Skeleton Representation Learning with Active Sampling and Adaptive Relabeling for Action RecognitionabstractSelf-supervised 3D skeleton representation learning has recently shown great potential for action recognition via contrastive learning. However, existing methods suffer from limited learning efficiency and the unreliability of representations, which is not conducive to action recognition. To this end, we propose an Active Sampling and Adaptive Relabeling (ASAR) contrastive learning method to achieve efficient and reliable learning of 3D skeleton representations. Specifically, the active sampling strategy is used to build a dictionary with informative samples for efficient representation learning. Additionally, the adaptive relabeling strategy is proposed to automatically modify the confidence scores of the extra positive samples and alleviate the unreliability of representations. Extensive experiments on NTU-60, NTU-120, and PKU-MMD datasets demonstrate the superiority of our approach. Hong Liu 0008, Tianyu Guo 0001, Jingwen Guo, Ti Wang, Yidi Li 0001 |
ICIP | 6 |
| 2022 | Multi-Modal Perception Attention Network with Self-Supervised Learning for Audio-Visual Speaker TrackingabstractMulti-modal fusion is proven to be an effective method to improve the accuracy and robustness of speaker tracking, especially in complex scenarios. However, how to combine the heterogeneous information and exploit the complementarity of multi-modal signals remains a challenging issue. In this paper, we propose a novel Multi-modal Perception Tracker (MPT) for speaker tracking using both audio and visual modalities. Specifically, a novel acoustic map based on spatial-temporal Global Coherence Field (stGCF) is first constructed for heterogeneous signal fusion, which employs a camera model to map audio cues to the localization space consistent with the visual cues. Then a multi-modal perception attention network is introduced to derive the perception weights that measure the reliability and effectiveness of intermittent audio and video streams disturbed by noise. Moreover, a unique cross-modal self-supervised learning method is presented to model the confidence of audio and visual observations by leveraging the complementarity and consistency between different modalities. Experimental results show that the proposed MPT achieves 98.6% and 78.3% tracking accuracy on the standard and occluded datasets, respectively, which demonstrates its robustness under adverse conditions and outperforms the current state-of-the-art methods. Yidi Li 0001, Hong Liu 0008, Hao Tang 0005 |
AAAI | 1 |
| 2020 | 3D Audio-Visual Speaker Tracking with A Novel Particle Filterabstract3D speaker tracking using co-located audio-visual sensors has received much attention recently. Though various methods have been attempted to this field, it is still challenging to obtain a reliable 3D tracking result since the position of colocated sensors are restricted to a small area. In this paper, a novel particle filter (PF) based method is proposed for 3D audio-visual speaker tracking. Compared with traditional PF based audio-visual speaker tracking method, our 3D audio-visual tracker has two main characteristics. In the prediction stage, we use audio-visual information at current frame to further adjust the direction of the particles after the particle state transition process, which can make the particles more concentrated around the speaker direction. In the update stage, the particle likelihood is calculated by fusing both the visual distance and audiovisual direction information. Specially, the distance likelihood is obtained according to the camera projection model and the adaptively estimated size of speaker face or head, and the direction likelihood is determined by audio-visual particle fitness. In this way, the particle likelihood can better represent the speaker presence probability in 3D space. Experimental results show that the proposed tracker outperforms other methods and provides a favorable speaker tracking performance both in 3D space and on the image plane. Hong Liu 0008, Yongheng Sun, Yidi Li 0001, Bing Yang 0004 |
ICPR | 3 |
| 2019 | 3D Audio-Visual Speaker Tracking with A Two-Layer Particle FilterabstractAudio-visual speaker tracking in 3D space is a challenging problem. Although the classical particle filter based methods have shown effectiveness in audio-visual speaker tracking, the performance degrades considerably when the measurements are disturbed by noise. To this end, a novel two-layer particle filter is proposed for 3D audio-visual speaker tracking. Firstly, two groups of particles, which are generated from the audio and video streams respectively, are propagated independently in the audio layer and visual layer. Then, the audio and visual likelihoods are combined in an adaptive sigmoid function, which can adjust particle weights according to the confidence of two modalities. Finally, an optimal particle set selected from two groups of particles is proposed to determine the speaker position and reset the particle positions in the next frame. Experiments on AV16.3 database show that our method outperforms the trackers using individual modalities and the existing approaches in the 3D space and on the image plane. Hong Liu 0008, Yidi Li 0001, Bing Yang 0004 |
ICIP | 2 |