EDBT 2026 Demo / reviewers in the wild / expert
Lei Liu 0049
dblp:21/2715-49
· DBLP profile ↗
36ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0003-2749-5528ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EVF-SAM: Early Vision-Language Fusion for text-prompted Segment Anything Model
Tianheng Cheng, Lianghui Zhu, Lei Liu 0049, Longjin Ran, Xiaoxin Chen 0001, Wenyu Liu 0001, Xinggang Wang |
Image Vis. Comput. | 4 |
| 2026 | RGBT tracking via supervised mutual guiding
Lei Liu 0049, Chenglong Li 0002, Jin Tang 0001, Changhe Li |
Pattern Recognit. | 1 |
| 2026 | Temporal multimodal knowledge distillation for modality-missing RGBT tracking
Rui Ruan, Yunlong Kang, Lei Liu 0049, Jingpeng Sun, Chenglong Li 0002, Jin Tang 0001 |
Pattern Recognit. | 3 |
| 2026 | Unveiling the Power of Multi-Modal Template Update in RGBT TrackingabstractTemplate update is essential for improving the adaptability of tracking algorithms to target appearance variations. While previous methods have leveraged the spatio-temporal complementarity of multi-modal templates for RGBT tracking, a comprehensive analysis of the template update mechanism remains underexplored. In this work, we propose a novel prototype-based framework that decomposes the multi-modal template update process from the perspective of prototype learning into four key components: multi-modal prototype, prototype integration, prototype evaluation, and prototype update algorithm. Our findings highlight that the multi-modal prototype is the most critical factor in enhancing tracking adaptability to appearance variations, leading to more robust target representations. While prototype integration is less crucial when the target representation is already robust, it still contributes to learning a more discriminative representation. Additionally, the accuracy of template updates is strongly influenced by prototype evaluation, which controls the accuracy of the update process. Finally, the prototype update algorithm, which determines when and how template updates occur, is key to maintaining tracking robustness. Building on these insights, we introduce the Multi-modal Prototype RGBT Tracker (MPTrack), which adapts dynamically to appearance variations through prototype learning. MPTrack combines a fixed template from the first frame with both modality-shared and modality-specific templates, forming a robust multi-modal prototype representation. It incorporates a prototype evaluation module that guides updates based on template reliability, and an adaptive update algorithm to manage templates effectively. Additionally, a prototype-guided cross-modal integration module enhances the discriminative power of multi-modal relation modeling. Experimental results on five challenging RGBT tracking benchmarks demonstrate that MPTrack consistently outperforms state-of-the-art methods, setting new performance records. The experimental data and source code will be made publicly available at: https://github.com/mmic-lcl/Datasets-and-benchmark-code. Lei Liu 0049, Chenglong Li 0002, Andong Lu, Yabin Zhu, Shoufei Han, Xinye Cai, Changhe Li |
IEEE Trans. Image Process. | 1 |
| 2025 | Reliable Imputed-Sample Assisted Vertical Federated LearningabstractVertical Federated Learning (VFL) is a well-known FL variant that enables multiple parties to collaboratively train a model without sharing their raw data. Existing VFL approaches focus on overlapping samples among different parties, while their performance is constrained by the limited number of these samples, leaving numerous non-overlapping samples unexplored. Some previous work has explored techniques for imputing missing values in samples, but often without adequate attention to the quality of the imputed samples. To address this issue, we propose a Reliable Imputed-Sample Assisted (RISA) VFL framework to effectively exploit non-overlapping samples by selecting reliable imputed samples for training VFL models. Specifically, after imputing non-overlapping samples, we introduce evidence theory to estimate the uncertainty of imputed samples, and only samples with low uncertainty are selected. In this way, high-quality non-overlapping samples are utilized to improve VFL model. Experiments on two widely used datasets demonstrate the significant performance gains achieved by the RISA, especially with the limited overlapping samples, e.g., a 48% accuracy gain on CIFAR-10 with only 1% overlapping samples. Yaopei Zeng, Lei Liu 0049, Shaoguo Liu, Hongjian Dou, Baoyuan Wu, Li Liu 0036 |
ICASSP | 2 |
| 2025 | GroundingSuite: Measuring Complex Multi-Granular Pixel GroundingabstractPixel grounding, encompassing tasks such as Referring Expression Segmentation (RES), has garnered considerable attention due to its immense potential for bridging the gap between vision and language modalities. However, advancements in this domain are currently constrained by limitations inherent in existing datasets, including limited object categories, insufficient textual diversity, and a scarcity of high-quality annotations. To mitigate these limitations, we introduce GroundingSuite, which comprises: (1) an automated data annotation framework leveraging multiple Vision-Language Model (VLM) agents; (2) a large-scale training dataset encompassing 9.56 million diverse referring expressions and their corresponding segmentations; and (3) a meticulously curated evaluation benchmark consisting of 3,800 images. The GroundingSuite training dataset facilitates substantial performance improvements, enabling models trained on it to achieve state-of-the-art results. Specifically, a cIoU of 68.9 on gRefCOCO and a gIoU of 55.3 on RefCOCOm. Moreover, the GroundingSuite annotation framework demonstrates superior efficiency compared to the current leading data annotation method, i.e., $4.5 \times$ faster than GLaMM. Lianghui Zhu, Tianheng Cheng, Lei Liu 0049, Longjin Ran, Xiaoxin Chen 0001, Wenyu Liu 0001, Xinggang Wang |
ICCV | 5 |
| 2025 | CPSea: Large-scale cyclic peptide-protein complex dataset for machine learning in cyclic peptide designabstractCyclic peptides exhibit better binding affinity and proteolytic stability compared to their linear counterparts. However, the development of cyclic peptide design models is hindered by the scarcity of data. To address this, we introduce **CPSea**(**C**yclic **P**eptide **Sea**), a dataset of 2.71 million cyclic peptide-receptor complexes, curated through systematic mining of the AlphaFold Database (AFDB). Our pipeline extracts compact domains from AFDB, identifies cyclization sites using the $\beta$-carbon (C$_\beta$) distance thresholds, and applies multi-stage filtering to ensure structure fidelity and binding compatibility. Compared with experimental data of cyclic peptides, CPSea shows similar distributions in metrics on structure fidelity and wet-lab compatibility. To our knowledge, CPSea is the largest cyclic peptide-receptor dataset to date, enabling end-to-end model training for the first time. The dataset also showcases the feasibility of simulating inter-chain interactions using intra-chain interactions, expanding available resources for machine-learning models on protein-protein interactions. The dataset and relevant scripts are accessible on GitHub ([https://github.com/YZY010418/CPSea](https://github.com/YZY010418/CPSea)). Ziyi Yang 0011, Hanyuan Xie, Yinjun Jia, Xiangzhe Kong, Jiqing Zheng, Ziting Zhang, Yang Liu 0003, Lei Liu 0049, Yanyan Lan |
NeurIPS | 8 |
| 2025 | Hierarchical semantics guided multi-scale correlation network for alignment-free red-green-blue and thermal salient object detection
Chengmei Han, Lei Liu 0049, Kunpeng Wang 0005 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Multi-stage network for single image deblurring based on dual-domain window mambaabstractMulti-stage methods have been proven effective and widely used in image deblurring research. These methods, usually designed based on Convolutional Neural Networks (CNNs) or Vision Transformers (ViTs), have limitations, including the inability to capture global contextual information and a quadratic increase in computational complexity as image resolution. Additionally, although current methods have incorporated frequency domain information, they do not sufficiently explore the interrelationships of different frequencies. To address these issues, we proposed a Multi-Stage Visual Dual-Domain Window Mamba (DDWMamba) approach to realize image deblurring, leveraging the benefits of state space models (SSMs) for image data. First, to achieve better deblurring effects, we used a multi-stage design approach in which each stage maintains the details and global information of the original resolution image. Second, we proposed a DDWMamba Block, which includes a Spatial Window Visual Mamba and a Frequency Window Visual Mamba, aiming to fully explore the correlations between different pixels in both the spatial and frequency domains. Finally, to implement a coarse-to-fine design approach in the multi-stage method and reduce model complexity, we set a window operation with different window sizes for each stage. DDWMamba is extensively evaluated on several benchmark datasets, and the model achieves superior performance compared to existing state-of-the-art deblurring methods. Lei Liu 0049, Bin Li 0053, Zongyu Ye, Wangmeng Zuo |
Neural Networks | 2 |
| 2025 | Testing non-commutativity of reduce functions with multi-column inputs
Xiangyu Mu, Chenlu Zhu, Lei Liu 0049 |
Sci. Comput. Program. | 6 |
| 2025 | Efficient RGBT Tracking via Multi-Path Mamba Fusion NetworkabstractRGBT tracking aims to fully exploit the complementary advantages of visible and infrared modalities to achieve robust tracking, thus the design of multimodal fusion network is crucial. However, existing methods typically adopt CNNs or Transformer networks to construct the fusion network, which poses a challenge in achieving a balance between performance and efficiency. To overcome this issue, we introduce an innovative visual state space (VSS) model, represented by Mamba, for RGBT tracking. In particular, we design a novel multi-path Mamba fusion network that achieves robust multimodal fusion capability while maintaining a linear overhead. First, we design a multi-path Mamba layer to sufficiently fuse two modalities in both global and local perspectives. Second, to alleviate the issue of inadequate VSS modeling in the channel dimension, we introduce a simple yet effective channel swapping layer. Extensive experiments conducted on four public RGBT tracking datasets demonstrate that our method surpasses existing state-of-the-art trackers. Notably, our fusion method achieves higher tracking performance compared to the well-known Transformer-based fusion approach (TBSI), while also achieving 92.8% and 80.5% reductions in parameter count and computational cost, respectively. Fanghua Hong, Andong Lu, Lei Liu 0049, Qunjing Wang |
IEEE Signal Process. Lett. | 4 |
| 2025 | Adaptive Interaction and Correction Attention Network for Audio-Visual MatchingabstractAudio-visual matching techniques aim to recognize and match information across different identities by learning a similarity metric across modalities. However, modal differences arise from insufficient cross-modal correlations and noise interference, which substantially hinder the performance of traditional deep metric learning methods in audio-visual matching tasks. To address the modal differences issue, we propose a novel Adaptive Interactive and Correction Attention Network (AICANet). This network efficiently captures deep information connections, generating modality-consistent feature embeddings within a unified metric framework. The core of AICANet is its two-pronged approach to reducing modal differences. First, we propose the Adaptive Interactive Attention (AIA) module, which flexibly establishes associations among cross-modal local features using dynamically generated pseudo-labels. Second, we propose the Adaptive Correction Attention (ACA) mechanism, which employs an adaptive threshold to de-interference effectively and accurately adjust the representation of local feature associations. Notably, the ACA mechanism is suitable for both intra-modal and inter-modal refined attention correction. Additionally, we design a relative distance stretching metric loss (LRDSM), which reinforces the similarity invariance of feature embeddings in a uniform space and enhances matching accuracy. Extensive tests on the VoxCeleb and VoxCeleb2 datasets demonstrate that AICANet outperforms leading existing algorithms across several evaluation metrics, validating its superior performance. The codes can be found at https://github.com/w1018979952/AICANet. Jiaxiang Wang 0001, Aihua Zheng, Lei Liu 0049, Chenglong Li 0002, Ran He 0001, Jin Tang 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Cross-Modal Object Tracking via Modality-Aware Fusion Network and a Large-Scale DatasetabstractVisual object tracking often faces challenges such as invalid targets and decreased performance in low-light conditions when relying solely on RGB image sequences. While incorporating additional modalities like depth and infrared data has proven effective, existing multimodal imaging platforms are complex and lack real-world applicability. In contrast, near-infrared (NIR) imaging, commonly used in surveillance cameras, can switch between RGB and NIR based on light intensity. However, tracking objects across these heterogeneous modalities poses significant challenges, particularly due to the absence of modality switch signals during tracking. To address these challenges, we propose an adaptive cross-modal object tracking algorithm called modality-aware fusion network (MAFNet). MAFNet efficiently integrates information from both RGB and NIR modalities using an adaptive weighting mechanism, effectively bridging the appearance gap and enabling a modality-aware target representation. It consists of two key components: an adaptive weighting module and a modality-specific representation module. The adaptive weighting module predicts fusion weights to dynamically adjust the contribution of each modality, while the modality-specific representation module captures discriminative features specific to RGB and NIR modalities. MAFNet offers great flexibility as it can effortlessly integrate into diverse tracking frameworks. With its simplicity, effectiveness, and efficiency, MAFNet outperforms state-of-the-art methods in cross-modal object tracking. To validate the effectiveness of our algorithm and overcome the scarcity of data in this field, we introduce CMOTB, a comprehensive and extensive benchmark dataset for cross-modal object tracking. CMOTB consists of 61 categories and 1000 video sequences, comprising a total of over 799K frames. We believe that our proposed method and dataset offer a strong foundation for advancing cross-modal object-tracking research. The dataset, toolkit, experimental data, and source code will be publicly available at: https://github.com/mmic-lcl/ Datasets-and-benchmark-code. Lei Liu 0049, Chenglong Li 0002, Jin Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Prior-free Balanced Replay: Uncertainty-guided Reservoir Sampling for Long-Tailed Continual LearningabstractEven in the era of large models, one of the well-known issues in continual learning (CL) is catastrophic forgetting, which is significantly challenging when the continual data stream exhibits a long-tailed distribution, termed as Long-Tailed Continual Learning (LTCL). Existing LTCL solutions generally require the label distribution of the data stream to achieve re-balance training. However, obtaining such prior information is often infeasible in real scenarios since the model should learn without pre-identifying the majority and minority classes. To this end, we propose a novel Prior-free Balanced Replay (PBR) framework to learn from long-tailed data stream with less forgetting. Concretely, motivated by our experimental finding that the minority classes are more likely to be forgotten due to the higher uncertainty, we newly design an uncertainty-guided reservoir sampling strategy to prioritize rehearsing minority data without using any prior information, which is based on the mutual dependence between the model and samples. Additionally, we incorporate two prior-free components to further reduce the forgetting issue: (1) Boundary constraint is to preserve uncertain boundary supporting samples for continually re-estimating task boundaries. (2) Prototype constraint is to maintain the consistency of learned class prototypes along with training. Our approach is evaluated on three standard long-tailed benchmarks, demonstrating superior performance to existing CL methods and previous SOTA LTCL approach in both task- and class-incremental learning settings, as well as ordered- and shuffled-LTCL settings. © 2024 ACM. Lei Liu 0049, Li Liu 0036, Yawen Cui |
ACM Multimedia | 1 |
| 2024 | RGBT Tracking based on modality feature enhancement
Sulan Zhai, Lei Liu 0049, Jin Tang 0001 |
Multim. Tools Appl. | 3 |
| 2024 | A method of test case set generation in the commutativity test of reduce functions
Xiangyu Mu, Lei Liu 0049, Hui Li 0037 |
Sci. Comput. Program. | 2 |
| 2024 | Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech RecognitionabstractCued Speech (CS) is a pure visual coding method used by hearing-impaired people that combines lip reading with several specific hand shapes to make the spoken language visible. Automatic CS recognition (ACSR) seeks to transcribe visual cues of speech into text, which can help hearing-impaired people to communicate effectively. The visual information of CS contains lip reading and hand cueing, thus the fusion of them plays an important role in ACSR. However, most previous fusion methods struggle to capture the global dependency present in long sequence inputs of multi-modal CS data. As a result, these methods generally fail to learn the effective cross-modal relationships that contribute to the fusion. Recently, attentionbased transformers have been a prevalent idea for capturing the global dependency over the long sequence in multi-modal fusion, but existing multi-modal fusion transformers suffer from both poor recognition accuracy and inefficient computation for the ACSR task. To address these problems, we develop a novel computation and parameter efficient multi-modal fusion transformer by proposing a novel Token-Importance-Aware Attention mechanism (TIAA), where a token utilization rate (TUR) is formulated to select the important tokens from the multi-modal streams. More precisely, TIAA firstly models the modality-specific fine-grained temporal dependencies over all tokens of each modality, and then learns the efficient cross-modal interaction for the modality-shared coarse-grained temporal dependencies over the important tokens of different modalities. Besides, a lightweight gated hidden projection is designed to control the feature flows of TIAA. The resulting model, named Economical Cued Speech Fusion Transformer (EcoCued), achieves state-of-the-art performance on all existing CS datasets (i.e., Mandarin Chinese, French, and British CS), compared with existing transformerbased fusion methods and ACSR fusion methods. Notably, our method dramatically reduces the computational complexity from O(T2) to O(T). We will release the source code and data as open source. Lei Liu 0049, Li Liu 0036, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | RGBT Tracking via Challenge-Based Appearance Disentanglement and InteractionabstractRGB and thermal source data suffer from both shared and specific challenges, and how to explore and exploit them plays a critical role in representing the target appearance in RGBT tracking. In this paper, we propose a novel approach, which performs target appearance representation disentanglement and interaction via both modality-shared and modality-specific challenge attributes, for robust RGBT tracking. In particular, we disentangle the target appearance representations via five challenge-based branches with different structures according to their properties, including three parameter-shared branches to model modality-shared challenges and two parameter-independent branches to model modality-specific challenges. Considering the complementary advantages between modality-specific cues, we propose a guidance interaction module to transfer discriminative features from one modality to another one to enhance the discriminative ability of weak modality. Moreover, we design an aggregation interaction module to combine all challenge-based target representations, which could form more discriminative target representations and fit the challenge-agnostic tracking process. These challenge-based branches are able to model the target appearance under certain challenges so that the target representations can be learned by a few parameters even in the situation of insufficient training data. In addition, to relieve labor costs and avoid label ambiguity, we design a generation strategy to generate training data with different challenge attributes. Comprehensive experiments demonstrate the superiority of the proposed tracker against the state-of-the-art methods on four benchmark datasets. Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003, Rui Ruan, Minghao Fan |
IEEE Trans. Image Process. | 1 |
| 2024 | Text-to-Image Vehicle Re-Identification: Multi-Scale Multi-View Cross-Modal Alignment Network and a Unified BenchmarkabstractVehicle Re-IDentification (Re-ID) aims to retrieve the most similar images with a given query vehicle image from a set of images captured by non-overlapping cameras, and plays a crucial role in intelligent transportation systems and has made impressive advancements in recent years. In real-world scenarios, we can often acquire the text descriptions of target vehicle through witness accounts, and then manually search the image queries for vehicle Re-ID, which is time-consuming and labor-intensive. To solve this problem, this paper introduces a new fine-grained cross-modal retrieval task called text-to-image vehicle re-identification, which seeks to retrieve target vehicle images based on the given text descriptions. To bridge the significant gap between language and visual modalities, we propose a novel Multi-scale multi-view Cross-modal Alignment Network (MCANet). In particular, we incorporate view masks and multi-scale features to align image and text features in a progressive way. In addition, we design the Masked Bidirectional InfoNCE (MB-InfoNCE) loss to enhance the training stability and make the best use of negative samples. To provide an evaluation platform for text-to-image vehicle re-identification, we create a Text-to-Image Vehicle Re-Identification dataset (T2I VeRi), which contains 2465 image-text pairs from 776 vehicles with an average sentence length of 26.8 words. Extensive experiments conducted on T2I VeRi demonstrate MCANet outperforms the current state-of-art (SOTA) method by 2.2% in rank-1 accuracy. Leqi Ding, Lei Liu 0049, Yan Huang 0008, Chenglong Li 0002, Cheng Zhang 0010, Wei Wang 0115, Liang Wang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Spatio-Temporal Structure Consistency for Semi-Supervised Medical Image ClassificationabstractIntelligent medical diagnosis has shown remarkable progress on the large-scale datasets with full annotations. However, very few labeled images are available due to significantly expensive annotations by experts. To efficiently leverage abundant unlabeled data, we propose a novel Spatio-Temporal Structure Consistent (STSC) learning framework to combine both spatial and temporal structure consistency. Specifically, a gram matrix is derived to capture the structural similarity among different training samples in the representation space. At the spatial level, our framework explicitly enforces the consistency of structural similarity among different samples under perturbations. At the temporal level, we desire to maintain the consistency of the structural similarity in different training iterations by digging out the stable sub-structures in a relation graph. Experiments on two medical image datasets (i.e., ISIC 2018 and ChestX-ray14) show that our method outperforms state-of-the-art Semi-Supervised Learning (SSL) methods. Furthermore, extensive qualitative analysis on the Gram matrices and heatmaps by Grad-CAM are presented to validate the effectiveness of our method. Wentao Lei, Lei Liu 0049, Li Liu 0036 |
ICASSP | 2 |
| 2023 | Cross-Modal Mutual Learning for Cued Speech RecognitionabstractAutomatic Cued Speech Recognition (ACSR) provides an intelligent human-machine interface for visual communications, where the Cued Speech (CS) system utilizes lip movements and hand gestures to code spoken language for hearing-impaired people. Previous ACSR approaches often utilize direct feature concatenation as the main fusion paradigm. However, the asynchronous modalities (i.e., lip, hand shape and hand position) in CS may cause interference for feature concatenation. To address this challenge, we propose a transformer based cross-modal mutual learning framework to prompt multi-modal interaction. Compared with the vanilla self-attention, our model forces modality-specific information of different modalities to pass through a modality-invariant codebook, concatenating linguistic representations with tokens of each modality. Then the shared linguistic knowledge is used to re-synchronize multi-modal sequences. Moreover, we establish a novel large-scale multi-speaker CS dataset for Mandarin Chinese. To our knowledge, this is the first work on ACSR for Mandarin Chinese. Extensive experiments are conducted for different languages (i.e., Chinese, French, and British English). Results demonstrate that our model exhibits superior recognition performance to the state-of-the-art by a large margin. Lei Liu 0049, Li Liu 0036 |
ICASSP | 1 |
| 2023 | Global Balanced Experts for Federated Long-Tailed LearningabstractFederated learning (FL) is a prevalent distributed machine learning approach that enables collaborative training of a global model across multiple devices without sharing local data. However, the presence of long-tailed data can negatively deteriorate the model’s performance in real-world FL applications. Moreover, existing re-balance strategies are less effective for the federated long-tailed issue when directly utilizing local label distribution as the class prior at the clients’ side. To this end, we propose a novel Global Balanced Multi-Expert (GBME) framework to optimize a balanced global objective, which does not require additional information beyond the standard FL pipeline. In particular, a proxy is derived from the accumulated gradients uploaded by the clients after local training, and is shared by all clients as the class prior for re-balance training. Such a proxy can also guide the client grouping to train a multi-expert model, where the knowledge from different clients can be aggregated via the ensemble of different experts corresponding to different client groups. To further strengthen the privacy-preserving ability, we present a GBME-p algorithm with a theoretical guarantee to prevent privacy leakage from the proxy. Extensive experiments on long-tailed decentralized datasets demonstrate the effectiveness of GBME and GBME-p, both of which show superior performance to state-of-the-art methods. The code is available at here. Yaopei Zeng, Lei Liu 0049, Li Liu 0036, Li Shen 0008, Shaoguo Liu, Baoyuan Wu |
ICCV | 2 |
| 2023 | Quality-Aware RGBT Tracking via Supervised Reliability Learning and Weighted Residual GuidanceabstractRGB and thermal infrared (TIR) data have different visual properties, which make their fusion essential for effective object tracking in diverse environments and scenes. Existing RGBT tracking methods commonly use attention mechanisms to generate reliability weights for multi-modal feature fusion. However, without explicit supervision, these weights may be unreliably estimated, especially in complex scenarios. To address this problem, we propose a novel Quality-Aware RGBT Tracker (QAT) for robust RGBT tracking. QAT learns reliable weights for each modality in a supervised manner and performs weighted residual guidance to extract and leverage useful features from both modalities. We address the issue of the lack of labels for reliability learning by designing an efficient three-branch network that generates reliable pseudo labels, and a simple binary classification scheme that estimates high-accuracy reliability weights, mitigating the effect of noisy pseudo labels. To propagate useful features between modalities while reducing the influence of noisy modal features on the migrated information, we design a weighted residual guidance module based on the estimated weights and residual connections. We evaluate our proposed QAT on five benchmark datasets, including GTOT, RGBT210, RGBT234, LasHeR, and VTUAV, and demonstrate its excellent performance compared to state-of-the-art methods. Experimental results show that QAT outperforms existing RGBT tracking methods in various challenging scenarios, demonstrating its efficacy in improving the reliability and accuracy of RGBT tracking. Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001 |
ACM Multimedia | 1 |
| 2023 | Cuing Without Sharing: A Federated Cued Speech Recognition Framework via Mutual Knowledge DistillationabstractCued Speech (CS) is a visual coding tool to encode spoken languages at the phonetic level, which combines lip-reading and hand gestures to effectively assist communication among people with hearing impairments. The Automatic CS Recognition (ACSR) task aims to recognize CS videos into linguistic texts, which involves both lips and hands as two distinct modalities conveying complementary information. However, the traditional centralized training approach poses potential privacy risks due to the use of facial and gesture videos in CS data. To address this issue, we propose a new Federated Cued Speech Recognition (FedCSR) framework to train an ACSR model over the decentralized CS data without sharing private information. In particular, a mutual knowledge distillation method is proposed to maintain cross-modal semantic consistency of the Non-IID CS data, which ensures learning a unified feature space for linguistic and visual information. On the server side, a globally shared linguistic model is trained to capture the long-term dependencies in the text sentences, which is aligned with the visual information from the local clients via visual-to-linguistic distillation. On the client side, the visual model of each client is trained with its own local data, assisted by linguistic-to-visual distillation treating the linguistic model as the teacher. To the best of our knowledge, this is the first approach to consider the federated ACSR task for privacy protection. Experimental results on the Chinese CS dataset with multiple cuers demonstrate that our approach outperforms both mainstream federated learning baselines and existing centralized state-of-the-art ACSR methods, achieving 9.7% performance improvement for character error rate (CER) and 15.0% for word error rate (WER). The Chinese CS dataset and our code will be open-sourced. Lei Liu 0049, Li Liu 0036 |
ACM Multimedia | 2 |
| 2023 | Siamese transformer RGBT tracking
Futian Wang, Lei Liu 0049, Chenglong Li 0002, Jing Tang 0001 |
Appl. Intell. | 3 |
| 2023 | Diag-IoU Loss for Object DetectionabstractExisting IoU-based loss functions have achieved promising performance for bounding box regression in object detection. However, they cannot fully reflect the relation between the predicted and target boxes in the case of box inclusions, and might thus deteriorate detection accuracy and efficiency. In this paper, we design a novel similarity measurement based on the box diagonal called Diag-IoU to well represent the divergence between the predicted and target boxes even in the case of box inclusions, and thus achieve superior localization accuracy and fast convergence. In particular, we equivalently represent a rectangular box with its box diagonal, which contains exclusive and informative geometrical factors, and define the Diag-IoU based on the similarities of a set of sampled point pairs from the predicted and target box diagonals. Based on the Diag-IoU, we design a general Diag-IoU loss, which can provide holistic information in measuring two boxes and thus differentiate the two boxes in the case of box inclusions. To validate the effectiveness of the proposed method, we apply the Diag-IoU loss to several representative object detectors, including YOLO v5s, Faster R-CNN, and FCOS. Extensive experiments on the synthetic data and two challenging object detection benchmark datasets, i.e., MS COCO and PASCAL VOC, demonstrate the superior performance of the proposed Diag-IoU loss compared to previous IoU-based losses as well as other metrics. Shuangqing Zhang, Chenglong Li 0002, Lei Liu 0049, Zhang Zhang 0001, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Cross-Modal Object Tracking: Modality-Aware Representations and a Unified BenchmarkabstractIn many visual systems, visual tracking often bases on RGB image sequences, in which some targets are invalid in low-light conditions, and tracking performance is thus affected significantly. Introducing other modalities such as depth and infrared data is an effective way to handle imaging limitations of individual sources, but multi-modal imaging platforms usually require elaborate designs and cannot be applied in many real-world applications at present. Near-infrared (NIR) imaging becomes an essential part of many surveillance cameras, whose imaging is switchable between RGB and NIR based on the light intensity. These two modalities are heterogeneous with very different visual properties and thus bring big challenges for visual tracking. However, existing works have not studied this challenging problem. In this work, we address the cross-modal object tracking problem and contribute a new video dataset, including 654 cross-modal image sequences with over 481K frames in total, and the average video length is more than 735 frames. To promote the research and development of cross-modal object tracking, we propose a new algorithm, which learns the modality-aware target representation to mitigate the appearance gap between RGB and NIR modalities in the tracking process. It is plug-and-play and could thus be flexibly embedded into different tracking frameworks. Extensive experiments on the dataset are conducted, and we demonstrate the effectiveness of the proposed algorithm in two representative tracking frameworks against 19 state-of-the-art tracking methods. Dataset, code, model and results are available at https://github.com/mmic-lcl/source-code. Chenglong Li 0002, Tianhao Zhu, Lei Liu 0049, Xiaonan Si, Zilin Fan, Sulan Zhai |
AAAI | 3 |
| 2022 | Attribute-Based Progressive Fusion Network for RGBT TrackingabstractRGBT tracking usually suffers from various challenge factors, such as fast motion, scale variation, illumination variation, thermal crossover and occlusion, to name a few. Existing works often study fusion models to solve all challenges simultaneously, and it requires fusion models complex enough and training data large enough, which are usually difficult to be constructed in real-world scenarios. In this work, we disentangle the fusion process via the challenge attributes, and thus propose a novel Attribute-based Progressive Fusion Network (APFNet) to increase the fusion capacity with a small number of parameters while reducing the dependence on large-scale training data. In particular, we design five attribute-specific fusion branches to integrate RGB and thermal features under the challenges of thermal crossover, illumination variation, scale variation, occlusion and fast motion respectively. By disentangling the fusion process, we can use a small number of parameters for each branch to achieve robust fusion of different modalities and train each branch using the small training subset with the corresponding attribute annotation. Then, to adaptive fuse features of all branches, we design an aggregation fusion module based on SKNet. Finally, we also design an enhancement fusion transformer to strengthen the aggregated feature and modality-specific features. Experimental results on benchmark datasets demonstrate the effectiveness of our APFNet against other state-of-the-art methods. Yun Xiao 0003, Chenglong Li 0002, Lei Liu 0049, Jin Tang 0001 |
AAAI | 4 |
| 2022 | The First Challenge on Moving Object Detection and Tracking in Satellite Videos: Methods and ResultsabstractIn this paper, we briefly summarize the first challenge on moving object detection and tracking in satellite videos (SatVideoDT). This challenge has three tracks related to satellite video analysis, including moving object detection (Track 1), single object tracking (Track 2), and multiple-object tracking (Track 3). 123, 89, and 70 participants successfully registered, while 37, 42, and 29 teams submitted their final results on the test datasets for Tracks 1-3, respectively. The top-performing methods and their results in each track are described with details. This challenge establishes a new benchmark for satellite video analysis. Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Chenguang Dai, Weilong Guo, Xiyu Qi, Kelong Tu, Shudan Zhu, Lai Chen, Bin Lin 0013, Chaocan Xue, Jinlei Zheng, Limei Qin, Ying Li 0017, Manqi Zhao, Lu Ruan 0003, Mingpeng Cui, Guanchen Ding, Guangwei Jiang, Zhenzhong Chen 0001, Kaiyang Cao, Lingyu Kong, Shaodong Chen, Zhicheng Zhao 0001, Qin Shen, Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003 |
ICPR | 34 |
| 2022 | EllipseIoU: A General Metric for Aerial Object Detection
Xinbo Yang, Chenglong Li 0002, Rui Ruan, Lei Liu 0049, Bin Luo 0001 |
PRCV (3) | 4 |
| 2022 | RGBT tracking based on cooperative low-rank graph model
Longfeng Shen, Xiaoxiao Wang 0003, Lei Liu 0049, Bin Hou, Yulei Jian, Jin Tang 0001, Bin Luo 0001 |
Neurocomputing | 3 |
| 2020 | Challenge-Aware RGBT Tracking
Chenglong Li 0002, Lei Liu 0049, Andong Lu, Jin Tang 0001 |
ECCV (22) | 2 |
| 2020 | Semi-Supervised Active Learning for COVID-19 Lung Ultrasound Multi-symptom ClassificationabstractUltrasound (US) is a non-invasive yet effective medical diagnostic imaging technique for the COVID-19 global pandemic. However, due to complex feature behaviors and expensive annotations of US images, it is difficult to apply Artificial Intelligence (AI) assisting approaches for the lung's multi-symptom (multi-label) classification. To overcome these difficulties, we propose a novel semi-supervised Two-Stream Active Learning (TSAL) method to model complicated features and reduce labeling costs in an iterative manner. The core component of TSAL is the multi-label learning mechanism, in which label correlation information is used to design a multi-label margin (MLM) strategy and a confidence validation for automatically selecting informative samples and confident labels. In this framework, a multi-symptom multi-label (MSML) classification network is proposed to learn discriminative features of lung symptoms, and a human-machine interaction (HMI) is exploited to confirm the final annotations that are used to fine-tune MSML. Moreover, a novel lung US dataset named COVID19-LUSMS is built, currently containing 71 clinical patients with 6,836 images sampled from 678 videos. Experimental evaluations show that TSAL can achieve superior performance to the baseline and the state-of-the-art using only 20% data. Qualitatively, visualization of the attention map confirms a good consistency between the model prediction and the clinical knowledge. Lei Liu 0049, Wentao Lei, Li Liu 0036, Yongfang Luo |
ICTAI | 1 |
| 2012 | Combination of Heterogeneous Features for Wrist Pulse Blood Flow Signal Diagnosis via Multiple Kernel LearningabstractWrist pulse signal is of great importance in the analysis of the health status and pathologic changes of a person. A number of feature extraction methods have been proposed to extract linear and nonlinear, and time and frequency features of wrist pulse signal. These features are heterogeneous in nature and are likely to contain complementary information, which highlights the need for the integration of heterogeneous features for pulse classification and diagnosis. In this paper, we propose a novel effective method to classify the wrist pulse blood flow signals by using the multiple kernel learning (MKL) algorithm to combine multiple types of features. In the proposed method, seven types of features are first extracted from the wrist pulse blood flow signals using the state-of-the-art pulse feature extraction methods, and are then fed to an efficient MKL method, SimpleMKL, to combine heterogeneous features for more effective classification. Experimental results show that the proposed method is promising in integrating multiple types of pulse features to further enhance the classification performance. Lei Liu 0049, Wangmeng Zuo, David Zhang 0001, Naimin Li |
IEEE Trans. Inf. Technol. Biomed. | 1 |
| 2011 | Adaptive Weighted Fusion of Local Kernel Classifiers for Effective Pattern Classification
Shixin Yang, Wangmeng Zuo, Lei Liu 0049, Yanlai Li, David Zhang 0001 |
ICIC (1) | 3 |
| 2009 | Spatially Smooth Subspace Face Recognition Using LOG and DOG Penalties
Wangmeng Zuo, Lei Liu 0049, Kuanquan Wang, David Zhang 0001 |
ISNN (3) | 2 |