EDBT 2026 Demo / reviewers in the wild / expert
Lingyan Liang
dblp:284/2563
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2026
0009-0002-3648-0678ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust Exemplar Prompt Learning via Bi-directional Visual-Semantic Alignment for Multi-Object TrackingabstractRecent multi-object tracking (MOT) approaches increasingly leverage pre-trained CLIP models to boost cross-domain generalization. A common strategy uses a predefined TrackBook—a closed-set of visual concepts—as textual prompts to guide learning of domain-invariant representations. However, these fixed prompts lack adaptive context, causing limited generalization. To address this limitation, this paper introduces a robust Exemplar Prompt Learning (EPL) framework via Bi-directional Visual-Semantic Alignment (BiVSA), termed EPL-MOT, which augments textual prompts with instance-aware contextual information derived during tracking. Specifically, an EPL module is designed to dynamically enrich textual prompts with contextual cues, enabling instance-specific adaptation without inducing category shift. Furthermore, a BiVSA module is proposed to deepen cross-modal interaction by incorporating bidirectional learnable prompts into both textual and visual branches. This facilitates progressive integration of global semantic features with local visual structures, resulting in a more effectively aligned visual-semantic space. Finally, to enhance robustness against distractors, a Category-guided Detection Query Generator (CDQG) is constructed, which incorporates base-class textual information to suppress irrelevant targets. Comprehensive evaluations on MOT17 and MOT20 demonstrate that the proposed EPL-MOT achieves competitive performance across both in-domain and cross-domain settings. Lingyan Liang, Gang Dong, Dongchao Wen, Kaihua Zhang 0001 |
ICMR | 1 |
| 2025 | Continuously Learning Video-level Object Tokens for Robust UAV trackingabstractDue to the dynamic changes in flight motion and viewpoint, the objects in unmanned aerial vehicle (UAV) tracking scenarios often suffer from drastic appearance variations. Existing UAV trackers often leverage a frame-level matching mechanism, which measures the appearance similarity between the object template and the search frame. The drastic object appearance variations degrade the learned model, leading to drift issue. To this end, this paper presents a video-level UAV tracking framework that focuses on Continuously Learning (CL) effective and efficient spatio-temporal object tokens for robust tracking, dubbed as CLTrack. Specifically, the CLTrack first learns a series of spatio-temporal object tokens via a dynamic filtering module (DFM), which encodes more consensus object appearance information from each frame. Afterwards, a spatio-temporal enhancement module (STEM) is designed via cascading a temporal and a spatial attention to fully interact with the selected tokens with stable long-range spatio-temporal context information of the tracked object. Finally, to ensure the learned model encodes the rich context information without catastrophic forgetting, a video-level tracking loss is designed to supervise feature learning from the whole video frames. Extensive experiments on three UAV benchmarks including UAV123, DTB70 and VisDrone2018 demonstrate that the proposed CLTrack achieves state-of-the-art performance. Shenglong Hu, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001 |
ICASSP | 4 |
| 2025 | Easy-to-hard Instance-level Feature Fusion for Co-saliency DetectionabstractExisting leading deep learning-based Co-saliency Detection (CoD) methods often learn the consensus features from the input image group without considering the complexity of each image. Despite the demonstrated success, the input images may contain hard samples with high complexity, e.g., those containing distractors that have similar appearance but different semantics to the co-salient objects. This is prone to mislead the learned model to treat these distractors as co-salient objects, leading to classification ambiguity. To address this issue, this paper presents an easy-to-hard instance-level feature Fusion framework for CoD, termed E2HCoD. The E2HCoD exploits the instance-level co-salient object consensus cues from the easy samples as reliable guidance to accurately fuse the co-salient object features in the hard samples. First, we design a Feature Filtering Module (FFM) that evaluates image complexity by integrating entropy, variance, texture, and edge density cues, allowing the model to select the easy samples with relatively easy backgrounds. Then, we develop an Easy-instance Embedding Branch (EEB), which accurately segments the co-salient object masks from the easy samples as the instance-level guidance to learn the accurate co-salient object consensus cues. Then, with the consensus knowledge from the easy samples as guidance, we construct an Easy-instance guided Fusion Branch (EFB), which fully interacts with the consensus features from the hard samples via a cross-attention mechanism, yielding the refined features that highlight the co-salient objects while suppressing the distractors. Finally, the refined features are fed into the decoder, generating a high-quality CoD prediction. Extensive experiments demonstrate that the proposed E2HCoD achieves state-of-the-art performance on CoSal2015, CoCA, and CoSOD3k. Chuang Ding, Zhidong Han, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001 |
ICASSP | 4 |
| 2025 | Spatio-Semantic Prompt guided Adaptive Segment Anything for Remote Sensing Change DetectionabstractExisting leading remote sensing change detection (RSCD) often takes a semantic-agnostic learning paradigm, which uses a binary ground-truth mask as supervision for model training. Despite the demonstrated success, due to the intrinsic characteristic of extremely complicated scene changes in RS images, this paradigm is prone to be misled by irrelevant semantic category changes, leading to a noisy CD mask prediction. To address this issue, this paper presents a Spatio-Semantic Prompt (SSP) guided adaptive Segment Anything Model (SAM) for RSCD, dubbed as SSP-SAM. The SSP-SAM introduces sparse textual and dense mask prompts into SAM to encode the task-specific semantic knowledge for RSCD. Specifically, we first encode the powerful textual semantic knowledge using Contrastive Language-Image Pre-training (CLIP) to determine the desired change semantic category. Then, we design a spatial dense prompt module that yields an attention map as prompt features to further refine the desired changed regions. Subsequently, we fine-tune the SAM through an adaptor to integrate the spatial-semantic prompt cues, yielding a coarse CD mask prediction. Finally, guided by the coarse CD mask, a multi-scale mask attention mechanism is adopted to learn the refined semantic representations of the changed targets, predicting the accurate CD mask. Extensive experiments on a variety of benchmark datasets demonstrate that the proposed SSP-SAM achieves state-of-the-art performance. Shenglong Hu, Zhidong Han, Gang Dong, Lingyan Liang, Dongchao Wen, Kaihua Zhang 0001 |
ICASSP | 4 |
| 2024 | Segment Anything Model Guided Semantic Knowledge Learning For Remote Sensing Change DetectionabstractExisting deep learning based remote sensing change detection (RSCD) methods only rely on binary ground-truth to guide the network learning while neglecting the useful semantic guidance. As a result, the network can be readily misled by irrelevant category changes, leading to degraded performance and slow convergence of the model. To this end, we propose a novel segment anything model (SAM) guided framework, termed as SAM-CD, which mines the rich semantic knowledge from the SAM for RSCD. Specifically, we first employ a transformer encoder to extract multi-scale global features from the bi-temporal images. Meanwhile, we obtain semantic prior masks from the bi-temporal images by providing the SAM with category-relevant text prompts. Then, using the semantic prior masks as constraints, we design a masked attention module (MAM) that generates local features related to the interested categories. Finally, the local and global features are fused and fed into a multi-layer perception (MLP) decoder to obtain the change map. The whole network is trained in an end-to-end manner that can readily encode the rich semantic knowledge of the changed targets to predict an accurate change map. Extensive experiments demonstrate that the proposed SAM-CD achieves state-of-the-art performance on a variety of benchmark datasets. Zixuan Sun, Huihui Song 0003, Kaihua Zhang 0001, Gang Dong, Lingyan Liang, Yaqian Zhao |
ICASSP | 5 |
| 2024 | Glance, Focus and Refinement Network for Remote Sensing Change DetectionabstractExisting change detection (CD) methods often directly fuse the multi-level features from bi-temporal remote sensing images without discriminatively considering each pixel's importance. Despite the demonstrated success, unselectively mixing the features degrades the model's performance to effectively capture the change targets due to the imbalance ratio between the change regions and the whole scene. To this end, this paper presents a glance, focus, and refinement network (GFRNet), which formulates CD as a continuous, step-by-step focusing process to mimic the human visual system. Specifically, the GFRNet first employs a transformer encoder to extract the global features from the bi-temporal images, where each feature takes a glance at the whole scene. Then, the GFRNet gradually pays attention to a cascade of salient regions, and ultimately progressively refines its focus on the desired areas of change. Comprehensive evaluations on two extensively utilized benchmark datasets, including LEVIR-CD and WHU-CD, demonstrate the superiority of our GFR-Net to a variety of state-of-the-art methods. Zixuan Sun, Yuhui Zheng, Kaihua Zhang 0001, Gang Dong, Lingyan Liang, Yaqian Zhao |
ICASSP | 6 |
| 2024 | Group-wise co-salient object detection via multi-view self-labeling novel class discovery
Gang Dong, Lingyan Liang, Yaqian Zhao, Kaihua Zhang 0001 |
Frontiers Comput. Sci. | 3 |
| 2023 | Group-Wise Co-Salient Object Detection with Siamese Transformers Via Brownian Distance Covariance MatchingabstractCo-salient object detection (CoSOD) aims to discover and segment foreground targets in a group of images with the same semantic category. Existing mainstream approaches often employ convolutional neural networks (CNNs) to learn the semantic-invariant features from a group of images. Despite demonstrated success, there exist two limitations: 1) The CNNs introduce the inductive bias of locality that are difficult to model long-range dependency, limiting their feature representation capability. 2) Their models lack discriminability to differentiate semantic differences between different groups since only one group of images with the same semantic category has been taken into account for model training. To address these issues, this paper presents a Siamese Transformer architecture for CoSOD that can fully mine the group-wise semantic contrast information for more discriminative feature learning. Specifically, the designed Siamese Transformer takes two groups of images as input for feature contrastive learning. Each group is processed by a Transformer branch with shared weights to capture the long-range interaction information. Besides, to model the complex non-linear interactions between these two branches, we further design a Brownian distance covariance (BDC) module that uses joint distribution to measure the inter- and intra-group semantic similarity. The BDC can be efficiently calculated in closed form that can fully characterize independence for effective feature contrastive learning. Extensive evaluations on the three largest and most challenging benchmark datasets (CoSal2015, CoCA, and CoSOD3k) demonstrate the superiority of our method over a variety of state-of-the-art methods. Lingyan Liang, Yaqian Zhao, Kaihua Zhang 0001 |
ICASSP | 3 |
| 2023 | Object-Aware Calibrated Depth-Guided Transformer for RGB-D Co-Salient Object DetectionabstractThe key role of RGB-D co-salient object detection is to effectively fuse the common information of RGB and depth signals. Existing works directly mix the information captured from both original depth maps and RGB images, but ignore one critical issue: due to the low contrast of the neighborhood objects in depth, the depth maps’ salient regions may correspond to the interference background regions in the RGB images, thereby leading to unsatisfying performance. To address this issue, we propose an Object-aware Calibrated Depth guided transformer (dubbed as OCDFormer) for RGB-D co-salient object detection. The OCDFormer mainly consists of two key designs: First, we design a depth calibration module via spectral clustering, which yields a group of calibrated depth maps that can highlight the co-object region while suppressing the interference regions. Second, we construct a cross-modal transformer, in which the common information from the RGB and the calibrated depth maps are fully captured by first injecting common tokens into the individual tokens, and then mixing them with an interaction-attention mechanism. Extensive evaluations demonstrate that our OCDFormer sets a new state-of-the-art on two public standard benchmarks including RGB-D CoSall5O and RGB-D CoSegl83. Lingyan Liang, Yaqian Zhao, Kaihua Zhang 0001 |
ICME | 2 |
| 2021 | Unsupervised Active Learning via Subspace LearningabstractUnsupervised active learning has been an active research topic in machine learning community, with the purpose of choosing representative samples to be labelled in an unsupervised manner. Previous works usually take the minimization of data reconstruction loss as the criterion to select representative samples which can better approximate original inputs. However, data are often drawn from low-dimensional subspaces embedded in an arbitrary high-dimensional space in many scenarios, thus it might severely bring in noise if attempting to precisely reconstruct all entries of one observation, leading to a suboptimal solution. In view of this, this paper proposes a novel unsupervised Active Learning model via Subspace Learning, called ALSL. In contrast to previous approaches, ALSL aims to discovery the low-rank structures of data, and then perform sample selection based on learnt low-rank representations. To this end, we devise two different strategies and propose two corresponding formulations to perform unsupervised active learning with and under low-rank sample representations respectively. Since the proposed formulations involve several non-smooth regularization terms, we develop a simple but effective optimization procedure to solve them. Extensive experiments are performed on five publicly available datasets, and experimental results demonstrate the proposed first formulation achieves comparable performance with the state-of-the-arts, while the second formulation significantly outperforms them, achieving a 13\% improvement over the second best baseline at most. Kaihang Mao, Lingyan Liang, Dongchun Ren, Ye Yuan 0001, Guoren Wang |
AAAI | 3 |
| 2021 | Auto-weighted centralised multi-task learning via integrating functional and structural connectivity for subjective cognitive decline diagnosis
Bai Ying Lei, Nina Cheng, Alejandro F. Frangi, Bihan Yu, Lingyan Liang, Wei Mai, Gaoxiong Duan, Xiucheng Nong, Jiahui Su, Tianfu Wang 0001, Lihua Zhao, Demao Deng, Zhiguo Zhang 0001 |
Medical Image Anal. | 6 |