EDBT 2026 Demo / reviewers in the wild / expert
Luping Ji
dblp:92/281
· DBLP profile ↗
57ranked-venue papers
8as first author
42since 2021 · last 2026
0000-0002-1200-5218ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 7 first-author · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain-Auxiliary Infrared Moving Small Target Detection by Learning to Overlook Domain DiscrepancyabstractCurrently, almost all traditional infrared small target detection methods work on the assumption that training and test sets always belong to the same domain, and training samples are sufficient. However, in real applications, a new detection task could often have no sufficient training samples from a special domain. In this situation, adopting the auxiliary data from big-sample domains is usually believed to be one of the most potential solutions. However, exceeding expectations, it is found that simply adding auxiliary samples cannot often be always effective, even causing performance decline, due to existing infrared domain shift. To overcome this unexpected problem, we propose the first infrared moving small target detection framework with domain-auxiliary supports by Learning to Overlook Domain Discrepancy (Loddis). This framework consists of three primary processing stages: correlation weakening, domain confusing, and target consistency contrastive learning. Breaking through traditional learning paradigm, through auxiliary data, it enables the model to focus more on targets themselves, and less on image backgrounds, minimizing the sensitivity to domain discrepancy. The extensive experiments on 6 different-domain datasets show the effectiveness and superiority of the proposed Loddis framework for infrared small target detection. Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001 |
AAAI | 2 |
| 2026 | SeViL: Semi-supervised Vision-Language Learning with Text Prompt Guiding for Moving Infrared Small Target DetectionabstractUnlike traditional object detection, moving infrared small target detection is highly challenging due to tiny target size and limited labeled samples. Currently, most existing methods mainly focus on the pure-vision features usually by fully-supervised learning, heavily relying on extensive high-cost manual annotations. Moreover, they almost have not concerned the potentials of multi-modal (e.g., vision and text) learning yet. To address these issues, inspired by prevalent vision-language models, we propose the first semi-supervised vision-language (SeViL) framework with adaptive text prompt guiding. Breaking through traditional pure-vision modality, it takes text prompts as prior knowledge to adaptively enhance target regions and then filter the low-quality pseudo-labels generated on unlabeled data. In the meanwhile, we employ an adaptive cross-modal masking strategy to align text and vision features, promoting cross-modal deep interactions. Remarkably, our extensive experiments on three public datasets (DAUB, ITSDT-15K and IRDST) verify that our new scheme could outperform other semi-supervised ones, and even achieve comparable performance to fully-supervised state-of-the-art (SOTA) methods, with only 10% labeled training samples. Luping Ji, Jianghong Huang, Sicheng Zhu |
AAAI | 2 |
| 2026 | Cross-domain Joint Learning with Prototype-guided Mixture-of-Experts for Infrared Moving Small Target DetectionabstractInfrared small target detection often faces significant domain gaps across datasets due to varying sensors and scene distributions. Currently, most existing methods are typically based on single-domain learning (i.e., training and test are on the same dataset), requiring training separate detectors when considering different datasets. However, they overlook the valuable public knowledge across domains and limit the applicability in multiple infrared scenarios. To break through single-domain learning, implementing only one universal detector simultaneously on multiple datasets, as the first exploration, we propose a cross-domain joint learning task framework with prototype-guided Mixture-of-Experts (CoMoE). Specifically, it designs a hyperspherical prototype learning to adaptively maintain both domain-specific prototypes and global prototypes, enhancing cross-domain feature representation. Meanwhile, a domain-aware Mixture-of-Experts with Top-K routing strategy is proposed to select the optimal domain experts. Moreover, to enhance cross-domain feature alignment, we design an adaptive cross-domain feature modulation with noise-guided contrastive learning. The extensive experiments on a newly constructed benchmark comprising three datasets verify the superiority of our CoMoE, even under limited data settings. It could often surpass general joint learning methods, and state-of-the-art (SOTA) single-domain ones. Luping Ji, Jianghong Huang, Sicheng Zhu, Mao Ye 0001 |
AAAI | 2 |
| 2026 | Multi-view Invariance Learning for 3D Scene Graph Pre-training via Collaborative Cross-Modal Regularizationabstract3D scene graph generation is a pivotal task in scene understanding. Its performance is easy to be constrained by the limited availability of annotated data. Currently, the existing solutions on point cloud pre-training usually emphasize on object-centric representations while neglecting the predicate feature learning. This limitation significantly hinders their relational reasoning capabilities, as inter-object relationships are fundamentally governed by predicate features. To enhance 3D Scene Graphs Pre-training, this paper proposes a task-specific Multi-view Invariance Learning framework with Collaborative Cross-modal Regularization. In detail, the inherent horizontal-rotation invariance of 3D objects and their semantic relationships are leveraged to construct a self-supervised paradigm for triplet feature learning. Moreover, our framework harnesses the cross-modal prior knowledge from the vision-language model to regularize model optimization. It could further achieve the semantic discrimination via unsupervised deep clustering. To resolve the knowledge discrepancies arising from the pre-trained model in fine-tuning, a predicate adapter equipped with knowledge filtering gate is devised to selectively aggregate the predicate features of pre-trained model. Extensive experiments demonstrate that our framework is effective in boosting 3D scene graph generation performance, surpassing state-of-the-art ones. Luping Ji, Ruijie Xiao, Jiayuan Sun |
AAAI | 2 |
| 2026 | Local-global collaborative feature learning with level-wise decoding for infrared small target detection
Luping Ji, Shengjia Chen, Jianghong Huang |
Comput. Vis. Image Underst. | 2 |
| 2026 | BeltCrack: The first sequential-image industrial conveyor belt crack detection dataset and its baseline with triple-domain feature learning
Jianghong Huang, Luping Ji, Mao Ye 0001 |
Pattern Recognit. | 2 |
| 2026 | Deformable Feature Alignment and Refinement for moving infrared small target detection
Dengyan Luo, Yanping Xiang, Luping Ji, Shuai Li 0005, Mao Ye 0001 |
Pattern Recognit. | 4 |
| 2026 | Long-Short Match for Lost Control in UAV Multi-Object TrackingabstractMulti-Object Tracking (MOT) in Unmanned Aerial Vehicles (UAV) aims to continuously and stably detect and track objects in videos captured by UAVs. In existing MOT tracking-by-detection schemes, the tracker with a fixed step size is always employed, and a fixed length of past tracking information is input to the tracker to guide position prediction. However, the limited prediction range of a single-scale tracker leads to frequent tracking losses, and limited historical information also reduces tracking accuracy. To address these limitations, we propose a novel Long-Short Match (LSMTrack) tracking method. The key idea is to use long and short trackers and maintain a long-term motion state to improve tracking performance, thus reducing the likelihood of entering the lost status. To this end, a new Mamba-based tracker and a long-short match strategy are proposed. For long and short trackers, the same architecture is used based on Mamba. Unlike the previous Mamba-based approach, the proposed tracker maintains a long-term state while updating the state and making position predictions in each time step, so we call it a step Mamba tracker. Meanwhile, we devise a long-short match strategy at the inference stage to integrate long and short trackers, and design a lost control operation which updates the long-term states using historical state values. In this way, the matching probability and the inference efficiency are guaranteed. Experimental results on two UAV MOT datasets confirm the state-of-the-art performance. Specifically, the best results are achieved in terms of two popular MOTA and IDF1 tracking evaluation metrics. Zi-Zhuang Zou, Mao Ye 0001, Luping Ji, Lihua Zhou, Song Tang 0001, Yan Gan, Shuai Li 0005 |
IEEE Trans. Multim. | 3 |
| 2025 | Motion Prior Knowledge Learning with Homogeneous Language Descriptions for Moving Infrared Small Target DetectionabstractDifferent from traditional object detection, pure vision is not enough to infrared small target detection, due to small target size and weak background contrast. For promoting detection performance, more target representations are needed. Currently, motion representations have been proved to be one of the most potential feature kinds for infrared small target detection. Existing methods have an obvious weakness, that besides vision features, they could only capture coarse motion representations from temporal domain. With vision features, fine motion representations could be more effective to enhance detection performance. To overcome this weakness, inspired by prevalent vision-language models, we propose the first vision-language framework with motion prior knowledge learning (MoPKL). Breaking through traditional pure-vision modality, it utilizes homogeneous language descriptions, formatted for moving targets, to directionally guide vision channel learning motion prior knowledge. With the facilitation of motion-vision alignment and motion-relation mining, the motion of infrared small targets is further refined by graph attention, to generate more fine motion representations. The extensive experiments on datasets ITSDT-15K and IRDST show that our framework is effective. It could often obviously outperform other methods. Shengjia Chen, Luping Ji, Mao Ye 0001 |
AAAI | 2 |
| 2025 | Queryable Prototype Multiple Instance Learning with Vision-Language Models for Incremental Whole Slide Image ClassificationabstractWhole Slide Image (WSI) classification has very significant applications in clinical pathology, e.g., tumor identification and cancer diagnosis. Currently, most research attention is focused on Multiple Instance Learning (MIL) using static datasets. One of the most obvious weaknesses of these methods is that they cannot efficiently preserve and utilize previously learned knowledge. With any new data arriving, classification models are required to be re-trained on both previous and current new data. To overcome this shortcoming and break through traditional vision modality, this paper proposes the first Vision-Language-based framework with Queryable Prototype Multiple Instance Learning (QPMIL-VL) specially designed for incremental WSI classification. This framework mainly consists of two information processing branches: one is for generating bag-level features by prototype-guided aggregation of instance features, while the other is for enhancing class features through a combination of class ensemble, tunable vector and class similarity loss. The experiments on four public WSI datasets demonstrate that our QPMIL-VL framework is effective for incremental WSI classification and often significantly outperforms other compared methods, achieving state-of-the-art (SOTA) performance. Jiaxiang Gou, Luping Ji, Pei Liu 0008, Mao Ye 0001 |
AAAI | 2 |
| 2025 | Mining In-distribution Attributes in Outliers for Out-of-distribution DetectionabstractOut-of-distribution (OOD) detection is indispensable for deploying reliable machine learning systems in real-world scenarios. Recent works, using auxiliary outliers in training, have shown good potential. However, they seldom concern the intrinsic correlations between in-distribution (ID) and OOD data. In this work, we discover an obvious correlation that OOD data usually possesses significant ID attributes. These attributes should be factored into the training process, rather than blindly suppressed as in previous approaches. Based on this insight, we propose a structured multi-view-based out-of-distribution detection learning (MVOL) framework, which facilitates rational handling of the intrinsic in-distribution attributes in outliers. We provide theoretical insights on the effectiveness of MVOL for OOD detection. Extensive experiments demonstrate the superiority of our framework to others. MVOL effectively utilizes both auxiliary OOD datasets and even wild datasets with noisy ID data. Luping Ji, Pei Liu 0008 |
AAAI | 2 |
| 2025 | Pseudo Visible Feature Fine-Grained Fusion for Thermal Object DetectionabstractThermal object detection is a critical task in various fields, such as surveillance and autonomous driving. Current state-of-the-art (SOTA) models always leverage a prior Thermal-To-Visible (T2V) translation model to obtain visible spectrum information, followed by a cross-modality aggregation module to fuse information from both modalities. However, this fusion approach does not fully exploit the complementary visible spectrum information beneficial for thermal detection. To address this issue, we propose a novel cross-modal fusion method called Pseudo Visible Feature Fine-Grained Fusion (PFGF). Specifically, a graph is constructed with nodes generated from multi-level thermal features and pseudo-visual latent features produced by the T2V model. Each level of features corresponds to a subgraph. An Inter-Mamba block is proposed to perform cross-modality fusion between nodes at the lowest level; while a Cascade Knowledge Integration (CKI) strategy is used to fuse low-level fused information to high-level subgraphs in a cascade manner. After several iterations of graph node updating, each subgraph outputs an aggregated feature to the detection head respectively. Unlike previous cross-modal fusion methods, our approach explicitly models high-level relationships between cross-modal data, effectively fusing different granularity information. Experimental results demonstrate that our method achieves SOTA detection performance. Code is available at https://github.com/liting1018/PFGF. Mao Ye 0001, Tianwen Wu, Nianxin Li, Shuaifeng Li, Song Tang 0001, Luping Ji |
CVPR | 7 |
| 2025 | Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational PathologyabstractHistopathology Whole-Slide Images (WSIs) provide an important tool to assess cancer prognosis in computational pathology (CPATH). While existing survival analysis (SA) approaches have made exciting progress, they are generally limited to adopting highly-expressive network architectures and only coarse-grained patient-level labels to learn visual prognostic representations from gigapixel WSIs. Such learning paradigm suffers from critical performance bottlenecks, when facing present scarce training data and standard multi-instance learning (MIL) framework in CPATH. To overcome it, this paper, for the first time, proposes a new Vision-Language-based SA (**VLSA**) paradigm. Concretely, (1) VLSA is driven by pathology VL foundation models. It no longer relies on high-capability networks and shows the advantage of *data efficiency*. (2) In vision-end, VLSA encodes textual prognostic prior and then employs it as *auxiliary signals* to guide the aggregating of visual prognostic features at instance level, thereby compensating for the weak supervision in MIL. Moreover, given the characteristics of SA, we propose i) *ordinal survival prompt learning* to transform continuous survival labels into textual prompts; and ii) *ordinal incidence function* as prediction target to make SA compatible with VL-based prediction. Notably, VLSA's predictions can be interpreted intuitively by our Shapley values-based method. The extensive experiments on five datasets confirm the effectiveness of our scheme. Our VLSA could pave a new way for SA in CPATH by offering weakly-supervised MIL an effective means to learn valuable prognostic clues from gigapixel WSIs. Our source code is available at https://github.com/liupei101/VLSA. Pei Liu 0008, Luping Ji, Jiaxiang Gou, Bo Fu 0007, Mao Ye 0001 |
ICLR | 2 |
| 2025 | SAM-Guided Semantic Knowledge Fusion for Visible-Infrared Object DetectionabstractVisible-infrared object detection has gained significant attention because of its applications in autonomous driving, video surveillance, and related fields. The effective fusion of multimodal information is fundamental to its success. The existing approaches concentrate on improving the pixel-level fusion mechanisms; detection performance has reached a plateau. We propose a new framework for SAM-guided semantic knowledge fusion ( SemFusion ). The core idea is to leverage semantic priors from large models while incorporating a lightweight cross-modal fusion strategy. Specifically, our method comprises two stages. In the first stage, the Flow-Guided RGB Feature Alignment (FGRA) module establishes object-aware correspondences between multimodalities based on SAM-generated masks. This ensures semantic-level feature matching by deformable convolution alignment. In the second stage, the Semantic Knowledge Distillation (SKD) strategy facilitates the transfer of large-model knowledge to the detection model through SAM feature, offset, and mask level distillations. For the detector model, three blocks are designed to augment any off-the-shelf detector. They are deformable cross-modal alignment, spatio-channel preliminary fusion, and mask-guided feature refinement. By alignment with SAM masks, semantic alignment and fusion can be achieved, breaking the pixel-level fusion barrier. Extensive experiments demonstrate that our method, as a plugin, exhibits superior performance on the DroneVehicle, VEDAI, and LLVIP datasets. Code is available at https://github.com/liting1018/SemFusion. Shuaifeng Li, Xiaolin Qin, Maoyuan Zhao, Luping Ji, Mao Ye 0001 |
ACM Multimedia | 6 |
| 2025 | Multimodal Causal Reasoning for UAV Object DetectionabstractUnmanned Aerial Vehicle (UAV) object detection faces significant challenges due to complex environmental conditions and different imaging conditions. These factors introduce significant changes in scale and appearance, particularly for small objects that occupy limited pixels and exhibit limited information, complicating detection tasks. To address these challenges, we propose a Multimodel Causal Reasoning framework based on YOLO backbone for UAV Object Detection (MCR-UOD). The key idea is to use the backdoor adjustment to discover the condition-invariant object representation for easy detection. Specifically, the YOLO backbone is first adjusted to incorporate the pre-trained vision-language model. The original category labels are replaced with semantic text prompts, and the detection head is replaced with text-image contrastive learning. Based on this backbone, our method consists of two parts. The first part, named language guided region exploration, discovers the regions with high probability of object existence using text embeddings based on vision-language model such as CLIP. Another part is the backdoor adjustment casual reasoning module, which constructs a confounder dictionary tailored to different imaging conditions to capture global image semantics and derives a prior probability distribution of shooting conditions. During causal inference, we use the confounder dictionary and the prior to intervene on local instance features, disentangling condition variations, and obtaining condition-invariant representations. Experimental results on several public datasets confirm the state-of-the-art performance of our approach. The code, data and models will be released upon publication of this paper. Nianxin Li, Mao Ye 0001, Lihua Zhou, Shuaifeng Li, Song Tang 0001, Luping Ji, Ce Zhu |
NeurIPS | 6 |
| 2025 | Adaptive graph attention networks with interactive learning for attributed graph clustering
Luping Ji, Lijun Wu 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Moving infrared dim and small target detection by mixed spatio-temporal encoding
Luping Ji, Shengjia Chen, Sicheng Zhu |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | WSOE: Weakly Supervised Outlier Exposure for Object-level Out-of-distribution detection
Luping Ji, Pei Liu 0008 |
Expert Syst. Appl. | 2 |
| 2025 | Spatial-temporal-channel collaborative feature learning with transformers for infrared small target detection
Sicheng Zhu, Luping Ji, Shengjia Chen |
Image Vis. Comput. | 2 |
| 2025 | Language-Driven Motion Prior Knowledge Learning for Moving Infrared Small Target DetectionabstractDifferent from traditional object detection, pure vision is often not enough to infrared small target detection (ISTD), due to the small target size and weak background contrast. For promoting performance, more target representations are often needed. Currently, motion representations have been proved to be one of the most potential feature patterns for infrared small targets. Besides vision features, existing methods have an obvious weakness that they could only capture coarse motion representations from the temporal domain. By vision features, fine motion representations could often be more effective to enhance detection performance. To overcome this weakness, and inspired by prevalent vision-language models (VLMs), the first vision-language framework with motion prior knowledge learning (MoPKL) was proposed in our previous work. To further extend it, we repropose an improved version, i.e., iMoPKL. Breaking through traditional pure-vision modality, it utilizes the homogeneous language descriptions, specially formatted for moving targets, to directionally guide vision channels to learn the motion prior knowledge of targets. In detail, it learns the distribution of target motion reconstruction corresponding to the language description as a type of prior knowledge. With the facilitation of language-driven motion alignment, the motion of infrared small targets could be further refined by motion-relation learning, to generate more fine motion representations. The extensive experiments on ITSDT-15K, DAUB-R, and IRDST-H show that our improvement version is effective. It could often obviously outperform the other methods, including our original MoPKL. Our source codes are available athttps://github.com/UESTC-nnLab/MoPKL Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001, Yongsheng Sang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Weakly Supervised Contrastive Learning With Quantity Prompts for Moving Infrared Small Target Detection
Luping Ji, Shengjia Chen, Sicheng Zhu, Jianghong Huang, Mao Ye 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Semi-Supervised Multiview Prototype Learning With Motion Reconstruction for Moving Infrared Small Target DetectionabstractMoving infrared small target detection is critical for various applications, e.g., remote sensing and military. Due to tiny target size and limited labeled data, accurately detecting targets is highly challenging. Currently, existing methods primarily focus on fully-supervised learning, which relies heavily on numerous annotated frames for training. However, annotating a large number of frames for each video is often expensive, time-consuming, and redundant, especially for low-quality infrared images. To break through traditional fully-supervised framework, we propose a new semi-supervised multi-view prototype (S2MVP) learning scheme that incorporates motion reconstruction. In our scheme, we design a bi-temporal motion perceptor based on bidirectional ConvGRU cells to effectively model the motion paradigms of targets by perceiving both forward and backward. Additionally, to explore the potential of unlabeled data, it generates the multi-view feature prototypes of targets as soft labels to guide feature learning by calculating cosine similarity. Imitating human visual system, it retains only the feature prototypes of recent frames. Moreover, it eliminates noisy pseudo-labels to enhance the quality of pseudo-labels through anomaly-driven pseudo-label filtering. Furthermore, we develop a target-aware motion reconstruction loss to provide additional supervision and prevent the loss of target details. To our best knowledge, the proposed S2MVP is the first work to utilize large-scale unlabeled video frames to detect moving infrared small targets. Although 10% labeled training samples are used, the experiments on three public benchmarks (DAUB, ITSDT-15K and IRDST) verify the superiority of our scheme compared to other methods. Source codes are available at https://github.com/UESTC-nnLab/S2MVP. Luping Ji, Jianghong Huang, Shengjia Chen, Sicheng Zhu, Mao Ye 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | MICPL: Motion-Inspired Cross-Pattern Learning for Small-Object Detection in Satellite VideosabstractFor small-object detection, vision patterns can only provide limited support to feature learning. Most prior schemes mainly depend on a single vision pattern to learn object features, seldom considering more latent motion patterns. In the real world, humans often efficiently perceive small objects through multipattern signals. Inspired by this observation, this article attempts to address small-object detection from a new prospective of latent pattern learning. To fulfill this purpose, it regards a real-world moving object as the spatiotemporal sequences of a static object to capture latent motion patterns. In view of this, we propose a motion-inspired cross-pattern learning (MICPL) scheme to capture the motion patterns for moving small-object scenarios. This scheme mainly consists of two crucial parts: motion pattern mining (MPM) and motion-vision adaption. The former is designed to effectively mine the motion pattern from time-dependent representation space. The latter is devised to correlate between motion patterns and vision semantics. In the meanwhile, we explore their cross-pattern interactions to guide MICPL to capture motion patterns effectively. Comparison experiments verify that, cooperated by motion pattern, even a simple detector could often refresh state-of-the-art (SOTA) results on moving small-object detection. Moreover, the experiments on two small-object-related tasks further prove the adaptivity and advantages of our cross-pattern feature learning scheme. Our source codes are available at https://github.com/UESTC-nnLab/MICPL. Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Dual Consistency Constraint-Based Self-Supervised Representation Learning for Heterogeneous Graphs With Missing AttributesabstractMissing attribute completion for unattributed nodes in heterogeneous graphs has received increasing attention, but previous works still suffer from the following issues: 1) they ignore the noise in the raw attributes, resulting in noise propagation and even inaccurate information generation during attribute completion, thus further influencing the representation learning; and 2) they ignore constraints on unattributed nodes when conducting consistency learning across augmented graph views, resulting in data inconsistency across views. To solve these issues, in this article, we propose a new dual consistency constraint-based self-supervised representation learning method for heterogeneous graphs with missing attributes. Specifically, we first investigate the representation completion and the within-view consistency loss to complete missing information in the representation space, and then, we investigate the cross-view consistency loss to ensure data consistency across views. We further reconstruct the masked data to avoid information loss due to the masking process. As a result, our method effectively filters out noise and inaccurate information by the representation completion process as well as achieves discriminative representation learning for heterogeneous graphs with missing attributes. Experimental results on various downstream tasks verify the superiority of our method. Yajie Lei, Yujie Mo, Luping Ji, Xiaofeng Zhu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Weakly-Supervised Residual Evidential Learning for Multi-Instance Uncertainty EstimationabstractUncertainty estimation (UE), as an effective means of quantifying predictive uncertainty, is crucial for safe and reliable decision-making, especially in high-risk scenarios. Existing UE schemes usually assume that there are completely-labeled samples to support fully-supervised learning. In practice, however, many UE tasks often have no sufficiently-labeled data to use, such as the Multiple Instance Learning (MIL) with only weak instance annotations. To bridge this gap, this paper, for the first time, addresses the weakly-supervised issue of *Multi-Instance UE* (MIUE) and proposes a new baseline scheme, *Multi-Instance Residual Evidential Learning* (MIREL). Particularly, at the fine-grained instance UE with only weak supervision, we derive a multi-instance residual operator through the Fundamental Theorem of Symmetric Functions. On this operator derivation, we further propose MIREL to jointly model the high-order predictive distribution at bag and instance levels for MIUE. Extensive experiments empirically demonstrate that our MIREL not only could often make existing MIL networks perform better in MIUE, but also could surpass representative UE methods by large margins, especially in instance-level UE tasks. Our source code is available at https://github.com/liupei101/MIREL. Pei Liu 0008, Luping Ji |
ICML | 2 |
| 2024 | Spatio-temporal fusion with motion masks for the moving small target detection from remote-sensing videos
Sicheng Zhu, Luping Ji, Jiewen Zhu, Shengjia Chen, Haohao Ren |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Multi-scale LBP fusion with the contours from deep CellNNs for texture classification
Mingzhe Chang, Luping Ji, Jiewen Zhu |
Expert Syst. Appl. | 2 |
| 2024 | TMP: Temporal Motion Perception with spatial auxiliary enhancement for moving Infrared dim-small target detection
Sicheng Zhu, Luping Ji, Jiewen Zhu, Shengjia Chen |
Expert Syst. Appl. | 2 |
| 2024 | AdvMIL: Adversarial multiple instance learning for the survival analysis on whole-slide images
Pei Liu 0008, Luping Ji, Bo Fu 0007 |
Medical Image Anal. | 2 |
| 2024 | Toward Dense Moving Infrared Small Target Detection: New Datasets and BaselineabstractAs an important research branch of infrared small target detection, dense target detection (e.g., drone swarm detection) has always been a topic worth exploring. Currently, existing datasets cover only one or several (sparse) targets, with almost no dataset available for the research on dense small target detection. To advance this kind of search, for the first time, we synthesize two special dense moving target datasets (DMIST-60 and DMIST-100) on DAUB data. They both contain far more than 50 infrared small targets per frame. In the meantime, for evaluating our new datasets and flourishing detection methodology research, we propose a linking-aware sliced network (LASNet) as the baseline of our datasets. It mainly consists of visual feature extraction, motion feature extraction and motion-affinity fusion. The comprehensive experiments on our synthesized datasets confirm: i) both datasets are practical and effective for dense moving infrared small target detection and ii) proposed LASNet could always obviously outperform other compared methods in both sparse and dense target scenarios. Our new datasets and source codes are currently available athttps://github.com/UESTC-nnLab/DMIST. Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001, Haohao Ren, Yongsheng Sang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | SSTNet: Sliced Spatio-Temporal Network With Cross-Slice ConvLSTM for Moving Infrared Dim-Small Target DetectionabstractInfrared dim-small target detection, as an important branch of object detection, has been attracting research attention in recent decades. Its challenges mainly lie in the small target sizes and dim contrast to background images. Recent research schemes on it mainly focus on improving the feature representation of spatio-temporal domains only in single-slice temporal scope. More cross-slice motion, i.e., past and future, is seldom considered to enhance target features. To use cross-slice motion context, this article proposes a sliced spatio-temporal network (SSTNet) with cross-slice enhancement for moving infrared dim-small target detection. In our scheme, a new cross-slice ConvLSTM node is designed to capture spatio-temporal motion features from both inner slice and inter-slices. Moreover, to improve infrared small target motion feature learning, we extend conventional loss function by adopting a new motion-coordination loss (MCL) term. On these, we propose a motion-coupling neck to assist feature extractor in facilitating the capturing and utilization of motion features from multiframes. To our best knowledge, our work is the first one to explore the cross-slice spatio-temporal motion modeling for infrared dim-small targets. Experiments verify that our SSTNet could refresh most state-of-the-art metrics on two public benchmarks (DAUB and IRDST). Our source codes are available athttps://github.com/UESTC-nnLab/SSTNet. Shengjia Chen, Luping Ji, Jiewen Zhu, Mao Ye 0001, Xiaoyong Yao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Triple-Domain Feature Learning With Frequency-Aware Memory Enhancement for Moving Infrared Small Target DetectionabstractAs a subfield of object detection, moving infrared small target detection (ISTD) presents significant challenges due to tiny target sizes and low contrast against backgrounds. Currently existing methods primarily rely on the features extracted only from spatiotemporal domain. Frequency domain has hardly been concerned yet, although it has been widely applied in image processing. To extend feature source domains and enhance feature representation, we propose a new triple-domain strategy (Tridos) with the frequency-aware memory enhancement on spatiotemporal domain for ISTD. In this scheme, it effectively detaches and enhances frequency features by a local-global frequency-aware module (LGFM) with Fourier transform (FT). Inspired by human visual system (HVS), our memory enhancement is designed to capture the spatial relationships of infrared targets among video frames. Furthermore, it encodes temporal dynamics motion features via differential learning and residual enhancing. In addition, we further design a residual compensation to reconcile possible cross-domain feature mismatches. To our best knowledge, proposed Tridos is the first work to explore infrared target feature learning comprehensively in spatiotemporal-frequency domains. The extensive experiments on three datasets (i.e., DAUB, ITSDT-15K, and IRDST) validate that our triple-domain infrared feature learning scheme could often be obviously superior to state-of-the-art (SOTA) ones. Source codes are available athttps://github.com/UESTC-nnLab/Tridos. Luping Ji, Shengjia Chen, Sicheng Zhu, Mao Ye 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Illumination Distribution-Aware Thermal Pedestrian DetectionabstractPedestrian detection is an important task in computer vision, which is also an important part of intelligent transportation systems. For privacy protection, thermal images are widely used in pedestrian detection problems. However, thermal pedestrian detection is challenging due to the significant effect of temperature variation on the illumination of images and that fine-grained illumination annotations are difficult to be acquired. The existing methods have attempted to exploit coarse-grained day/night labels, which however even hampers the model performance. In this work, we introduce a novel idea of regressing conditional thermal-visible feature distribution, dubbed as Illumination Distribution-Aware adaptation (IDA). The key idea is to predict the conditional visible feature distribution given a thermal image, subject to their pre-computed joint distribution. Specifically, we first estimate the thermal-visible feature joint distribution by constructing feature co-occurrence matrices, offering a conditional probability distribution for any given thermal image. With this pairing information, we then form a conditional probability distribution regression task for model optimization. Critically, as a model agnostic strategy, this allows the visible feature knowledge to be transferred to the thermal counterpart implicitly for learning more discriminating feature representation. Experiment results show that our method outperforms the prior art methods, which use extra illumination annotations. Besides, as a plug-in, our method can averagely reduce about 2% MR on KAIST dataset, and improve about 1% mAP on FLIR-aligned and Autonomous Vehicles datasets without extra calculation for test. Code is available athttps://github.com/HaMeow-lst1/IDA. Mao Ye 0001, Luping Ji, Song Tang 0001, Yan Gan, Xiatian Zhu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Pseudo-Bag Mixup Augmentation for Multiple Instance Learning-Based Whole Slide Image ClassificationabstractGiven the special situation of modeling gigapixel images, multiple instance learning (MIL) has become one of the most important frameworks for Whole Slide Image (WSI) classification. In current practice, most MIL networks often face two unavoidable problems in training: i) insufficient WSI data and ii) the sample memorization inclination inherent in neural networks. These problems may hinder MIL models from adequate and efficient training, suppressing the continuous performance promotion of classification models on WSIs. Inspired by the basic idea of Mixup, this paper proposes a new Pseudo-bag Mixup (PseMix) data augmentation scheme to improve the training of MIL models. This scheme generalizes the Mixup strategy for general images to special WSIs via pseudo-bags so as to be applied in MIL-based WSI classification. Cooperated by pseudo-bags, our PseMix fulfills the critical size alignment and semantic alignment in Mixup strategy. Moreover, it is designed as an efficient and decoupled method, neither involving time-consuming operations nor relying on MIL model predictions. Comparative experiments and ablation studies are specially designed to evaluate the effectiveness and advantages of our PseMix. Experimental results show that PseMix could often assist state-of-the-art MIL networks to refresh their classification performance on WSIs. Besides, it could also boost the generalization performance of MIL models in special test scenarios, and promote their robustness to patch occlusion and label noise. Our source code is available at https://github.com/liupei101/PseMix. Pei Liu 0008, Luping Ji |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Shared Coupling-Bridge Scheme for Weakly Supervised Local Feature LearningabstractLocal feature learning is believed to be of important significance in classic vision tasks such as visual localization, image matching and 3D reconstruction. Limited by training samples, weakly-supervised strategy has become one of widely-concerned effective schemes for local feature learning. Currently, it still has some weaknesses needing further improvement, mainly including the discrimination power of extracted local descriptors, the localization accuracy of detected keypoints, and the efficiency of weakly-supervised local feature learning. Focusing on promoting the performance of sparse local feature learning with camera pose supervision, this article pertinently proposes a Shared Coupling-bridge scheme with four light-weight yet effective improvements for weakly-supervised local feature (SCFeat) learning. It mainly contains: i) theFeature-Fusion-ResUNet Backbone(F2R-Backbone) for local descriptors learning, ii) a shared coupling-bridge normalization to improve the decoupling training of description network and detection network, iii) an improved detection network with peakiness measurement to detect keypoints and iv) a new reward factor of fundamental matrix error to further optimize feature detection training. Extensive experiments prove that our SCFeat scheme is effective and has wide task adaptability. It could often obtain a state-of-the-art performance on classic image matching and visual localization. Even in terms of 3D reconstruction, it could still achieve competitive results. Jiayuan Sun, Luping Ji, Jiewen Zhu |
IEEE Trans. Multim. | 2 |
| 2023 | AugTarget Data Augmentation for Infrared Small Target DetectionabstractSample shortage has always been a frequently-faced problem for the machine-learning models in infrared small target detection. As one of main limitations, it is hampering the further promotion of target detection performance. In this paper, we propose a simple and effective data augmentation scheme, AugTarget, to address this shortage issue of small target samples. Our scheme mainly consists of two crucial algorithms: target augmentation and batch augmentation. The former is designed to generate sufficient targets, by random target representation. The latter is devised to diversify training samples. Moreover, the initially-generated image samples of small targets are further enriched by randomly aggregating the feature representation of different images. The experiments on public datasets demonstrate that our AugTarget could bring an obvious improvement to mean intersection over union (mIoU). Cooperated by the augmentation of our AugTarget, the state-of-the-art (SOTA) method, AGPC could even achieve a distinct performance promotion by 3.03%, 2.17% and 2.63% on MDFA, SIRST-Aug and Merged datasets, respectively. In addition, the experimental results on three baseline models also show the universality & adaptivity of AugTarget to different dataset augmentation. Our source codes are available at https://github.com/UESTC-nnLab/AugTarget. Shengjia Chen, Jiewen Zhu, Luping Ji, Hongjun Pan |
ICASSP | 3 |
| 2023 | Sanet: Spatial Attention Network with Global Average Contrast Learning for Infrared Small Target DetectionabstractInfrared small target detection has always been a challenging theme, due to small target size, unconspicuous contour and texture, even low vision contrast to background. Because of these causes, some popular object detection methods, such as Faster-RCNN and YOLOV, could often lose effectiveness. Aiming to promote the comprehensive performance of detection, this paper proposes a Spatial Attention Network (SANet) with global average contrast learning specially for infrared small target. Different from the other detection strategies by pixel-level segmentation, our scheme extends traditional contrast methods to target detection framework of deep learning, so as to achieve robust performance. In feature extraction, a group of cross stage partial networks (CSPNet) is designed to capture the local semantic information, and a cluster of spatial attention modules with global average contrast (SAG) is devised to obtain global spatial semantics. Moreover, a series of selective kernel convolution (SKConv) is adopted to effectively fuse semantic and spatial features. For robust feature representation, an Spatial Pyramid Pooling (SPP) scheme is utilized in our detection model. The experiments on two public datasets show that our detection model could often obviously outperform current state-of-the-art ones. The source code is available at https://github.com/UESTC-nnLab/SANet. Jiewen Zhu, Shengjia Chen, Lexiao Li, Luping Ji |
ICASSP | 4 |
| 2023 | DSCA: A dual-stream network with cross-attention on whole-slide image pyramids for cancer prognosis
Pei Liu 0008, Bo Fu 0007, Rui Yang 0033, Luping Ji |
Expert Syst. Appl. | 5 |
| 2022 | Visual Sound Source Separation with Partial Supervision LearningabstractRecent deep learning approaches have achieved impressive performance in visually-guided sound source separation tasks. However, due to the lack of real-world mixed/separated audio sample pairs, most methods seriously rely on the "Mix-and-Separate" manner to learn sound source separation, often unsuitable for real-world mixtures. To address this issue, we utilize a semi-supervised learning technique — preserving audio-visual consistency — to improve the separation performance of real-world scenarios. In this way, our network is trained jointly by artificial and real-world mixtures. To the best of our knowledge, this could be the first attempt to improve real-world generalization. We also design a category-guided audiovisual fusion module to learn audio-visual matching. Comparative experiments are performed on two publicly-available datasets, MUSIC and AudioSet. Experiment results demonstrate that our method could often outperform other state-of-the-art ones in visual sound separation. Huasen Wang, Lingling Gao, Qianchao Tan, Luping Ji |
ICIP | 4 |
| 2022 | Geometry Attention Transformer with position-aware LSTMs for image captioning
Yulin Shen 0001, Luping Ji |
Expert Syst. Appl. | 3 |
| 2022 | Prototype-Based Multisource Domain AdaptationabstractUnsupervised domain adaptation aims to transfer knowledge from labeled source domain to unlabeled target domain. Recently, multisource domain adaptation (MDA) has begun to attract attention. Its performance should go beyond simply mixing all source domains together for knowledge transfer. In this article, we propose a novel prototype-based method for MDA. Specifically, for solving the problem that the target domain has no label, we use the prototype to transfer the semantic category information from source domains to target domain. First, a feature extraction network is applied to both source and target domains to obtain the extracted features from which the domain-invariant features and domain-specific features will be disentangled. Then, based on these two kinds of features, the named inherent class prototypes and domain prototypes are estimated, respectively. Then a prototype mapping to the extracted feature space is learned in the feature reconstruction process. Thus, the class prototypes for all source and target domains can be constructed in the extracted feature space based on the previous domain prototypes and inherent class prototypes. By forcing the extracted features are close to the corresponding class prototypes for all domains, the feature extraction network is progressively adjusted. In the end, the inherent class prototypes are used as a classifier in the target domain. Our contribution is that through the inherent class prototypes and domain prototypes, the semantic category information from source domains is transformed into the target domain by constructing the corresponding class prototypes. In our method, all source and target domains are aligned twice at the feature level for better domain-invariant features and more closer features to the class prototypes, respectively. Several experiments on public data sets also prove the effectiveness of our method. Lihua Zhou, Mao Ye 0001, Ce Zhu, Luping Ji |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Multi-view subspace clustering via partition fusion
Juncheng Lv, Zhao Kang 0001, Boyu Wang 0004, Luping Ji, Zenglin Xu |
Inf. Sci. | 4 |
| 2020 | Recurrent convolutions of binary-constraint Cellular Neural Network for texture recognition
Luping Ji, Mingzhe Chang, Yulin Shen 0001 |
Neurocomputing | 1 |
| 2018 | Median local ternary patterns optimized with rotation-invariant uniform-three mapping for noisy texture classification
Luping Ji, Xiaorong Pu, Guisong Liu |
Pattern Recognit. | 1 |
| 2018 | Training-Based Gradient LBP Feature Models for Multiresolution Texture ClassificationabstractLocal binary pattern (LBP) is a simple, yet efficient coding model for extracting texture features. To improve texture classification, this paper designs a median sampling regulation, defines a group of gradient LBP (gLBP) descriptors, proposes a training-based feature model mapping method, and then develops a texture classification frame using the multiresolution feature fusion of four gLBP descriptors. Cooperated by median sampling, four descriptors encode a pixel respectively by central gradient, radial gradient, magnitude gradient and tangent gradient to generate initial gLBP patterns. The feature mapping models of gLBP descriptors are constructed by the maximal relative-variation rate (mr2) of rotation-invariant patterns, and then prestored as mapping lookup files. By mapping, initial patterns can be transformed into low-dimensional ones. And then it generates multiresolution texture features via the joint and concatenation of gLBP descriptors on different sampling parameters. A trained nearest neighbor classifier with chi-square distance is applied to classify textures by feature histograms. The experimental results of simulation on five public texture databases show that the proposed method is reliable and efficient in texture classification. In comparison with nine other similar approaches, including two state-of-the-art ones, the proposed method runs faster than most of them and also outperforms all of them in terms of classification accuracy and noise robustness. It achieves higher accuracy and has also better robustness to the Salt&Pepper and Gaussian noise added artificially into texture images. Luping Ji, Guisong Liu, Xiaorong Pu |
IEEE Trans. Cybern. | 1 |
| 2015 | One-dimensional pairwise CNN for the global alignment of two DNA sequences
Luping Ji, Xiaorong Pu, Hong Qu 0002, Guisong Liu |
Neurocomputing | 1 |
| 2015 | Computing k shortest paths using modified pulse-coupled neural network
Guisong Liu, Hong Qu 0002, Luping Ji |
Neurocomputing | 4 |
| 2015 | Facial expression recognition from image sequences using twofold random forest classifier
Xiaorong Pu, Luping Ji, Zhihu Zhou |
Neurocomputing | 4 |
| 2015 | Computing k shortest paths from a source node to each other node
Guisong Liu, Hong Qu 0002, Luping Ji, Alexander Takacs |
Soft Comput. | 4 |
| 2012 | Parameter selection of support vector machines and genetic algorithm based on change area search
Luping Ji, Jianping Li 0002, Mingtian Zhou |
Neural Comput. Appl. | 3 |
| 2011 | Feature selection and parameter optimization for support vector machines: A new approach based on genetic algorithm with feature chromosomes
Luping Ji, Mingtian Zhou |
Expert Syst. Appl. | 3 |
| 2009 | Constrained ZIP code segmentation by a PCNN-based thinning algorithm
Lifeng Shang, Zhang Yi 0001, Luping Ji |
Neurocomputing | 3 |
| 2008 | A mixed noise image filtering method using weighted-linking PCNNs
Luping Ji, Zhang Yi 0001 |
Neurocomputing | 1 |
| 2008 | An improved pulse coupled neural network for image processing
Luping Ji, Zhang Yi 0001, Lifeng Shang |
Neural Comput. Appl. | 1 |
| 2008 | Fingerprint orientation field estimation using ridge projection
Luping Ji, Zhang Yi 0001 |
Pattern Recognit. | 1 |
| 2007 | Binary Image Thinning Using Autowaves Generated by PCNN
Lifeng Shang, Zhang Yi 0001, Luping Ji |
Neural Process. Lett. | 3 |
| 2007 | Binary Fingerprint Image Thinning Using Template-Based PCNNsabstractThis correspondence presents a coarse-to-fine binary-image-thinning algorithm by proposing a template-based pulse-coupled neural-network model. Under the control of coupled templates, this algorithm iteratively skeletonizes a binary image by changing the load signals of pulse neurons. A direction-constraining scheme for avoiding fingerprint ridge spikes has been discussed. Experiments show that this algorithm is effective for fingerprint thinning, as well as other common images. Moreover, this algorithm can be coupled with a fingerprint identification system to improve the recognition performance. Luping Ji, Zhang Yi 0001, Lifeng Shang, Xiaorong Pu |
IEEE Trans. Syst. Man Cybern. Part B | 1 |