EDBT 2026 Demo / reviewers in the wild / expert
Shengjia Chen
dblp:276/5017
· DBLP profile ↗
23ranked-venue papers
12as first author
19since 2021 · last 2026
0000-0002-8083-7905ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Domain-Auxiliary Infrared Moving Small Target Detection by Learning to Overlook Domain DiscrepancyabstractCurrently, almost all traditional infrared small target detection methods work on the assumption that training and test sets always belong to the same domain, and training samples are sufficient. However, in real applications, a new detection task could often have no sufficient training samples from a special domain. In this situation, adopting the auxiliary data from big-sample domains is usually believed to be one of the most potential solutions. However, exceeding expectations, it is found that simply adding auxiliary samples cannot often be always effective, even causing performance decline, due to existing infrared domain shift. To overcome this unexpected problem, we propose the first infrared moving small target detection framework with domain-auxiliary supports by Learning to Overlook Domain Discrepancy (Loddis). This framework consists of three primary processing stages: correlation weakening, domain confusing, and target consistency contrastive learning. Breaking through traditional learning paradigm, through auxiliary data, it enables the model to focus more on targets themselves, and less on image backgrounds, minimizing the sensitivity to domain discrepancy. The extensive experiments on 6 different-domain datasets show the effectiveness and superiority of the proposed Loddis framework for infrared small target detection. Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001 |
AAAI | 1 |
| 2026 | Local-global collaborative feature learning with level-wise decoding for infrared small target detection
Luping Ji, Shengjia Chen, Jianghong Huang |
Comput. Vis. Image Underst. | 3 |
| 2025 | Motion Prior Knowledge Learning with Homogeneous Language Descriptions for Moving Infrared Small Target DetectionabstractDifferent from traditional object detection, pure vision is not enough to infrared small target detection, due to small target size and weak background contrast. For promoting detection performance, more target representations are needed. Currently, motion representations have been proved to be one of the most potential feature kinds for infrared small target detection. Existing methods have an obvious weakness, that besides vision features, they could only capture coarse motion representations from temporal domain. With vision features, fine motion representations could be more effective to enhance detection performance. To overcome this weakness, inspired by prevalent vision-language models, we propose the first vision-language framework with motion prior knowledge learning (MoPKL). Breaking through traditional pure-vision modality, it utilizes homogeneous language descriptions, formatted for moving targets, to directionally guide vision channel learning motion prior knowledge. With the facilitation of motion-vision alignment and motion-relation mining, the motion of infrared small targets is further refined by graph attention, to generate more fine motion representations. The extensive experiments on datasets ITSDT-15K and IRDST show that our framework is effective. It could often obviously outperform other methods. Shengjia Chen, Luping Ji, Mao Ye 0001 |
AAAI | 1 |
| 2025 | Moving infrared dim and small target detection by mixed spatio-temporal encoding
Luping Ji, Shengjia Chen, Sicheng Zhu |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Spatial-temporal-channel collaborative feature learning with transformers for infrared small target detection
Sicheng Zhu, Luping Ji, Shengjia Chen |
Image Vis. Comput. | 3 |
| 2025 | Language-Driven Motion Prior Knowledge Learning for Moving Infrared Small Target DetectionabstractDifferent from traditional object detection, pure vision is often not enough to infrared small target detection (ISTD), due to the small target size and weak background contrast. For promoting performance, more target representations are often needed. Currently, motion representations have been proved to be one of the most potential feature patterns for infrared small targets. Besides vision features, existing methods have an obvious weakness that they could only capture coarse motion representations from the temporal domain. By vision features, fine motion representations could often be more effective to enhance detection performance. To overcome this weakness, and inspired by prevalent vision-language models (VLMs), the first vision-language framework with motion prior knowledge learning (MoPKL) was proposed in our previous work. To further extend it, we repropose an improved version, i.e., iMoPKL. Breaking through traditional pure-vision modality, it utilizes the homogeneous language descriptions, specially formatted for moving targets, to directionally guide vision channels to learn the motion prior knowledge of targets. In detail, it learns the distribution of target motion reconstruction corresponding to the language description as a type of prior knowledge. With the facilitation of language-driven motion alignment, the motion of infrared small targets could be further refined by motion-relation learning, to generate more fine motion representations. The extensive experiments on ITSDT-15K, DAUB-R, and IRDST-H show that our improvement version is effective. It could often obviously outperform the other methods, including our original MoPKL. Our source codes are available athttps://github.com/UESTC-nnLab/MoPKL Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001, Yongsheng Sang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | Weakly Supervised Contrastive Learning With Quantity Prompts for Moving Infrared Small Target Detection
Luping Ji, Shengjia Chen, Sicheng Zhu, Jianghong Huang, Mao Ye 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Semi-Supervised Multiview Prototype Learning With Motion Reconstruction for Moving Infrared Small Target DetectionabstractMoving infrared small target detection is critical for various applications, e.g., remote sensing and military. Due to tiny target size and limited labeled data, accurately detecting targets is highly challenging. Currently, existing methods primarily focus on fully-supervised learning, which relies heavily on numerous annotated frames for training. However, annotating a large number of frames for each video is often expensive, time-consuming, and redundant, especially for low-quality infrared images. To break through traditional fully-supervised framework, we propose a new semi-supervised multi-view prototype (S2MVP) learning scheme that incorporates motion reconstruction. In our scheme, we design a bi-temporal motion perceptor based on bidirectional ConvGRU cells to effectively model the motion paradigms of targets by perceiving both forward and backward. Additionally, to explore the potential of unlabeled data, it generates the multi-view feature prototypes of targets as soft labels to guide feature learning by calculating cosine similarity. Imitating human visual system, it retains only the feature prototypes of recent frames. Moreover, it eliminates noisy pseudo-labels to enhance the quality of pseudo-labels through anomaly-driven pseudo-label filtering. Furthermore, we develop a target-aware motion reconstruction loss to provide additional supervision and prevent the loss of target details. To our best knowledge, the proposed S2MVP is the first work to utilize large-scale unlabeled video frames to detect moving infrared small targets. Although 10% labeled training samples are used, the experiments on three public benchmarks (DAUB, ITSDT-15K and IRDST) verify the superiority of our scheme compared to other methods. Source codes are available at https://github.com/UESTC-nnLab/S2MVP. Luping Ji, Jianghong Huang, Shengjia Chen, Sicheng Zhu, Mao Ye 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | MICPL: Motion-Inspired Cross-Pattern Learning for Small-Object Detection in Satellite VideosabstractFor small-object detection, vision patterns can only provide limited support to feature learning. Most prior schemes mainly depend on a single vision pattern to learn object features, seldom considering more latent motion patterns. In the real world, humans often efficiently perceive small objects through multipattern signals. Inspired by this observation, this article attempts to address small-object detection from a new prospective of latent pattern learning. To fulfill this purpose, it regards a real-world moving object as the spatiotemporal sequences of a static object to capture latent motion patterns. In view of this, we propose a motion-inspired cross-pattern learning (MICPL) scheme to capture the motion patterns for moving small-object scenarios. This scheme mainly consists of two crucial parts: motion pattern mining (MPM) and motion-vision adaption. The former is designed to effectively mine the motion pattern from time-dependent representation space. The latter is devised to correlate between motion patterns and vision semantics. In the meanwhile, we explore their cross-pattern interactions to guide MICPL to capture motion patterns effectively. Comparison experiments verify that, cooperated by motion pattern, even a simple detector could often refresh state-of-the-art (SOTA) results on moving small-object detection. Moreover, the experiments on two small-object-related tasks further prove the adaptivity and advantages of our cross-pattern feature learning scheme. Our source codes are available at https://github.com/UESTC-nnLab/MICPL. Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Spatio-temporal fusion with motion masks for the moving small target detection from remote-sensing videos
Sicheng Zhu, Luping Ji, Jiewen Zhu, Shengjia Chen, Haohao Ren |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | TMP: Temporal Motion Perception with spatial auxiliary enhancement for moving Infrared dim-small target detection
Sicheng Zhu, Luping Ji, Jiewen Zhu, Shengjia Chen |
Expert Syst. Appl. | 4 |
| 2024 | Toward Dense Moving Infrared Small Target Detection: New Datasets and BaselineabstractAs an important research branch of infrared small target detection, dense target detection (e.g., drone swarm detection) has always been a topic worth exploring. Currently, existing datasets cover only one or several (sparse) targets, with almost no dataset available for the research on dense small target detection. To advance this kind of search, for the first time, we synthesize two special dense moving target datasets (DMIST-60 and DMIST-100) on DAUB data. They both contain far more than 50 infrared small targets per frame. In the meantime, for evaluating our new datasets and flourishing detection methodology research, we propose a linking-aware sliced network (LASNet) as the baseline of our datasets. It mainly consists of visual feature extraction, motion feature extraction and motion-affinity fusion. The comprehensive experiments on our synthesized datasets confirm: i) both datasets are practical and effective for dense moving infrared small target detection and ii) proposed LASNet could always obviously outperform other compared methods in both sparse and dense target scenarios. Our new datasets and source codes are currently available athttps://github.com/UESTC-nnLab/DMIST. Shengjia Chen, Luping Ji, Sicheng Zhu, Mao Ye 0001, Haohao Ren, Yongsheng Sang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | SSTNet: Sliced Spatio-Temporal Network With Cross-Slice ConvLSTM for Moving Infrared Dim-Small Target DetectionabstractInfrared dim-small target detection, as an important branch of object detection, has been attracting research attention in recent decades. Its challenges mainly lie in the small target sizes and dim contrast to background images. Recent research schemes on it mainly focus on improving the feature representation of spatio-temporal domains only in single-slice temporal scope. More cross-slice motion, i.e., past and future, is seldom considered to enhance target features. To use cross-slice motion context, this article proposes a sliced spatio-temporal network (SSTNet) with cross-slice enhancement for moving infrared dim-small target detection. In our scheme, a new cross-slice ConvLSTM node is designed to capture spatio-temporal motion features from both inner slice and inter-slices. Moreover, to improve infrared small target motion feature learning, we extend conventional loss function by adopting a new motion-coordination loss (MCL) term. On these, we propose a motion-coupling neck to assist feature extractor in facilitating the capturing and utilization of motion features from multiframes. To our best knowledge, our work is the first one to explore the cross-slice spatio-temporal motion modeling for infrared dim-small targets. Experiments verify that our SSTNet could refresh most state-of-the-art metrics on two public benchmarks (DAUB and IRDST). Our source codes are available athttps://github.com/UESTC-nnLab/SSTNet. Shengjia Chen, Luping Ji, Jiewen Zhu, Mao Ye 0001, Xiaoyong Yao |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Triple-Domain Feature Learning With Frequency-Aware Memory Enhancement for Moving Infrared Small Target DetectionabstractAs a subfield of object detection, moving infrared small target detection (ISTD) presents significant challenges due to tiny target sizes and low contrast against backgrounds. Currently existing methods primarily rely on the features extracted only from spatiotemporal domain. Frequency domain has hardly been concerned yet, although it has been widely applied in image processing. To extend feature source domains and enhance feature representation, we propose a new triple-domain strategy (Tridos) with the frequency-aware memory enhancement on spatiotemporal domain for ISTD. In this scheme, it effectively detaches and enhances frequency features by a local-global frequency-aware module (LGFM) with Fourier transform (FT). Inspired by human visual system (HVS), our memory enhancement is designed to capture the spatial relationships of infrared targets among video frames. Furthermore, it encodes temporal dynamics motion features via differential learning and residual enhancing. In addition, we further design a residual compensation to reconcile possible cross-domain feature mismatches. To our best knowledge, proposed Tridos is the first work to explore infrared target feature learning comprehensively in spatiotemporal-frequency domains. The extensive experiments on three datasets (i.e., DAUB, ITSDT-15K, and IRDST) validate that our triple-domain infrared feature learning scheme could often be obviously superior to state-of-the-art (SOTA) ones. Source codes are available athttps://github.com/UESTC-nnLab/Tridos. Luping Ji, Shengjia Chen, Sicheng Zhu, Mao Ye 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | AugTarget Data Augmentation for Infrared Small Target DetectionabstractSample shortage has always been a frequently-faced problem for the machine-learning models in infrared small target detection. As one of main limitations, it is hampering the further promotion of target detection performance. In this paper, we propose a simple and effective data augmentation scheme, AugTarget, to address this shortage issue of small target samples. Our scheme mainly consists of two crucial algorithms: target augmentation and batch augmentation. The former is designed to generate sufficient targets, by random target representation. The latter is devised to diversify training samples. Moreover, the initially-generated image samples of small targets are further enriched by randomly aggregating the feature representation of different images. The experiments on public datasets demonstrate that our AugTarget could bring an obvious improvement to mean intersection over union (mIoU). Cooperated by the augmentation of our AugTarget, the state-of-the-art (SOTA) method, AGPC could even achieve a distinct performance promotion by 3.03%, 2.17% and 2.63% on MDFA, SIRST-Aug and Merged datasets, respectively. In addition, the experimental results on three baseline models also show the universality & adaptivity of AugTarget to different dataset augmentation. Our source codes are available at https://github.com/UESTC-nnLab/AugTarget. Shengjia Chen, Jiewen Zhu, Luping Ji, Hongjun Pan |
ICASSP | 1 |
| 2023 | Sanet: Spatial Attention Network with Global Average Contrast Learning for Infrared Small Target DetectionabstractInfrared small target detection has always been a challenging theme, due to small target size, unconspicuous contour and texture, even low vision contrast to background. Because of these causes, some popular object detection methods, such as Faster-RCNN and YOLOV, could often lose effectiveness. Aiming to promote the comprehensive performance of detection, this paper proposes a Spatial Attention Network (SANet) with global average contrast learning specially for infrared small target. Different from the other detection strategies by pixel-level segmentation, our scheme extends traditional contrast methods to target detection framework of deep learning, so as to achieve robust performance. In feature extraction, a group of cross stage partial networks (CSPNet) is designed to capture the local semantic information, and a cluster of spatial attention modules with global average contrast (SAG) is devised to obtain global spatial semantics. Moreover, a series of selective kernel convolution (SKConv) is adopted to effectively fuse semantic and spatial features. For robust feature representation, an Spatial Pyramid Pooling (SPP) scheme is utilized in our detection model. The experiments on two public datasets show that our detection model could often obviously outperform current state-of-the-art ones. The source code is available at https://github.com/UESTC-nnLab/SANet. Jiewen Zhu, Shengjia Chen, Lexiao Li, Luping Ji |
ICASSP | 2 |
| 2023 | Improving semantic segmentation with knowledge reasoning network
Shengjia Chen, Xiwei Yang, Zhixin Li 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | Knowledge Reasoning for Semantic SegmentationabstractThe convolution operation suffers from a limited receptive field, while global modeling is fundamental to dense prediction tasks, such as semantic segmentation. However, most existing methods treat the recognition of each region separately and overlook crucial global semantic relations between regions in one scene. These methods cannot segment the semantic regions accurately due to the lack of global-level supervision or guidance of external knowledge. To overcome the limitation of the traditional method, we propose a Knowledge Reasoning Net (KRNet) that consists of two crucial modules: (1) a prior knowledge mapping module that incorporates external knowledge by graph convolutional network to guide learning semantic representations and (2) a knowledge reasoning module that correlates these representations with a graph built on the external knowledge and explores their interactions via the knowledge reasoning. Experiments on Cityscapes and ADE datasets demonstrate the effectiveness of our proposed methods on semantic segmentation. Shengjia Chen, Zhixin Li 0001, Xiwei Yang |
ICASSP | 1 |
| 2021 | Joint deep separable convolution network and border regression reinforcement for object detection
Yu Quan, Zhixin Li 0001, Shengjia Chen, Canlong Zhang, Huifang Ma |
Neural Comput. Appl. | 3 |
| 2020 | Image Captioning with Internal and External KnowledgeabstractAutomatically generating a human-like description for a given image is a potential research in artificial intelligence, which has attracted a great of attention recently. Most of the existing attention methods explore the mapping relationships between words in sentence and regions in image, such unpredictable matching manner sometimes causes inharmonious alignments that may reduce the quality of generated captions. In this paper, we make our efforts to reason about more accurate and meaningful captions. We first propose word attention to improve the correctness of visual attention when generating sequential descriptions word-by-word. The special word attention emphasizes on word importance when focusing on different regions of the input image, and makes full use of the internal annotation knowledge to assist the calculation of visual attention. Then, in order to reveal those incomprehensible intentions that cannot be expressed straightforwardly by machines, we inject external knowledge extracted from knowledge graph into the encoder-decoder framework to facilitate meaningful captioning. We validate our model on two freely available captioning benchmarks: Microsoft COCO dataset and Flickr30k dataset. The results demonstrate that our approach achieves state-of-the-art performance and outperforms many of the existing approaches. Feicheng Huang, Zhixin Li 0001, Shengjia Chen, Canlong Zhang, Huifang Ma |
CIKM | 3 |
| 2020 | Improving Object Detection with Relation Mining NetworkabstractDue to the deteriorated quality of feature in the propagation process of the neural network, it may be hard for traditional detector to identify a small object by just utilizing information within one region proposal. To overcome the limitation of the traditional object detector, we proposed a graph based relation mining network, to capture the relation information from labels and images. The semantic relation network is proposed to mine the global semantic relation in labels, and the spatial relation network is proposed to capture the local spatial relation in images. The feature representation is further improved by aggregating the outputs of the two networks. Instead of directly disseminating visual features in the network, the relation mining network explores more advanced feature information. Experiments on the PASCAL VOC and MS COCO datasets demonstrate that key relation information significantly improve the performance of object detection with better ability to detect small objects and reasonable bounding box. The results on COCO dataset demonstrate our method can detect objects robustly, increasing the detection performance of small objects from average precision and average recall by 4.7% and 7.6% respectively in performance relative to Faster R-CNN. Shengjia Chen, Zhixin Li 0001, Feicheng Huang, Canlong Zhang, Huifang Ma |
ICDM | 1 |
| 2020 | Object Detection Using Dual Graph NetworkabstractMost object detection methods focus only on the local information near the region proposal and ignore the object's global semantic relation and local spatial relation information, resulting in limited performance. To capture and explore these important relations, we propose a detection method based on a graph convolutional network (GCN). Two independent relation graph networks are used to obtain the global semantic information of the object in labels and the local spatial information in images. Semantic relation networks can implicitly acquire global knowledge, and by constructing a directed graph on the dataset, each node is represented by the word embedding of labels and then sent to the GCN to obtain high-level semantic representation. The spatial relation network encodes the relation by the positional relation module and the visual connection module, and enriches the object features through local key information from objects. The feature representation is further improved by aggregating the outputs of the two networks. Instead of directly disseminating visual features in the network, the dual-graph network explores more advanced feature information, giving the detector the ability to obtain key relations in labels and region proposals. Experiments on the PASCAL VOC and MS COCO datasets demonstrate that key relation information significantly improve the performance of detection with better ability to detect small objects and reasonable boduning box. The results on COCO dataset demonstrate our method obtains around 32.3% improvement on AP in terms of small objects. Shengjia Chen, Zhixin Li 0001, Feicheng Huang, Canlong Zhang, Huifang Ma |
ICPR | 1 |
| 2020 | Relation R-CNN: A Graph Based Relation-Aware Network for Object DetectionabstractDue to the deteriorated quality of feature in the propagation process of the neural network, it may be hard for traditional detector to identify a small object by just utilizing information within one region proposal. To overcome the limitation of the traditional object detector, we proposed a graph based relation-aware network, to capture the relation information from labels, and images. The semantic relation network is proposed to mine the global semantic relation in labels, and the spatial relation network is proposed to capture the local spatial relation in images. The feature representation is further improved by aggregating the outputs of the two networks. Instead of directly disseminating visual features in the network, the relation-aware network explores more advanced feature information. Experiments on the PASCAL VOC, and MS COCO datasets demonstrate that key relation information significantly improve the performance of object detection with better ability to detect small objects, and reasonable bounding box. The results on COCO dataset demonstrate our method can detect objects robustly, increasing the detection performance of small objects from average precision, and average recall by 31.8%, and 32.3% respectively in performance relative to Faster R-CNN. Shengjia Chen, Zhixin Li 0001, Zhenjun Tang |
IEEE Signal Process. Lett. | 1 |