VLDB 2026 Research / reviewers in the wild / expert
Xingqun Jiang
dblp:247/1119
· DBLP profile ↗
11ranked-venue papers
0as first author
10since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AMPL: An adaptive meta-prompt learner for few-shot image classification
Zhiping Wu, Lian Huai, Zeyu Shangguan, Lei Wang 0001, Jing Huo, Wenbin Li 0006, Yang Gao 0001, Xingqun Jiang |
Neural Networks | 9 |
| 2026 | Unsupervised few-shot learning with object-aware and attribute-consistent augmentation
Zhiping Wu, Lian Huai, Zeyu Shangguan, Wenbin Li 0006, Yang Gao 0001, Xingqun Jiang |
Pattern Recognit. | 7 |
| 2026 | From Abstract Events to Grounded Cues: Cue-Guided Vision-Language Anomaly DetectionabstractVideo anomaly detection (VAD) must be reliable under large appearance variation and weak supervision, yet provide explanations grounded in human-interpretable evidence. Vision-Language Models (VLMs) often suffer from brittle direct visual-text alignment which is unstable across domains, while deep models are accurate but have limited interpretability. We address this by introducing event-related but more concrete cues as intermediate representations, making the mapping from frames to evidence more stable than directly mapping frames to event labels. Building on this idea, we propose a two-stage cooperative framework: a VLM discovers cues and generates cue-guided pseudo frame-level labels, and a Symbolic Learning Model (SLM) learns from them to produce segment-level cue presence estimates and anomaly scores via cross-modal matching. In inference, cue estimates, SLM scores and video segments are combined by VLM to output final decisions with cue-based rationales. Experiments show competitive performance. Code is available athttps://github.com/AllenYLJiang/Atoms-to-Events-Categorical-Evidence-Composition-for-Video-Anomaly-Detection. Yalong Jiang, Lian Huai, Yuyu Liu, Xingqun Jiang |
IEEE Signal Process. Lett. | 4 |
| 2025 | WeakMedSAM: Weakly-Supervised Medical Image Segmentation via SAM With Sub-Class Exploration and Prompt Affinity MiningabstractWe have witnessed remarkable progress in foundation models in vision tasks. Currently, several recent works have utilized the segmenting anything model (SAM) to boost the segmentation performance in medical images, where most of them focus on training an adaptor for fine-tuning a large amount of pixel-wise annotated medical images following a fully supervised manner. In this paper, to reduce the labeling cost, we investigate a novel weakly-supervised SAM-based segmentation model, namely WeakMedSAM. Specifically, our proposed WeakMedSAM contains two modules: 1) to mitigate severe co-occurrence in medical images, a sub-class exploration module is introduced to learn accurate feature representations. 2) to improve the quality of the class activation maps, our prompt affinity mining module utilizes the prompt capability of SAM to obtain an affinity map for random-walk refinement. Our method can be applied to any SAM-like backbone, and we conduct experiments with SAMUS and EfficientSAM. The experimental results on three popularly-used benchmark datasets, i.e., BraTS 2019, AbdomenCT-1K, and MSD Cardiac dataset, show the promising results of our proposed WeakMedSAM. Our code is available at https://github.com/wanghr64/WeakMedSAM. Lian Huai, Wenbin Li 0006, Lei Qi 0001, Xingqun Jiang, Yinghuan Shi |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Decoupled DETR for Few-Shot Object Detection
Zeyu Shangguan, Lian Huai, Yuyu Liu, Xingqun Jiang |
ACCV (8) | 5 |
| 2024 | Test-Time Linear Out-of-Distribution DetectionabstractOut-of-Distribution (OOD) detection aims to address the excessive confidence prediction by neural networks by triggering an alert when the input sample deviates significantly from the training distribution (in-distribution), indicating that the output may not be reliable. Current OOD detection approaches explore all kinds of cues to identify OOD data, such as finding irregular patterns in the feature space, logit space, gradient space, or the raw image space. Surprisingly, we observe a linear trend between the OOD score produced by current OOD detection algorithms and the network features on several datasets. We conduct a thorough investigation, theoretically and empirically, to analyze and understand the meaning of such a linear trend in OOD detection. This paper proposes a Robust Test-time Linear method (RTL) to utilize such linear trends like a ‘free lunch’ when we have a batch of data to perform OOD detection. By using a simple linear regression as a test time adaptation, we can make a more precise OOD prediction. We further propose an online variant of the proposed method, which achieves promising performance and is more practical for real applications. Theoretical analysis is given to prove the effectiveness of our methods. Extensive experiments on several OOD datasets show the efficacy of RTL for OOD detection tasks, significantly improving the results of base OOD detectors. Project will be available at https://github.com/kfan21/RTL. Xingyu Qiu, Yikai Wang 0002, Lian Huai, Zeyu Shangguan, Shuang Gou, Fengjian Liu, Yuqian Fu, Yanwei Fu 0001, Xingqun Jiang |
CVPR | 11 |
| 2024 | Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector
Yuqian Fu, Yu Wang 0002, Yixuan Pan, Lian Huai, Xingyu Qiu, Zeyu Shangguan, Yanwei Fu 0001, Luc Van Gool, Xingqun Jiang |
ECCV (58) | 10 |
| 2023 | Few-shot Object Detection with Refined Contrastive LearningabstractDue to the scarcity of sampling data in reality, few-shot object detection (FSOD) has drawn more and more attention because of its ability to quickly train new detection concepts with less data. However, there are still failure identifications due to the difficulty in distinguishing confusable classes. We also notice that the high standard deviation of average precision reveals the inconsistent detection performance. To this end, we propose a novel FSOD method with Refined Contrastive Learning (FSRC). A pre-determination component is introduced to find out the Resemblance Group from novel classes which contains confusable classes. Afterwards, Refined Contrastive Learning (RCL) is pointedly performed on this group of classes in order to increase the inter-class distances among them. In the meantime, the detection results distribute more uniformly which further improve the performance. Experimental results based on PASCAL VOC and COCO datasets demonstrate our proposed method outperforms the current state-of-the-art research. Zeyu Shangguan, Lian Huai, Xingqun Jiang |
ICTAI | 4 |
| 2022 | edge-SR: Super-Resolution For The MassesabstractClassic image scaling (e.g. bicubic) can be seen as one convolutional layer and a single upscaling filter. Its implementation is ubiquitous in all display devices and image processing software. In the last decade deep learning systems have been introduced for the task of image super-resolution (SR), using several convolutional layers and numerous filters. These methods have taken over the benchmarks of image quality for upscaling tasks. Would it be possible to replace classic upscalers with deep learning architectures on edge devices such as display panels, tablets, laptop computers, etc.? On one hand, the current trend in Edge–AI chips shows a promising future in this direction, with rapid development of hardware that can run deep–learning tasks efficiently. On the other hand, in image SR only few architectures have pushed the limit to extreme small sizes that can actually run on edge devices at real-time. We explore possible solutions to this problem with the aim to fill the gap between classic upscalers and small deep learning configurations. As a transition from classic to deep–learning upscaling we propose edge–SR (eSR), a set of one–layer architectures that use interpretable mechanisms to upscale images. Certainly, a one–layer architecture cannot reach the quality of deep learning systems. Nevertheless, we find that for high speed requirements, eSR becomes better at trading–off image quality and runtime performance. Filling the gap between classic and deeplearning architectures for image upscaling is critical for massive adoption of this technology. It is equally important to have an interpretable system that can reveal the inner strategies to solve this problem and guide us to future improvements and better understanding of larger networks. Pablo Navarrete Michelini, Yunhua Lu, Xingqun Jiang |
WACV | 3 |
| 2021 | Back-Projection PipelineabstractWe propose a simple extension of residual networks that works simultaneously in multiple resolutions for the problem of image super-resolution. Our network design is inspired by the iterative back-projection algorithm and seeks the more difficult task of learning how to enhance images. Compared to similar approaches, we propose a novel solution to make back-projections run in multiple resolutions by using a data pipeline workflow. Features are updated at multiple scales in each layer of the network. The update dynamic through these layers includes interactions between different resolutions in a way that is causal in scale, and it is represented by a system of ODEs, as opposed to a single ODE in the case of ResNets. Pablo Navarrete Michelini, Yunhua Lu, Xingqun Jiang |
ICIP | 4 |
| 2019 | A Tour of Convolutional Networks Guided by Linear InterpretersabstractConvolutional networks are large linear systems divided into layers and connected by non-linear units. These units are the "articulations" that allow the network to adapt to the input. To understand how a network manages to solve a problem we must look at the articulated decisions in entirety. If we could capture the actions of non-linear units for a particular input, we would be able to replay the whole system back and forth as if it was always linear. It would also reveal the actions of non-linearities because the resulting linear system, a Linear Interpreter, depends on the input image. We introduce a hooking layer, called a LinearScope, which allows us to run the network and the linear interpreter in parallel. Its implementation is simple, flexible and efficient. From here we can make many curious inquiries: how do these linear systems look like? When the rows and columns of the transformation matrix are images, how do they look like? What type of basis do these linear transformations rely on? The answers depend on the problems presented, through which we take a tour to some popular architectures used for classification, super-resolution (SR) and image-to-image translation (I2I). For classification we observe that popular networks use a pixel-wise vote per class strategy and heavily rely on bias parameters. For SR and I2I we find that CNNs use wavelet-type basis similar to the human visual system. For I2I we reveal copy-move and template-creation strategies to generate outputs. Pablo Navarrete Michelini, Yunhua Lu, Xingqun Jiang |
ICCV | 4 |