EDBT 2026 Demo / reviewers in the wild / expert
Xingjia Pan
dblp:210/4520
· DBLP profile ↗
19ranked-venue papers
4as first author
16since 2021 · last 2024
0000-0003-3699-8936ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ${A^{2}Pt}$: Anti-Associative Prompt Tuning for Open Set Visual RecognitionabstractMulti-modality pre-trained models (PTMs) have considerably boosted the performance on a broad range of computer vision topics. Still, they have not been explored purposefully in open set recognition (OSR) scenarios when applying PTMs to downstream recognition tasks. Directly fine/prompt tuning PTMs on closed-set classification tasks will inevitably suffer from data bias and always learn more or less target class-irrelevant cooccurring contextual information, which leads to over-confident predictions on unknown samples. In this paper, we propose a simple yet effective approach, termed Anti-Associative Prompt Tuning(A2Pt), toward learning compact and accurate class-related representation with few class-irrelevant associations from context using multi-modal priors. Specifically, a cross-modal guided activation module is adopted to refine the class-aware representation and suppress the associations from co-occurring contexts by involving text-modal information. We further design an anti-association calibration module to obtain compact class-aware and class-irrelevant representations, respectively, by introducing two additional object functions. Extensive experiments on publicly available benchmarks, including CIFAR series, Tiny-ImageNet, and ImageNet-21K-P, show that the proposed(A2Pt)achieves substantial and consistent performance gains compared with both SOTA OSR and PTM prompt tuning approaches. Hairui Ren, Fan Tang, Xingjia Pan, Juan Cao 0001, Weiming Dong, Zhiwen Lin, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2024 | TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL), which trains object localization models using solely image category annotations, remains a challenging problem. Existing approaches based on convolutional neural networks (CNNs) tend to miss full object extent while activating discriminative object parts. Based on our analysis, this is caused by CNN's intrinsic characteristics, which experiences difficulty to capture object semantics at long distances. In this article, we introduce the vision transformer to WSOL, with the aim to capture long-range semantic dependency of features by leveraging transformer's cascaded self-attention mechanism. We propose the token semantic coupled attention map (TS-CAM) method, which first decomposes class-aware semantics and then couples the semantics with attention maps for semantic-aware activation. To capture object semantics at long distances and avoid partial activation, TS-CAM performs spatial embedding by partitioning an image to a set of patch tokens. To incorporate object category information to patch tokens, TS-CAM reallocates category-related semantics to each patch token. The patch tokens are finally coupled with attention maps which are semantic-agnostic to perform semantic-aware object localization. By introducing semantic tokens to produce semantic-aware attention maps, we further explore the capability of TS-CAM for multicategory object localization. Experiments show that TS-CAM outperforms its CNN-CAM counterpart by 11.6% and 28.9% on ILSVRC and CUB-200-2011 datasets, respectively, improving the state-of-the-art with large margins. TS-CAM also demonstrates superiority for multicategory object localization on the Pascal VOC dataset. The code is available at github.com/yuanyao366/ts-cam-extension. Fang Wan 0001, Wei Gao 0050, Xingjia Pan, Zhiliang Peng, Qi Tian 0001, Qixiang Ye |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Inter-image Contrastive Consistency for Multi-Person Pose EstimationabstractMulti-person pose estimation (MPPE) has achieved impressive progress in recent years. However, due to the large variance of appearances among images or occlusions, the model can hardly learn consistent patterns enough, which leads to severe location jitter and missing issues. In this study, we propose a novel framework, termed Inter-image Contrastive consistency (ICON), to strengthen the keypoint consistency among images for MPPE. Concretely, we consider two-fold consistency constraints, which include single keypoint contrastive consistency (SKCC) and pair relation contrastive consistency (PRCC). The SKCC learns to strengthen the consistency of individual keypoints across images in the same category to improve the category-specific robustness. Only with SKCC, the model can effectively reduce location errors caused by large appearance variations, but remains challenging with extreme postures (e.g., occlusions) due to lack of relational guidance. Therefore, PRCC is proposed to strengthen the consistency of pair-wise joint relation between images to preserve the instructive relation. Cooperating with SKCC, PRCC further improves structure aware robustness in handling extreme postures. Extensive experiments on kinds of architectures across three datasets (i.e., MS-COCO, MPII, CrowdPose) show the proposed ICON achieves substantial improvements over baselines. Furthermore, ICON under the semi-supervised setup can obtain comparable results with the fully-supervised methods using only 30% labeled data. Xixia Xu, Yingguo Gao, Xingjia Pan, Qi Zou 0001 |
AAAI | 3 |
| 2023 | CrossRectify: Leveraging disagreement for semi-supervised object detection
Chengcheng Ma, Xingjia Pan, Qixiang Ye, Fan Tang, Weiming Dong, Changsheng Xu |
Pattern Recognit. | 2 |
| 2023 | Semantic-Context Graph Network for Point-Based 3D Object DetectionabstractPoint-based indoor 3D object detection has received increasing attention with the large demand for augmented reality, autonomous driving, and robot technology in the industry. However, the detection precision suffers from inputs with semantic ambiguity, i.e., shape symmetries, occlusion, and texture missing, which would lead that different objects appearing similar from different viewpoints and then confusing the detection model. Typical point-based detectors relieve this problem via learning proposal representations with both geometric and semantic information, while the entangled representation may cause a reduction in both semantic and spatial discrimination. In this paper, we focus on alleviating the confusion from entanglement and then enhancing the proposal representation by considering the proposal’s semantics and the context in one scene. A semantic-context graph network (SCGNet) is proposed, which mainly includes two modules: a category-aware proposal recoding module (CAPR) and a proposal context aggregation module (PCAg). To produce semantically clear features from entanglement representation, the CAPR module learns a high-level semantic embedding for each category to extract discriminative semantic clues. In view of further enhancing the proposal representation and leveraging the semantic clues, the PCAg module builds a graph to mine the most relevant context in the scene. With few bells and whistles, the SCGNet achieves SOTA performance and obtains consistent gains when applying to different backbones (0.9% ~ 2.4% on ScanNet V2 and 1.6% ~ 2.2% on SUN RGB-D for [email protected]). Code is available at https://github.com/dsw-jlu-rgzn/SCGNet. Shuwei Dong, Xiaoyu Kong, Xingjia Pan, Fan Tang, Weiming Dong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | SPA2Net: Structure-Preserved Attention Activated Network for Weakly Supervised Object LocalizationabstractBy exploring the localizable representations in deep CNN, weakly supervised object localization (WSOL) methods could determine the position of the object in each image just trained by the classification task. However, the partial activation problem caused by the discriminant function makes the network unable to locate objects accurately. To alleviate this problem, we propose Structure-Preserved Attention Activated Network (SPA2Net), a simple and effective one-stage WSOL framework to explore the ability of structure preservation of deep features. Different from traditional WSOL approaches, we decouple the object localization task from the classification branch to reduce their mutual influence by involving a localization branch which is online refined by a self-supervised structural-preserved localization mask. Specifically, we employ the high-order self-correlation as structural prior to enhance the perception of spatial interaction within convolutional features. By succinctly combining the structural prior with spatial attention, activations by SPA2Net will spread from part to the whole object during training. To avoid the structure-missing issue caused by the classification network, we furthermore utilize the restricted activation loss (RAL) to distinguish the difference between foreground and background in the channel dimension. In conjunction with the self-supervised localization branch, SPA2Net can directly predict the class-irrelevant localization map while prompting the network to pay more attention to the target region for accurate localization. Extensive experiments on two publicly available benchmarks, including CUB-200-2011 and ILSVRC, show that our SPA2Net achieves substantial and consistent performance gains compared with baseline approaches. The code and models are available at https://github.com/MsterDC/SPA2Net. Dong Chen 0044, Xingjia Pan, Fan Tang, Weiming Dong, Changsheng Xu |
IEEE Trans. Image Process. | 2 |
| 2023 | SMNet: Synchronous Multi-Scale Low Light Enhancement Network With Local and Global ConcernabstractLimited by objectively poor lighting conditions and hardware devices, low-light images with low visual quality and low visibility are inevitable in the real world. Accurate local details and reasonable global information play their essential and distinct roles in low-light image enhancement: local details contribute to fine textures, while global information is critical for a proper understanding of the global brightness level. In this paper, we focus on integrating local and global aspects to achieve high-quality low-light image enhancement by proposing the synchronous multi-scale low-light enhancement network (SMNet). A synchronous multi-scale representation learning structure and a global feature recalibration module are adopted in SMNet. Different from the traditional multi-scale feature learning architecture, SMNet carries out the multi-scale representation learning in a synchronous way: we first calculate the rough contextual representations in a top-down manner and then learn multi-scale representations in a bottom-up way to generate representations with rich local details. To acquire global brightness information, a global feature recalibration module (GFRM) is applied after the synchronous multi-scale representations to perceive and exploit proper global information by global pooling and projection to recalibrate channel weights globally. The synchronous multi-scale representation and GFRM compose the basic local-and-global block. Experimental results on mainstream real-world dataset LOL and synthetic dataset MIT-Adobe FiveK show that the proposed SMNet not only leads the way on objective metrics (0.41/2.31 improvement of PSNR on two datasets) but is also superior in subjective comparisons compared with typical SoTA methods. The code had already been uploaded tohttps://github.com/linshideng/SMNet. Shideng Lin, Fan Tang, Weiming Dong, Xingjia Pan, Changsheng Xu |
IEEE Trans. Multim. | 4 |
| 2022 | StyTr2: Image Style Transfer with TransformersabstractThe goal of image style transfer is to render an image with artistic features guided by a style reference while maintaining the original content. Owing to the locality in convolutional neural networks (CNNs), extracting and maintaining the global information of input images is difficult. Therefore, traditional neural style transfer methods face biased content representation. To address this critical issue, we take long-range dependencies of input images into account for image style transfer by proposing a transformer-based approach called StyTr2. In contrast with visual transformers for other vision tasks, StyTr2 contains two different transformer encoders to generate domain-specific sequences for content and style, respectively. Following the encoders, a multi-layer transformer decoder is adopted to stylize the content sequence according to the style sequence. We also analyze the deficiency of existing positional encoding methods and propose the content-aware positional encoding (CAPE), which is scale-invariant and more suitable for image style transfer tasks. Qualitative and quantitative experiments demonstrate the effectiveness of the proposed StyTr2 compared with state-of-the-art CNN-based and flow-based approaches. Code and models are available at https://github.com/diyiiyiii/StyTR-2. Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Xingjia Pan, Changsheng Xu |
CVPR | 5 |
| 2022 | Expanding Low-Density Latent Regions for Open-Set Object DetectionabstractModern object detectors have achieved impressive progress under the close-set setup. However, open-set object detection (OSOD) remains challenging since objects of unknown categories are often misclassified to existing known classes. In this work, we propose to identify unknown objects by separating high/low-density regions in the latent space, based on the consensus that unknown objects are usually distributed in low-density latent regions. As traditional threshold-based methods only maintain limited low-density regions, which cannot cover all unknown objects, we present a novel Openset Detector (OpenDet) with expanded low-density regions. To this aim, we equip Open-Det with two learners, Contrastive Feature Learner (CFL) and Unknown Probability Learner (UPL). CFL performs instance-level contrastive learning to encourage compact features of known classes, leaving more low-density regions for unknown classes; UPL optimizes unknown probability based on the uncertainty of predictions, which further divides more low-density regions around the cluster of known classes. Thus, unknown objects in low-density regions can be easily identified with the learned unknown probability. Extensive experiments demonstrate that our method can significantly improve the OSOD performance, e.g., OpenDet reduces the Absolute Open-Set Errors by 25%-35% on six OSOD benchmarks. Code is available at: https://github.com/csuhan/opendet2. Jiaming Han, Yuqiang Ren, Jian Ding 0001, Xingjia Pan, Gui-Song Xia |
CVPR | 4 |
| 2022 | SIOD: Single Instance Annotated Per Category Per Image for Object DetectionabstractObject detection under imperfect data receives great attention recently. Weakly supervised object detection (WSOD) suffers from severe localization issues due to the lack of instance-level annotation, while semi-supervised object detection (SSOD) remains challenging led by the inter-image discrepancy between labeled and unlabeled data. In this study, we propose the Single Instance annotated Object Detection (SIOD), requiring only one instance annotation for each existing category in an image. Degraded from inter-task (WSOD) or inter-image (SSOD) discrepancies to the intra-image discrepancy, SIOD provides more reliable and rich prior knowledge for mining the rest of unlabeled instances and trades off the annotation cost and performance. Under the SIOD setting, we propose a simple yet effective framework, termed Dual-Mining (DMiner), which consists of a Similarity-based Pseudo Label Generating module (SPLG) and a Pixel-level Group Contrastive Learning module (PGCL). SPLG firstly mines latent instances from feature representation space to alleviate the annotation missing problem. To avoid being misled by inaccurate pseudo labels, we propose PGCL to boost the tolerance to false pseudo labels. Extensive experiments on MS COCO verify the feasibility of the SIOD setting and the superiority of the proposed method, which obtains consistent and significant improvements compared to baseline methods and achieves comparable results with fully supervised object detection (FSOD) methods with only 40% instances annotated. Code is available at https://github.com/solicucu/SIOD. Hanjun Li 0004, Xingjia Pan, Fan Tang, Wei-Shi Zheng 0001 |
CVPR | 2 |
| 2022 | SeqTR: A Simple Yet Universal Network for Visual Grounding
Yiyi Zhou, Yunhang Shen, Gen Luo, Xingjia Pan, Mingbao Lin, Chao Chen 0026, Liujuan Cao, Xiaoshuai Sun, Rongrong Ji |
ECCV (35) | 5 |
| 2022 | A Comparative Study of CNN- and Transformer-Based Visual Style Transfer
HuaPeng Wei, Yingying Deng, Fan Tang, Xingjia Pan, Weiming Dong |
J. Comput. Sci. Technol. | 4 |
| 2021 | Unveiling the Potential of Structure Preserving for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) remains an open problem given the deficiency of finding object extent information using a classification network. Although prior works struggled to localize objects through various spatial regularization strategies, we argue that how to extract object structural information from the trained classification network is neglected. In this paper, we propose a two-stage approach, termed structure-preserving activation (SPA), toward fully leveraging the structure information incorporated in convolutional features for WSOL. First, a restricted activation module (RAM) is designed to alleviate the structure-missing issue caused by the classification network on the basis of the observation that the unbounded classification map and global average pooling layer drive the network to focus only on object parts. Second, we designed a post-process approach, termed self-correlation map generating (SCG) module to obtain structure-preserving localization maps on the basis of the activation maps acquired from the first stage. Specifically, we utilize the high-order self-correlation (HSC) to extract the inherent structural information retained in the learned model and then aggregate HSC of multiple points for precise object localization. Extensive experiments on two publicly available benchmarks including CUB-2002011 and ILSVRC show that the proposed SPA achieves substantial and consistent performance gains compared with baseline approaches. Code and models are available at github.com/Panxjia/SPA CVPR2021. Xingjia Pan, Yingguo Gao, Zhiwen Lin, Fan Tang, Weiming Dong, Haolei Yuan, Feiyue Huang, Changsheng Xu |
CVPR | 1 |
| 2021 | TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationabstractWeakly supervised object localization (WSOL) is a challenging problem when given image category labels but requires to learn object localization models. Optimizing a convolutional neural network (CNN) for classification tends to activate local discriminative regions while ignoring complete object extent, causing the partial activation issue. In this paper, we argue that partial activation is caused by the intrinsic characteristics of CNN, where the convolution operations produce local receptive fields and experience difficulty to capture long-range feature dependency among pixels. We introduce the token semantic coupled attention map (TS-CAM) to take full advantage of the self-attention mechanism in visual transformer for long-range dependency extraction. TS-CAM first splits an image into a sequence of patch tokens for spatial embedding, which produce attention maps of long-range visual dependency to avoid partial activation. TS-CAM then re-allocates category-related semantics for patch tokens, enabling each of them to be aware of object categories. TS-CAM finally couples the patch tokens with the semantic-agnostic attention map to achieve semantic-aware localization. Experiments on the ILSVRC/CUB-200-2011 datasets show that TS-CAM outperforms its CNN-CAM counterparts by 7.1%/27.1% for WSOL, achieving state-of-the-art performance. Code is available at https://github.com/vasgaowei/TS-CAM Wei Gao 0050, Fang Wan 0001, Xingjia Pan, Zhiliang Peng, Qi Tian 0001, Zhenjun Han, Bolei Zhou, Qixiang Ye |
ICCV | 3 |
| 2021 | Distributed Attention for Grounded Image CaptioningabstractWe study the problem of weakly supervised grounded image captioning. That is, given an image, the goal is to automatically generate a sentence describing the context of the image with each noun word grounded to the corresponding region in the image. This task is challenging due to the lack of explicit fine-grained region word alignments as supervision. Previous weakly supervised methods mainly explore various kinds of regularization schemes to improve attention accuracy. However, their performances are still far from the fully supervised ones. One main issue that has been ignored is that the attention for generating visually groundable words may only focus on the most discriminate parts and can not cover the whole object. To this end, we propose a simple yet effective method to alleviate the issue, termed as partial grounding problem in our paper. Specifically, we design a distributed attention mechanism to enforce the network to aggregate information from multiple spatially different regions with consistent semantics while generating the words. Therefore, the union of the focused region proposals should form a visual region that encloses the object of interest completely. Extensive experiments have demonstrated the superiority of our proposed method compared with the state-of-the-arts. Nenglun Chen, Xingjia Pan, Runnan Chen, Lei Yang 0048, Zhiwen Lin, Yuqiang Ren, Haolei Yuan, Feiyue Huang, Wenping Wang 0001 |
ACM Multimedia | 2 |
| 2021 | Content-Based Visual Summarization for Image CollectionsabstractWith the surge of images in the information era, people demand an effective and accurate way to access meaningful visual information. Accordingly, effective and accurate communication of information has become indispensable. In this article, we propose a content-based approach that automatically generates a clear and informative visual summarization based on design principles and cognitive psychology to represent image collections. We first introduce a novel method to make representative and nonredundant summarizations of image collections, thereby ensuring data cleanliness and emphasizing important information. Then, we propose a tree-based algorithm with a two-step optimization strategy to generate the final layout that operates as follows: (1) an initial layout is created by constructing a tree randomly based on the grouping results of the input image set; (2) the layout is refined through a coarse adjustment in a greedy manner, followed by gradient back propagation drawing on the training procedure of neural networks. We demonstrate the usefulness and effectiveness of our method via extensive experimental results and user studies. Our visual summarization algorithm can precisely and efficiently capture the main content of image collections better than alternative methods or commercial tools. Xingjia Pan, Fan Tang, Weiming Dong, Chongyang Ma, Yiping Meng, Feiyue Huang, Tong-Yee Lee, Changsheng Xu |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Dynamic Refinement Network for Oriented and Densely Packed Object DetectionabstractObject detection has achieved remarkable progress in the past decade. However, the detection of oriented and densely packed objects remains challenging because of following inherent reasons: (1) receptive fields of neurons are all axis-aligned and of the same shape, whereas objects are usually of diverse shapes and align along various directions; (2) detection models are typically trained with generic knowledge and may not generalize well to handle specific objects at test time; (3) the limited dataset hinders the development on this task. To resolve the first two issues, we present a dynamic refinement network that consists of two novel components, i.e., a feature selection module (FSM) and a dynamic refinement head (DRH). Our FSM enables neurons to adjust receptive fields in accordance with the shapes and orientations of target objects, whereas the DRH empowers our model to refine the prediction dynamically in an object-aware manner. To address the limited availability of related benchmarks, we collect an extensive and fully annotated dataset, namely, SKU110K-R, which is relabeled with oriented bounding boxes based on SKU110K. We perform quantitative evaluations on several publicly available benchmarks including DOTA, HRSC2016, SKU110K, and our own SKU110K-R dataset. Experimental results show that our method achieves consistent and substantial gains compared with baseline approaches. Our source code and dataset will be released to encourage follow-up research. Xingjia Pan, Yuqiang Ren, Kekai Sheng, Weiming Dong, Haolei Yuan, Chongyang Ma, Changsheng Xu |
CVPR | 1 |
| 2020 | Self-Supervised Feature Augmentation for Large Image Object DetectionabstractInput scale plays an important role in modern detection frameworks, and an optimal training scale for images exists empirically. However, the optimal one usually cannot be reached in facing extremely large images under the memory constraint. In this study, we explore the scale effect inside the object detection pipeline and find that feature upsampling with the introduction of high-resolution information benefits the detection. Compared with direct input upscaling, feature upsampling trades a small performance loss for a large amount of memory savings. From these observations, we propose a self-supervised feature augmentation network, which takes downsampled images as inputs and aims to generate comparable features with the ones when feeding upscaled images to networks. We present a guided feature upsampling module, which takes downsampled images as inputs, to learn upscaled feature representations with the supervision of real large features acquired from upscaled images. In a self-supervised learning manner, we can introduce detailed information of images to the network. For an efficient feature upsampling, we design a residualized sub-pixel convolution block based on a sub-pixel convolution layer, which involves considerable information in upsampling process. Experiments on Mapillary Vistas Dataset (MVD), Cityscapes, and COCO are conducted to demonstrate the effectiveness of our method. On the MVD and Cityscapes detection benchmarks, in which the images are extremely large, our method surpasses current approaches. On COCO, the proposed method obtains comparable results to existing methods but with higher efficiency. Xingjia Pan, Fan Tang, Weiming Dong, Zhichao Song, Yiping Meng, Pengfei Xu 0013, Oliver Deussen, Changsheng Xu |
IEEE Trans. Image Process. | 1 |
| 2018 | Automatic Building Rooftop Extraction From Aerial Images via Hierarchical RGB-D PriorsabstractAccurate building rooftop extraction from high-resolution aerial images is of crucial importance in a wide range of applications. Owing to the varying appearance and large-scale range of scene objects, especially for building rooftops in different scales and heights, single-scale or individual prior-based extraction technique is insufficient in pursuing efficient, generic, and accurate extraction results. The trend toward integrating multiscale or several cue techniques appears to be the best way; thus, such integration is the focus of this paper. We first propose a novel salient rooftop detector integrating four correlative RGB-D priors (depth cue, uniqueness prior, shape prior, and transition surface prior) for improved rooftop extraction to address the preceding complex issues mentioned. Then, these correlative cues are computed from image layers created by our multilevel segmentation and further fused into the state-of-the-art high-order conditional random field (CRF) framework to locate the rooftop. Finally, an iterative optimization strategy is applied for high-quality solving, which can robustly handle varying appearance of building rooftops. Performance evaluations in the SZTAKI-INRIA benchmark data sets show that our method outperforms the traditional color-based algorithm and the original high-order CRF algorithm and its variants. The proposed algorithm is also evaluated and found to produce consistently satisfactory results for various large-scale, real-world data sets. Shibiao Xu, Xingjia Pan, Er Li, Baoyuan Wu, Shuhui Bu, Weiming Dong, Shiming Xiang, Xiaopeng Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |