Tianxiang Pan

dblp:195/8239 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0005-8126-2743ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Semantic Hierarchical Prompt Tuning for Parameter-Efficient Fine-Tuning
abstract
As the scale of vision models continues to grow, Visual Prompt Timing (VPT) has emerged as a parameter-efficient transfer learning technique, noted for its superior performance compared to full fine-tuning. However, indiscriminately applying prompts to every layer without considering their inherent correlations, can cause significant disturbances, leading to suboptimal transferability. Additionally, VPT disrupts the original self-attention structure, affecting the aggregation of visual features, and lacks a mechanism for explicitly mining discriminative visual features, which are crucial for classification. To address these issues, we propose a Semantic Hierarchical Prompt (SHIP) fine-tuning strategy. We adaptively construct semantic hierarchies and use semantic-independent and semantic-shared prompts to learn hierarchical representations. We also integrate attribute prompts and a prompt matching loss to enhance feature discrimination and employ decoupled attention for robustness and reduced inference costs. SHIP significantly improves performance, achieving a 4.9% gain in accuracy over VPT with a ViT-B/16 backbone on VTAB-1k tasks. Our code is available at https://github.com/haoweiz23/SHIP.
Haowei Zhu, Tianxiang Pan, Jun-Hai Yong, Bin Wang 0021
ICASSP4
2025 ReCon: Region-Controllable Data Augmentation with Rectification and Alignment for Object Detection
abstract
The scale and quality of datasets are crucial for training robust perception models. However, obtaining large-scale annotated data is both costly and time-consuming. Generative models have emerged as a powerful tool for data augmentation by synthesizing samples that adhere to desired distributions. However, current generative approaches often rely on complex post-processing or extensive fine-tuning on massive datasets to achieve satisfactory results, and they remain prone to content–position mismatches and semantic leakage. To overcome these limitations, we introduce ReCon, a novel augmentation framework that enhances the capacity of structure-controllable generative models for object detection. ReCon integrates region-guided rectification into the diffusion sampling process, using feedback from a pre-trained perception model to rectify misgenerated regions within diffusion sampling process. We further propose region-aligned cross-attention to enforce spatial–semantic alignment between image regions and their textual cues, thereby improving both semantic consistency and overall image fidelity. Extensive experiments demonstrate that ReCon substantially improve the quality and trainability of generated data, achieving consistent performance gains across various datasets, backbone architectures, and data scales.
Haowei Zhu, Tianxiang Pan, Jun-Hai Yong, Bin Wang 0021
NeurIPS2
2024 W2P: Switching from Weak Supervision to Partial Supervision for Semantic Segmentation
abstract
Current weakly-supervised semantic segmentation (WSSS) techniques concentrate on enhancing class activation maps (CAMs) with image-level annotations. Yet, the emphasis on producing these pseudo-labels often overshadows the pivotal role of training the segmentation model itself. This paper underscores the significant influence of noisy pseudo-labels on segmentation network performance, particularly in boundary region. To address above issues, we introduce a novel paradigm: Weak to Partial Supervision (W2P). At its core, W2P categorizes the pseudo-labels from WSSS into two unique supervisions: trustworthy clean labels and uncertain noisy labels. Next, our proposed partially-supervised framework adeptly employs these clean labels to rectify the noisy ones, thereby promoting the continuous enhancement of the segmentation model. To further optimize boundary segmentation, we incorporate a noise detection mechanism that specifically preserves boundary regions while eliminating noise. During the noise refinement phase, we adopt a boundary-conscious noise correction technique to extract comprehensive boundaries from noisy areas. Furthermore, we devise a boundary generation approach that assists in predicting intricate boundary zones. Evaluations on the PASCAL VOC 2012 and MS COCO 2014 datasets confirm our method's impressive segmentation capabilities across various pseudo-labels.
Tianxiang Pan, Jun-Hai Yong, Bin Wang 0021
AAAI2
2024 Boundary-Enhanced Instance Segmentation
abstract
Despite significant progress in instance segmentation, recent solutions still fall short of boundary accuracy especially for overlapping instances of the same category. In this paper, we propose a novel boundary-enhanced instance segmentation (BEIS) framework that explicitly models the feature relationships across object boundaries for high-quality instance segmentation. Specifically, BEIS generates boundary-enhanced features using both intra-mask and cross-image boundary discrimination learning. The intra-mask boundary discrimination learning (IBDL) employs pixel-level discrimination learning to disentangle pixel representations along boundaries. The cross-image boundary discrimination learning (CBDL) learns a boundary-aware feature bank from training data to further boost the performance. Thus, CBDL can take advantage of boundary relations across images to enhance the quality of segmented boundaries. To focus on hard-to-segment boundaries, we propose an adaptive sampling strategy to automatically construct discriminative pairs in regions with high possibilities of confusion. Extensive experiments show BEIS outperforms on various datasets.
Tianxiang Pan, Yu-Wing Tai, Bin Wang 0021
ECAI2
2023 Low-Confidence Samples Mining for Semi-supervised Object Detection
abstract
Reliable pseudo labels from unlabeled data play a key role in semi-supervised object detection (SSOD). However, the state-of-the-art SSOD methods all rely on pseudo labels with high confidence, which ignore valuable pseudo labels with lower confidence. Additionally, the insufficient excavation for unlabeled data results in an excessively low recall rate thus hurting the network training. In this paper, we propose a novel Low-confidence Samples Mining (LSM) method to utilize low confidence pseudo labels efficiently. Specifically, we develop an additional pseudo information mining (PIM) branch on account of low-resolution feature maps to extract reliable large area instances, the IoUs of which are higher than small area ones. Owing to the complementary predictions between PIM and the main branch, we further design self-distillation (SD) to compensate for both in a mutually learning manner. Meanwhile, the extensibility of the above approaches enables our LSM to apply to Faster-RCNN and Deformable-DETR respectively. On the MS-COCO benchmark, our method achieves 3.54% mAP improvement over state-of-the-art methods under 5% labeling ratios.
Guandu Liu, Tianxiang Pan, Jun-Hai Yong, Bin Wang 0021
IJCAI3
2022 Semi-supervised Object Detection with Adaptive Class-Rebalancing Self-Training
abstract
While self-training achieves state-of-the-art results in semi-supervised object detection (SSOD), it severely suffers from foreground-background and foreground-foreground imbalances in SSOD. In this paper, we propose an Adaptive Class-Rebalancing Self-Training (ACRST) with a novel memory module called CropBank to alleviate these imbalances and generate unbiased pseudo-labels. Besides, we observe that both self-training and data-rebalancing procedures suffer from noisy pseudo-labels in SSOD. Therefore, we contribute a simple yet effective two-stage pseudo-label filtering scheme to obtain accurate supervision. Our method achieves competitive performance on MS-COCO and VOC benchmarks. When using only 1% labeled data of MS-COCO, our method achieves 17.02 mAP improvement over the supervised method and 5.32 mAP gains compared with state-of-the-arts.
Tianxiang Pan, Bin Wang 0021
AAAI2
2020 P2MAT-NET: Learning medial axis transform from sparse point clouds
Baorong Yang, Junfeng Yao, Bin Wang 0021, Jianwei Hu 0003, Yiling Pan, Tianxiang Pan, Wenping Wang 0001, Xiaohu Guo
Comput. Aided Geom. Des.6
2019 Low Shot Box Correction for Weakly Supervised Object Detection
abstract
Weakly supervised object detection (WSOD) has been widely studied but the accuracy of state-of-art methods remains far lower than strongly supervised methods. One major reason for this huge gap is the incomplete box detection problem which arises because most previous WSOD models are structured on classification networks and therefore tend to recognize the most discriminative parts instead of complete bounding boxes. To solve this problem, we define a low-shot weakly supervised object detection task and propose a novel low-shot box correction network to address it. The proposed task enables to train object detectors on a large data set all of which have image-level annotations, but only a small portion or few shots have box annotations. Given the low-shot box annotations, we use a novel box correction network to transfer the incomplete boxes into complete ones. Extensive empirical evidence shows that our proposed method yields state-of-art detection accuracy under various settings on the PASCAL VOC benchmark.
Tianxiang Pan, Bin Wang 0021, Guiguang Ding, Jungong Han, Jun-Hai Yong
IJCAI1
2018 Shadow Detection Using Robust Texture Learning
Tianxiang Pan, Bin Wang 0021, Guiguang Ding, Jun-Hai Yong
BMVC1
2018 CropNet: Real-Time Thumbnailing
abstract
We present a deep learning-based thumbnail generation method called CropNet in this paper. Unlike previous deep learning-based methods, such as Fast-AT, which can utilize detectors introduced in object detection frameworks and generate thousands of proposals, our detector is straightforward and concise, thereby ensuring that the final cropping window is computed by its center and width, with the input aspect ratio. To achieve this goal, CropNet learns specific filters to estimate the center position and utilizes a cascade structure of filters and single neuron for width inference. In addition, CropNet optimizes the center and width jointly for optimal results. We collect a data set of more than 29,000 thumbnail annotations to train CropNet and perform cross-validation between existing data sets. Experiments show that CropNet outperforms existing techniques. Our result is achieved at a test-time speed of 17 ms per image, which is six times faster than the fastest method at present.
Huarong Chen, Bin Wang 0021, Tianxiang Pan, Liwang Zhou, Hua Zeng
ACM Multimedia3
2017 Fully Convolutional Neural Networks with Full-Scale-Features for Semantic Segmentation
abstract
In this work, we propose a novel method to involve full-scale-features into the fully convolutional neural networks (FCNs) for Semantic Segmentation. Current works on FCN has brought great advances in the task of semantic segmentation, but the receptive field, which represents region areas of input volume connected to any output neuron, limits the available information of output neuron's prediction accuracy. We investigate how to involve the full-scale or full-image features into FCNs to enrich the receptive field. Specially, the full-scale feature network (FFN) extends the full-connected network and makes an end-to-end unified training structure. It has two appealing properties. First, the introduction of full-scale-features is beneficial for prediction. We build a unified extracting network and explore several fusion functions for concatenating features. Amounts of experiments have been carried out to prove that full-scale-features makes fair accuracy raising. Second, FFN is applicable to many variants of FCN which could be regarded as a general strategy to improve the segmentation accuracy. Our proposed method is evaluated on PASCAL VOC 2012, and achieves a state-of-art result.
Tianxiang Pan, Bin Wang 0021, Guiguang Ding, Jun-Hai Yong
AAAI1