Yanchun Xie

dblp:184/9463 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Reproducibility Companion Paper: u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model
Jinjin Xu, Xilu Wang 0001, Liwu Xu, Yuzhe Yang 0001, Xiang Li 0179, Fanyi Wang, Yanchun Xie, Yi-Jie Huang, Yunfan Hu
ICMR7
2025 Open-Set Image Tagging with Multi-Grained Text Supervision
abstract
This paper introduces the Recognize Anything Plus Model (RAM++), an open-set image tagging model effectively leveraging multi-grained text supervision. Previous approaches (e.g., CLIP) primarily utilize global text supervision paired with images, leading to sub-optimal performance in recognizing multiple individual semantic tags. In contrast, RAM++ seamlessly integrates individual tag supervision with global text supervision, all within a unified alignment framework. This integration not only ensures efficient recognition of predefined tag categories, but also enhances generalization capabilities for diverse open-set categories. Furthermore, RAM++ employs large language models (LLMs) to convert semantically constrained tag supervision into more expansive tag description supervision, thereby enriching the scope of open-set visual description concepts. Comprehensive evaluations on various image recognition benchmarks demonstrate RAM++ exceeds existing state-of-the-art (SOTA) open-set image tagging models on most aspects. Specifically, for predefined commonly used tag categories, RAM++ showcases 10.2 mAP and 15.4 mAP enhancements over CLIP on OpenImages and ImageNet. For open-set categories beyond predefined, RAM++ records improvements of 5.0 mAP and 6.4 mAP over CLIP and RAM respectively on OpenImages. For diverse human-object interaction phrases, RAM++ achieves 7.8 mAP and 4.7 mAP improvements on the HICO benchmark.
Yi-Jie Huang, Youcai Zhang, Rui Feng 0001, Yuejie Zhang, Yanchun Xie, Lei Zhang 0001
ACM Multimedia7
2025 Toward Realistic Hierarchical Object Detection: Problem, Benchmark, and Solution
abstract
With the continuous advancement of deep learning, object detection has made remarkable progress in accurately identifying a wide range of object categories, even within increasingly complex scenes. However, as the number of categories grows, visual concepts naturally organize into a label hierarchy. We contend that existing hierarchical classification and detection methods predominantly prioritize fine-grained prediction, potentially leading to inconsistencies with realistic human perception. From this perspective, we investigate the Hierarchical Object Detection (HOD) problem to better align with real-world perception. To address the lack of benchmarks in the field, we build a large-scale HOD benchmark termed RHOD with open-source datasets, comprising 740 categories. To better align the hierarchical object detectors towards realistic perception, we propose a new evaluation metric named Hierarchical Average Precision (HAP). Furthermore, we present a novel hierarchical object detection method that includes two components, Tree Soft Labeling (TSL) and Hierarchical Extension and Suppression (HES). Our method mitigates the issue of overconfidence in fine-grained predictions, which has been prevalent in previous approaches. We evaluate a range of existing methods on the RHOD benchmark, including plain, hierarchical, and open-vocabulary models. Additionally, we perform comprehensive experiments to assess the performance of our proposed method. The experimental results show that our method achieves state-of-the-art performance on the RHOD benchmark.
Juexiao Feng, Yuhong Yang 0008, Mengyao Lyu, Tianxiang Hao 0001, Yi-Jie Huang, Yanchun Xie, Jungong Han, Liuyu Xiang, Guiguang Ding
IEEE Trans. Circuits Syst. Video Technol.6
2024 Debiased Novel Category Discovering and Localization
abstract
In recent years, object detection in deep learning has experienced rapid development. However, most existing object detection models perform well only on closed-set datasets, ignoring a large number of potential objects whose categories are not defined in the training set. These objects are often identified as background or incorrectly classified as pre-defined categories by the detectors. In this paper, we focus on the challenging problem of Novel Class Discovery and Localization (NCDL), aiming to train detectors that can detect the categories present in the training data, while also actively discover, localize, and cluster new categories. We analyze existing NCDL methods and identify the core issue: object detectors tend to be biased towards seen objects, and this leads to the neglect of unseen targets. To address this issue, we first propose an Debiased Region Mining (DRM) approach that combines class-agnostic Region Proposal Network (RPN) and class-aware RPN in a complementary manner. Additionally, we suggest to improve the representation network through semi-supervised contrastive learning by leveraging unlabeled data. Finally, we adopt a simple and efficient mini-batch K-means clustering method for novel class discovery. We conduct extensive experiments on the NCDL benchmark, and the results demonstrate that the proposed DRM approach significantly outperforms previous methods, establishing a new state-of-the-art.
Juexiao Feng, Yuhong Yang 0008, Yanchun Xie, Yandong Guo, Liuyu Xiang, Guiguang Ding
AAAI3
2024 u-LLaVA: Unifying Multi-Modal Tasks via Large Language Model
abstract
Recent advancements in multi-modal large language models (MLLMs) have led to substantial improvements in visual understanding, primarily driven by sophisticated modality alignment strategies. However, predominant approaches prioritize global or regional comprehension, with less focus on fine-grained, pixel-level tasks. To address this gap, we introduce u-LLaVA, an innovative unifying multi-task framework that integrates pixel, regional, and global features to refine the perceptual faculties of MLLMs. We commence by leveraging an efficient modality alignment approach, harnessing both image and video datasets to bolster the model’s foundational understanding across diverse visual contexts. Subsequently, a joint instruction tuning method with task-specific projectors and decoders for end-to-end downstream training is presented. Furthermore, this work contributes a novel mask-based multi-task dataset comprising 277K samples, crafted to challenge and assess the fine-grained perception capabilities of MLLMs. The overall framework is simple, effective, and achieves state-of-the-art performance across multiple benchmarks. We make model, data, and code publicly accessible at https://github.com/OPPOMKLab/u-LLaVA.
Jinjin Xu, Liwu Xu, Yuzhe Yang 0001, Xiang Li 0179, Fanyi Wang, Yanchun Xie, Yi-Jie Huang
ECAI6
2022 Progressive multi-scale fusion network for RGB-D salient object detection
abstract
Salient object detection (SOD) aims at locating the most significant object within a given image. In recent years, great progress has been made in applying SOD on many vision tasks. The depth map could provide additional spatial prior and boundary cues to boost the performance. Combining the depth information with image data obtained from standard visual cameras has been widely used in recent SOD works, however, introducing depth information in a suboptimal fusion strategy may have negative influence in the performance of SOD. In this paper, we discuss about the advantages of the so-called progressive multi-scale fusion method and propose a mask-guided feature aggregation module (MGFA). The proposed framework can effectively combine the two features of different modalities and, furthermore, alleviate the impact of erroneous depth features, which are inevitably caused by the variation of depth quality. We further introduce a mask-guided refinement module (MGRM) to complement the high-level semantic features and reduce the irrelevant features from multi-scale fusion, leading to an overall refinement of detection. Experiments on five challenging benchmarks demonstrate that the proposed method outperforms 11 state-of-the-art methods under different evaluation metrics.
Guangyu Ren, Yanchun Xie, Tianhong Dai, Tania Stathaki
Comput. Vis. Image Underst.2
2020 Feature Representation Matters: End-to-End Learning for Reference-Based Image Super-Resolution
Yanchun Xie, Jimin Xiao, Mingjie Sun, Kaizhu Huang
ECCV (4)1
2020 Adaptive ROI generation for video object segmentation using reinforcement learning
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yanchun Xie, Jiashi Feng
Pattern Recognit.4
2020 Correlation Filter Selection for Visual Tracking Using Reinforcement Learning
abstract
Correlation filter has been proven to be an effective tool for a number of approaches in visual tracking, particularly for seeking a good balance between tracking accuracy and speed. However, correlation filter-based models are susceptible to wrong updates stemming from inaccurate tracking results. To date, very little effort has been devoted towards handling the correlation filter update problem. In this paper, we propose a novel approach to address the correlation filter update problem. In our approach, we update and maintain multiple correlation filter models in parallel, and we use deep reinforcement learning for the selection of an optimal correlation filter model among them. To facilitate the decision process in an efficient manner, we propose a decision-net to deal with target appearance modeling, which is trained through hundreds of challenging videos using proximal policy optimization and a lightweight learning network. An exhaustive evaluation of the proposed approach on the OTB100 and OTB2013 benchmarks shows that the approach is effective enough to achieve the average success rate of 62.3% and the average precision score of 81.2%, both exceeding the performance of traditional correlation filter-based trackers.
Yanchun Xie, Jimin Xiao, Kaizhu Huang, Jeyan Thiyagalingam, Yao Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 IAN: The Individual Aggregation Network for Person Search
Jimin Xiao, Yanchun Xie, Tammam Tillo, Kaizhu Huang, Yunchao Wei, Jiashi Feng
Pattern Recognit.2
2018 Siamese network ensemble for visual tracking
Chenru Jiang, Jimin Xiao, Yanchun Xie, Tammam Tillo, Kaizhu Huang
Neurocomputing3
2016 3D video super-resolution using fully convolutional neural networks
abstract
Large amount of redundant information and huge data size have been a serious problem for multiview video systems. To address this problem, one popular solution is mixed-resolution, where only few viewpoints are kept with full resolution and other views are kept with lower resolution. In this paper, we propose a super-resolution (SR) method, where the low-resolution viewpoints in the 3D video are up-sampled using a fully convolutional neural network. By simply projecting the neighboring high resolution image to the position of the low resolution image, we learn the relationship of high and low resolution patches, and reconstruct the low resolution images into high resolution ones using the projected image information. We propose to use a fully convolutional neural network to establish a mapping between those images. The network is barely trained on 17 pairs of multiview images, and tested on other multiview images and video sequences. It is observed that our proposed method outperforms existing methods objectively and subjectively, with more than 1 dB average gain achieved. Meanwhile, our network training procedure is efficient, with less than 3 hours using one Titan X GPU.
Yanchun Xie, Jimin Xiao, Tammam Tillo, Yunchao Wei, Yao Zhao 0001
ICME1