Yufei Xu

dblp:43/7400 · DBLP profile ↗
← Back
22ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0002-9931-5138ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Computer networks · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use
abstract
Junjie Ye, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zehui Chen, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Siyu Yuan, Tao Gui, Qi Zhang, Xuanjing Huang, Jiecao Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Junjie Ye 0005, Zhengyin Du, Xuesong Yao, Weijian Lin, Yufei Xu, Zaiyuan Wang, Sining Zhu, Zhiheng Xi, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001, Jiecao Chen
ACL (1)5
2025 Simple but Effective: Sub-Volume Contrastive Learning for Class-Imbalanced Semi-Supervised 3D Medical Image Segmentation
abstract
Medical image segmentation is essential for precise anatomical delineation and clinical decision-making. However, fully supervised methods are limited by the substantial cost of acquiring pixel-level annotations, particularly for 3D volumetric data. Semi-supervised learning (SSL) alleviates this challenge by leveraging unlabeled data, yet it remains hindered by severe class imbalance, where dominant structures disproportionately occupy the voxel space, leading to feature degradation and unreliable pseudo-labels. To address this issue, we propose a simple but effective SSL framework, namely Sub-Volume Contrastive Learning (SuVCL), to enhance feature discriminability in imbalanced 3D medical image segmentation. Our approach incorporates localized contrastive learning through sub-volume sampling, which captures small but semantically informative regions to retain fine-grained structural details while mitigating computational overhead. Furthermore, we introduce a balanced memory bank mechanism, which dynamically maintains class-specific feature representations with adaptive updates guided by class-predictive confidence. Extensive experimental evaluations demonstrate that our method substantially enhances segmentation performance for minority classes, demonstrating substantial performance gains over existing SOTAs.
Xianrun Xu, Baoyao Yang, Wanyun Li, Jingsong Lin, Yufei Xu
ACM Multimedia5
2025 A novel hybrid model for short-term freeway network traffic state prediction under heavy rainstorm
Jian Wang 0085, Yanxun Chen, Yufei Xu
Expert Syst. Appl.4
2024 HandRefiner: Refining Malformed Hands in Generated Images by Diffusion-based Conditional Inpainting
abstract
Diffusion models have achieved remarkable success in generating realistic images but suffer from generating accurate human hands, such as incorrect finger counts or irregular shapes. This difficulty arises from the complex task of learning the physical structure and pose of hands from training images, which involves extensive deformations and occlusions. For correct hand generation, our paper introduces a lightweight post-processing solution called HandRefiner. HandRefiner employs a conditional inpainting approach to rectify malformed hands while leaving other parts of the image untouched. We leverage the hand mesh reconstruction model that consistently adheres to the correct number of fingers and hand shape, while also being capable of fitting the desired hand pose in the generated image. Given a generated failed image due to malformed hands, we utilize ControlNet modules to re-inject such correct hand information. Additionally, we uncover a phase transition phenomenon within ControlNet as we vary the control strength. It enables us to take advantage of more readily available synthetic data without suffering from the domain gap between realistic and synthetic hands. Experiments demonstrate that HandRefiner can significantly improve the generation quality quantitatively and qualitatively. The code is available at https://github.com/wenquanlu/HandRefiner.
Wenquan Lu, Yufei Xu, Jing Zhang 0037, Dacheng Tao
ACM Multimedia2
2024 When ControlNet Meets Inexplicit Masks: A Case Study of ControlNet on its Contour-following Ability
abstract
ControlNet excels at creating content that closely matches precise contours in user-provided masks. However, when these masks contain noise, as a frequent occurrence with non-expert users, the output would include unwanted artifacts. This paper first highlights the crucial role of controlling the impact of these inexplicit masks with diverse deterioration levels through in-depth analysis. Subsequently, to enhance controllability with inexplicit masks, an advanced Shape-aware ControlNet consisting of a deterioration estimator and a shape-prior modulation block is devised. The deterioration estimator assesses the deterioration factor of the provided masks. Then this factor is used in a modulation block to adaptively adjust the model's contour-following ability, which helps it dismiss the noise part in the inexplicit masks. Extensive experiments prove its effectiveness in encouraging ControlNet to interpret inaccurate spatial conditions robustly rather than blindly following the given contours, suitable for diverse kinds of conditions. We showcase application scenarios like modifying shape priors and composable shape-controllable generation. Codes are available at github.
Wenjie Xuan, Yufei Xu, Shanshan Zhao 0001, Juhua Liu, Bo Du 0001, Dacheng Tao
ACM Multimedia2
2024 ViTPose++: Vision Transformer for Generic Body Pose Estimation
abstract
In this paper, we show the surprisingly good properties of plain vision transformers for body pose estimation from various aspects, namely simplicity in model structure, scalability in model size, flexibility in training paradigm, and transferability of knowledge between models, through a simple baseline model dubbed ViTPose. ViTPose employs the plain and non-hierarchical vision transformer as an encoder to encode features and a lightweight decoder to decode body keypoints in either a top-down or a bottom-up manner. It can be scaled to 1B parameters by taking the advantage of the scalable model capacity and high parallelism, setting a new Pareto front for throughput and performance. Besides, ViTPose is very flexible regarding the attention type, input resolution, and pre-training and fine-tuning strategy. Based on the flexibility, a novel ViTPose++ model is proposed to deal with heterogeneous body keypoint categories via knowledge factorization, i.e., adopting task-agnostic and task-specific feed-forward networks in the transformer. We also demonstrate that the knowledge of large ViTPose models can be easily transferred to small ones via a simple knowledge token. Our largest single model ViTPose-G sets a new record on the MS COCO test set without model ensemble. Furthermore, our ViTPose++ model achieves state-of-the-art performance simultaneously on a series of body pose estimation tasks, including MS COCO, AI Challenger, OCHuman, MPII for human keypoint detection, COCO-Wholebody for whole-body keypoint detection, as well as AP-10K and APT-36K for animal keypoint detection, without sacrificing inference speed.
Yufei Xu, Jing Zhang 0037, Qiming Zhang 0001, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Vision Transformer With Quadrangle Attention
abstract
Window-based attention has become a popular choice in vision transformers due to its superior performance, lower computational complexity, and less memory footprint. However, the design of hand-crafted windows, which is data-agnostic, constrains the flexibility of transformers to adapt to objects of varying sizes, shapes, and orientations. To address this issue, we propose a novel quadrangle attention (QA) method that extends the window-based attention to a general quadrangle formulation. Our method employs an end-to-end learnable quadrangle regression module that predicts a transformation matrix to transform default windows into target quadrangles for token sampling and attention calculation, enabling the network to model various targets with different shapes and orientations and capture rich context information. We integrate QA into plain and hierarchical vision transformers to create a new architecture named QFormer, which offers minor code modifications and negligible extra computational cost. Extensive experiments on public benchmarks demonstrate that QFormer outperforms existing representative vision transformers on various vision tasks, including classification, object detection, semantic segmentation, and pose estimation. The code will be made publicly available at QFormer.
Qiming Zhang 0001, Jing Zhang 0037, Yufei Xu, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 CLAMP: Prompt-based Contrastive Learning for Connecting Language and Animal Pose
abstract
Animal pose estimation is challenging for existing image-based methods because of limited training data and large intra- and inter-species variances. Motivated by the progress of visual-language research, we propose that pre-trained language models (e.g., CLIP) can facilitate animal pose estimation by providing rich prior knowledge for describing animal keypoints in text. However, we found that building effective connections between pre-trained language models and visual animal keypoints is non-trivial since the gap between text-based descriptions and keypoint-based visual features about animal pose can be significant. To address this issue, we introduce a novel prompt-based Contrastive learning scheme for connecting Language and AniMal Pose (CLAMP) effectively. The CLAMP attempts to bridge the gap by adapting the text prompts to the animal keypoints during network training. The adaptation is decomposed into spatialaware and feature-aware processes, and two novel contrastive losses are devised correspondingly. In practice, the CLAMP enables the first cross-modal animal pose estimation paradigm. Experimental results show that our method achieves state-of-the-art performance under the supervised, few-shot, and zero-shot settings, outperforming image-based methods by a large margin. The code is available at https://github.com/xuzhang1199/CLAMP.
Wen Wang 0009, Zhe Chen 0013, Yufei Xu, Jing Zhang 0037, Dacheng Tao
CVPR4
2023 Transformer-Based Context Condensation for Boosting Feature Pyramids in Object Detection
abstract
Abstract Current object detectors typically have a feature pyramid (FP) module for multi-level feature fusion (MFF) which aims to mitigate the gap between features from different levels and form a comprehensive object representation to achieve better detection performance. However, they usually require heavy cross-level connections or iterative refinement to obtain better MFF result, making them complicated in structure and inefficient in computation. To address these issues, we propose a novel and efficient context modeling mechanism that can help existing FPs deliver better MFF results while reducing the computational costs effectively. In particular, we introduce a novel insight that comprehensive contexts can be decomposed and condensed into two types of representations for higher efficiency. The two representations include a locally concentrated representation and a globally summarized representation, where the former focuses on extracting context cues from nearby areas while the latter extracts general contextual representations of the whole image scene as global context cues. By collecting the condensed contexts, we employ a Transformer decoder to investigate the relations between them and each local feature from the FP and then refine the MFF results accordingly. As a result, we obtain a simple and light-weight Transformer-based Context Condensation (TCC) module, which can boost various FPs and lower their computational costs simultaneously. Extensive experimental results on the challenging MS COCO dataset show that TCC is compatible to four representative FPs and consistently improves their detection accuracy by up to 7.8% in terms of average precision and reduce their complexities by up to around 20% in terms of GFLOPs, helping them achieve state-of-the-art performance more efficiently. Code will be released at https://github.com/zhechen/TCC .
Zhe Chen 0013, Jing Zhang 0037, Yufei Xu, Dacheng Tao
Int. J. Comput. Vis.3
2023 ViTAEv2: Vision Transformer Advanced by Exploring Inductive Bias for Image Recognition and Beyond
Qiming Zhang 0001, Yufei Xu, Jing Zhang 0037, Dacheng Tao
Int. J. Comput. Vis.2
2023 Advancing Plain Vision Transformer Toward Remote Sensing Foundation Model
abstract
Large-scale vision foundation models have made significant progress in visual tasks on natural images, with vision transformers (ViTs) being the primary choice due to their good scalability and representation ability. However, large-scale models in remote sensing (RS) have not yet been sufficiently explored. In this article, we resort to plain ViTs with about 100 million parameters and make the first attempt to propose large vision models tailored to RS tasks and investigate how such large models perform. To handle the large sizes and objects of arbitrary orientations in RS images, we propose a new rotated varied-size window attention to replace the original full attention in transformers, which can significantly reduce the computational cost and memory footprint while learning better object representation by extracting rich context from the generated diverse windows. Experiments on detection tasks show the superiority of our model over all state-of-the-art models, achieving 81.24% mean average precision (mAP) on the DOTA-V1.0 dataset. The results of our models on downstream classification and segmentation tasks also show competitive performance compared to existing advanced methods. Further experiments show the advantages of our models in terms of computational complexity and data efficiency in transferring. The code and models will be released athttps://github.com/ViTAE-Transformer/Remote-Sensing-RVSA.
Di Wang 0023, Qiming Zhang 0001, Yufei Xu, Jing Zhang 0037, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 RegionCL: Exploring Contrastive Region Pairs for Self-supervised Representation Learning
Yufei Xu, Qiming Zhang 0001, Jing Zhang 0037, Dacheng Tao
ECCV (33)1
2022 VSA: Learning Varied-Size Window Attention in Vision Transformers
Qiming Zhang 0001, Yufei Xu, Jing Zhang 0037, Dacheng Tao
ECCV (25)2
2022 ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation
abstract
Although no specific domain knowledge is considered in the design, plain vision transformers have shown excellent performance in visual recognition tasks. However, little effort has been made to reveal the potential of such simple structures for pose estimation tasks. In this paper, we show the surprisingly good capabilities of plain vision transformers for pose estimation from various aspects, namely simplicity in model structure, scalability in model size, flexibility in training paradigm, and transferability of knowledge between models, through a simple baseline model called ViTPose. Specifically, ViTPose employs plain and non-hierarchical vision transformers as backbones to extract features for a given person instance and a lightweight decoder for pose estimation. It can be scaled up from 100M to 1B parameters by taking the advantages of the scalable model capacity and high parallelism of transformers, setting a new Pareto front between throughput and performance. Besides, ViTPose is very flexible regarding the attention type, input resolution, pre-training and finetuning strategy, as well as dealing with multiple pose tasks. We also empirically demonstrate that the knowledge of large ViTPose models can be easily transferred to small ones via a simple knowledge token. Experimental results show that our basic ViTPose model outperforms representative methods on the challenging MS COCO Keypoint Detection benchmark, while the largest model sets a new state-of-the-art. The code and models are available at https://github.com/ViTAE-Transformer/ViTPose.
Yufei Xu, Jing Zhang 0037, Qiming Zhang 0001, Dacheng Tao
NeurIPS1
2022 APT-36K: A Large-scale Benchmark for Animal Pose Estimation and Tracking
abstract
Animal pose estimation and tracking (APT) is a fundamental task for detecting and tracking animal keypoints from a sequence of video frames. Previous animal-related datasets focus either on animal tracking or single-frame animal pose estimation, and never on both aspects. The lack of APT datasets hinders the development and evaluation of video-based animal pose estimation and tracking methods, limiting the applications in real world, e.g., understanding animal behavior in wildlife conservation. To fill this gap, we make the first step and propose APT-36K, i.e., the first large-scale benchmark for animal pose estimation and tracking. Specifically, APT-36K consists of 2,400 video clips collected and filtered from 30 animal species with 15 frames for each video, resulting in 36,000 frames in total. After manual annotation and careful double-check, high-quality keypoint and tracking annotations are provided for all the animal instances. Based on APT-36K, we benchmark several representative models on the following three tracks: (1) supervised animal pose estimation on a single frame under intra- and inter-domain transfer learning settings, (2) inter-species domain generalization test for unseen animals, and (3) animal pose estimation with animal tracking. Based on the experimental results, we gain some empirical insights and show that APT-36K provides a useful animal pose estimation and tracking benchmark, offering new challenges and opportunities for future research. The code and dataset will be made publicly available at https://github.com/pandorgan/APT-36K.
Yuxiang Yang 0001, Yufei Xu, Jing Zhang 0037, Long Lan, Dacheng Tao
NeurIPS3
2022 DUT: Learning Video Stabilization by Simply Watching Unstable Videos
abstract
Previous deep learning-based video stabilizers require a large scale of paired unstable and stable videos for training, which are difficult to collect. Traditional trajectory-based stabilizers, on the other hand, divide the task into several sub-tasks and tackle them subsequently, which are fragile in textureless and occluded regions regarding the usage of hand-crafted features. In this paper, we attempt to tackle the video stabilization problem in a deep unsupervised learning manner, which borrows the divide-and-conquer idea from traditional stabilizers while leveraging the representation power of DNNs to handle the challenges in real-world scenarios. Technically, DUT is composed of a trajectory estimation stage and a trajectory smoothing stage. In the trajectory estimation stage, we first estimate the motion of keypoints, initialize and refine the motion of grids via a novel multi-homography estimation strategy and a motion refinement network, respectively, and get the grid-based trajectories via temporal association. In the trajectory smoothing stage, we devise a novel network to predict dynamic smoothing kernels for trajectory smoothing, which can well adapt to trajectories with different dynamic patterns. We exploit the spatial and temporal coherence of keypoints and grid vertices to formulate the training objectives, resulting in an unsupervised training scheme. Experiment results on public benchmarks show that DUT outperforms state-of-the-art methods both qualitatively and quantitatively. The source code is available at https://github.com/Annbless/DUTCode.
Yufei Xu, Jing Zhang 0037, Stephen J. Maybank, Dacheng Tao
IEEE Trans. Image Process.1
2021 Out-of-boundary View Synthesis Towards Full-Frame Video Stabilization
abstract
Warping-based video stabilizers smooth camera trajectory by constraining each pixel's displacement and warp stabilized frames from unstable ones accordingly. However, since the view outside the boundary is not available during warping, the resulting holes around the boundary of the stabilized frame must be discarded (i.e., cropping) to maintain visual consistency, and thus does leads to a tradeoff between stability and cropping ratio. In this paper, we make a first attempt to address this issue by proposing a new Out-of-boundary View Synthesis (OVS) method. By the nature of spatial coherence between adjacent frames and within each frame, OVS extrapolates the out-of-boundary view by aligning adjacent frames to each reference one. Technically, it first calculates the optical flow and propagates it to the outer boundary region according to the affinity, and then warps pixels accordingly. OVS can be integrated into existing warping-based stabilizers as a plug-and-play module to significantly improve the cropping ratio of the stabilized results. In addition, stability is improved because the jitter amplification effect caused by cropping and resizing is reduced. Experimental results on the NUS benchmark show that OVS can improve the performance of five representative state-of-the-art methods in terms of objective metrics and subjective visual quality.1
Yufei Xu, Jing Zhang 0037, Dacheng Tao
ICCV1
2021 ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
abstract
Transformers have shown great potential in various computer vision tasks owing to their strong capability in modeling long-range dependency using the self-attention mechanism. Nevertheless, vision transformers treat an image as 1D sequence of visual tokens, lacking an intrinsic inductive bias (IB) in modeling local visual structures and dealing with scale variance. Alternatively, they require large-scale training data and longer training schedules to learn the IB implicitly. In this paper, we propose a new Vision Transformer Advanced by Exploring intrinsic IB from convolutions, i.e., ViTAE. Technically, ViTAE has several spatial pyramid reduction modules to downsample and embed the input image into tokens with rich multi-scale context by using multiple convolutions with different dilation rates. In this way, it acquires an intrinsic scale invariance IB and is able to learn robust feature representation for objects at various scales. Moreover, in each transformer layer, ViTAE has a convolution block in parallel to the multi-head self-attention module, whose features are fused and fed into the feed-forward network. Consequently, it has the intrinsic locality IB and is able to learn local features and global dependencies collaboratively. Experiments on ImageNet as well as downstream tasks prove the superiority of ViTAE over the baseline transformer and concurrent works. Source code and pretrained models will be available at https://github.com/Annbless/ViTAE.
Yufei Xu, Qiming Zhang 0001, Jing Zhang 0037, Dacheng Tao
NeurIPS1
2012 Design of fault tolerant wireless sensor networks satisfying survivability and lifetime requirements
Ataul Bari, Arunita Jaekel, Jin Jiang 0001, Yufei Xu
Comput. Commun.4
2010 H∞ filter design for a class of networked control systems via T-S fuzzy model approach
abstract
This paper is concerned with H∞filter design for a class of networked control systems (NCSs) with multiple state-delays via Takagi-Sugeno (T-S) fuzzy model. The transfer delays and packet loss which are induced by the limited bandwidth of communication networks, are considered. The focus of this paper is on the analysis and design of a full-order H∞filter such that the filtering error dynamics is stochastically stable and a prescribed H∞attenuation level is guaranteed. Sufficient conditions are established for the existence of the desired filter in terms of linear matrix inequalities (LMIs). An example is given to illustrate the effectiveness and applicability of the proposed design method.
Zehui Mao, Bin Jiang 0001, Yufei Xu
FUZZ-IEEE3
2010 Adaptive Fault-Tolerant Tracking Control of Near-Space Vehicle Using Takagi-Sugeno Fuzzy Models
abstract
Based on the adaptive-control technique, this paper deals with the problem of fault-tolerant tracking control for near-space-vehicle (NSV) attitude dynamics. First, Takagi–Sugeno (T–S) fuzzy models are used to describe the NSV attitude dynamics; then, an actuator-fault model is developed. Next, an adaptive fault-tolerant tracking-control scheme is proposed based on the online estimation of actuator faults, in which a compensation control term is introduced in order to reduce the effect of actuator faults. Compared with some existing results of fault-tolerant control (FTC) in nonlinear systems, the technique presented in this paper is not dependent on fault detection and isolation (FDI) mechanism and is easy to implement in aerospace-engineering applications. Finally, simulation results are given to illustrate the effectiveness and potential of the proposed FTC scheme.
Bin Jiang 0001, Zhifeng Gao, Peng Shi 0001, Yufei Xu
IEEE Trans. Fuzzy Syst.4
2008 Integrated Placement and Routing of Relay Nodes for Fault-Tolerant Hierarchical Sensor Networks
abstract
Two-tiered sensor networks have gained popularity in recent years, due to their ability to facilitate load-balanced data gathering, fault-tolerance as well as increased network connectivity and coverage. Using higher-powered relay nodes as cluster heads can lead to further improvements in network performance. It is important to determine an appropriate placement scheme for such relay nodes, in order to achieve specified coverage and connectivity requirements with as few relay nodes as possible. A significant amount of work has been done in this area in recent years. However, existing placement strategies typically do not consider energy dissipation due to routing and are not capable of optimizing the routing scheme and placement concurrently. In this paper, we propose an integrated integer linear program (ILP) formulation that determines the minimum number of relay nodes, along with their locations and a suitable communication strategy such that i) all sensor nodes are able to connect to at least ksrelay nodes, ii) the upper tier relay node network is at least Kr-connected and iii) the network has a guaranteed lifetime. We also present an intersection based scheme for creating the initial set of potential relay node positions, which are used by our ILP, and evaluate its performance under different conditions. Experimental results on networks with hundreds of sensor nodes show that our approach leads to significant improvements over existing energy-unaware placement schemes.
Ataul Bari, Yufei Xu, Arunita Jaekel
ICCCN2