EDBT 2026 Demo / reviewers in the wild / expert
Xuetao Feng
dblp:24/6348
· DBLP profile ↗
17ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0002-1309-8954ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 9 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IPFormer: Instance Prompt-guided Transformer for Multi-modal Multi-shot Video UnderstandingabstractVideo Large Language Models (VideoLLMs), which adopt large language models for video understanding, have been demonstrated for single-shot videos. However, they usually struggle in multi-shot videos with frequent shot changes, varying camera angles, etc., which makes VideoLLMs hardly answer questions about multiple instances or shots over the whole video. We attribute this challenge to two issues: 1) the lack of multi-shot multi-instance annotations of existing datasets, and 2) the negligence of instance-aware modeling of current VideoLLMs. Therefore, we first introduce a new dataset termed MultiClip-Bench, featuring dense descriptions and question-answering pairs tailored for multi-shot and multi-instance scenarios. Moreover, since the existing VideoLLMs neglect the explicit modeling of instance-related features, we propose a novel Instance Prompt-guided Transformer, named IPFormer, to achieve instance-aware videounderstanding. In the IPFormer, we design a simple but effective instance-aware feature injection module, which encodes instance features as instance prompts via an attention-based connector. By this means, IPFormer can aggregate instance-specific information across multiple shots. Extensive experiments not only show that our dataset and model significantly improve multi-shot video understanding. but also show that our MultiClip-Bench can provide valuable training data and benchmarks for various video understanding tasks. Yujia Liang, Jile Jiao, Xuetao Feng, Xinchen Liu, Zixuan Ye |
AAAI | 3 |
| 2026 | MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and UnderstandingabstractGenerating lifelike human motions from descriptive texts has experienced remarkable research focus in recent years, propelled by the emerging requirements of digital humans. Despite impressive advances, existing approaches are often constrained by limited control modalities, task specificity, and focus solely on body motion representations. In this paper, we present MotionGPT-2, a unified Large Motion-Language Model (LMLM) that addresses these limitations. MotionGPT-2 accommodates multiple motion-relevant tasks and supports multimodal control conditions through pre-trained Large Language Models (LLMs). It quantizes multimodal inputs—such as text and single-frame poses—into discrete, LLM-interpretable tokens, seamlessly integrating them into the LLM’s vocabulary. These tokens are then organized into unified prompts, guiding the LLM to generate motion outputs through a pretraining-then-finetuning paradigm. We also show that the proposed MotionGPT-2 is highly adaptable to the challenging 3D holistic motion generation task, enabled by the innovative motion discretization framework, Part-Aware VQVAE, which facilitates fine-grained representations of body and hand movements. Extensive experiments and visualizations validate the effectiveness of our method, demonstrating the adaptability of MotionGPT-2 across motion generation, motion captioning, and generalized motion completion tasks. Wanli Ouyang, Jile Jiao, Xuetao Feng, Dan Xu 0002, Shixiang Tang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Mamba-3VL: Taming State Space Model for 3D Vision Language Learning
Zhongang Qi, Jile Jiao, Xuetao Feng, Yujia Liang, Ying Shan |
ICCV | 6 |
| 2024 | Visual-Augmented Dynamic Semantic Prototype for Generative Zero-Shot LearningabstractGenerative Zero-shot learning (ZSL) learns a generator to synthesize visual samples for unseen classes, which is an effective way to advance ZSL. However, existing generative methods rely on the conditions of Gaussian noise and the predefined semantic prototype, which limit the generator only optimized on specific seen classes rather than characterizing each visual instance, resulting in poor generalizations (e.g., overfitting to seen classes). To address this issue, we propose a novel Visual-Augmented Dynamic Semantic prototype method (termed VADS) to boost the generator to learn accurate semantic-visual mapping by fully exploiting the visual-augmented knowledge into semantic conditions. In detail, VADS consists of two modules: (1) Visual-aware Domain Knowledge Learning module (VDKL) learns the local bias and global prior of the visual features (referred to as domain visual knowledge), which replace pure Gaussian noise to provide richer prior noise information; (2) VisionOriented Semantic Updation module (VOSU) updates the semantic prototype according to the visual representations of the samples. Ultimately, we concatenate their output as a dynamic semantic prototype, which serves as the condition of the generator. Extensive experiments demonstrate that our VADS achieves superior CZSL and GZSL performances on three prominent datasets and outperforms other state-of-the-art methods with averaging increases by 6.4%, 5.9% and 4.2% on SUN, CUB and AWA2, respectively. Wenjin Hou, Shiming Chen 0002, Shuhuang Chen, Ziming Hong, Xuetao Feng, Salman Khan 0001, Fahad Shahbaz Khan, Xinge You |
CVPR | 6 |
| 2023 | Dual-Tuning: Joint Prototype Transfer and Structure Regularization for Compatible Feature LearningabstractVisual retrieval system faces frequent model update and deployment. It is a heavy workload to re-extract features of the whole database every time. Feature compatibility enables the learned new visual features to be directly compared with the old features stored in the database. In this way, when updating the deployed model, we can bypass the inflexible and time-consuming feature re-extraction process. However, the old feature space that needs to be compatible is not ideal and faces outlier samples. Besides, the new and old models may be supervised by different losses, which will further causes distribution discrepancy problem between these two feature spaces. In this article, we propose a global optimization Dual-Tuning method to obtain feature compatibility against different networks and losses. A feature-level prototype loss is proposed to explicitly align two types of embedding features, by transferring global prototype information. Furthermore, we design a component-level mutual structural regularization to implicitly optimize the feature intrinsic structure. Experiments are conducted on six datasets, including person ReID datasets, face recognition datasets, and million-scale ImageNet and Place365. Experimental results demonstrate that our Dual-Tuning is able to obtain feature compatibility without sacrificing performance. Jile Jiao, Yihang Lou, Shengsen Wu, Jun Liu 0036, Xuetao Feng, Ling-Yu Duan |
IEEE Trans. Multim. | 6 |
| 2023 | Consistent Discrepancy Learning for Intra-Camera Supervised Person Re-IdentificationabstractSince annotating pedestrians across different views is extremely costly, intra-camera supervised person re-identification (ReID) aims to learn a ReID model from the intra-view labeled data. Under this setting, the most challenge lies in learning a view-invariant feature embedding in the absence of the cross-view annotations. Previous works focus on assigning a pseudo identity label for each image based on the feature similarity and learn view-invariant features by classification loss. However, because of the cross-view variations in lighting, background, etc., the pseudo labels are often noisy, and therefore not reliable for classification. In this paper, we explore learning a consistent discrepancy for pairwise images. Our main idea is that the discrepancy between pedestrian images should be consistent across different views regardless of view change so that it mainly depicts the identity difference. Due to the lack of cross-view annotations, we project images into different views and obtain likelihood prototypes for cross-view learning. These likelihood prototypes are used to measure the discrepancies between pairwise images under different views. And then, we propose an intra-view discrepancy preservation module to enforce the discrepancy to be view-consistent so as to encourage the model to distinguish the images based on the identities regardless of view change. Extensive experiments on multiple datasets show that our method outperforms existing related methods by clear margins and our method is comparable to supervised counterparts. Code will be made publicly available. Yi-Xing Peng, Jile Jiao, Xuetao Feng, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Neural Surface Reconstruction of Dynamic Scenes with Monocular RGB-D CameraabstractWe propose Neural-DynamicReconstruction (NDR), a template-free method to recover high-fidelity geometry and motions of a dynamic scene from a monocular RGB-D camera. In NDR, we adopt the neural implicit function for surface representation and rendering such that the captured color and depth can be fully utilized to jointly optimize the surface and deformations. To represent and constrain the non-rigid deformations, we propose a novel neural invertible deforming network such that the cycle consistency between arbitrary two frames is automatically satisfied. Considering that the surface topology of dynamic scene might change over time, we employ a topology-aware strategy to construct the topology-variant correspondence for the fused frames. NDR also further refines the camera poses in a global optimization manner. Experiments on public datasets and our collected dataset demonstrate that NDR outperforms existing monocular dynamic reconstruction methods. Hongrui Cai, Wanquan Feng, Xuetao Feng, Juyong Zhang |
NeurIPS | 3 |
| 2021 | Person30K: A Dual-Meta Generalization Network for Person Re-IdentificationabstractRecently, person re-identification (ReID) has vastly benefited from the surging waves of data-driven methods. However, these methods are still not reliable enough for real-world deployments, due to the insufficient generalization capability of the models learned on existing benchmarks that have limitations in multiple aspects, including limited data scale, capture condition variations, and appearance diversities. To this end, we collect a new dataset named Person30K with the following distinct features: 1) a very large scale containing 1.38 million images of 30K identities, 2) a large capture system containing 6,497 cameras deployed at 89 different sites, 3) abundant sample diversities including varied backgrounds and diverse person poses. Furthermore, we propose a domain generalization ReID method, dual-meta generalization network (DMG-Net), to exploit the merits of meta-learning in both the training procedure and the metric space learning. Concretely, we design a "learning then generalization evaluation" metatraining procedure and a meta-discrimination loss to enhance model generalization and discrimination capabilities. Comprehensive experiments validate the effectiveness of our DMG-Net. Jile Jiao, Ce Wang 0007, Jun Liu 0036, Yihang Lou, Xuetao Feng, Ling-Yu Duan |
CVPR | 6 |
| 2021 | Occluded Person Re-Identification with Single-scale Global RepresentationsabstractOccluded person re-identification (ReID) aims at re-identifying occluded pedestrians from occluded or holistic images taken across multiple cameras. Current state-of-the-art (SOTA) occluded ReID models rely on some auxiliary modules, including pose estimation, feature pyramid and graph matching modules, to learn multi-scale and/or part-level features to tackle the occlusion challenges. This unfortunately leads to complex ReID models that (i) fail to generalize to challenging occlusions of diverse appearance, shape or size, and (ii) become ineffective in handling non-occluded pedestrians. However, real-world ReID applications typically have highly diverse occlusions and involve a hybrid of occluded and non-occluded pedestrians. To address these two issues, we introduce a novel ReID model that learns discriminative single-scale global-level pedestrian features by enforcing a novel exponentially sensitive yet bounded distance loss on occlusion-based augmented data. We show for the first time that learning single-scale global features without using these auxiliary modules is able to outperform the SOTA multi-scale and/or part-level feature-based models. Further, our simple model can achieve new SOTA performance in both occluded and non-occluded ReID, as shown by extensive results on three occluded and two general ReID benchmarks. Additionally, we create a large-scale occluded person ReID dataset with various occlusions in different scenes, which is significantly larger and contains more diverse occlusions and pedestrian dressings than existing occluded ReID datasets, providing a more faithful occluded ReID benchmark. The dataset is available at: https://git.io/OPReID Guansong Pang, Jile Jiao, Xiao Bai 0001, Xuetao Feng, Chunhua Shen |
ICCV | 5 |
| 2021 | BV-Person: A Large-scale Dataset for Bird-view Person Re-identificationabstractPerson Re-IDentification (ReID) aims at re-identifying persons from non-overlapping cameras. Existing person ReID studies focus on horizontal-view ReID tasks, in which the person images are captured by the cameras from a (nearly) horizontal view. In this work we introduce a new ReID task, bird-view person ReID, which aims at searching for a person in a gallery of horizontal-view images with the query images taken from a bird's-eye view, i.e., an elevated view of an object from above. The task is important because there are a large number of video surveillance cameras capturing persons from such an elevated view at public places. However, it is a challenging task in that the images from the bird view (i) provide limited person appearance information and (ii) have a large discrepancy compared to the persons in the horizontal view. We aim to facilitate the development of person ReID from this line by introducing a large-scale real-world dataset for this task. The proposed dataset, named BV-Person, contains 114k images of 18k identities in which nearly 20k images of 7.4k identities are taken from the bird's-eye view. We further introduce a novel model for this new ReID task. Large-scale experiments are performed to evaluate our model and 11 current state-of-the-art ReID models on BV-Person to establish performance benchmarks from multiple perspectives. The empirical results show that our model consistently and substantially outperforms the state-of-the-art models on all five datasets derived from BV-Person. Our model also achieves state-of-the-art performance on two general ReID datasets. The BV-Person dataset is available at: https://git.io/BVPerson Guansong Pang, Lei Wang 0001, Jile Jiao, Xuetao Feng, Chunhua Shen |
ICCV | 5 |
| 2021 | High-Performance Discriminative Tracking with TransformersabstractEnd-to-end discriminative trackers improve the state of the art significantly, yet the improvement in robustness and efficiency is restricted by the conventional discriminative model, i.e., least-squares based regression. In this paper, we present DTT, a novel single-object discriminative tracker, based on an encoder-decoder Transformer architecture. By self- and encoder-decoder attention mechanisms, our approach is able to exploit the rich scene information in an end-to-end manner, effectively removing the need for hand-designed discriminative models. In online tracking, given a new test frame, dense prediction is performed at all spatial positions. Not only location, but also bounding box of the target object is obtained in a robust fashion, streamlining the discriminative tracking pipeline. DTT is conceptually simple and easy to implement. It yields state-of-the-art performance on four popular benchmarks including GOT-10k, LaSOT, NfS, and TrackingNet while running at over 50 FPS, confirming its effectiveness and efficiency. We hope DTT may provide a new perspective for single-object visual tracking. Ming Tang 0001, Linyu Zheng, Guibo Zhu, Jinqiao Wang, Xuetao Feng, Hanqing Lu |
ICCV | 7 |
| 2021 | Multi-initialization Optimization Network for Accurate 3D Human Pose and Shape Estimationabstract3D human pose and shape recovery from a monocular RGB image is a challenging task. Existing learning based methods highly depend on weak supervision signals, e.g. 2D and 3D joint location, due to the lack of in-the-wild paired 3D supervision. However, considering the 2D-to-3D ambiguities existed in these weak supervision labels, the network is easy to get stuck in local optima when trained with such labels. In this paper, we reduce the ambituity by optimizing multiple initializations. Specifically, we propose a three-stage framework named Multi-Initialization Optimization Network (MION). In the first stage, we strategically select different coarse 3D reconstruction candidates which are compatible with the 2D keypoints of input sample. Each coarse reconstruction can be regarded as an initialization leads to one optimization branch. In the second stage, we design a mesh refinement transformer (MRT) to respectively refine each coarse reconstruction result via a self-attention mechanism. Finally, a Consistency Estimation Network (CEN) is proposed to find the best result from mutiple candidates by evaluating if the visual evidence in RGB image matches a given 3D reconstruction. Experiments demonstrate that our Multi-Initialization Optimization Network outperforms existing 3D mesh based methods on multiple public benchmarks. Zhiwei Liu 0004, Xiangyu Zhu 0001, Lu Yang 0006, Ming Tang 0001, Zhen Lei 0001, Guibo Zhu, Xuetao Feng, Yan Wang 0068, Jinqiao Wang |
ACM Multimedia | 8 |
| 2015 | Robust pose normalization for face recognition under varying viewsabstractUnconstrained face recognition under varying views is one of the most challenging tasks, since the difference in appearances caused by poses may be even larger than that due to identity. In this paper, we exploit and analyze a novel pose normalization scheme for facial images under varying views via robust 3D shape reconstruction from single, unconstrained photos in the wild. Specifically, to address the problem of ambiguous 2D-to-3D landmark correspondence and imperfect landmark detector, for each input 2D face, the 3D shape is suggested to be learned by iteratively refining the 3D landmarks and the weighting coefficients of each landmark. Experimental results on both LFW and a large-scale self-collected face databases demonstrate that the proposed approach performs better than the existing representative technologies. Xuetao Feng, Lujin Gong, Wonjun Hwang, Jae-Joon Han |
ICIP | 2 |
| 2015 | Combining nonuniform sampling, hybrid super vector, and random forest with discriminative decision trees for action recognitionabstractTrajectory-based features have become popular for action recognition and achieve the state-of-the-art results on a variety of datasets. In this paper, we propose a novel framework to improve the performance of action recognition. Specifically, we first apply the nonuniform sampling method to efficiently select features for given actions. The proposed hybrid super vector, namely fisher vector (FV) combined with vector of locally aggregated descriptors (VLAD), is then employed to encode sampled trajectories. A random forest with discriminative decision trees, where every tree node is a discriminative classifier, is finally applied to predict action labels. We have achieved 88.2% in average accuracy on the UCF101 dataset, which outperforms the best results that have been reported in the literature. Kuanhong Xu, Ya Lu, Xuetao Feng, Jae-Joon Han |
ICIP | 4 |
| 2011 | Robust facial expression tracking based on composite constraints AAMabstractFacial expression tracking is a challenging task because the head pose may change in a large range and the expression is highly non-rigid. It can be formulized as an energy minimization problem. The two most important issues are the construction of the cost function and the selection of the initial value. In this paper, we present a fast and robust expression tracking algorithm called Composite Constraints AAM. Firstly, a novel cost function is proposed to enhance the convergence by combining multiple constraints in a unified framework. Secondly, widely used local features are strictly tested with face videos, and an efficient motion estimation method is presented to provide a good initial value to the iterative optimization process. Experimental result demonstrates that our system can track the head pose and facial expression with very high stability in real time speed. Xuetao Feng, Xiaolu Shen, Mingcai Zhou, Jung-Bae Kim |
ICIP | 1 |
| 2008 | On edge structure based adaptive observation model for facial feature trackingabstractFacial feature tracking is a crucial and challenging task in computer vision. Recently online-learning methods have become increasingly popular on account of their strong ability to adapt to variations and have achieved good results in tracking. However, all previous work used only raw intensity to build the model, which is very sensitive to condition changes. In this work, we present a real time, fully automatic facial feature detection and tracking approach using adaptive observation models based on edge structure, which is more reliable especially when the lighting state alters during tracking. Experimental results demonstrate that using edge map measures in observation modeling can improve the accuracy and robustness of tracking. Yangsheng Wang, Xuetao Feng, Mingcai Zhou |
ICPR | 3 |
| 2006 | Boosting Gabor Feature Classifier for Face Recognition Using Random SubspaceabstractGabor feature has been widely viewed as a good representation method for face recognition. AdaBoost is an excellent machine learning technique. Learning Gabor feature based classifier using AdaBoost is one of the best face recognition algorithms. However, dimensionality of Gabor feature space usually is very high, which makes the training program need huge memory or else take a very long time to run. In this paper, we propose a method which not only can solve the problem but also can improve recognition accuracy. Several subspaces with moderate size are randomly generated from original high dimensional Gabor feature space. Then strong classifier is trained in every random subspace (T. Kam Ho, 1998) respectively and the outputs of multiple classifiers are combined in the final decision. Experimental results demonstrate that the method saves a great amount of training time, and achieves an exciting recognition rate of 97.91% on the FERET Fb test set Yong Gao 0002, Yangsheng Wang, Xuetao Feng, Xiaoxu Zhou |
ICASSP (2) | 3 |