EDBT 2026 Demo / reviewers in the wild / expert
Jile Jiao
dblp:119/5841
· DBLP profile ↗
9ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0003-2644-717XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IPFormer: Instance Prompt-guided Transformer for Multi-modal Multi-shot Video UnderstandingabstractVideo Large Language Models (VideoLLMs), which adopt large language models for video understanding, have been demonstrated for single-shot videos. However, they usually struggle in multi-shot videos with frequent shot changes, varying camera angles, etc., which makes VideoLLMs hardly answer questions about multiple instances or shots over the whole video. We attribute this challenge to two issues: 1) the lack of multi-shot multi-instance annotations of existing datasets, and 2) the negligence of instance-aware modeling of current VideoLLMs. Therefore, we first introduce a new dataset termed MultiClip-Bench, featuring dense descriptions and question-answering pairs tailored for multi-shot and multi-instance scenarios. Moreover, since the existing VideoLLMs neglect the explicit modeling of instance-related features, we propose a novel Instance Prompt-guided Transformer, named IPFormer, to achieve instance-aware videounderstanding. In the IPFormer, we design a simple but effective instance-aware feature injection module, which encodes instance features as instance prompts via an attention-based connector. By this means, IPFormer can aggregate instance-specific information across multiple shots. Extensive experiments not only show that our dataset and model significantly improve multi-shot video understanding. but also show that our MultiClip-Bench can provide valuable training data and benchmarks for various video understanding tasks. Yujia Liang, Jile Jiao, Xuetao Feng, Xinchen Liu, Zixuan Ye |
AAAI | 2 |
| 2026 | MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and UnderstandingabstractGenerating lifelike human motions from descriptive texts has experienced remarkable research focus in recent years, propelled by the emerging requirements of digital humans. Despite impressive advances, existing approaches are often constrained by limited control modalities, task specificity, and focus solely on body motion representations. In this paper, we present MotionGPT-2, a unified Large Motion-Language Model (LMLM) that addresses these limitations. MotionGPT-2 accommodates multiple motion-relevant tasks and supports multimodal control conditions through pre-trained Large Language Models (LLMs). It quantizes multimodal inputs—such as text and single-frame poses—into discrete, LLM-interpretable tokens, seamlessly integrating them into the LLM’s vocabulary. These tokens are then organized into unified prompts, guiding the LLM to generate motion outputs through a pretraining-then-finetuning paradigm. We also show that the proposed MotionGPT-2 is highly adaptable to the challenging 3D holistic motion generation task, enabled by the innovative motion discretization framework, Part-Aware VQVAE, which facilitates fine-grained representations of body and hand movements. Extensive experiments and visualizations validate the effectiveness of our method, demonstrating the adaptability of MotionGPT-2 across motion generation, motion captioning, and generalized motion completion tasks. Wanli Ouyang, Jile Jiao, Xuetao Feng, Dan Xu 0002, Shixiang Tang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Mamba-3VL: Taming State Space Model for 3D Vision Language Learning
Zhongang Qi, Jile Jiao, Xuetao Feng, Yujia Liang, Ying Shan |
ICCV | 5 |
| 2023 | Dual-Tuning: Joint Prototype Transfer and Structure Regularization for Compatible Feature LearningabstractVisual retrieval system faces frequent model update and deployment. It is a heavy workload to re-extract features of the whole database every time. Feature compatibility enables the learned new visual features to be directly compared with the old features stored in the database. In this way, when updating the deployed model, we can bypass the inflexible and time-consuming feature re-extraction process. However, the old feature space that needs to be compatible is not ideal and faces outlier samples. Besides, the new and old models may be supervised by different losses, which will further causes distribution discrepancy problem between these two feature spaces. In this article, we propose a global optimization Dual-Tuning method to obtain feature compatibility against different networks and losses. A feature-level prototype loss is proposed to explicitly align two types of embedding features, by transferring global prototype information. Furthermore, we design a component-level mutual structural regularization to implicitly optimize the feature intrinsic structure. Experiments are conducted on six datasets, including person ReID datasets, face recognition datasets, and million-scale ImageNet and Place365. Experimental results demonstrate that our Dual-Tuning is able to obtain feature compatibility without sacrificing performance. Jile Jiao, Yihang Lou, Shengsen Wu, Jun Liu 0036, Xuetao Feng, Ling-Yu Duan |
IEEE Trans. Multim. | 2 |
| 2023 | Consistent Discrepancy Learning for Intra-Camera Supervised Person Re-IdentificationabstractSince annotating pedestrians across different views is extremely costly, intra-camera supervised person re-identification (ReID) aims to learn a ReID model from the intra-view labeled data. Under this setting, the most challenge lies in learning a view-invariant feature embedding in the absence of the cross-view annotations. Previous works focus on assigning a pseudo identity label for each image based on the feature similarity and learn view-invariant features by classification loss. However, because of the cross-view variations in lighting, background, etc., the pseudo labels are often noisy, and therefore not reliable for classification. In this paper, we explore learning a consistent discrepancy for pairwise images. Our main idea is that the discrepancy between pedestrian images should be consistent across different views regardless of view change so that it mainly depicts the identity difference. Due to the lack of cross-view annotations, we project images into different views and obtain likelihood prototypes for cross-view learning. These likelihood prototypes are used to measure the discrepancies between pairwise images under different views. And then, we propose an intra-view discrepancy preservation module to enforce the discrepancy to be view-consistent so as to encourage the model to distinguish the images based on the identities regardless of view change. Extensive experiments on multiple datasets show that our method outperforms existing related methods by clear margins and our method is comparable to supervised counterparts. Code will be made publicly available. Yi-Xing Peng, Jile Jiao, Xuetao Feng, Wei-Shi Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Person30K: A Dual-Meta Generalization Network for Person Re-IdentificationabstractRecently, person re-identification (ReID) has vastly benefited from the surging waves of data-driven methods. However, these methods are still not reliable enough for real-world deployments, due to the insufficient generalization capability of the models learned on existing benchmarks that have limitations in multiple aspects, including limited data scale, capture condition variations, and appearance diversities. To this end, we collect a new dataset named Person30K with the following distinct features: 1) a very large scale containing 1.38 million images of 30K identities, 2) a large capture system containing 6,497 cameras deployed at 89 different sites, 3) abundant sample diversities including varied backgrounds and diverse person poses. Furthermore, we propose a domain generalization ReID method, dual-meta generalization network (DMG-Net), to exploit the merits of meta-learning in both the training procedure and the metric space learning. Concretely, we design a "learning then generalization evaluation" metatraining procedure and a meta-discrimination loss to enhance model generalization and discrimination capabilities. Comprehensive experiments validate the effectiveness of our DMG-Net. Jile Jiao, Ce Wang 0007, Jun Liu 0036, Yihang Lou, Xuetao Feng, Ling-Yu Duan |
CVPR | 2 |
| 2021 | Occluded Person Re-Identification with Single-scale Global RepresentationsabstractOccluded person re-identification (ReID) aims at re-identifying occluded pedestrians from occluded or holistic images taken across multiple cameras. Current state-of-the-art (SOTA) occluded ReID models rely on some auxiliary modules, including pose estimation, feature pyramid and graph matching modules, to learn multi-scale and/or part-level features to tackle the occlusion challenges. This unfortunately leads to complex ReID models that (i) fail to generalize to challenging occlusions of diverse appearance, shape or size, and (ii) become ineffective in handling non-occluded pedestrians. However, real-world ReID applications typically have highly diverse occlusions and involve a hybrid of occluded and non-occluded pedestrians. To address these two issues, we introduce a novel ReID model that learns discriminative single-scale global-level pedestrian features by enforcing a novel exponentially sensitive yet bounded distance loss on occlusion-based augmented data. We show for the first time that learning single-scale global features without using these auxiliary modules is able to outperform the SOTA multi-scale and/or part-level feature-based models. Further, our simple model can achieve new SOTA performance in both occluded and non-occluded ReID, as shown by extensive results on three occluded and two general ReID benchmarks. Additionally, we create a large-scale occluded person ReID dataset with various occlusions in different scenes, which is significantly larger and contains more diverse occlusions and pedestrian dressings than existing occluded ReID datasets, providing a more faithful occluded ReID benchmark. The dataset is available at: https://git.io/OPReID Guansong Pang, Jile Jiao, Xiao Bai 0001, Xuetao Feng, Chunhua Shen |
ICCV | 3 |
| 2021 | BV-Person: A Large-scale Dataset for Bird-view Person Re-identificationabstractPerson Re-IDentification (ReID) aims at re-identifying persons from non-overlapping cameras. Existing person ReID studies focus on horizontal-view ReID tasks, in which the person images are captured by the cameras from a (nearly) horizontal view. In this work we introduce a new ReID task, bird-view person ReID, which aims at searching for a person in a gallery of horizontal-view images with the query images taken from a bird's-eye view, i.e., an elevated view of an object from above. The task is important because there are a large number of video surveillance cameras capturing persons from such an elevated view at public places. However, it is a challenging task in that the images from the bird view (i) provide limited person appearance information and (ii) have a large discrepancy compared to the persons in the horizontal view. We aim to facilitate the development of person ReID from this line by introducing a large-scale real-world dataset for this task. The proposed dataset, named BV-Person, contains 114k images of 18k identities in which nearly 20k images of 7.4k identities are taken from the bird's-eye view. We further introduce a novel model for this new ReID task. Large-scale experiments are performed to evaluate our model and 11 current state-of-the-art ReID models on BV-Person to establish performance benchmarks from multiple perspectives. The empirical results show that our model consistently and substantially outperforms the state-of-the-art models on all five datasets derived from BV-Person. Our model also achieves state-of-the-art performance on two general ReID datasets. The BV-Person dataset is available at: https://git.io/BVPerson Guansong Pang, Lei Wang 0001, Jile Jiao, Xuetao Feng, Chunhua Shen |
ICCV | 4 |
| 2017 | A Convolutional Neural Network Based Two-Stage Document DeblurringabstractBlurring often happens when capturing documents with hand held cameras, which has negative effects on the Optical Character Recognition systems. In this paper, we propose a Convolutional Neural Network (CNN) based two-stage deblurring method. The method can deal with both real motion blur and focal blur situations, while it does not require exact estimation of the blur kernel. To achieve this, the whole blur kernel space is divided into several degradative sub-spaces. Firstly, a CNN classifier is trained to predict which sub-space the blurry image belongs to at the patch level. Then, several patches voting for the specific blur kernel sub-space is developed. Given the strong learning ability of CNN, only one CNN model corresponding to a degradative kernel sub-space is trained to restore the sharp images in the image restoration step. Experimental results show that the proposed approach performs well on the real blurring document images. In addition, we demonstrate that the proposed method could also handle the spatially-varying blurring. Jile Jiao, Jun Sun 0004, Satoshi Naoi |
ICDAR | 1 |