EDBT 2026 Demo / reviewers in the wild / expert
Xinchen Liu
dblp:36/2028
· DBLP profile ↗
38ranked-venue papers
9as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | IPFormer: Instance Prompt-guided Transformer for Multi-modal Multi-shot Video UnderstandingabstractVideo Large Language Models (VideoLLMs), which adopt large language models for video understanding, have been demonstrated for single-shot videos. However, they usually struggle in multi-shot videos with frequent shot changes, varying camera angles, etc., which makes VideoLLMs hardly answer questions about multiple instances or shots over the whole video. We attribute this challenge to two issues: 1) the lack of multi-shot multi-instance annotations of existing datasets, and 2) the negligence of instance-aware modeling of current VideoLLMs. Therefore, we first introduce a new dataset termed MultiClip-Bench, featuring dense descriptions and question-answering pairs tailored for multi-shot and multi-instance scenarios. Moreover, since the existing VideoLLMs neglect the explicit modeling of instance-related features, we propose a novel Instance Prompt-guided Transformer, named IPFormer, to achieve instance-aware videounderstanding. In the IPFormer, we design a simple but effective instance-aware feature injection module, which encodes instance features as instance prompts via an attention-based connector. By this means, IPFormer can aggregate instance-specific information across multiple shots. Extensive experiments not only show that our dataset and model significantly improve multi-shot video understanding. but also show that our MultiClip-Bench can provide valuable training data and benchmarks for various video understanding tasks. Yujia Liang, Jile Jiao, Xuetao Feng, Xinchen Liu, Zixuan Ye |
AAAI | 4 |
| 2026 | JanusMM: A Benchmark for Self-Deprecation Understanding in Real-World Multimodal ConversationsabstractXinyi Xu, Bingguang Hao, Yongyi Xiong, Zimo Chen, Xinchen Liu, Hongxin Guo, Xuelong Wang, Silin Zhou, Shihan Dou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Bingguang Hao, Yongyi Xiong, Zimo Chen, Xinchen Liu, Hongxin Guo, Xuelong Wang, Silin Zhou, Shihan Dou |
ACL (1) | 5 |
| 2026 | DSA: A Dual-Safety Approach for Test-Time Temporal Risk Removal of Video Generation
Qi Liu 0081, Kun Liu 0016, Xinchen Liu, Xiaoyan Gu 0001, Zhaochun Ren, Yongdong Zhang 0001, Wu Liu 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video GenerationabstractText-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale videos with accurate captions for HOI. To address this issue, we introduce HOIGen-1M, the first large-scale dataset for HOI Generation, consisting of over one million high-quality videos collected from diverse sources. In particular, to guarantee the high quality of videos, we first design an efficient framework to automatically curate HOI videos using the powerful multimodal large language models (MLLMs), and then the videos are further cleaned by human annotators. Moreover, to obtain accurate textual captions for HOI videos, we design a novel video description method based on a Mixture-of-Multimodal-Experts (MoME) strategy that not only generates expressive captions but also eliminates the hallucination by individual MLLM. Further-more, due to the lack of an evaluation framework for gen-erated HOI videos, we propose two new metrics to assess the quality of generated videos in a coarse-to-fine manner. Extensive experiments reveal that current T2V models struggle to generate high-quality HOI videos and confirm that our HOIGen-1M dataset is instrumental for improving HOI video generation. Kun Liu 0016, Qi Liu 0081, Xinchen Liu, Yongdong Zhang 0001, Jiebo Luo 0001, Xiaodong He 0001, Wu Liu 0005 |
CVPR | 3 |
| 2025 | Scaling Down Text Encoders of Text-to-Image Diffusion ModelsabstractText encoders in diffusion models have rapidly evolved, transitioning from CLIP to T5-XXL. Although this evolution has significantly enhanced the models’ ability to understand complex prompts and generate text, it also leads to a substantial increase in the number of parameters. Despite T5 series encoders being trained on the C4 natural language corpus, which includes a significant amount of non-visual data, diffusion models with T5 encoder do not respond to those non-visual prompts, indicating redundancy in representational power. Therefore, it raises an important question: "Do we really need such a large text encoder?" In pursuit of an answer, we employ vision-based knowledge distillation to train a series of T5 encoder models. To fully inherit T5-XXL’s capabilities, we constructed our dataset based on three criteria: image quality, semantic understanding, and text-rendering. Our results demonstrate the scaling down pattern that the distilled T5-base model can generate images of comparable quality to those produced by T5-XXL, while being 50 times smaller in size. This reduction in model size significantly lowers the GPU requirements for running state-of-the-art models such as FLUX and SD3, making high-quality text-to-image generation more accessible. Our code is available at https: //lifuwang-66.github.io/ScalingDownTE/. Daqing Liu, Xinchen Liu, Xiaodong He 0001 |
CVPR | 3 |
| 2024 | HumanNeRF-SE: A Simple yet Effective Approach to Animate HumanNeRF with Diverse PosesabstractWe present HumanNeRF-SE, a simple yet effective method that synthesizes diverse novel pose images with sim-ple input. Previous HumanNeRF works require a large number of optimizable parameters to fit the human images. Instead, we reload these approaches by combining explicit and implicit human representations to design both general-ized rigid deformation and specific non-rigid deformation. Our key insight is that explicit shape can reduce the sam-pling points used to fit implicit representation, and frozen blending weights from SMPL constructing a generalized rigid deformation can effectively avoid overfitting and im-prove pose generalization performance. Our architecture involving both explicit and implicit representation is sim-ple yet effective. Experiments demonstrate our model can synthesize images under arbitrary poses with few-shot input and increase the speed of synthesizing images by 15 times through a reduction in computational complexity without using any existing acceleration modules. Compared to the state-of-the-art HumanNeRF studies, HumanNeRF-SE achieves better performance with fewer learnable parame-ters and less training time. Caoyuan Ma, Yu-Lun Liu 0001, Zhixiang Wang 0001, Wu Liu 0005, Xinchen Liu, Zheng Wang 0007 |
CVPR | 5 |
| 2024 | CLaM: An Open-Source Library for Performance Evaluation of Text-driven Human Motion GenerationabstractText-driven human motion generation, which creates motion sequences based on textual descriptions, has attracted great attention in the communities of multimedia and artificial intelligence. By parsing and comprehending textual information and converting it into specific human movements, it realizes a direct transformation from human semantics to motion sequences. New text-driven human motion generators are springing up to achieve better performance. However, the absence of well-trained evaluators that can effectively estimate the consistency between the text prompts and motions generated by existing generators remains a challenge. To address the above issues, we propose an open-source library with a powerful Contrastive Language-and-Motion (CLaM) pre-training evaluator, which can be employed for evaluating a variety of text-driven human motion generation algorithms. We perform a thorough performance evaluation of the existing algorithms on various metrics, such as R-Precision. As a by-product, we build a large-scale HumanML3D-synthesis dataset, which consists of 14,616 motion sequences and 547,102 textual descriptions, which is ten times larger than the widely-used HumanML3D dataset. The source codes and models for CLaM are available at~https://github.com/SheldongChen/CLaM/. Xiaodong Chen 0011, Kunlang He, Wu Liu 0005, Xinchen Liu, Zhengjun Zha, Tao Mei 0001 |
ACM Multimedia | 4 |
| 2024 | It Takes Two: Accurate Gait Recognition in the Wild via Cross-granularity AlignmentabstractExisting studies for gait recognition primarily utilized sequences of either binary silhouette or human parsing to encode the shapes and dynamics of persons during walking. Silhouettes exhibit accurate segmentation quality and robustness to environmental variations, but their low information entropy may result in sub-optimal performance. In contrast, human parsing provides fine-grained part segmentation with higher information entropy, but the segmentation quality may deteriorate due to the complex environments. To discover the advantages of silhouette and parsing and overcome their limitations, this paper proposes a novel cross-granularity alignment gait recognition method, named XGait, to unleash the power of gait representations of different granularity. To achieve this goal, the XGait first contains two branches of backbone encoders to map the silhouette sequences and the parsing sequences into two latent spaces, respectively. Moreover, to explore the complementary knowledge across the features of two representations, we design the Global Cross-granularity Module (GCM) and the Part Cross-granularity Module (PCM) after the two encoders. In particular, the GCM aims to enhance the quality of parsing features by leveraging global features from silhouettes, while the PCM aligns the dynamics of human parts between silhouette and parsing features using the high information entropy in parsing sequences. In addition, to effectively guide the alignment of two representations with different granularity at the part level, an elaborate-designed learnable division mechanism is proposed for the parsing features. Finally, comprehensive experiments on two large-scale gait datasets not only show the superior performance of XGait with the Rank-1 accuracy of 80.5% on Gait3D and 88.3% CCPG but also reflect the robustness of the learned features even under challenging conditions like occlusions and cloth changes Jinkai Zheng, Xinchen Liu, Boyue Zhang 0004, Chenggang Yan 0001, Jiyong Zhang 0001, Wu Liu 0005, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2024 | SPSO-Pruner: a network pruning method on YOLOv5 for fewer categories scenarios
Xinchen Liu, Linyao Xie, Hongbo Wang 0001, Osman Salem |
Multim. Tools Appl. | 2 |
| 2024 | SigFormer: Sparse Signal-guided Transformer for Multi-modal Action SegmentationabstractMulti-modal human action segmentation is a critical and challenging task with a wide range of applications. Nowadays, the majority of approaches concentrate on the fusion of dense signals (i.e., RGB, optical flow, and depth maps). However, the potential contributions of sparse IoT sensor signals, which can be crucial for achieving accurate recognition, have not been fully explored. To make up for this, we introduce a S parse s i gnal- g uided Transformer ( SigFormer ) to combine both dense and sparse signals. We employ mask attention to fuse localized features by constraining cross-attention within the regions where sparse signals are valid. However, since sparse signals are discrete, they lack sufficient information about the temporal action boundaries. Therefore, in SigFormer, we propose to emphasize the boundary information at two stages to alleviate this problem. In the first feature extraction stage, we introduce an intermediate bottleneck module to jointly learn both category and boundary features of each dense modality through the inner loss functions. After the fusion of dense modalities and sparse signals, we then devise a two-branch architecture that explicitly models the interrelationship between action category and temporal boundary. Experimental results demonstrate that SigFormer outperforms the state-of-the-art approaches on a multi-modal action segmentation dataset from real industrial environments, reaching an outstanding F1 score of 0.958. The codes and pre-trained models have been made available at https://github.com/LIUQI-creat/SigFormer . Qi Liu 0081, Xinchen Liu, Kun Liu 0016, Xiaoyan Gu 0001, Wu Liu 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Overcoming catastrophic forgetting with classifier expander
Xinchen Liu, Hongbo Wang 0001, Yingjian Tian, Linyao Xie |
ACML | 1 |
| 2023 | An Adaptive Detector for Few Shot Object Detection
Jiming Yan, Hongbo Wang 0001, Xinchen Liu |
ICONIP (1) | 3 |
| 2023 | FastReID: A Pytorch Toolbox for General Instance Re-identificationabstractGeneral Instance Re-identification is a very important task in computer vision, which can be widely used in many practical applications, such as person/vehicle re-identification, face recognition, wildlife protection, commodity tracing, snapshots, and so on. To meet the increasing application demand for general instance re-identification, we present FastReID as a widely used software system. In FastReID, the highly modular and extensible design makes it easy for the researcher to achieve new research ideas. Friendly manageable system configuration and engineering deployment functions allow practitioners to quickly deploy models into productions. We have implemented some state-of-the-art projects, including person re-id, partial re-id, cross-domain re-id, and vehicle re-id. Moreover, we plan to release these pre-trained models on multiple benchmark datasets. FastReID is by far the most general and high-performance toolbox that supports single and multiple GPU servers, it can reproduce our project results very easily. The source codes and models have been released at https://github.com/JDAI-CV/fast-reid. Lingxiao He, Xingyu Liao, Wu Liu 0005, Xinchen Liu, Peng Cheng 0002, Tao Mei 0001 |
ACM Multimedia | 4 |
| 2023 | HCMA '23: 4th International Workshop on Human-Centric Multimedia AnalysisabstractUnderstanding human interactions within diverse media contexts has emerged as a fundamental challenge. The explosive growth of multimedia data not only provides opportunities for human-centirc analysis but also increases the complexity of processing multimodal data. To address this pivotal challenge and explore its multifaceted dimensions, the Fourth International Workshop on Human-Centric Multimedia Analysis is concentrated on the tasks of human-centric analysis with multimedia and multimodal information. By delving into the nuances of human behavior within multimedia, this workshop aims to uncover novel insights, showcase innovative methodologies, and discuss future directions. With a spotlight on cutting-edge research and a focus on real-world applications, the workshop seeks to equip researchers and practitioners with the tools and knowledge to navigate the intricacies of human-centric multimedia analysis. Jingkuan Song, Wu Liu 0005, Xinchen Liu, Dingwen Zhang, Chaowei Fang, Hongyuan Zhu 0002, Wenbing Huang 0001, John R. Smith, Xin Wang 0019 |
ACM Multimedia | 3 |
| 2023 | Parsing is All You Need for Accurate Gait Recognition in the WildabstractBinary silhouettes and keypoint-based skeletons have dominated human gait recognition studies for decades since they are easy to extract from video frames. Despite their success in gait recognition for in-the-lab environments, they usually fail in real-world scenarios due to their low information entropy for gait representations. To achieve accurate gait recognition in the wild, this paper presents a novel gait representation, named Gait Parsing Sequence (GPS). GPSs are sequences of fine-grained human segmentation, i.e., human parsing, extracted from video frames, so they have much higher information entropy to encode the shapes and dynamics of fine-grained human parts during walking. Moreover, to effectively explore the capability of the GPS representation, we propose a novel human parsing-based gait recognition framework, named ParsingGait. ParsingGait contains a Convolutional Neural Network (CNN)-based backbone and two light-weighted heads. The first head extracts global semantic features from GPSs, while the other one learns mutual information of part-level features through Graph Convolutional Networks to model the detailed dynamics of human walking. Furthermore, due to the lack of suitable datasets, we build the first parsing-based dataset for gait recognition in the wild, named Gait3D-Parsing, by extending the large-scale and challenging Gait3D dataset. Based on Gait3D-Parsing, we comprehensively evaluate our method and existing gait recognition methods. Specifically, ParsingGait achieves a 17.5% Rank-1 increase compared with the state-of-the-art silhouette-based method. In addition, by replacing silhouettes with GPSs, current gait recognition methods achieve about 12.5% ~ 19.2% improvements in Rank-1 accuracy. The experimental results show a significant improvement in accuracy brought by the GPS representation and the superiority of ParsingGait. Jinkai Zheng, Xinchen Liu, Shuai Wang 0003, Chenggang Yan 0001, Wu Liu 0005 |
ACM Multimedia | 2 |
| 2022 | Gait Recognition in the Wild with Dense 3D Representations and A BenchmarkabstractExisting studies for gait recognition are dominated by 2D representations like the silhouette or skeleton of the human body in constrained scenes. However, humans live and walk in the unconstrained 3D space, so projecting the 3D human body onto the 2D plane will discard a lot of crucial information like the viewpoint, shape, and dynamics for gait recognition. Therefore, this paper aims to explore dense 3D representations for gait recognition in the wild, which is a practical yet neglected problem. In particular, we propose a novel framework to explore the 3D Skinned Multi-Person Linear (SMPL) model of the human body for gait recognition, named SMPLGait. Our framework has two elaborately-designed branches of which one extracts appearance features from silhouettes, the other learns knowledge of 3D viewpoints and shapes from the 3D SMPL model. In addition, due to the lack of suitable datasets, we build the first large-scale 3D representation-based gait recognition dataset, named Gait3D. It contains 4,000 subjects and over 25,000 sequences extracted from 39 cameras in an unconstrained indoor scene. More importantly, it provides 3D SMPL models recovered from video frames which can provide dense 3D information of body shape, viewpoint, and dynamics. Based on Gait3D, we comprehensively compare our method with existing gait recognition approaches, which reflects the superior performance of our framework and the potential of 3D representations for gait recognition in the wild. The code and dataset are available at: https://gait3d.github.io. Jinkai Zheng, Xinchen Liu, Wu Liu 0005, Lingxiao He, Chenggang Yan 0001, Tao Mei 0001 |
CVPR | 2 |
| 2022 | Part-level Action Parsing via a Pose-guided Coarse-to-Fine FrameworkabstractAction recognition from videos, i.e., classifying a video into one of the pre-defined action types, has been a popular topic in the communities of artificial intelligence, multimedia, and signal processing. However, existing methods usually consider an input video as a whole and learn models, e.g., Convolutional Neural Networks (CNNs), with coarse video-level class labels. These methods can only output an action class for the video, but cannot provide fine-grained and explainable cues to answer why the video shows a specific action. Therefore, researchers start to focus on a new task, Part-level Action Parsing (PAP), which aims to not only predict the video-level action but also recognize the frame-level fine-grained actions or interactions of body parts for each person in the video. To this end, we propose a coarse-to-fine framework for this challenging task. In particular, our framework first predicts the video-level class of the input video, then localizes the body parts and predicts the part-level action. Moreover, to balance the accuracy and computation in part-level action parsing, we propose to recognize the part-level actions by segment-level features. Furthermore, to overcome the ambiguity of body parts, we propose a pose-guided positional embedding method to accurately localize body parts. Through comprehensive experiments on a large-scale dataset, i.e., Kinetics-TPS, our framework achieves state-of-the-art performance and outperforms existing methods over 31.10% ROC score. Xiaodong Chen 0011, Xinchen Liu, Wu Liu 0005, Kun Liu 0016, Yongdong Zhang 0001, Tao Mei 0001 |
ISCAS | 2 |
| 2022 | MAPLE: Masked Pseudo-Labeling autoEncoder for Semi-supervised Point Cloud Action RecognitionabstractRecognizing human actions from point cloud videos has attracted tremendous attention from both academia and industry due to its wide applications like automatic driving, robotics, and so on. However, current methods for point cloud action recognition usually require a huge amount of data with manual annotations and a complex backbone network with high computation cost, which makes it impractical for real-world applications. Therefore, this paper considers the task of semi-supervised point cloud action recognition. We propose a Masked Pseudo-Labeling autoEncoder (MAPLE) framework to learn effective representations with much fewer annotations for point cloud action recognition. In particular, we design a novel and efficient Decoupled spatial-temporal TransFormer (DestFormer) as the backbone of MAPLE. In DestFormer, the spatial and temporal dimensions of the 4D point cloud videos are decoupled to achieve an efficient self-attention for learning both long-term and short-term features. Moreover, to learn discriminative features from fewer annotations, we design a masked pseudo-labeling autoencoder structure to guide the DestFormer to reconstruct features of masked frames from the available frames. More importantly, for unlabeled data, we exploit the pseudo-labels from the classification head as the supervision signal for the reconstruction of features from the masked frames. Finally, comprehensive experiments demonstrate that MAPLE achieves superior results on three public benchmarks and outperforms the state-of-the-art method by 8.08% accuracy on the MSR-Action3D dataset. Xiaodong Chen 0011, Wu Liu 0005, Xinchen Liu, Yongdong Zhang 0001, Jungong Han, Tao Mei 0001 |
ACM Multimedia | 3 |
| 2022 | Delving into the Frequency: Temporally Consistent Human Motion Transfer in the Fourier SpaceabstractHuman motion transfer refers to synthesizing photo-realistic and temporally coherent videos that enable one person to imitate the motion of others. However, current synthetic videos suffer from the temporal inconsistency in sequential frames that significantly degrades the video quality, yet is far from solved by existing methods in the pixel domain. Recently, some works on DeepFake detection try to distinguish the natural and synthetic images in the frequency domain because of the frequency insufficiency of image synthesizing methods. Nonetheless, there is no work to study the temporal inconsistency of synthetic videos from the aspects of the frequency-domain gap between natural and synthetic videos. Therefore, in this paper, we propose to delve into the frequency space for temporally consistent human motion transfer. First of all, we make the first comprehensive analysis of natural and synthetic videos in the frequency domain to reveal the frequency gap in both the spatial dimension of individual frames and the temporal dimension of the video. To close the frequency gap between the natural and synthetic videos, we propose a novel Frequency-based human MOtion TRansfer framework, named FreMOTR, which can effectively mitigate the spatial artifacts and the temporal inconsistency of the synthesized videos. FreMOTR explores two novel frequency-based regularization modules: 1) the Frequency-domain Appearance Regularization (FAR) to improve the appearance of the person in individual frames and 2) Temporal Frequency Regularization (TFR) to guarantee the temporal consistency between adjacent frames. Finally, comprehensive experiments demonstrate that the FreMOTR not only yields superior performance in temporal consistency metrics but also improves the frame-level visual quality of synthetic videos. In particular, the temporal consistency metrics are improved by nearly 30% than the state-of-the-art model. Guang Yang 0031, Wu Liu 0005, Xinchen Liu, Xiaoyan Gu 0001, Juan Cao 0001, Jintao Li 0001 |
ACM Multimedia | 3 |
| 2022 | REMOT: A Region-to-Whole Framework for Realistic Human Motion TransferabstractHuman Video Motion Transfer (HVMT) aims to, given an image of a source person, generate his/her video that imitates the motion of the driving person. Existing methods for HVMT mainly exploit Generative Adversarial Networks (GANs) to perform the warping operation based on the flow estimated from the source person image and each driving video frame. However, these methods always generate obvious artifacts due to the dramatic differences in poses, scales, and shifts between the source person and the driving person. To overcome these challenges, this paper presents a novel REgion-to-whole human MOtion Transfer (REMOT) framework based on GANs. To generate realistic motions, the REMOT adopts a progressive generation paradigm: it first generates each body part in the driving pose without flow-based warping, then composites all parts into a complete person of the driving motion. Moreover, to preserve the natural global appearance, we design a Global Alignment Module to align the scale and position of the source person with those of the driving person based on their layouts. Furthermore, we propose a Texture Alignment Module to keep each part of the person aligned according to the similarity of the texture. Finally, through extensive quantitative and qualitative experiments, our REMOT achieves state-of-the-art results on two public benchmarks. Quanwei Yang, Xinchen Liu, Wu Liu 0005, Hongtao Xie 0001, Xiaoyan Gu 0001, Lingyun Yu 0002, Yongdong Zhang 0001 |
ACM Multimedia | 2 |
| 2022 | HCMA'22: 3rd International Workshop on Human-Centric Multimedia AnalysisabstractThe Third International Workshop on Human-Centric Multimedia Analysis concentrates on the tasks of human-centric analysis with multimedia and multimodal information. It involves multiple tasks such as face detection and recognition, human body pattern analysis, person re-identification, human action detection, etc. Today, multiple multimedia sensing technologies and large-scale computing infrastructures are emerging at a rapid velocity a wide variety of big multi-modality data for human-centric analysis, which provides rich knowledge to help tackle these challenges. Researchers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as intelligent surveillance, retailing, fashion design, and services. Therefore, this workshop aims to provide a platform to bridge the gap between the communities of human analysis and multimedia. Dingwen Zhang, Chaowei Fang, Wu Liu 0005, Xinchen Liu, Jingkuan Song, Hongyuan Zhu 0002, Wenbing Huang 0001, John R. Smith |
ACM Multimedia | 4 |
| 2022 | Gait Recognition in the Wild with Multi-hop Temporal SwitchabstractExisting studies for gait recognition are dominated by in-the-lab scenarios. Since people live in real-world senses, gait recognition in the wild is a more practical problem that has recently attracted the attention of the community of multimedia and computer vision. Current methods that obtain state-of-the-art performance on in-the-lab benchmarks achieve much worse accuracy on the recently proposed in-the-wild datasets because these methods can hardly model the varied temporal dynamics of gait sequences in unconstrained scenes. Therefore, this paper presents a novel multi-hop temporal switch method to achieve effective temporal modeling of gait patterns in real-world scenes. Concretely, we design a novel gait recognition network, named Multi-hop Temporal Switch Network (MTSGait), to learn spatial features and multi-scale temporal features simultaneously. Different from existing methods that use 3D convolutions for temporal modeling, our MTSGait models the temporal dynamics of gait sequences by 2D convolutions. By this means, it achieves high efficiency with fewer model parameters and reduces the difficulty in optimization compared with 3D convolution-based models. Based on the specific design of the 2D convolution kernels, our method can eliminate the misalignment of features among adjacent frames. In addition, a new sampling strategy, i.e., non-cyclic continuous sampling, is proposed to make the model learn more robust temporal features. Finally, the proposed method achieves superior performance on two public gait in-the-wild datasets, i.e., GREW and Gait3D, compared with state-of-the-art methods. Jinkai Zheng, Xinchen Liu, Xiaoyan Gu 0001, Yaoqi Sun, Chuang Gan 0001, Jiyong Zhang 0001, Wu Liu 0005, Chenggang Yan 0001 |
ACM Multimedia | 2 |
| 2022 | SPU-Net: Self-Supervised Point Cloud Upsampling by Coarse-to-Fine Reconstruction With Self-Projection OptimizationabstractThe task of point cloud upsampling aims to acquire dense and uniform point sets from sparse and irregular point sets. Although significant progress has been made with deep learning models, state-of-the-art methods require ground-truth dense point sets as the supervision, which makes them limited to be trained under synthetic paired training data and not suitable to be under real-scanned sparse data. However, it is expensive and tedious to obtain large numbers of paired sparse-dense point sets as supervision from real-scanned sparse data. To address this problem, we propose a self-supervised point cloud upsampling network, named SPU-Net, to capture the inherent upsampling patterns of points lying on the underlying object surface. Specifically, we propose a coarse-to-fine reconstruction framework, which contains two main components: point feature extraction and point feature expansion, respectively. In the point feature extraction, we integrate the self-attention module with the graph convolution network (GCN) to capture context information inside and among local regions simultaneously. In the point feature expansion, we introduce a hierarchically learnable folding strategy to generate upsampled point sets with learnable 2D grids. Moreover, to further optimize the noisy points in the generated point sets, we propose a novel self-projection optimization associated with uniform and reconstruction terms as a joint loss to facilitate the self-supervised point cloud upsampling. We conduct various experiments on both synthetic and real-scanned datasets, and the results demonstrate that we achieve comparable performances to state-of-the-art supervised methods. Xinhai Liu, Xinchen Liu, Yu-Shen Liu, Zhizhong Han |
IEEE Trans. Image Process. | 2 |
| 2021 | Explainable Person Re-Identification with Attribute-guided Metric DistillationabstractDespite the great progress of person re-identification (ReID) with the adoption of Convolutional Neural Networks, current ReID models are opaque and only outputs a scalar distance between two persons. There are few methods providing users semantically understandable explanations for why two persons are the same one or not. In this paper, we propose a post-hoc method, named Attribute-guided Metric Distillation (AMD), to explain existing ReID models. This is the first method to explore attributes to answer: 1) what and where the attributes make two persons different, and 2) how much each attribute contributes to the difference. In AMD, we design a pluggable interpreter network for target models to generate quantitative contributions of attributes and visualize accurate attention maps of the most discriminative attributes. To achieve this goal, we propose a metric distillation loss by which the interpreter learns to decompose the distance of two persons into components of attributes with knowledge distilled from the target model. Moreover, we propose an attribute prior loss to make the interpreter generate attribute-guided attention maps and to eliminate biases caused by the imbalanced distribution of attributes. This loss can guide the interpreter to focus on the exclusive and discriminative attributes rather than the large-area but common attributes of two persons. Comprehensive experiments show that the interpreter can generate effective and intuitive explanations for varied models and generalize well under cross-domain settings. As a by-product, the accuracy of target models can be further improved with our interpreter.1 Xiaodong Chen 0011, Xinchen Liu, Wu Liu 0005, Xiao-Ping Zhang 0002, Yongdong Zhang 0001, Tao Mei 0001 |
ICCV | 2 |
| 2021 | TraND: Transferable Neighborhood Discovery for Unsupervised Cross-Domain Gait RecognitionabstractGait, i.e., the movement pattern of human limbs during locomotion, is a promising biometrie for identification of persons. Despite significant improvement in gait recognition with deep learning, existing studies still neglect a more practical but challenging scenario - unsupervised cross-domain gait recognition which aims to learn a model on a labeled dataset then adapt it to an unlabeled dataset. Due to the domain shift and class gap, directly applying a model trained on one source dataset to other target datasets usually obtains very poor results. Therefore, this paper proposes a Transferable Neighborhood Discovery (TraND) framework to bridge the domain gap for unsupervised cross-domain gait recognition. To learn effective prior knowledge for gait representation, we first adopt a backbone network pre- trained on the labeled source data in a supervised manner. Then we design an end-to-end trainable approach to automatically discover the confident neighborhoods of unlabeled samples in the latent space. During training, the class consistency indicator is adopted to select confident neighborhoods of samples based on their entropy measurements. Moreover, we explore a high- entropy-first neighbor selection strategy, which can effectively transfer prior knowledge to the target domain. Our method achieves the state-of-the-art results on two public datasets, i.e., CASIA-B and OU-LP. Jinkai Zheng, Xinchen Liu, Chenggang Yan 0001, Jiyong Zhang 0001, Wu Liu 0005, Xiao-Ping Zhang 0002, Tao Mei 0001 |
ISCAS | 2 |
| 2021 | HUMA'21: 2nd International Workshop on Human-centric Multimedia AnalysisabstractThe Second International Workshop on Human-centric Multimedia Analysis is focused on human-centric analysis using multimedia information. The human-centric multimedia analysis is one of the fundamental and challenging problems of multimedia understanding. It involves various human-centric analysis tasks like face recognition, human pose estimation, person re-identification, human action recognition, person tracking, human-computer interaction, etc. Nowadays, various multimedia sensing devices and large-scale computing infrastructures are generating a wide variety of multi-modality data at a rapid velocity, which supplies rich knowledge to tackle these challenges for human-centric analysis. Researchers and engineers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as smart city, retailing, intelligent manufacturing, and public services. To this end, our workshop aims to provide a platform to promote exchanges and integration for the fields of human analysis and multimedia. Wu Liu 0005, Xinchen Liu, Jingkuan Song, Dingwen Zhang, Wenbing Huang 0001, Junbo Guo, John R. Smith |
ACM Multimedia | 2 |
| 2020 | A Cross-modality and Progressive Person Search SystemabstractThis demonstration presents an instant and progressive cross-modality person search system, called 'CMPS'. Through the system, users can instantly find the lost children or elderly persons by simply describing their appearance through speech. Unlike most existing person search applications which have to cost much time to find the probe images, CMPS will save more valuable time in the early stage of losing. The proposed CMPS is one of the first attempts towards instant and progressive person search leveraging the audio, text, and visual modalities together. In detail, the system first takes the speech that describes the appearance of a person as the input to obtain a textual description by speech-to-text conversion. Then the cross-modal search is performed by matching the textual embedding with the visual representations of images in the learned latent space. The searched images can be used as candidates for query expansion. If the candidates are not right, the user can quickly adjust their description through speech. Once a right image is found, the user can directly click it as a new query. Finally the system will give the complete track of the lost person by once-click. On the built CUHK-PEDES-AUDIOS dataset, the system can achieve 82.46% rank-1 accuracy in real-time speed. Our code of CMPS is available at https://github.com/SheldongChen/Search-People-With-Audio. Xiaodong Chen 0011, Wu Liu 0005, Xinchen Liu, Yongdong Zhang 0001, Tao Mei 0001 |
ACM Multimedia | 3 |
| 2020 | Beyond the Parts: Learning Multi-view Cross-part Correlation for Vehicle Re-identificationabstractVehicle re-identification (Re-Id) is a challenging task due to the inter-class similarity, the intra-class difference, and the cross-view misalignment of vehicle parts. Although recent methods achieve great improvement by learning detailed features from keypoints or bounding boxes of parts, vehicle Re-Id is still far from being solved. Different from existing methods, we propose a Parsing-guided Cross-part Reasoning Network, named as PCRNet, for vehicle Re-Id. The PCRNet explores vehicle parsing to learn discriminative part-level features, model the correlation among vehicle parts, and achieve precise part alignment for vehicle Re-Id. To accurately segment vehicle parts, we first build a large-scale Multi-grained Vehicle Parsing (MVP) dataset from surveillance images. With the parsed parts, we extract regional features for each part and build a part-neighboring graph to explicitly model the correlation among parts. Then, the graph convolutional networks (GCNs) are adopted to propagate local information among parts, which can discover the most effective local features of varied viewpoints. Moreover, we propose a self-supervised part prediction loss to make the GCNs generate features of invisible parts from visible parts under different viewpoints. By this means, the same vehicle from different viewpoints can be matched with the well-aligned and robust feature representations. Through extensive experiments, our PCRNet significantly outperforms the state-of-the-art methods on three large-scale vehicle Re-Id datasets. Xinchen Liu, Wu Liu 0005, Jinkai Zheng, Chenggang Yan 0001, Tao Mei 0001 |
ACM Multimedia | 1 |
| 2020 | MetaSearch: Incremental Product Search via Deep Meta-LearningabstractWith the advancement of image processing and computer vision technology, content-based product search is applied in a wide variety of common tasks, such as online shopping, automatic checkout systems, and intelligent logistics. Given a product image as a query, existing product search systems mainly perform the retrieval process using predefined databases with fixed product categories. However, real-world applications often require inserting new categories or updating existing products in the product database. When using existing product search methods, the image feature extraction models must be retrained and database indexes must be rebuilt to accommodate the updated data, and these operations incur high costs for data annotation and training time. To this end, we propose a few-shot incremental product search framework with meta-learning, which requires very few annotated images and has a reasonable training time. In particular, our framework contains a multipooling-based product feature extractor that learns a discriminative representation for each product, and we also design a meta-learning-based feature adapter to guarantee the robustness of the few-shot features. Furthermore, when expanding new categories in batches during a product search, we reconstruct the few-shot features by using an incremental weight combiner to accommodate the incremental search task. Through extensive experiments, we demonstrate that the proposed framework achieves excellent performance for new products while still guaranteeing the high search accuracy of the base categories after gradually expanding new product categories without forgetting. Qi Wang 0079, Xinchen Liu, Wu Liu 0005, Anan Liu, Wenyin Liu, Tao Mei 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Social Relation Recognition From Videos via Multi-Scale Spatial-Temporal ReasoningabstractDiscovering social relations, e.g., kinship, friendship, etc., from visual contents can make machines better interpret the behaviors and emotions of human beings. Existing studies mainly focus on recognizing social relations from still images while neglecting another important media--video. On one hand, the actions and storylines in videos provide more important cues for social relation recognition. On the other hand, the key persons may appear at arbitrary spatial-temporal locations, even not in one same image from beginning to the end. To overcome these challenges, we propose a Multi-scale Spatial-Temporal Reasoning (MSTR) framework to recognize social relations from videos. For the spatial representation, we not only adopt a temporal segment network to learn global action and scene information, but also design a Triple Graphs model to capture visual relations between persons and objects. For the temporal domain, we propose a Pyramid Graph Convolutional Network to perform temporal reasoning with multi-scale receptive fields, which can obtain both long-term and short-term storylines in videos. By this means, MSTR can comprehensively explore the multi-scale actions and storylines in spatial-temporal dimensions for social relation reasoning in videos. Extensive experiments on a new large-scale Video Social Relation dataset demonstrate the effectiveness of the proposed framework. Xinchen Liu, Wu Liu 0005, Jingwen Chen 0001, Lianli Gao, Chenggang Yan 0001, Tao Mei 0001 |
CVPR | 1 |
| 2019 | Multi-Granularity Reasoning for Social Relation Recognition From ImagesabstractDiscovering social relations in images can make machines better interpret the behavior of human beings. However, automatically recognizing social relations in images is a challenging task due to the significant gap between the domains of visual content and social relation. Existing studies separately process various features such as faces expressions, body appearance, and contextual objects, thus they cannot comprehensively capture the multi-granularity semantics, such as scenes, regional cues of persons, and interactions among persons and objects. To bridge the domain gap, we propose a Multi-Granularity Reasoning framework for social relation recognition from images. The global knowledge and mid-level details are learned from the whole scene and the regions of persons and objects, respectively. Most importantly, we explore the fine-granularity pose keypoints of persons to discover the interactions among persons and objects. Specifically, the pose-guided Person-Object Graph and Person-Pose Graph are proposed to model the actions from persons to object and the interactions between paired persons, respectively. Based on the graphs, social relation reasoning is performed by graph convolutional networks. Finally, the global features and reasoned knowledge are integrated as a comprehensive representation for social relation recognition. Extensive experiments on two public datasets show the effectiveness of the proposed framework. Xinchen Liu, Wu Liu 0005, Anfu Zhou, Huadong Ma, Tao Mei 0001 |
ICME | 2 |
| 2019 | BraidNet: Braiding Semantics and Details for Accurate Human ParsingabstractThis paper focuses on fine-grained human parsing in images. This is a very challenging task due to the diverse person appearance, semantic ambiguity of different body parts and clothing, and extremely small parsing targets. Although existing approaches can achieve significant improvement by pyramid feature learning, multi-level supervision, and joint learning with pose estimation, human parsing is still far from being solved. Different from existing approaches, we propose a Braiding Network, named as BraidNet, to learn complementary semantics and details for fine-grained human parsing. The BraidNet contains a two-stream braid-like architecture. The first stream is a semantic abstracting net with a deep yet narrow structure which can learn semantic knowledge by a hierarchy of fully convolution layers to overcome the challenges of diverse person appearance. To capture low-level details of small targets, the detail-preserving net is designed to exploit a shallow yet wide network without down-sampling, which can retain sufficient local structures for small objects. Moreover, we design a group of braiding modules across the two sub-nets, by which complementary information can be exchanged during end-to-end training. Besides, in the end of BraidNet, a Pairwise Hard Region Embedding strategy is propose to eliminate the semantic ambiguity of different body parts and clothing. Extensive experiments show that the proposed BraidNet achieves better performance than the state-of-the-art methods for fine-grained human parsing. Xinchen Liu, Wu Liu 0005, Jingkuan Song, Tao Mei 0001 |
ACM Multimedia | 1 |
| 2019 | PVSS: A Progressive Vehicle Search System for Video Surveillance Networks
Xinchen Liu, Huadong Ma, Shuangqun Li |
J. Comput. Sci. Technol. | 1 |
| 2018 | PROVID: Progressive and Multimodal Vehicle Reidentification for Large-Scale Urban SurveillanceabstractCompared with person reidentification, which has attracted concentrated attention, vehicle reidentification is an important yet frontier problem in video surveillance and has been neglected by the multimedia and vision communities. Since most existing approaches mainly consider the general vehicle appearance for reidentification while overlooking the distinct vehicle identifier, such as the license plate number, they attain suboptimal performance. In this paper, we propose PROVID, a PROgressive Vehicle re-IDentification framework based on deep neural networks. In particular, our framework not only utilizes the multimodality data in large-scale video surveillance, such as visual features, license plates, camera locations, and contextual information, but also considers vehicle reidentification in two progressive procedures: coarse-to-fine search in the feature domain, and near-to-distant search in the physical space. Furthermore, to evaluate our progressive search framework and facilitate related research, we construct the VeRi dataset, which is the most comprehensive dataset from real-world surveillance videos. It not only provides large numbers of vehicles with varied labels and sufficient cross-camera recurrences but also contains license plate numbers and contextual information. Extensive experiments on the VeRi dataset demonstrate both the accuracy and efficiency of our progressive vehicle reidentification framework. Xinchen Liu, Wu Liu 0005, Tao Mei 0001, Huadong Ma |
IEEE Trans. Multim. | 1 |
| 2017 | Beyond Human-level License Plate Super-resolution with Progressive Vehicle Search and Domain Priori GANabstractIn this paper, we address the challenging problem of vehicle license plate image super-resolution. Different from existing image super-resolution approaches only resorted to one single image, we propose to leverage complementary information from multiple images to recover the license plate numbers. To achieve this goal, we design a principled license plate images super-resolution framework which is composed of two components: progressive vehicle search and Domain Priori GAN (DP-GAN). Particularly, we design a null space based progressive vehicle search approach to retrieve the relevant images captured by different cameras given one vehicle with a low-resolution license plate. To handle the extremely varied license plate images caused by different sensors, times, depths, and viewpoints, we also propose a DP-GAN framework to generate multiple spatial correspondences and high-resolution plate images. In the generator network of DP-GAN, a license plate synthesis pipeline is exploited to generate the nearly canonical license plates. In the discriminator network, a spatial split layer is designed to simultaneously preserve the global and local manufacture standards of the license plate. Finally, a multiple images super-resolution GAN is exploited to combine all the synthetic license plates into one high-resolution image. Different from previous super-resolution criteria mainly focus on pixel-level detail recovery condition, we leverage the downstream tasks, i.e. license plate recognition and vehicle search as criteria. The results on a new collected real-world dataset demonstrate that the proposed method achieves the beyond human-level license plate super-resolution performance for automatic license plate recognition and vehicle search. Wu Liu 0005, Xinchen Liu, Huadong Ma, Peng Cheng 0002 |
ACM Multimedia | 2 |
| 2016 | A Deep Learning-Based Approach to Progressive Vehicle Re-identification for Urban Surveillance
Xinchen Liu, Wu Liu 0005, Tao Mei 0001, Huadong Ma |
ECCV (2) | 1 |
| 2016 | Large-scale vehicle re-identification in urban surveillance videosabstractVehicle, as a significant object class in urban surveillance, attracts massive focuses in computer vision field, such as detection, tracking, and classification. Among them, vehicle re-identification (Re-Id) is an important yet frontier topic, which not only faces the challenges of enormous intra-class and subtle inter-class differences of vehicles in multicameras, but also suffers from the complicated environments in urban surveillance scenarios. Besides, the existing vehicle related datasets all neglect the requirements of vehicle Re-Id: 1) massive vehicles captured in real-world traffic environment; and 2) applicable recurrence rate to give cross-camera vehicle search for vehicle Re-Id. To facilitate vehicle Re-Id research, we propose a large-scale benchmark dataset for vehicle Re-Id in the real-world urban surveillance scenario, named “VeRi”. It contains over 40,000 bounding boxes of 619 vehicles captured by 20 cameras in unconstrained traffic scene. Moreover, each vehicle is captured by 2~18 cameras in different viewpoints, illuminations, and resolutions to provide high recurrence rate for vehicle Re-Id. Finally, we evaluate six competitive vehicle Re-Id methods on VeRi and propose a baseline which combines the color, texture, and highlevel semantic information extracted by deep neural network. Xinchen Liu, Wu Liu 0005, Huadong Ma, Huiyuan Fu |
ICME | 1 |
| 2001 | Multicasts on WDM All-Optical Multistage Interconnection NetworksabstractWavelength-division multiplexing (WDM) optical networks provide huge bandwidth by allowing multiple data streams to be transmitted simultaneously along the same optical fiber, with each stream assigned a distinct wavelength. A key issue of WDM optical networks is the minimization of the number of wavelengths for realizing a routing request. Let W be the number of wavelengths supported by a WDM optical network. For a routing request R which needs l wavelengths, if l/spl les/W then R can be realized in one round of routing. However, when l>W, multiple rounds of routing for R are required. In this case, it is important to minimize the number of routing rounds. Multicast transmits a data stream from one input to multiple outputs (one-to-many), a fundamental communication pattern in many applications. We study the problem of minimizing the number of wavelengths and the number of routing rounds for realizing a set R={(u, /spl nu/)} of multicasts, where each output /spl nu/ receives a data stream from exactly one input u, on an n-dimensional WDM all-optical multistage interconnection networks (MINs). For a network with wavelength converters, we show that any set of multicasts can be realized by 2/sup [(n-1)/(k+1)]/ wavelengths in k rounds of routing. For one round of routing, the upper bound 2/sup [(n-1)/2]/ is tight to the lower bound. We also give algorithms for multicasts on a network without wavelength converters. Computer simulation results show that any set of multicasts can be realized in at most two rounds of routing on a network of practical size. Xinchen Liu, Qian-Ping Gu |
ICPADS | 1 |