Wu Liu 0005

dblp:31/4112-5 · DBLP profile ↗
← Back
135ranked-venue papers
15as first author
80since 2021 · last 2026
0000-0003-1633-7575ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 114 · 13 first-author · 65 since 2021Artificial intelligence and machine learning · 41 · 3 first-author · 27 since 2021Computer networks · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents
abstract
Recent advances in vision-language models (VLMs) and reinforcement learning (RL) have driven progress in GUI automation. However, most existing methods rely on static, one-shot visual inputs and passive perception, lacking the ability to adaptively determine when, whether, and how to observe the interface. We present GUI-Eyes, a reinforcement learning framework for active visual perception in GUI tasks. To acquire more informative observations, the agent learns to make strategic decisions on both whether and how to invoke visual tools, such as cropping or zooming, within a two-stage reasoning process. To support this behavior, we introduce a progressive perception strategy that decomposes the decision-making into coarse exploration and fine-grained grounding, coordinated by a two-level policy. In addition, we design a spatially continuous reward function tailored to tool usage, which integrates both location proximity and region overlap to provide dense supervision and alleviate the reward sparsity common in GUI environments. On the ScreenSpot-Pro benchmark, GUI-Eyes-3B achieves 44.8% grounding accuracy using only 3k labeled samples, significantly outperforming both supervised and RL-based baselines. These results highlight that tool-aware active perception, enabled by staged policy reasoning and fine-grained reward feedback, is critical for building robust and data-efficient GUI agents.
Jiawei Shao, Dakuan Lu, Haoyi Hu, Xiangcheng Liu, Hantao Yao, Wu Liu 0005
AAAI7
2026 GASim: A Graph-Accelerated Hybrid Framework for Social Simulation
abstract
Large-scale social simulators are essential for studying complex social patterns.Prior work explores hybrid methods to scale up simulations, combining large language models (LLM)-based agents with numerical agentbased models (ABM).However, this incurs high latency due to expensive memory retrieval and sequential ABM execution.To address this challenge, we propose GASim, a graph-accelerated hybrid multi-agent framework for large-scale social simulations.For core agents driven by LLM, GASim introduces Graph-Optimized Memory (GOM) to replace intensive LLM-based retrieval pipelines with lightweight propagation over a sparse memory graph.For the majority of ordinary agents, GASim employs Graph Message Passing (GMP), substituting sequential ABM execution with parallel updates by fine-grained feature aggregation and Graph Attention Network.We further introduce Entropy-Driven Grouping (EDG) that coordinates this hybrid partitioning, leveraging information entropy to dynamically identify emergent core agents situated in information-diverse neighborhoods.Extensive experiments show that GASim not only delivers a substantial 9.94× end-to-end speedup over the traditional hybrid framework but also consumes less than 20% of baseline tokens, significantly reducing costs while preserving strong alignment with real-world public opinion trends.Our code is available at https://github.com/Jasmine0201/GASim.
Yanhui Sun, Hantao Yao, Allen He, Yongdong Zhang 0001, Wu Liu 0005
ACL (1)6
2026 Enhancing Federated Class-Incremental Learning via Spatial-Temporal Statistics Aggregation
abstract
The growing presence of mobile and IoT devices has led to massive decentralized and evolving data, driving the rise of Federated Learning (FL) to enable collaborative training without data sharing. However, traditional FL assumes static data distributions, which is unrealistic for dynamic real-world environments. To address this challenge, Federated Class-Incremental Learning (FCIL) has emerged as a promising framework that enables flexible adaptation to newly introduced classes over time. Existing FCIL methods typically integrate old knowledge preservation into local client training. However, these methods cannot avoid spatial-temporal client drift caused by data heterogeneity and often incur significant computational and communication overhead, limiting practical deployment. To address these challenges simultaneously, we propose a novel approach, Spatial-Temporal Statistics Aggregation (STSA), which provides a unified framework to aggregate feature statistics both spatially (across clients) and temporally (across stages). The aggregated feature statistics are unaffected by data heterogeneity and can be used to update the classifier in closed form at each stage. Additionally, we introduce STSA-E, a communication-efficient variant that enables the server to approximate global second-order feature statistics using first-order statistics uploaded from clients. Theoretical analysis shows that it achieves similar performance to STSA with much lower communication overhead. Extensive experiments on three widely used FCIL datasets, with varying degrees of data heterogeneity, show that our method outperforms state-of-the-art FCIL methods in terms of performance, flexibility, and both communication and computation efficiency. The code is available at https://github.com/Yuqin-G/STSA.
Zenghao Guan, Guojun Zhu, Yucan Zhou, Wu Liu 0005, Weiping Wang 0005, Jiebo Luo 0001, Xiaoyan Gu 0001
WWW4
2026 Multi-Granularity Multi-Modal Knowledge Graph Representation Learning via Subgraph-Aware Adaptive Fusion and Hierarchical Relation Modeling
Peining Li, Meiyu Liang, Junping Du 0001, Zhe Xue, Guanhua Ye, Wu Liu 0005, Lei Shi 0030
WWW7
2026 AdaptiveMamba: Comprehensive visual representation learning with adaptive semantic perception
Chaojie Chen, Xingcai Wu, Wu Liu 0005, Qi Wang 0079
Knowl. Based Syst.3
2026 Bitscaling: Streamlining neural network compression via predictive multi-scale growth of mixed-precision networks
Yuehao Li 0003, Haifang Jian, Hongchang Wang, Linghe Zhang, Wu Liu 0005
Neural Networks6
2026 DSA: A Dual-Safety Approach for Test-Time Temporal Risk Removal of Video Generation
Qi Liu 0081, Kun Liu 0016, Xinchen Liu, Xiaoyan Gu 0001, Zhaochun Ren, Yongdong Zhang 0001, Wu Liu 0005
IEEE Trans. Circuits Syst. Video Technol.7
2026 How to Understand Named Entities: Using Commonsense for News Captioning
abstract
News captioning aims to describe an image with its news article body as input. It greatly relies on a set of detected named entities, including real-world people, organizations, and places. This article exploits commonsense knowledge to understand named entities for news captioning. By “understand,” we mean correlating the news content with commonsense in the wild, which helps an agent to (1) distinguish semantically similar named entities and (2) describe named entities using words outside of training corpora. Our approach consists of three modules: (a) Filter Module aims to clarify the commonsense concerning a named entity from two aspects: what does it mean ? and what is it related to ?, which divide the commonsense into explanatory knowledge and relevant knowledge , respectively. (b) Distinguish Module aggregates explanatory knowledge from node-degree , dependency , and distinguish three aspects to distinguish semantically similar named entities. (c) Enrich Module attaches relevant knowledge to named entities to enrich the entity description by commonsense information (e.g., identity and social position). Finally, all of information is integrated into the large multimodal model to generate the news caption. Extensive experiments on two challenging datasets (i.e., GoodNews and NYTimes) demonstrate the superiority of our method. Ablation studies and visualization further validate its effectiveness in understanding named entities.
Shenyuan Zhang, Ning Xu 0003, Yanhui Wang 0001, Tongle Ma, Wu Liu 0005, Jinlin Guo, Anan Liu
ACM Trans. Multim. Comput. Commun. Appl.5
2025 HOIGen-1M: A Large-scale Dataset for Human-Object Interaction Video Generation
abstract
Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale videos with accurate captions for HOI. To address this issue, we introduce HOIGen-1M, the first large-scale dataset for HOI Generation, consisting of over one million high-quality videos collected from diverse sources. In particular, to guarantee the high quality of videos, we first design an efficient framework to automatically curate HOI videos using the powerful multimodal large language models (MLLMs), and then the videos are further cleaned by human annotators. Moreover, to obtain accurate textual captions for HOI videos, we design a novel video description method based on a Mixture-of-Multimodal-Experts (MoME) strategy that not only generates expressive captions but also eliminates the hallucination by individual MLLM. Further-more, due to the lack of an evaluation framework for gen-erated HOI videos, we propose two new metrics to assess the quality of generated videos in a coarse-to-fine manner. Extensive experiments reveal that current T2V models struggle to generate high-quality HOI videos and confirm that our HOIGen-1M dataset is instrumental for improving HOI video generation.
Kun Liu 0016, Qi Liu 0081, Xinchen Liu, Yongdong Zhang 0001, Jiebo Luo 0001, Xiaodong He 0001, Wu Liu 0005
CVPR8
2025 A Semantic Knowledge Complementarity based Decoupling Framework for Semi-supervised Class-imbalanced Medical Image Segmentation
abstract
The limited data annotations have made semi-supervised learning (SSL) increasingly popular in medical image analysis. However, the use of pseudo labels in SSL degrades the performance of decoders that heavily rely on high-accuracy annotations. This issue is particularly pronounced in class-imbalanced multi-organ segmentation tasks, where small organs may be under-segmented or even ignored. In this paper, we propose SKCDF, a semantic knowledge complementarity based decoupling framework for multi-organ segmentation in class-imbalanced medical images. SKCDF decouples the data flow based on the responsibilities of the encoder and decoder during model training to make the model effectively learn semantic features, while mitigating the negative impact of unlabeled data on the semantic segmentation task. We also design a semantic knowledge complementarity module that adopts labeled data to guide the generation of pseudo labels and enriches the semantic features of labeled data with unlabeled data, which improves the quality of generated pseudo labels and the robustness of the overall model. Furthermore, we design an auxiliary balanced segmentation head based training strategy to further enhance the segmentation performance of small organs. Experimental results on the Synapse and AMOS datasets show that our method significantly outperforms existing methods.
Zheng Zhang 0038, Guanchun Yin, Bo Zhang 0032, Wu Liu 0005, Xiuzhuang Zhou, Wendong Wang 0003
CVPR4
2025 MotionPro: A Precise Motion Controller for Image-to-Video Generation
abstract
Animating images with interactive motion control has garnered popularity for image-to-video (I2V) generation. Modern approaches typically rely on large Gaussian kernels to extend motion trajectories as condition without explicitly defining movement region, leading to coarse motion control and failing to disentangle object and camera moving. To alleviate these, we present MotionPro, a precise motion controller that novelly leverages region-wise trajectory and motion mask to regulate fine-grained motion synthesis and identify target motion category (i.e., object or camera moving), respectively. Technically, MotionPro first estimates the flow maps on each training video via a tracking model, and then samples the region-wise trajectories to simulate inference scenario. Instead of extending flow through large Gaussian kernels, our region-wise trajectory approach enables more precise control by directly utilizing trajectories within local regions, thereby effectively characterizing fine-grained movements. A motion mask is simultaneously derived from the predicted flow maps to capture the holistic motion dynamics of the movement regions. To pursue natural motion control, MotionPro further strengthens video denoising by incorporating both region-wise trajectories and motion mask through feature modulation. More remarkably, we meticulously construct a benchmark, i.e., MC-Bench, with 1.1K user-annotated image-trajectory pairs, for the evaluation of both fine-grained and object-level I2V motion control. Extensive experiments conducted on WebVid-10M and MC-Bench demonstrate the effectiveness of MotionPro. Please refer to our project page for more results: https://zhw-zhang.github.io/MotionPro-page/.
Fuchen Long, Zhaofan Qiu, Yingwei Pan, Wu Liu 0005, Ting Yao 0003, Tao Mei 0001
CVPR5
2025 PlugMark: A Plug-In Zero-Watermarking Framework for Diffusion Models
Pengzhen Chen, Yanwei Liu 0001, Xiaoyan Gu 0001, Enci Liu, Zhuoyi Shang, Xiangyang Ji, Wu Liu 0005
ICCV7
2025 High-Fidelity Object Removal through Boosting Diffusion Processes
abstract
The remarkable image understanding and generation capabilities of diffusion models have made image editing a highly promising area of research. As a significant subtask within the field, object removal aims to remove objects from a specified region and fill the missing pixels with visually coherent and semantically sound content. Despite the great progress made in deep generative models, research in this area still faces several challenges: i. high expense of model training induced by the data scarcity and the difficulity in large-scale real data collection ii. current training-free methods are unable to drastically change the behavior of the attention layer that has been set during the pre-training phase for text-guided object inpainting. In this paper, we introduce HybridRemover, a two-stage diffusion scheme. We decouple the task into two subtasks: one to remove the specified objects from the target region and one to perform image restoration of the target region. By fine-tuning the SD-Inpainting model with a very small amount of data, we transform it into a model that focuses only on the complete removal of the object without considering the surrounding effects, and the overall repair of the image is taken care of by the SD-Inpainting model cascaded behind it. As a result of our efforts, our method achieved state-of-the-art performance in object removal tasks. Even when a strong perspective distortion gets involved, our method delivers exceptional results.
Haihui Fan, Xiaoyan Gu 0001, Wei Zhang 0031, Wu Liu 0005
ISCAS5
2025 Edit-by-Example: Adaptive Exemplar-Based Image Editing
abstract
Recent advances in diffusion-based image editing models have demonstrated remarkable success. However, these models primarily rely on high-quality textual prompts to guide image manipulation, creating a significant barrier for non-expert users. In this demonstration, we present an exemplar-based image editing framework named Edit-by-Example, which eliminates the reliance on textual prompts, requires only a single pair of before-and-after images to encapsulate the desired editing effect that can readily be applied on the user-provided query image without any model fine-tuning. Technically, our framework comprises two components: an Adaptive Editing Policy Module (AEPM) and a Generation Module (GM). The AEPM jointly analyzes cross-image relationships in exemplar pairs and query image content to derive optimal editing directions, while GM executes these policies through an off-the-shelf image editor with optional semantic alignment verification. We introduce EEdBench, a comprehensive benchmark for exemplar-based image editing containing 1,500 test cases across 15 categories. Experiments demonstrate that our framework outperforms existing prompt-free methods in editing direction accuracy (S-Visual) and fidelity (FID).
Yaojie Li, Zhaofan Qiu, Yingwei Pan, Wu Liu 0005, Ting Yao 0003, Tao Mei 0001
ACM Multimedia5
2025 Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Retrieval
abstract
In recent years, pre-trained multimodal large models have attracted widespread attention due to their outstanding performance in various multimodal applications. Nonetheless, the extensive computational resources and vast datasets required for their training present significant hurdles for deployment in environments with limited computational resources. Many existing methods attempt to compress pre-trained multimodal large models through knowledge distillation, typically focusing on a single optimization objective. While such methods successfully reduce model parameters, they often incur significant performance degradation. Moreover, single-scale optimization fails to ensure comprehensive learning of the teacher model's knowledge across different aspects. In this work, we propose, for the first time, a dynamic self-adaptive multiscale distillation (DSMD) from pre-trained multi-modal large model for efficient cross-modal retrieval method, considering multiple scales from the perspectives of fine granularity, global structure, and hard negative sample mining. Furthermore, we design a dynamic loss balancer, eliminating the need to manually tune objective weights during distillation. This dynamic mechanism ensures that all objectives are optimized in a balanced and adaptive manner throughout the training process. Experiments demonstrate that our multiscale distillation framework achieves significant performance improvements over traditional single-scale distillation methods. Additionally, our proposed dynamic balancer effectively stabilizes the distillation process, ensuring consistent optimization across objectives. The distilled student model achieves 90% of the teacher model's performance while using only 10% of its parameters. Notably, our model also achieves state-of-the-art performance on cross-modal retrieval tasks, outperforming existing approaches. Codes are available at https://github.com/chrisx599/DSMD.
Zhengyang Liang, Meiyu Liang, Yawen Li 0001, Wu Liu 0005, Yingxia Shao, Kangkang Lu 0002
ACM Multimedia5
2025 Asymmetric Pre-aligned Anchor Contrastive Enhanced Diffusion Hashing Model for Incomplete Multimodal Retrieval
abstract
Multimodal hashing stands as an efficient approach for multimodal retrieval, yet it frequently grapples with the challenge of misaligned representation spaces across different modalities. This misalignment can degrade the consistency and discrimination of multimodal representations, complicating the learning of effective representations for image and text pairs. Particularly, the task becomes arduous when the system must handle incomplete data while ensuring accurate and relevant retrieval outcomes. To address these challenges, we propose the Asymmetric Pre-aligned Anchor Contrastive Enhanced Diffusion Hashing Model (AADH) for Incomplete Multimodal Retrieval. Our model is specifically tailored to robustly manage multimodal incomplete data scenarios. Initially, we develop an Asymmetric Pre-alignment Strategy that utilizes asymmetric contrastive learning to preliminarily align the semantic disparities between various modalities. Subsequently, we propose an innovative Anchor Contrastive Reinforcement Diffusion Hashing Model, which integrates image and text modalities to varying extents during the reverse diffusion process. It constructs an anchor space that not only facilitates the learning of incomplete multimodal hashing representations through anchor contrastive learning but also leverages inter-modal and intra-modal contrastive learning to enhance the representations. Moreover, we effectively bridge the modality gap between different modal hash codes by employing the anchor space to constrain the representations of different modal hashes. By adjusting the initial noise of the diffusion model, we indirectly expand the data volume, which in turn bolsters the model's robustness. Our extensive experimental results across multiple datasets demonstrate that the proposed AADH model achieves state-of-the-art (SOTA) results.
Meiyu Liang, Juncheng Zheng, Kangkang Lu 0002, Yawen Li 0001, Junping Du 0001, Zhe Xue, Wu Liu 0005
ACM Multimedia9
2025 Statistics Caching Test-Time Adaptation for Vision-Language Models
abstract
Test-time adaptation (TTA) for Vision-Language Models (VLMs) aims to enhance performance on unseen test data. However, existing methods struggle to achieve robust and continuous knowledge accumulation during test time. To address this, we propose Statistics Caching test-time Adaptation (SCA), a novel cache-based approach. Unlike traditional feature-caching methods prone to forgetting, SCA continuously accumulates task-specific knowledge from all encountered test samples. By formulating the reuse of past features as a least squares problem, SCA avoids storing raw features and instead maintains compact, incrementally updated feature statistics. This design enables efficient online adaptation without the limitations of fixed-size caches, ensuring that the accumulated knowledge grows persistently over time. Furthermore, we introduce adaptive strategies that leverage the VLM's prediction uncertainty to reduce the impact of noisy pseudo-labels and dynamically balance multiple prediction sources, leading to more robust and reliable performance. Extensive experiments demonstrate that SCA achieves compelling performance while maintaining competitive computational efficiency.
Zenghao Guan, Yucan Zhou, Wu Liu 0005, Xiaoyan Gu 0001
NeurIPS3
2025 LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature Representation
abstract
Speech-driven 3D facial animation has attracted increasing interest since its potential to generate expressive and temporally synchronized digital humans. While recent works have begun to explore emotion-aware animation, they still depend on explicit one-hot encodings to represent identity and emotion with given emotion and identity labels, which limits their ability to generalize to unseen speakers. Moreover, the emotional cues inherently present in speech are often neglected, limiting the naturalness and adaptability of generated animations. In this work, we propose LSF-Animation, a novel framework that eliminates the reliance on explicit emotion and identity feature representations. Specifically, LSF-Animation implicitly extracts emotion information from speech and captures the identity features from a neutral facial mesh, enabling improved generalization to unseen speakers and emotional states without requiring manual labels. Furthermore, we introduce a Hierarchical Interaction Fusion Block (HIFB), which employs a fusion token to integrate dual transformer features and more effectively integrate emotional, motion-related and identity-related cues. Extensive experiments conducted on the 3DMEAD dataset demonstrate that our method surpasses recent state-of-the-art approaches in terms of emotional expressiveness, identity generalization, and animation realism. The source code will be released at: https://github.com/Dogter521/LSF-Animation.
Chuanqing Zhuang, Chenxi Jin, Zhengda Lu, Yiqun Wang 0001, Wu Liu 0005, Jun Xiao 0005
SIGGRAPH Asia6
2025 TalkingAvatar: Learning 3D talking human avatar via NeRF
Lingyun Yu 0002, Chuanbin Liu 0001, Wu Liu 0005, Quanwei Yang, Meng Shao
Neurocomputing4
2025 Single-Frame Supervision for Spatio-Temporal Video Grounding
abstract
Spatio-Temporal Video Grounding (STVG) aims at localizing the spatio-temporal tube of a specific object in an untrimmed video given a free-form natural language query. As the annotation of tubes is labor intensive, researchers are motivated to explore weakly supervised approaches in recent works, which usually results in significant performance degradation. To achieve a less expensive STVG method with acceptable accuracy, this work investigates the "single-frame supervision" paradigm that requires a single frame labeled with a bounding box within the temporal boundary of the fully supervised counterpart as the supervisory signal. Based on the characteristics of the STVG problem, we propose a Two-Stage Multiple Instance Learning (T-SMILE) method, which creates pseudo labels by expanding the annotated frame to its contextual frames, thereby establishing a fully-supervised problem to facilitate further model training. The innovations of the proposed method are three-folded, including 1) utilizing multiple instance learning to dynamically select instances in positive bags for the recognition of starting and ending timestamps, 2) learning highly discriminative query features by incorporating spatial prior constraints in cross-attention, and 3) designing a curriculum learning-based strategy that iterative assigns dynamic weights to spatial and temporal branches, thereby gradually adapting to the learning branch with larger difficulty. To facilitate future research on this task, we also contribute a large-scale benchmark containing 12,469 videos on complex scenes with single-frame annotation. The extensive experiments on two benchmarks demonstrate that T-SMILE significantly outperforms all weakly-supervised methods. Remarkably, it also performs better than some fully-supervised methods associated with much more annotation labor costs.
Kun Liu 0016, Mengxue Qu, Yang Liu 0235, Yunchao Wei, Wenming Zhe, Yao Zhao 0001, Wu Liu 0005
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 DRC: Discrete Representation Classifier With Salient Features via Fixed-Prototype
abstract
Image classification models including convolutional neural networks (CNN) and vision transformers (ViT) commonly employ a fully connected (FC) layer as the classifier. However, the fully connected nature of FC brings large amounts of weight parameters, limits the efficiency of inference, tends to over-fit the training data, and struggles to learn distinct class weights. To solve these problems, we propose a discrete representation classifier (DRC), a generic parameter-free classifier that offers efficiency, robustness, and more discriminative categorization. Specifically, the DRC discards numerous unimportant features and focuses solely on the salient features which are reinforced during training and presented in short discrete form during inference. Unlike the way of learning pseudo-prototypes (weights) from data laden with complex patterns and noises in FC, the DRC introducing discriminative fixed-prototypes which are almost uniformly distributed across the high-dimensional feature space, thus helps the model to learn more distinct boundaries between categories. Further leveraging the advantage of DRC’s focus on salient features, we propose Salient-CAM, which is able to locate the most important region in image without the need for weighting feature maps. The experiments demonstrate that simply replacing the model’s classifier from FC to DRC can lead to a significant acceleration in the whole model’s inference and a more robust classification. Additionally, the proposed Salient-CAM exhibits excellent object localization ability in complex natural scenes.
Qinglei Li, Qi Wang 0079, Yongbin Qin, Xingcai Wu, Shiming Chen 0002, Wu Liu 0005, Yong-Jin Liu 0001, Jiebo Luo 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Learning a Dynamic Neural Human via Poses Guided Dual Spaces Feature
abstract
Learning human representations from video is becoming increasingly important in various applications. However, due to the limited information in videos and the complexity of human deformation, existing methods cannot faithfully reconstruct the image representation of humans, including clothing folds and light and shadow. Our method is built upon a deformation-based approach, which uses pose-guided joint learning to derive human representations in both canonical space and observation space, thereby enhancing the model’s performance in human details. We conducted several experiments on publicly available datasets using our approach, achieving highly realistic reconstruction results that are difficult to distinguish from real frames. Our approach also showed improved overall evaluation metrics for video frames that were not visible in the original view angle.
Caoyuan Ma, Runqi Wang, Wu Liu 0005, Ziqiao Zhou, Zheng Wang 0007
AVSS3
2024 HumanNeRF-SE: A Simple yet Effective Approach to Animate HumanNeRF with Diverse Poses
abstract
We present HumanNeRF-SE, a simple yet effective method that synthesizes diverse novel pose images with sim-ple input. Previous HumanNeRF works require a large number of optimizable parameters to fit the human images. Instead, we reload these approaches by combining explicit and implicit human representations to design both general-ized rigid deformation and specific non-rigid deformation. Our key insight is that explicit shape can reduce the sam-pling points used to fit implicit representation, and frozen blending weights from SMPL constructing a generalized rigid deformation can effectively avoid overfitting and im-prove pose generalization performance. Our architecture involving both explicit and implicit representation is sim-ple yet effective. Experiments demonstrate our model can synthesize images under arbitrary poses with few-shot input and increase the speed of synthesizing images by 15 times through a reduction in computational complexity without using any existing acceleration modules. Compared to the state-of-the-art HumanNeRF studies, HumanNeRF-SE achieves better performance with fewer learnable parame-ters and less training time.
Caoyuan Ma, Yu-Lun Liu 0001, Zhixiang Wang 0001, Wu Liu 0005, Xinchen Liu, Zheng Wang 0007
CVPR4
2024 Norma: A Noise Robust Memory-Augmented Framework for Whole Slide Image Classification
Yu Bai 0020, Bo Zhang 0032, Zheng Zhang 0038, Zibo Ma, Wu Liu 0005, Xiuzhuang Zhou, Xiangyang Gong, Wendong Wang 0003
ECCV (51)6
2024 CLaM: An Open-Source Library for Performance Evaluation of Text-driven Human Motion Generation
abstract
Text-driven human motion generation, which creates motion sequences based on textual descriptions, has attracted great attention in the communities of multimedia and artificial intelligence. By parsing and comprehending textual information and converting it into specific human movements, it realizes a direct transformation from human semantics to motion sequences. New text-driven human motion generators are springing up to achieve better performance. However, the absence of well-trained evaluators that can effectively estimate the consistency between the text prompts and motions generated by existing generators remains a challenge. To address the above issues, we propose an open-source library with a powerful Contrastive Language-and-Motion (CLaM) pre-training evaluator, which can be employed for evaluating a variety of text-driven human motion generation algorithms. We perform a thorough performance evaluation of the existing algorithms on various metrics, such as R-Precision. As a by-product, we build a large-scale HumanML3D-synthesis dataset, which consists of 14,616 motion sequences and 547,102 textual descriptions, which is ten times larger than the widely-used HumanML3D dataset. The source codes and models for CLaM are available at~https://github.com/SheldongChen/CLaM/.
Xiaodong Chen 0011, Kunlang He, Wu Liu 0005, Xinchen Liu, Zhengjun Zha, Tao Mei 0001
ACM Multimedia3
2024 ADDG: An Adaptive Domain Generalization Framework for Cross-Plane MRI Segmentation
abstract
Multi-planar magnetic resonance imaging (MRI) can provide comprehensive 3D structural information for disease diagnosis. Compared to multi-source MRI, multi-planar MRI scans target areas in the human body from different directions. This atypical difference between directions may lead to poor performance of traditional domain generalization methods, especially when MRI from different planes also comes from different sources. In this paper, we propose ADDG, an Adaptive Domain Generalization framework for accurate cross-plane MRI segmentation. ADDG significantly mitigates the impact of information loss caused by slice spacing by injecting 3D shape prior to the segmentation target and capturing domain-agnostic feature differences from heterogeneous data sources through an adaptive data partitioning strategy. In addition, we propose a mesh deformation-based organ segmentation network to simultaneously delineate 2D boundary and 3D volume of organ, which could guide more accurate mesh deformation. We also develop an organ-specific mesh template and employ Loop subdivision for generating smoother 3D organ mesh. Furthermore, we design a flexible meta-learning paradigm to adaptively partition data domains based on invariant learning, which can learn domain-agnostic features from multi-source data to enhance the overall generalization ability. Experimental results show that ADDG outperforms several medical image segmentation, single-view 3D shape reconstruction, and domain generalization methods.
Zibo Ma, Bo Zhang 0032, Zheng Zhang 0038, Wu Liu 0005, Wufan Wang, Hui Gao 0002, Wendong Wang 0003
ACM Multimedia4
2024 It Takes Two: Accurate Gait Recognition in the Wild via Cross-granularity Alignment
abstract
Existing studies for gait recognition primarily utilized sequences of either binary silhouette or human parsing to encode the shapes and dynamics of persons during walking. Silhouettes exhibit accurate segmentation quality and robustness to environmental variations, but their low information entropy may result in sub-optimal performance. In contrast, human parsing provides fine-grained part segmentation with higher information entropy, but the segmentation quality may deteriorate due to the complex environments. To discover the advantages of silhouette and parsing and overcome their limitations, this paper proposes a novel cross-granularity alignment gait recognition method, named XGait, to unleash the power of gait representations of different granularity. To achieve this goal, the XGait first contains two branches of backbone encoders to map the silhouette sequences and the parsing sequences into two latent spaces, respectively. Moreover, to explore the complementary knowledge across the features of two representations, we design the Global Cross-granularity Module (GCM) and the Part Cross-granularity Module (PCM) after the two encoders. In particular, the GCM aims to enhance the quality of parsing features by leveraging global features from silhouettes, while the PCM aligns the dynamics of human parts between silhouette and parsing features using the high information entropy in parsing sequences. In addition, to effectively guide the alignment of two representations with different granularity at the part level, an elaborate-designed learnable division mechanism is proposed for the parsing features. Finally, comprehensive experiments on two large-scale gait datasets not only show the superior performance of XGait with the Rank-1 accuracy of 80.5% on Gait3D and 88.3% CCPG but also reflect the robustness of the learned features even under challenging conditions like occlusions and cloth changes
Jinkai Zheng, Xinchen Liu, Boyue Zhang 0004, Chenggang Yan 0001, Jiyong Zhang 0001, Wu Liu 0005, Yongdong Zhang 0001
ACM Multimedia6
2024 Watermarking Vision-Language Models
Shan Wan, Wu Liu 0005, Yijun Liu 0004, Feiniu Yuan, Chunli Meng
MMAsia2
2024 Prompting Industrial Anomaly Segment with Large Vision-Language Models
Jinheng Zhou, Wu Liu 0005, Guang Yang 0031, Feiniu Yuan
MMAsia2
2024 Bridging the Gap: Multi-Level Cross-Modality Joint Alignment for Visible-Infrared Person Re-Identification
abstract
Visible-Infrared person Re-IDentification (VI-ReID) is a challenging cross-modality image retrieval task that aims to match pedestrians’ images across visible and infrared cameras. To solve the modality gap, existing mainstream methods adopt a learning paradigm converting the image retrieval task into an image classification task with cross-entropy loss and auxiliary metric learning losses. These losses follow the strategy of adjusting the distribution of extracted embeddings to reduce the intra-class distance and increase the inter-class distance. However, such objectives do not precisely correspond to the final test setting of the retrieval task, resulting in a new gap at the optimization level. By rethinking these keys of VI-ReID, we propose a simple and effective method, the Multi-level Cross-modality Joint Alignment (MCJA), bridging both the modality and objective-level gap. For the former, we design the Visible-Infrared Modality Coordinator in the image space and propose the Modality Distribution Adapter in the feature space, effectively reducing modality discrepancy of the feature extraction process. For the latter, we introduce a new Cross-Modality Retrieval loss. It is the first work to constrain from the perspective of the ranking list in the VI-ReID, aligning with the goal of the testing stage. Moreover, to strengthen the robustness and cross-modality retrieval ability, we further introduce a Multi-Spectral Enhanced Ranking strategy for the testing phase. Based on the global feature only, our method outperforms existing methods by a large margin, achieving the remarkable rank-1 of 89.51% and mAP of 87.58% on the most challenging single-shot setting and all-search mode of the SYSU-MM01 dataset.
Tengfei Liang, Yi Jin 0001, Wu Liu 0005, Tao Wang 0011, Songhe Feng, Yidong Li
IEEE Trans. Circuits Syst. Video Technol.3
2024 A Generative-Based Image Fusion Strategy for Visible-Infrared Person Re-Identification
abstract
Cross-modality person re-identification task is a challenging task aiming to recognize images of the same identity between different modalities. To alleviate the cross-modality discrepancies between images, existing approaches mainly guide models to mine modality invariant features. Although those approaches are effective, they lose the modality-specific features that include important information beneficial to VI-ReID. Therefore, some approaches are using generative adversarial networks to compensate for modality information. However, the quality of images generated by these methods is usually poor, and most of them focus only on the learning of modality-sharable features. To solve these problems, this paper proposes a generative-based cross-modality image fusion strategy (GC-IFS), which can generate high-quality cross-modality paired images and fuse the information of the two modalities. Firstly, considering the importance of the identity discriminative information of the generated image, we propose a contrastive-learning image generation (CLIG) network to generate cross-modality paired images. Meanwhile, to fully integrate and utilize the information of the two modalities and eliminate the influence of cross-modality discrepancies, we design a part-based dual multi-modality feature fusion (P-DMFF) module to extract the unified feature representation. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate that our strategy outperforms the state-of-the-art methods for the VI-ReID task.
Jia Qi, Tengfei Liang, Wu Liu 0005, Yidong Li, Yi Jin 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 CDKM: Common and Distinct Knowledge Mining Network With Content Interaction for Dense Captioning
abstract
The dense captioning task aims at detecting multiple salient regions of an image and describing them separately in natural language. Although significant advancements in the field of dense captioning have been made, there are still some limitations to existing methods in recent years. On the one hand, most dense captioning methods lack strong target detection capabilities and struggle to cover all relevant content when dealing with target-intensive images. On the other hand, current transformer-based methods are powerful but neglect the acquisition and utilization of contextual information, hindering the visual understanding of local areas. To address these issues, we propose a common and distinct knowledge-mining network with content interaction for the task of dense captioning. Our network has a knowledge mining mechanism that improves the detection of salient targets by capturing common and distinct knowledge from multi-scale features. We further propose a content interaction module that combines region features into a unique context based on their correlation. Our experiments on various benchmarks have shown that the proposed method outperforms the current state-of-the-art methods.
Hongyu Deng, Yushan Xie, Qi Wang 0079, Weijian Ruan, Wu Liu 0005, Yong-Jin Liu 0001
IEEE Trans. Multim.6
2024 Learning Monocular Regression of 3D People in Crowds via Scene-Aware Blending and De-Occlusion
abstract
In this study, we address the challenge of estimating 3D body pose, shape, and depth relationships from single RGB images in crowded scenes. The difficulty lies in the limited availability of in-the-wild training samples, which feature densely populated scenes. To mitigate this issue, we introduce a synthesis-based approach that fuses multiple human samples into a single composite scene. Our innovative scene-aware blending technique maintains human-scene consistency by positioning individuals within plausible locations and adjusting their scales to conform to 3D settings. Furthermore, our method enables flexible per-subject occlusion management during the blending process, bolstering the robustness of 3D human body representations through a novel de-occlusion training scheme. We present a one-stage model, CBD, designed to learn monocular regression of 3D people in crowds by leveraging blending and de-occlusion techniques. Our quantitative and qualitative evaluations on four benchmark datasets reveal that CBD surpasses existing state-of-the-art approaches in terms of 3D human pose and mesh regression accuracy, thereby establishing it as a promising solution for monocular 3D human mesh recovery in densely populated scenes.
Yu Sun 0030, Lubing Xu, Qian Bao, Wu Liu 0005, Wenpeng Gao, Yili Fu 0001
IEEE Trans. Multim.4
2024 SigFormer: Sparse Signal-guided Transformer for Multi-modal Action Segmentation
abstract
Multi-modal human action segmentation is a critical and challenging task with a wide range of applications. Nowadays, the majority of approaches concentrate on the fusion of dense signals (i.e., RGB, optical flow, and depth maps). However, the potential contributions of sparse IoT sensor signals, which can be crucial for achieving accurate recognition, have not been fully explored. To make up for this, we introduce a S parse s i gnal- g uided Transformer ( SigFormer ) to combine both dense and sparse signals. We employ mask attention to fuse localized features by constraining cross-attention within the regions where sparse signals are valid. However, since sparse signals are discrete, they lack sufficient information about the temporal action boundaries. Therefore, in SigFormer, we propose to emphasize the boundary information at two stages to alleviate this problem. In the first feature extraction stage, we introduce an intermediate bottleneck module to jointly learn both category and boundary features of each dense modality through the inner loss functions. After the fusion of dense modalities and sparse signals, we then devise a two-branch architecture that explicitly models the interrelationship between action category and temporal boundary. Experimental results demonstrate that SigFormer outperforms the state-of-the-art approaches on a multi-modal action segmentation dataset from real industrial environments, reaching an outstanding F1 score of 0.958. The codes and pre-trained models have been made available at https://github.com/LIUQI-creat/SigFormer .
Qi Liu 0081, Xinchen Liu, Kun Liu 0016, Xiaoyan Gu 0001, Wu Liu 0005
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Learning to Segment Every Referring Object Point by Point
abstract
Referring Expression Segmentation (RES) can facilitate pixel-level semantic alignment between vision and language. Most of the existing RES approaches require massive pixel-level annotations, which are expensive and exhaustive. In this paper, we propose a new partially supervised training paradigm for RES, i.e., training using abundant referring bounding boxes and only a few (e.g., 1%) pixel-level referring masks. To maximize the transferability from the REC model, we construct our model based on the point-based sequence prediction model. We propose the co-content teacher-forcing to make the model explicitly associate the point coordinates (scale values) with the referred spatial features, which alleviates the exposure bias caused by the limited segmentation masks. To make the most of referring bounding box annotations, we further propose the resampling pseudo points strategy to select more accurate pseudo-points as supervision. Extensive experiments show that our model achieves 52.06% in terms of accuracy (versus 58.93% in fully supervised setting) on Re-fCOCO+@testA, when only using 1% of the mask annotations. Code is available at https://github.com/qumengxue/Partial-RES.git.
Mengxue Qu, Yu Wu 0011, Yunchao Wei, Wu Liu 0005, Xiaodan Liang, Yao Zhao 0001
CVPR4
2023 TRACE: 5D Temporal Regression of Avatars with Dynamic Cameras in 3D Environments
abstract
Although the estimation of 3D human pose and shape (HPS) is rapidly progressing, current methods still cannot reliably estimate moving humans in global coordinates, which is critical for many applications. This is particularly challenging when the camera is also moving, entangling human and camera motion. To address these issues, we adopt a novel 5D representation (space, time, and identity) that enables end-to-end reasoning about people in scenes. Our method, called TRACE, introduces several novel architectural components. Most importantly, it uses two new “maps” to reason about the 3D trajectory of people over time in camera, and world, coordinates. An additional memory unit enables persistent tracking of people even during long occlusions. TRACE is the first one-stage method to jointly recover and track 3D humans in global coordinates from dynamic cameras. By training it end-to-end, and using full image information, TRACE achieves state-of-the-art performance on tracking and HPS benchmarks. The code11https://www.yusun.work/TRACE/TRACE.html and dataset22https://github.com/Arthur151/DynaCam are released for research purposes.
Yu Sun 0030, Qian Bao, Wu Liu 0005, Tao Mei 0001, Michael J. Black
CVPR3
2023 CoCa: A Connectivity-Aware Cascade Framework for Histology Gland Segmentation
abstract
Gland segmentation is crucial for computer-aided diagnosis of adenocarcinoma. However, Topologically Critical Areas (TCAs), such as background tissues between two adjacent glands, can easily cause under- or over-connection of gland topological structures that may lead to the opposite diagnostic of the malignancy degree. Therefore, we provide a novel perspective for gland segmentation by incorporating gland connectivity information to locate critical errors within TCAs. We propose a Connectivity-Aware Cascade framework (CoCa) that explicitly encodes gland connectivity information into the network to locate all connectivity errors during training and then leverage attention operations to focus on these errors. Since under- or over-connected glands can change the Betti number (e.g., number of connected components) of glands, we design a Connectivity Refinement Module (CRM) to compare the Betti number of each gland to locate connectivity errors. We propose CoCa-Net to mine the topological relations among different biomedical entities to guide gland prediction. We also use contrastive learning to separate pixel embeddings of different classes within TCAs through our connectivity-aware hard example sampling strategy. Extensive experiments on the GlaS and CRAG datasets demonstrate the effectiveness of CoCa over state-of-the-art methods.
Yu Bai 0020, Bo Zhang 0032, Zheng Zhang 0038, Wu Liu 0005, Xiangyang Gong, Wendong Wang 0003
ACM Multimedia4
2023 FastReID: A Pytorch Toolbox for General Instance Re-identification
abstract
General Instance Re-identification is a very important task in computer vision, which can be widely used in many practical applications, such as person/vehicle re-identification, face recognition, wildlife protection, commodity tracing, snapshots, and so on. To meet the increasing application demand for general instance re-identification, we present FastReID as a widely used software system. In FastReID, the highly modular and extensible design makes it easy for the researcher to achieve new research ideas. Friendly manageable system configuration and engineering deployment functions allow practitioners to quickly deploy models into productions. We have implemented some state-of-the-art projects, including person re-id, partial re-id, cross-domain re-id, and vehicle re-id. Moreover, we plan to release these pre-trained models on multiple benchmark datasets. FastReID is by far the most general and high-performance toolbox that supports single and multiple GPU servers, it can reproduce our project results very easily. The source codes and models have been released at https://github.com/JDAI-CV/fast-reid.
Lingxiao He, Xingyu Liao, Wu Liu 0005, Xinchen Liu, Peng Cheng 0002, Tao Mei 0001
ACM Multimedia3
2023 HCMA '23: 4th International Workshop on Human-Centric Multimedia Analysis
abstract
Understanding human interactions within diverse media contexts has emerged as a fundamental challenge. The explosive growth of multimedia data not only provides opportunities for human-centirc analysis but also increases the complexity of processing multimodal data. To address this pivotal challenge and explore its multifaceted dimensions, the Fourth International Workshop on Human-Centric Multimedia Analysis is concentrated on the tasks of human-centric analysis with multimedia and multimodal information. By delving into the nuances of human behavior within multimedia, this workshop aims to uncover novel insights, showcase innovative methodologies, and discuss future directions. With a spotlight on cutting-edge research and a focus on real-world applications, the workshop seeks to equip researchers and practitioners with the tools and knowledge to navigate the intricacies of human-centric multimedia analysis.
Jingkuan Song, Wu Liu 0005, Xinchen Liu, Dingwen Zhang, Chaowei Fang, Hongyuan Zhu 0002, Wenbing Huang 0001, John R. Smith, Xin Wang 0019
ACM Multimedia2
2023 Factorized Omnidirectional Representation based Vision GNN for Anisotropic 3D Multimodal MR Image Segmentation
abstract
Anisotropy arises due to the influence of scanning equipment and parameters, resulting in a distance between slices that is often much greater than the actual distance represented by a single pixel within each slice. This can lead to inefficiency or ineffectiveness in 3D convolution. To address the anisotropy issue, we propose FOrViG, an asymmetric vision graph neural network (GNN) framework that captures the correlation between different slices by constructing a graph for multi-slice images and aggregating information from adjacent nodes. This allows FOrViG to efficiently extract 3D spatial scale information, and effectively identify feature nodes associated with small lesions efficiently, thereby improving the accuracy of lesion segmentation on anisotropic 3D multimodal MR images. As far as we know, this is the first study that adopts GNN to address anisotropy issues. Additionally, we also design a factorized omnidirectional representation method and a supervised multi-perspective contrastive learning strategy to enhance the capability of FOrViG in learning multi-scale omnidirectional presentation information, graphics construction, and distinguishing foreground from background. Extensive experiments on the PI-CAI dataset demonstrate that FOrViG significantly outperforms several state-of-the-art 3D segmentation algorithms.
Bo Zhang 0032, Yunpeng Tan, Zheng Zhang 0038, Wu Liu 0005, Hui Gao 0002, Zhijun Xi, Wendong Wang 0003
ACM Multimedia4
2023 Parsing is All You Need for Accurate Gait Recognition in the Wild
abstract
Binary silhouettes and keypoint-based skeletons have dominated human gait recognition studies for decades since they are easy to extract from video frames. Despite their success in gait recognition for in-the-lab environments, they usually fail in real-world scenarios due to their low information entropy for gait representations. To achieve accurate gait recognition in the wild, this paper presents a novel gait representation, named Gait Parsing Sequence (GPS). GPSs are sequences of fine-grained human segmentation, i.e., human parsing, extracted from video frames, so they have much higher information entropy to encode the shapes and dynamics of fine-grained human parts during walking. Moreover, to effectively explore the capability of the GPS representation, we propose a novel human parsing-based gait recognition framework, named ParsingGait. ParsingGait contains a Convolutional Neural Network (CNN)-based backbone and two light-weighted heads. The first head extracts global semantic features from GPSs, while the other one learns mutual information of part-level features through Graph Convolutional Networks to model the detailed dynamics of human walking. Furthermore, due to the lack of suitable datasets, we build the first parsing-based dataset for gait recognition in the wild, named Gait3D-Parsing, by extending the large-scale and challenging Gait3D dataset. Based on Gait3D-Parsing, we comprehensively evaluate our method and existing gait recognition methods. Specifically, ParsingGait achieves a 17.5% Rank-1 increase compared with the state-of-the-art silhouette-based method. In addition, by replacing silhouettes with GPSs, current gait recognition methods achieve about 12.5% ~ 19.2% improvements in Rank-1 accuracy. The experimental results show a significant improvement in accuracy brought by the GPS representation and the superiority of ParsingGait.
Jinkai Zheng, Xinchen Liu, Shuai Wang 0003, Chenggang Yan 0001, Wu Liu 0005
ACM Multimedia6
2023 RIO: A Benchmark for Reasoning Intention-Oriented Objects in Open Environments
abstract
Intention-oriented object detection aims to detect desired objects based on specific intentions or requirements. For instance, when we desire to "lie down and rest", we instinctively seek out a suitable option such as a "bed" or a "sofa" that can fulfill our needs. Previous work in this area is limited either by the number of intention descriptions or by the affordance vocabulary available for intention objects. These limitations make it challenging to handle intentions in open environments effectively. To facilitate this research, we construct a comprehensive dataset called Reasoning Intention-Oriented Objects (RIO). In particular, RIO is specifically designed to incorporate diverse real-world scenarios and a wide range of object categories. It offers the following key features: 1) intention descriptions in RIO are represented as natural sentences rather than a mere word or verb phrase, making them more practical and meaningful; 2) the intention descriptions are contextually relevant to the scene, enabling a broader range of potential functionalities associated with the objects; 3) the dataset comprises a total of 40,214 images and 130,585 intention-object pairs. With the proposed RIO, we evaluate the ability of some existing models to reason intention-oriented objects in open environments.
Mengxue Qu, Yu Wu 0011, Wu Liu 0005, Xiaodan Liang, Jingkuan Song, Yao Zhao 0001, Yunchao Wei
NeurIPS3
2023 Bird-Count: a multi-modality benchmark and system for bird population counting in the wild
Hongchang Wang, Huaxiang Lu, Huimin Guo, Haifang Jian, Chuang Gan 0001, Wu Liu 0005
Multim. Tools Appl.6
2023 Cross-Modality Transformer With Modality Mining for Visible-Infrared Person Re-Identification
abstract
The visible-infrared person re-identification (VI-ReID) is a challenging ReID task, which aims to retrieve and match the same identity's images between the heterogeneous visible and infrared modalities. Thus, the core of this task is to bridge the huge gap between these two modalities. The existing methods mainly face the problem of insufficient perception of modality information, and can not learn good discriminative modality-invariant embeddings for identities, which limits their performance. To solve these problems, we propose a new cross-modality transformer-based method (CMTR) for this visible-infrared person re-identification task, which can explicitly mine the information of each modality and generate better discriminative features based on it. Specifically, to capture inherent characteristics of modalities, we design the novel modality embeddings, which are fused with token embeddings to encode modality information directly. Moreover, to enhance representation of modality embeddings and adjust the distribution of embeddings, we further propose a modality-aware enhancement loss based on the learned modality information, which contains two components to reduce intra-class distance and enlarging inter-class distance simultaneously. To our knowledge, this is the first exploration of applying pure transformer network to the cross-modality re-identification task. We implement extensive experiments on the public SYSU-MM01 and RegDB datasets, and compared with previous methods, our method achieves good performance with more compact and powerful embeddings for the cross-modality retrieval.
Tengfei Liang, Yi Jin 0001, Wu Liu 0005, Yidong Li
IEEE Trans. Multim.3
2023 Semantic Embedding Guided Attention with Explicit Visual Feature Fusion for Video Captioning
abstract
Video captioning, which bridges vision and language, is a fundamental yet challenging task in computer vision. To generate accurate and comprehensive sentences, both visual and semantic information is quite important. However, most existing methods simply concatenate different types of features and ignore the interactions between them. In addition, there is a large semantic gap between visual feature space and semantic embedding space, making the task very challenging. To address these issues, we propose a framework named semantic embedding guided attention with Explicit visual Feature Fusion for vidEo CapTioning, EFFECT for short, in which we design an explicit visual-feature fusion (EVF) scheme to capture the pairwise interactions between multiple visual modalities and fuse multimodal visual features of videos in an explicit way. Furthermore, we propose a novel attention mechanism called semantic embedding guided attention (SEGA ), which cooperates with the temporal attention to generate a joint attention map. Specifically, in SEGA, the semantic word embedding information is leveraged to guide the model to pay more attention to the most correlated visual features at each decoding stage. In this way, the semantic gap between visual and semantic space is alleviated to some extent. To evaluate the proposed model, we conduct extensive experiments on two widely used datasets, i.e., MSVD and MSR-VTT. The experimental results demonstrate that our approach achieves state-of-the-art results in terms of four evaluation metrics.
Tian-Zi Niu, Xin Luo 0006, Wu Liu 0005, Xin-Shun Xu
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Introduction to the Special Issue on Trustworthy Multimedia Computing and Applications in Urban Scenes
abstract
Special Issue Part 1 (Issue 3) and Part 2 (Issue 4) of AIEDAM are based on a workshop on Learning and Creativity held at the 2002 conference on Artificial Intelligence in Design, AID '02 (www.cad.strath.ac.uk/AID02_workshop/Workshop_webpage.html; Gero, ...
Wu Liu 0005, Hailin Shi, Yunchao Wei, Dan Zeng 0001, Nicu Sebe, Jiebo Luo 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Putting People in their Place: Monocular Regression of 3D People in Depth
abstract
Given an image with multiple people, our goal is to directly regress the pose and shape of all the people as well as their relative depth. Inferring the depth of a person in an image, however, is fundamentally ambiguous without knowing their height. This is particularly problematic when the scene contains people of very different sizes, e.g. from infants to adults. To solve this, we need several things. First, we develop a novel method to infer the poses and depth of multiple people in a single image. While previous work that estimates multiple people does so by reasoning in the image plane, our method, called BEV, adds an additional imaginary Bird's-Eye-View representation to explicitly reason about depth. BEV reasons simultaneously about body centers in the image and in depth and, by combing these, estimates 3D body position. Unlike prior work, BEV is a single-shot method that is end-to-end differentiable. Second, height varies with age, making it impossible to resolve depth without also estimating the age of people in the image. To do so, we exploit a 3D body model space that lets BEV infer shapes from infants to adults. Third, to train BEV, we need a new dataset. Specifically, we create a “Relative Human” (RH) dataset that includes age labels and relative depth relationships between the people in the images. Extensive experiments on RH and AGORA demonstrate the effectiveness of the model and training scheme. BEV out-performs existing methods on depth reasoning, child shape estimation, and robustness to occlusion. The code11https://github.com/Arthur151/ROMP and dataset22https://github.com/Arthur151/Relative_Human are released for research purposes.
Yu Sun 0030, Wu Liu 0005, Qian Bao, Yili Fu 0001, Tao Mei 0001, Michael J. Black
CVPR2
2022 Gait Recognition in the Wild with Dense 3D Representations and A Benchmark
abstract
Existing studies for gait recognition are dominated by 2D representations like the silhouette or skeleton of the human body in constrained scenes. However, humans live and walk in the unconstrained 3D space, so projecting the 3D human body onto the 2D plane will discard a lot of crucial information like the viewpoint, shape, and dynamics for gait recognition. Therefore, this paper aims to explore dense 3D representations for gait recognition in the wild, which is a practical yet neglected problem. In particular, we propose a novel framework to explore the 3D Skinned Multi-Person Linear (SMPL) model of the human body for gait recognition, named SMPLGait. Our framework has two elaborately-designed branches of which one extracts appearance features from silhouettes, the other learns knowledge of 3D viewpoints and shapes from the 3D SMPL model. In addition, due to the lack of suitable datasets, we build the first large-scale 3D representation-based gait recognition dataset, named Gait3D. It contains 4,000 subjects and over 25,000 sequences extracted from 39 cameras in an unconstrained indoor scene. More importantly, it provides 3D SMPL models recovered from video frames which can provide dense 3D information of body shape, viewpoint, and dynamics. Based on Gait3D, we comprehensively compare our method with existing gait recognition approaches, which reflects the superior performance of our framework and the potential of 3D representations for gait recognition in the wild. The code and dataset are available at: https://gait3d.github.io.
Jinkai Zheng, Xinchen Liu, Wu Liu 0005, Lingxiao He, Chenggang Yan 0001, Tao Mei 0001
CVPR3
2022 SiRi: A Simple Selective Retraining Mechanism for Transformer-Based Visual Grounding
Mengxue Qu, Yu Wu 0011, Wu Liu 0005, Qiqi Gong, Xiaodan Liang, Olga Russakovsky, Yao Zhao 0001, Yunchao Wei
ECCV (35)3
2022 CAViT: Contextual Alignment Vision Transformer for Video Object Re-identification
Jinlin Wu, Lingxiao He, Wu Liu 0005, Yang Yang 0062, Zhen Lei 0001, Tao Mei 0001, Stan Z. Li
ECCV (14)3
2022 Genre-Conditioned Long-Term 3D Dance Generation Driven by Music
abstract
Dancing to music is an artistic behavior of humans, however, letting machines generate dances from music is still challenging. Most existing works have been made progress in tackling the problem of motion prediction conditioned by music, yet they rarely consider the importance of the musical genre. In this paper, we focus on generating long-term 3D dance from music with a specific genre. Specifically, we construct a pure transformer-based architecture to correlate motion features and music features. To utilize the genre information, we propose to embed the genre categories into the transformer decoder so that it can guide every frame. Moreover, different from previous inference schemes, we introduce the motion queries to output the dance sequence in parallel that significantly improves the efficiency. Extensive experiments on AIST++[1] dataset show that our model outperforms state-of-the-art methods with a much faster inference speed.
Yuhang Huang 0006, Junjie Zhang 0002, Qian Bao, Dan Zeng 0001, Zhineng Chen, Wu Liu 0005
ICASSP7
2022 Learning Monocular Mesh Recovery of Multiple Body Parts Via Synthesis
abstract
In this paper, we focus on simultaneously recovering the 3D mesh of multiple body parts from a single RGB image. One of the main challenges is that available datasets with full-body 3D annotations are very limited. This results in poor generalization ability of existing learning-based methods. Existing optimization-based methods iteratively fit the 3D mesh to the 2d pose, which is very time-consuming. To address these limitations, we propose to integrate multiple 3D single-body-part datasets to create a highly diverse whole-body 3D motion space for learning from controllable synthetics. Compared with the learning-based approaches, the proposed method greatly alleviates the reliance on training data. Compared with the optimization-based approaches, the proposed method is a hundred times faster. Our proposed method also outperforms previous state-of-the-art methods on CMU Panoptic dataset.
Yu Sun 0030, Qian Bao, Wu Liu 0005, Wenpeng Gao, Yili Fu 0001
ICASSP4
2022 Part-level Action Parsing via a Pose-guided Coarse-to-Fine Framework
abstract
Action recognition from videos, i.e., classifying a video into one of the pre-defined action types, has been a popular topic in the communities of artificial intelligence, multimedia, and signal processing. However, existing methods usually consider an input video as a whole and learn models, e.g., Convolutional Neural Networks (CNNs), with coarse video-level class labels. These methods can only output an action class for the video, but cannot provide fine-grained and explainable cues to answer why the video shows a specific action. Therefore, researchers start to focus on a new task, Part-level Action Parsing (PAP), which aims to not only predict the video-level action but also recognize the frame-level fine-grained actions or interactions of body parts for each person in the video. To this end, we propose a coarse-to-fine framework for this challenging task. In particular, our framework first predicts the video-level class of the input video, then localizes the body parts and predicts the part-level action. Moreover, to balance the accuracy and computation in part-level action parsing, we propose to recognize the part-level actions by segment-level features. Furthermore, to overcome the ambiguity of body parts, we propose a pose-guided positional embedding method to accurately localize body parts. Through comprehensive experiments on a large-scale dataset, i.e., Kinetics-TPS, our framework achieves state-of-the-art performance and outperforms existing methods over 31.10% ROC score.
Xiaodong Chen 0011, Xinchen Liu, Wu Liu 0005, Kun Liu 0016, Yongdong Zhang 0001, Tao Mei 0001
ISCAS3
2022 MAPLE: Masked Pseudo-Labeling autoEncoder for Semi-supervised Point Cloud Action Recognition
abstract
Recognizing human actions from point cloud videos has attracted tremendous attention from both academia and industry due to its wide applications like automatic driving, robotics, and so on. However, current methods for point cloud action recognition usually require a huge amount of data with manual annotations and a complex backbone network with high computation cost, which makes it impractical for real-world applications. Therefore, this paper considers the task of semi-supervised point cloud action recognition. We propose a Masked Pseudo-Labeling autoEncoder (MAPLE) framework to learn effective representations with much fewer annotations for point cloud action recognition. In particular, we design a novel and efficient Decoupled spatial-temporal TransFormer (DestFormer) as the backbone of MAPLE. In DestFormer, the spatial and temporal dimensions of the 4D point cloud videos are decoupled to achieve an efficient self-attention for learning both long-term and short-term features. Moreover, to learn discriminative features from fewer annotations, we design a masked pseudo-labeling autoencoder structure to guide the DestFormer to reconstruct features of masked frames from the available frames. More importantly, for unlabeled data, we exploit the pseudo-labels from the classification head as the supervision signal for the reconstruction of features from the masked frames. Finally, comprehensive experiments demonstrate that MAPLE achieves superior results on three public benchmarks and outperforms the state-of-the-art method by 8.08% accuracy on the MSR-Action3D dataset.
Xiaodong Chen 0011, Wu Liu 0005, Xinchen Liu, Yongdong Zhang 0001, Jungong Han, Tao Mei 0001
ACM Multimedia2
2022 Keypoint-Guided Modality-Invariant Discriminative Learning for Visible-Infrared Person Re-identification
abstract
The visible-infrared person re-identification (VI-ReID) task aims to retrieve images of pedestrians across cameras with different modalities. In this task, the major challenges arise from two aspects: intra-class variations among images of the same identity, and cross-modality discrepancies between visible and infrared images. Existing methods mainly focus on the latter, attempting to alleviate the impact of modality discrepancy, which ignore the former issue of identity variations and achieve limited discrimination. To address both aspects, we propose a Keypoint-guided Modality-invariant Discriminative Learning (KMDL) method, which can simultaneously adapt to intra-ID variations and bridge the cross-modality gap. By introducing human keypoints, our method makes further exploration in the image space, feature space and loss constraints to solve the above issues. Specifically, considering the modality discrepancy in original images, we first design a Hue Jitter Augmentation (HJA) strategy, introducing the hue disturbance to alleviate color dependence in the input stage. To obtain discriminative fine-grained representation for retrieval, we design the Global-Keypoint Graph Module (GKGM) in feature space, which can directly extract keypoint-aligned features and mine relationships within global and keypoint embeddings. Based on these semantic local embeddings, we further propose the Keypoint-Aware Center (KAC) loss that can effectively adjust the feature distribution under the supervision of ID and keypoint to learn discriminative representation for the matching. Extensive experiments on SYSU-MM01 and RegDB datasets demonstrate the effectiveness of our KMDL method.
Tengfei Liang, Yi Jin 0001, Wu Liu 0005, Songhe Feng, Tao Wang 0011, Yidong Li
ACM Multimedia3
2022 Towards Causality Inference for Very Important Person Localization
abstract
Very Important Person Localization (VIPLoc) aims at detecting certain individuals in a given image, who are more attractive than others in the image. Existing uncontrolled VIPLoc benchmark assumes that the image has one single VIP, which is not suitable for actual application scenarios when multiple VIPs or no VIPs appear in the image. In this paper, we re-built a complex uncontrolled conditions (CUC) dataset to make the VIPLoc closer to the actual situation, containing no, single, and multiple VIPs. Existing methods use the hand-designed and deep learning strategies to extract the features of persons and analyze the differences between VIPs and other persons from the perspective of statistics. They are not explainable as to why the VIP located this output for that input. Thus, there exist the severe performance degradation when we use these models in real-world VIPLoc. Specifically, we establish a causal inference framework that unpacks the causes of previous methods and derives a new principled solution for VIPLoc. It treats the scene as confounding factor, allowing the ever-elusive confounding effects to be eliminated and the essential determinants to be uncovered. Through extensive experiments, our method outperforms the state-of-the-art methods on public VIPLoc datasets and the re-built CUC dataset.
Xiao Wang 0029, Zheng Wang 0007, Wu Liu 0005, Xin Xu 0007, Qijun Zhao, Shin'ichi Satoh 0001
ACM Multimedia3
2022 WOC: A Handy Webcam-based 3D Online Chatroom
abstract
We develop WOC, a webcam-based 3D virtual online chatroom for multi-person interaction, which captures the 3D motion of users and drives their individual 3D virtual avatars in real-time. Compared to the existing wearable equipment-based solution, WOC offers convenient and low-cost 3D motion capture with a single camera. To promote the immersive chat experience, WOC provides high-fidelity virtual avatar manipulation, which also supports the user-defined characters. With the distributed data flow service, the system delivers highly synchronized motion and voice for all users. Deployed on the website and no installation required, users can freely experience the virtual online chat at https://yanch.cloud/.
Chuanhang Yan, Yu Sun 0030, Qian Bao, Jinhui Pang, Wu Liu 0005, Tao Mei 0001
ACM Multimedia5
2022 Delving into the Frequency: Temporally Consistent Human Motion Transfer in the Fourier Space
abstract
Human motion transfer refers to synthesizing photo-realistic and temporally coherent videos that enable one person to imitate the motion of others. However, current synthetic videos suffer from the temporal inconsistency in sequential frames that significantly degrades the video quality, yet is far from solved by existing methods in the pixel domain. Recently, some works on DeepFake detection try to distinguish the natural and synthetic images in the frequency domain because of the frequency insufficiency of image synthesizing methods. Nonetheless, there is no work to study the temporal inconsistency of synthetic videos from the aspects of the frequency-domain gap between natural and synthetic videos. Therefore, in this paper, we propose to delve into the frequency space for temporally consistent human motion transfer. First of all, we make the first comprehensive analysis of natural and synthetic videos in the frequency domain to reveal the frequency gap in both the spatial dimension of individual frames and the temporal dimension of the video. To close the frequency gap between the natural and synthetic videos, we propose a novel Frequency-based human MOtion TRansfer framework, named FreMOTR, which can effectively mitigate the spatial artifacts and the temporal inconsistency of the synthesized videos. FreMOTR explores two novel frequency-based regularization modules: 1) the Frequency-domain Appearance Regularization (FAR) to improve the appearance of the person in individual frames and 2) Temporal Frequency Regularization (TFR) to guarantee the temporal consistency between adjacent frames. Finally, comprehensive experiments demonstrate that the FreMOTR not only yields superior performance in temporal consistency metrics but also improves the frame-level visual quality of synthetic videos. In particular, the temporal consistency metrics are improved by nearly 30% than the state-of-the-art model.
Guang Yang 0031, Wu Liu 0005, Xinchen Liu, Xiaoyan Gu 0001, Juan Cao 0001, Jintao Li 0001
ACM Multimedia2
2022 REMOT: A Region-to-Whole Framework for Realistic Human Motion Transfer
abstract
Human Video Motion Transfer (HVMT) aims to, given an image of a source person, generate his/her video that imitates the motion of the driving person. Existing methods for HVMT mainly exploit Generative Adversarial Networks (GANs) to perform the warping operation based on the flow estimated from the source person image and each driving video frame. However, these methods always generate obvious artifacts due to the dramatic differences in poses, scales, and shifts between the source person and the driving person. To overcome these challenges, this paper presents a novel REgion-to-whole human MOtion Transfer (REMOT) framework based on GANs. To generate realistic motions, the REMOT adopts a progressive generation paradigm: it first generates each body part in the driving pose without flow-based warping, then composites all parts into a complete person of the driving motion. Moreover, to preserve the natural global appearance, we design a Global Alignment Module to align the scale and position of the source person with those of the driving person based on their layouts. Furthermore, we propose a Texture Alignment Module to keep each part of the person aligned according to the similarity of the texture. Finally, through extensive quantitative and qualitative experiments, our REMOT achieves state-of-the-art results on two public benchmarks.
Quanwei Yang, Xinchen Liu, Wu Liu 0005, Hongtao Xie 0001, Xiaoyan Gu 0001, Lingyun Yu 0002, Yongdong Zhang 0001
ACM Multimedia3
2022 HCMA'22: 3rd International Workshop on Human-Centric Multimedia Analysis
abstract
The Third International Workshop on Human-Centric Multimedia Analysis concentrates on the tasks of human-centric analysis with multimedia and multimodal information. It involves multiple tasks such as face detection and recognition, human body pattern analysis, person re-identification, human action detection, etc. Today, multiple multimedia sensing technologies and large-scale computing infrastructures are emerging at a rapid velocity a wide variety of big multi-modality data for human-centric analysis, which provides rich knowledge to help tackle these challenges. Researchers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as intelligent surveillance, retailing, fashion design, and services. Therefore, this workshop aims to provide a platform to bridge the gap between the communities of human analysis and multimedia.
Dingwen Zhang, Chaowei Fang, Wu Liu 0005, Xinchen Liu, Jingkuan Song, Hongyuan Zhu 0002, Wenbing Huang 0001, John R. Smith
ACM Multimedia3
2022 Gait Recognition in the Wild with Multi-hop Temporal Switch
abstract
Existing studies for gait recognition are dominated by in-the-lab scenarios. Since people live in real-world senses, gait recognition in the wild is a more practical problem that has recently attracted the attention of the community of multimedia and computer vision. Current methods that obtain state-of-the-art performance on in-the-lab benchmarks achieve much worse accuracy on the recently proposed in-the-wild datasets because these methods can hardly model the varied temporal dynamics of gait sequences in unconstrained scenes. Therefore, this paper presents a novel multi-hop temporal switch method to achieve effective temporal modeling of gait patterns in real-world scenes. Concretely, we design a novel gait recognition network, named Multi-hop Temporal Switch Network (MTSGait), to learn spatial features and multi-scale temporal features simultaneously. Different from existing methods that use 3D convolutions for temporal modeling, our MTSGait models the temporal dynamics of gait sequences by 2D convolutions. By this means, it achieves high efficiency with fewer model parameters and reduces the difficulty in optimization compared with 3D convolution-based models. Based on the specific design of the 2D convolution kernels, our method can eliminate the misalignment of features among adjacent frames. In addition, a new sampling strategy, i.e., non-cyclic continuous sampling, is proposed to make the model learn more robust temporal features. Finally, the proposed method achieves superior performance on two public gait in-the-wild datasets, i.e., GREW and Gait3D, compared with state-of-the-art methods.
Jinkai Zheng, Xinchen Liu, Xiaoyan Gu 0001, Yaoqi Sun, Chuang Gan 0001, Jiyong Zhang 0001, Wu Liu 0005, Chenggang Yan 0001
ACM Multimedia7
2022 TICNet: A Target-Insight Correlation Network for Object Tracking
abstract
Recently, the correlation filter (CF) and Siamese network have become the two most popular frameworks in object tracking. Existing CF trackers, however, are limited by feature learning and context usage, making them sensitive to boundary effects. In contrast, Siamese trackers can easily suffer from the interference of semantic distractors. To address the above problems, we propose an end-to-end target-insight correlation network (TICNet) for object tracking, which aims at breaking the above limitations on top of a unified network. TICNet is an asymmetric dual-branch network involving a target-background awareness model (TBAM), a spatial-channel attention network (SCAN), and a distractor-aware filter (DAF) for end-to-end learning. Specifically, TBAM aims to distinguish a target from the background in the pixel level, yielding a target likelihood map based on color statistics to mine distractors for DAF learning. SCAN consists of a basic convolutional network, a channel-attention network, and a spatial-attention network, aiming to generate attentive weights to enhance the representation learning of the tracker. Especially, we formulate a differentiable DAF and employ it as a learnable layer in the network, thus helping suppress distracting regions in the background. During testing, DAF, together with TBAM, yields a response map for the final target estimation. Extensive experiments on seven benchmarks demonstrate that TICNet outperforms the state-of-the-art methods while running at real-time speed.
Weijian Ruan, Mang Ye, Yi Wu 0001, Wu Liu 0005, Jun Chen 0001, Chao Liang 0001, Ge Li 0002, Chia-Wen Lin
IEEE Trans. Cybern.4
2022 FasterPose: A Faster Simple Baseline for Human Pose Estimation
abstract
The performance of human pose estimation depends on the spatial accuracy of keypoint localization. Most existing methods pursue the spatial accuracy through learning the high-resolution (HR) representation from input images. By the experimental analysis, we find that the HR representation leads to a sharp increase of computational cost, while the accuracy improvement remains marginal compared with the low-resolution (LR) representation. In this article, we propose a design paradigm for cost-effective network with LR representation for efficient pose estimation, named FasterPose. Whereas the LR design largely shrinks the model complexity, how to effectively train the network with respect to the spatial accuracy is a concomitant challenge. We study the training behavior of FasterPose and formulate a novel regressive cross-entropy (RCE) loss function for accelerating the convergence and promoting the accuracy. The RCE loss generalizes the ordinary cross-entropy loss from the binary supervision to a continuous range, thus the training of pose estimation network is able to benefit from the sigmoid function. By doing so, the output heatmap can be inferred from the LR features without loss of spatial accuracy, while the computational cost and model size has been significantly reduced. Compared with the previously dominant network of pose estimation, our method reduces 58% of the FLOPs and simultaneously gains 1.3% improvement of accuracy. Extensive experiments show that FasterPose yields promising results on the common benchmarks, i.e., COCO and MPII, consistently validating the effectiveness and efficiency for practical utilization, especially the low-latency and low-energy-budget applications in the non-GPU scenarios.
Hanbin Dai, Hailin Shi, Wu Liu 0005, Linfang Wang, Yinglu Liu, Tao Mei 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Group-aware Label Transfer for Domain Adaptive Person Re-identification
abstract
Unsupervised Domain Adaptive (UDA) person re-identification (ReID) aims at adapting the model trained on a labeled source-domain dataset to a target-domain dataset without any further annotations. Most successful UDA-ReID approaches combine clustering-based pseudo-label prediction with representation learning and perform the two steps in an alternating fashion. However, offline interaction between these two steps may allow noisy pseudo labels to substantially hinder the capability of the model. In this paper, we propose a Group-aware Label Transfer (GLT) algorithm, which enables the online interaction and mutual promotion of pseudo-label prediction and representation learning. Specifically, a label transfer algorithm simultaneously uses pseudo labels to train the data while refining the pseudo labels as an online clustering algorithm. It treats the online label refinery problem as an optimal transport problem, which explores the minimum cost for assigning M samples to N pseudo labels. More importantly, we introduce a group-aware strategy to assign implicit attribute group IDs to samples. The combination of the online label refining algorithm and the group-aware strategy can better correct the noisy pseudo label in an online fashion and narrow down the search space of the target identity. The effectiveness of the proposed GLT is demonstrated by the experimental results (Rank-1 accuracy) for Market1501→DukeMTMC (82.0%) and DukeMTMC→Market1501 (92.2%), remarkably closing the gap between unsupervised and supervised performance on person re-identification.1
Kecheng Zheng, Wu Liu 0005, Lingxiao He, Tao Mei 0001, Jiebo Luo 0001, Zhengjun Zha
CVPR2
2021 Explainable Person Re-Identification with Attribute-guided Metric Distillation
abstract
Despite the great progress of person re-identification (ReID) with the adoption of Convolutional Neural Networks, current ReID models are opaque and only outputs a scalar distance between two persons. There are few methods providing users semantically understandable explanations for why two persons are the same one or not. In this paper, we propose a post-hoc method, named Attribute-guided Metric Distillation (AMD), to explain existing ReID models. This is the first method to explore attributes to answer: 1) what and where the attributes make two persons different, and 2) how much each attribute contributes to the difference. In AMD, we design a pluggable interpreter network for target models to generate quantitative contributions of attributes and visualize accurate attention maps of the most discriminative attributes. To achieve this goal, we propose a metric distillation loss by which the interpreter learns to decompose the distance of two persons into components of attributes with knowledge distilled from the target model. Moreover, we propose an attribute prior loss to make the interpreter generate attribute-guided attention maps and to eliminate biases caused by the imbalanced distribution of attributes. This loss can guide the interpreter to focus on the exclusive and discriminative attributes rather than the large-area but common attributes of two persons. Comprehensive experiments show that the interpreter can generate effective and intuitive explanations for varied models and generalize well under cross-domain settings. As a by-product, the accuracy of target models can be further improved with our interpreter.1
Xiaodong Chen 0011, Xinchen Liu, Wu Liu 0005, Xiao-Ping Zhang 0002, Yongdong Zhang 0001, Tao Mei 0001
ICCV3
2021 Learning Instance-level Spatial-Temporal Patterns for Person Re-identification
abstract
Person re-identification (Re-ID) aims to match pedestrians under dis-joint cameras. Most Re-ID methods formulate it as visual representation learning and image search, and its accuracy is consequently affected greatly by the search space. Spatial-temporal information has been proven to be efficient to filter irrelevant negative samples and significantly improve Re-ID accuracy. However, existing spatial-temporal person Re-ID methods are still rough and do not exploit spatial-temporal information sufficiently. In this paper, we propose a novel Instance-level and Spatial-Temporal Disentangled Re-ID method (InSTD), to improve Re-ID accuracy. In our proposed framework, personalized information such as moving direction is explicitly considered to further narrow down the search space. Besides, the spatial-temporal transferring probability is disentangled from joint distribution to marginal distribution, so that outliers can also be well modeled. Abundant experimental analyses are presented, which demonstrates the superiority and provides more insights into our method. The proposed method achieves mAP of 90.8% on Market-1501 and 89.1% on DukeMTMC-reID, improving from the baseline 82.2% and 72.7%, respectively. Besides, in order to provide a better benchmark for person re-identification, we release a cleaned data list of DukeMTMC-reID with this paper: https://github.com/RenMin1991/cleaned-DukeMTMC-reID/
Lingxiao He, Xingyu Liao, Wu Liu 0005, Yunlong Wang 0003, Tieniu Tan
ICCV4
2021 Monocular, One-stage, Regression of Multiple 3D People
abstract
This paper focuses on the regression of multiple 3D people from a single RGB image. Existing approaches predominantly follow a multi-stage pipeline that first detects people in bounding boxes and then independently regresses their 3D body meshes. In contrast, we propose to Regress all meshes in a One-stage fashion for Multiple 3D People (termed ROMP). The approach is conceptually simple, bounding box-free, and able to learn a per-pixel representation in an end-to-end manner. Our method simultaneously predicts a Body Center heatmap and a Mesh Parameter map, which can jointly describe the 3D body mesh on the pixel level. Through a body-center-guided sampling process, the body mesh parameters of all people in the image are easily extracted from the Mesh Parameter map. Equipped with such a fine-grained representation, our one-stage framework is free of the complex multi-stage process and more robust to occlusion. Compared with state-of-the-art methods, ROMP achieves superior performance on the challenging multi-person benchmarks, including 3DPW and CMU Panoptic. Experiments on crowded/occluded datasets demonstrate the robustness under various types of occlusion. The code, released at https://github.com/Arthur151/ROMP, is the first real-time implementation of monocular multi-person 3D mesh regression.
Yu Sun 0030, Qian Bao, Wu Liu 0005, Yili Fu 0001, Michael J. Black, Tao Mei 0001
ICCV3
2021 Neural Architecture Search for Joint Human Parsing and Pose Estimation
abstract
Human parsing and pose estimation are crucial for the understanding of human behaviors. Since these tasks are closely related, employing one unified model to perform two tasks simultaneously allows them to benefit from each other. However, since human parsing is a pixel-wise classification process while pose estimation is usually a regression task, it is non-trivial to extract discriminative features for both tasks while modeling their correlation in the joint learning fashion. Recent studies have shown that Neural Architecture Search (NAS) has the ability to allocate efficient feature connections for specific tasks automatically. With the spirit of NAS, we propose to search for an efficient network architecture (NPPNet) to tackle two tasks at the same time. On the one hand, to extract task-specific features for the two tasks and lay the foundation for the further searching of feature interaction, we propose to search their encoder-decoder architectures, respectively. On the other hand, to ensure two tasks fully communicate with each other, we propose to embed NAS units in both multi-scale feature interaction and high-level feature fusion to establish optimal connections between two tasks. Experimental results on both parsing and pose estimation benchmark datasets have demonstrated that the searched model achieves state-of-the-art performances on both tasks.1
Dan Zeng 0001, Yuhang Huang 0006, Qian Bao, Junjie Zhang 0002, Chi Su, Wu Liu 0005
ICCV6
2021 TraND: Transferable Neighborhood Discovery for Unsupervised Cross-Domain Gait Recognition
abstract
Gait, i.e., the movement pattern of human limbs during locomotion, is a promising biometrie for identification of persons. Despite significant improvement in gait recognition with deep learning, existing studies still neglect a more practical but challenging scenario - unsupervised cross-domain gait recognition which aims to learn a model on a labeled dataset then adapt it to an unlabeled dataset. Due to the domain shift and class gap, directly applying a model trained on one source dataset to other target datasets usually obtains very poor results. Therefore, this paper proposes a Transferable Neighborhood Discovery (TraND) framework to bridge the domain gap for unsupervised cross-domain gait recognition. To learn effective prior knowledge for gait representation, we first adopt a backbone network pre- trained on the labeled source data in a supervised manner. Then we design an end-to-end trainable approach to automatically discover the confident neighborhoods of unlabeled samples in the latent space. During training, the class consistency indicator is adopted to select confident neighborhoods of samples based on their entropy measurements. Moreover, we explore a high- entropy-first neighbor selection strategy, which can effectively transfer prior knowledge to the target domain. Our method achieves the state-of-the-art results on two public datasets, i.e., CASIA-B and OU-LP.
Jinkai Zheng, Xinchen Liu, Chenggang Yan 0001, Jiyong Zhang 0001, Wu Liu 0005, Xiao-Ping Zhang 0002, Tao Mei 0001
ISCAS5
2021 MSO: Multi-Feature Space Joint Optimization Network for RGB-Infrared Person Re-Identification
abstract
The RGB-infrared cross-modality person re-identification (ReID) task aims to recognize the images of the same identity between the visible modality and the infrared modality. Existing methods mainly use a two-stream architecture to eliminate the discrepancy between the two modalities in the final common feature space, which ignore the single space of each modality in the shallow layers. To solve it, in this paper, we present a novel multi-feature space joint optimization (MSO) network, which can learn modality-sharable features in both the single-modality space and the common space. Firstly, based on the observation that edge information is modality-invariant, we propose an edge features enhancement module to enhance the modality-sharable features in each single-modality space. Specifically, we design a perceptual edge features (PEF) loss after the edge fusion strategy analysis. According to our knowledge, this is the first work that proposes explicit optimization in the single-modality feature space on cross-modality ReID task. Moreover, to increase the difference between cross-modality distance and class distance, we introduce a novel cross-modality contrastive-center (CMCC) loss into the modality-joint constraints in the common feature space. The PEF loss and CMCC loss jointly optimize the model in an end-to-end manner, which markedly improves the network's performance. Extensive experiments demonstrate that the proposed model significantly outperforms state-of-the-art methods on both the SYSU-MM01 and RegDB datasets.
Yajun Gao, Tengfei Liang, Yi Jin 0001, Xiaoyan Gu 0001, Wu Liu 0005, Yidong Li, Congyan Lang
ACM Multimedia5
2021 HUMA'21: 2nd International Workshop on Human-centric Multimedia Analysis
abstract
The Second International Workshop on Human-centric Multimedia Analysis is focused on human-centric analysis using multimedia information. The human-centric multimedia analysis is one of the fundamental and challenging problems of multimedia understanding. It involves various human-centric analysis tasks like face recognition, human pose estimation, person re-identification, human action recognition, person tracking, human-computer interaction, etc. Nowadays, various multimedia sensing devices and large-scale computing infrastructures are generating a wide variety of multi-modality data at a rapid velocity, which supplies rich knowledge to tackle these challenges for human-centric analysis. Researchers and engineers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as smart city, retailing, intelligent manufacturing, and public services. To this end, our workshop aims to provide a platform to promote exchanges and integration for the fields of human analysis and multimedia.
Wu Liu 0005, Xinchen Liu, Jingkuan Song, Dingwen Zhang, Wenbing Huang 0001, Junbo Guo, John R. Smith
ACM Multimedia1
2021 Consistency-Constancy Bi-Knowledge Learning for Pedestrian Detection in Night Surveillance
abstract
Pedestrian detection in the night surveillance is a challenging yet not largely explored task. As the success of the detector in the daytime surveillance and the convenient acquisition of all-weather data, we learn knowledge from these data to benefit pedestrian detection in night surveillance. We find two key properties of surveillance: distribution cross-time consistency and background cross-frame constancy. This paper proposes a consistency-constancy bi-knowledge learning (CCBL) for pedestrian detection in night surveillance, which is able to simultaneously achieve the night pedestrian detection's useful knowledge, coming from day and night surveillance. Firstly, based on the robustness of the existing detector in day surveillance, we obtain pedestrians' distribution in the daytime scene using the detector's detection results in the daytime scene. Based on the consistency of pedestrians' distribution during the day and night in the same scene, the pedestrian distribution from daytime is used as the consistency-knowledge for pedestrian detection in night surveillance. Secondly, the background as a constant knowledge of the surveillance scene is extractable and contributes to the division of the foreground, which contains most of the pedestrian regions and helps in pedestrian detection for night surveillance. Finally, we add bi-knowledge representation to promote each other and merge them together as the final pedestrian representation. Through extensive experiments, our CCBL significantly outperforms the state-of-the-art methods on public pedestrian detection datasets. In the NightSurveillance dataset, CCBL reduced the average missed detection rate by 3.04% compared to the existing best method.
Xiao Wang 0029, Zheng Wang 0007, Wu Liu 0005, Xin Xu 0007, Jing Chen 0003, Chia-Wen Lin
ACM Multimedia3
2021 Boosting End-to-end Multi-Object Tracking and Person Search via Knowledge Distillation
abstract
Multi-Object Tracking (MOT) and Person Search both demand to localize and identify specific targets from raw image frames. Existing methods can be classified into two categories, namely two-step strategy and end-to-end strategy. Two-step approaches have high accuracy but suffer from costly computations, while end-to-end methods show greater efficiency with limited performance. In this paper, we dissect the gap between two-step and end-to-end strategy and propose a simple yet effective end-to-end framework with knowledge distillation. Our proposed framework is simple in concept and easy to benefit from external datasets. Experimental results demonstrate that our model performs competitively with other sophisticated two-step and end-to-end methods in multi-object tracking and person search.
Wei Zhang 0255, Lingxiao He, Xingyu Liao, Wu Liu 0005, Qi Li 0005, Zhenan Sun
ACM Multimedia5
2021 Hierarchical multi-view context modelling for 3D object classification and retrieval
Anan Liu, Heyu Zhou, Weizhi Nie, Zhenguang Liu, Wu Liu 0005, Hongtao Xie 0001, Zhendong Mao 0001, Xuanya Li, Dan Song 0006
Inf. Sci.5
2021 Coupled-dynamic learning for vision and language: Exploring Interaction between different tasks
Ning Xu 0003, Hongshuo Tian, Yanhui Wang 0001, Weizhi Nie, Dan Song 0006, Anan Liu, Wu Liu 0005
Pattern Recognit.7
2021 A Real-Time Action Representation With Temporal Encoding and Deep Compression
abstract
Deep neural networks have achieved remarkable success for video-based action recognition. However, most of existing approaches cannot be deployed in practice due to the high computational cost. To address this challenge, we propose a new real-time convolutional architecture, called Temporal Convolutional 3D Network (T-C3D), for action representation. T-C3D learns video action representations in a hierarchical multi-granularity manner while obtaining a high process speed. Specifically, we propose a residual 3D Convolutional Neural Network (CNN) to capture complementary information on the appearance of a single frame and the motion between consecutive frames. Based on this CNN, we develop a new temporal encoding method to explore the temporal dynamics of the whole video. Furthermore, we integrate deep compression techniques with T-C3D to further accelerate the deployment of models via reducing the size of the model. By these means, heavy calculations can be avoided when doing the inference, which enables the method to deal with videos beyond real-time speed while keeping promising performance. We validate our approach by studying its action representation performance on four benchmarks over three different tasks. Our method achieves clear improvements on UCF101 action recognition benchmark against the state-of-the-art real-time methods by 5.4% in terms of accuracy and 2 times faster in terms of inference speed with a less than 5MB storage model. The source code and the pre-trained models are publicly available at https://github.com/tc3d.
Kun Liu 0016, Wu Liu 0005, Huadong Ma, Mingkui Tan, Chuang Gan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Pose-Guided Tracking-by-Detection: Robust Multi-Person Pose Tracking
abstract
Multi-person pose tracking task aims to estimate and track person keypoints in videos. Most of the previous methods follow the general track-by-detection strategy that ignores the consistent pose information during the whole framework. Thus, they often suffer from missing detections or inaccurate human association in challenging scenes with motion blur or person occlusion. To handle those problems, we propose a pose-guided tracking-by-detection framework that fuses pose information into both video human detection and human association procedures. In the video human detection stage, we adopt the pose-guided person location prediction exploiting the temporal information to make up missing detections. Technically, pose heatmaps are utilized to cope with the person-specific intra-class distractors. Furthermore, in the human association stage, we propose an appearance discriminative model based on the hierarchical pose-guided graph convolutional networks (PoseGCN). The PoseGCN-based model exploits human structural relations to boost person representation. Extensive experiments show the superiority of our method on the challenging pose tracking benchmark. Our proposed method ranks first on the PoseTrack leaderboard.11http://posetrack.net/leaderboard.php till the submission date (22-Aug-2019) of this paper. Our code has been publicly available at https://github.com/human-centric982/PGPT.
Qian Bao, Wu Liu 0005, Yuhao Cheng, Boyan Zhou, Tao Mei 0001
IEEE Trans. Multim.2
2021 Universal Cross-Domain 3D Model Retrieval
abstract
Recent advances in 3D modeling technologies such as 3D scanning, reconstruction and printing produce an explosive increasing of 3D models, consequently 3D model management becomes urgent to facilitate related applications such as CAD, VR/AR and autonomous driving. However, we usually lack the labels of the recently emerging 3D models and even have no prior knowledge toward the label set relationship between new datasets and existing labeled datasets, which makes the management challenging. In this paper, a universal cross-domain 3D model retrieval framework is proposed for utilizing the labeled 2D images or 3D models to manage unlabeled 3D models with no prior knowledge about label sets. Specifically, a sample-level weighting mechanism is adopted to automatically detect the samples from the common label set for both domains. Then, both the domain-level and class-level alignments are performed for domain adaptation. Finally, the adapted features are used for 3D model retrieval. We conduct experiments on the cross-domain 3D model retrieval dataset NTU-PSB (PSB-NTU) and image-based 3D model retrieval dataset MI3DOR, and the results validate the superiority and effectiveness of the proposed method.
Dan Song 0006, Tianbao Li 0001, Wenhui Li 0001, Weizhi Nie, Wu Liu 0005, Anan Liu
IEEE Trans. Multim.5
2021 Hierarchical Soft Quantization for Skeleton-Based Human Action Recognition
abstract
In daily life, human beings rely on hands and body parts to complete particular actions cooperatively. These selected body parts and their cooperative relationships are essential cues to distinguish these actions. However, most existing action recognition methods, which try to model the body appearance or spatial relations in skeleton sequences, often ignore the essential cooperation relationship among joints. Differently, in this paper, we propose a spatio-temporal hierarchical soft quantization method to extract the congenerous motion features, which reflect the cooperation relations among joints and body parts. Specifically, we design a hierarchical network with multiple soft quantization layers to extract congenerous features. The hierarchical network not only models the spatial hierarchy of skeleton structure for joint, part, and body, but also extracts the temporal hierarchy with sliding windows for frame, fragment, and sequence. Moreover, the features in each layer are visually explainable, which reflect the cooperation among body parts. The trainable parameters in the network are also significantly reduced, which reduces computational cost. Extensive experiments conducted on four benchmarks demonstrate that our method can provide competitive results compared with state-of-the-arts. The visualized congenerous features also validate that our approach can effectively perceive the essential cooperation relations.
Jianyu Yang 0002, Wu Liu 0005, Junsong Yuan 0001, Tao Mei 0001
IEEE Trans. Multim.2
2021 Correlation Discrepancy Insight Network for Video Re-identification
abstract
Video-based person re-identification (ReID) aims at re-identifying a specified person sequence from videos that were captured by disjoint cameras. Most existing works on this task ignore the quality discrepancy across frames by using all video frames to develop a ReID method. Additionally, they adopt only the person self-characteristic as the representation, which cannot adapt to cross-camera variation effectively. To that end, we propose a novel correlation discrepancy insight network for video-based person ReID, which consists of an unsupervised correlation insight model (CIM) for video purification and a discrepancy description network (DDN) for person representation. Concretely, CIM is constructed by using kernelized correlation filters to encode person half-parts, which evaluates the frame quality by the cross correlation across frames for selecting discriminative video fragments. Furthermore, DDN exploits the selected video fragments to generate a discrepancy descriptor using a compression network, which aims at employing the discrepancies with other persons’ to facilitate the representation of the target person rather than only using the self-characteristic. Due to the advantage in handling cross-domain variation, the discrepancy descriptor is expected to provide a new pattern for the object representation in cross-camera tasks. Experimental results on three public benchmarks demonstrate that the proposed method outperforms several state-of-the-art methods.
Weijian Ruan, Chao Liang 0001, Yi Yu 0001, Zheng Wang 0007, Wu Liu 0005, Jun Chen 0001, Jiayi Ma 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2020 When Pedestrian Detection Meets Nighttime Surveillance: A New Benchmark
abstract
Pedestrian detection at nighttime is a crucial and frontier problem in surveillance, but has not been well explored by the computer vision and artificial intelligence communities. Most of existing methods detect pedestrians under favorable lighting conditions (e.g. daytime) and achieve promising performances. In contrast, they often fail under unstable lighting conditions (e.g. nighttime). Night is a critical time for criminal suspects to act in the field of security. The existing nighttime pedestrian detection dataset is captured by a car camera, specially designed for autonomous driving scenarios. The dataset for nighttime surveillance scenario is still vacant. There are vast differences between autonomous driving and surveillance, including viewpoint and illumination. In this paper, we build a novel pedestrian detection dataset from the nighttime surveillance aspect: NightSurveillance1. As a benchmark dataset for pedestrian detection at nighttime, we compare the performances of state-of-the-art pedestrian detectors and the results reveal that the methods cannot solve all the challenging problems of NightSurveillance. We believe that NightSurveillance can further advance the research of pedestrian detection, especially in the field of surveillance security at nighttime.
Xiao Wang 0029, Jun Chen 0001, Zheng Wang 0007, Wu Liu 0005, Shin'ichi Satoh 0001, Chao Liang 0001, Chia-Wen Lin
IJCAI4
2020 KTAN: Knowledge Transfer Adversarial Network
abstract
Knowledge distillation was pioneered to transfer the generalization ability of a large teacher deep network to a light-weight student network. The student network can retain the high quality of the teacher network, yet exhibiting low computational complexity and storage requirement, which is attractive for deploying a deep convolution neural network on a resource-constrained mobile device. However, most of the existing methods focus on transferring the probability distribution of a softmax layer in a teacher network and neglect the intermediate representations. However, we find that the intermediate representation is critical for a student network to better understand the transferred generalization as compared to the probability distribution only. In this paper, therefore, we propose such a knowledge transfer adversarial network method which holistically considers both intermediate representations and probability distributions of a teacher network. To transfer the knowledge of intermediate representations, we set high-level teacher feature maps as a target, toward which the method trains student feature maps. Furthermore, to support various structures of a student network, we arrange a novel teacher-to-student layer. Finally, the proposed method employs an adversarial learning process. Specifically, it includes a discriminator network to fully exploit the spatial correlation of feature maps during the training process of a student network. The experimental results demonstrate that the proposed method can significantly improve the performance of a student network on two important vision tasks, image classification and object detection.
Peiye Liu, Wu Liu 0005, Huadong Ma, Zhewei Jiang, Mingoo Seok
IJCNN2
2020 Effective and Efficient: Toward Open-world Instance Re-identification
abstract
Instance Re-identification (ReID) system facilitates various applications that require painful and boring video watching. Its efficiency and effectiveness accelerate the process of video analysis. In this tutorial, we summarize ReID technologies and provide an overview. We'll introduce fundamental technologies, existing challenges, trends, etc. This tutorial would be useful for multimedia content analysis and system-level multimedia retrieval, especially for an effective and efficient open-world ReID system for the practical, large-scale, and open-set domain.
Zheng Wang 0007, Wu Liu 0005, Yusuke Matsui 0001, Shin'ichi Satoh 0001
ACM Multimedia2
2020 Pose-native Network Architecture Search for Multi-person Human Pose Estimation
abstract
Multi-person pose estimation has achieved great progress in recent years, even though, the precise prediction for occluded and invisible hard keypoints remains challenging. Most of the human pose estimation networks are equipped with an image classification-based pose encoder for feature extraction and a handcrafted pose decoder for high-resolution representations. However, the pose encoder might be sub-optimal because of the gap between image classification and pose estimation. The widely used multi-scale feature fusion in pose decoder is still coarse and cannot provide sufficient high-resolution details for hard keypoints. Neural Architecture Search (NAS) has shown great potential in many visual tasks to automatically search efficient networks. In this work, we present the Pose-native Network Architecture Search (PoseNAS) to simultaneously design a better pose encoder and pose decoder for pose estimation. Specifically, we directly search a data-oriented pose encoder with stacked searchable cells, which can provide an optimum feature extractor for the pose specific task. In the pose decoder, we exploit scale-adaptive fusion cells to promote rich information exchange across the multi-scale feature maps. Meanwhile, the pose decoder adopts a Fusion-and-Enhancement manner to progressively boost the high-resolution representations that are non-trivial for the precious prediction of hard keypoints. With the exquisitely designed search space and search strategy, PoseNAS can simultaneously search all modules in an end-to-end manner. PoseNAS achieves state-of-the-art performance on three public datasets, MPII, COCO, and PoseTrack, with small-scale parameters compared with the existing methods. Our best model obtains 76.7% mAP and 75.9% mAP on the COCO validation set and test set with only 33.6M parameters. Code and implementation are available at https://github.com/for-code0216/PoseNAS.
Qian Bao, Wu Liu 0005, Ling-Yu Duan, Tao Mei 0001
ACM Multimedia2
2020 A Cross-modality and Progressive Person Search System
abstract
This demonstration presents an instant and progressive cross-modality person search system, called 'CMPS'. Through the system, users can instantly find the lost children or elderly persons by simply describing their appearance through speech. Unlike most existing person search applications which have to cost much time to find the probe images, CMPS will save more valuable time in the early stage of losing. The proposed CMPS is one of the first attempts towards instant and progressive person search leveraging the audio, text, and visual modalities together. In detail, the system first takes the speech that describes the appearance of a person as the input to obtain a textual description by speech-to-text conversion. Then the cross-modal search is performed by matching the textual embedding with the visual representations of images in the learned latent space. The searched images can be used as candidates for query expansion. If the candidates are not right, the user can quickly adjust their description through speech. Once a right image is found, the user can directly click it as a new query. Finally the system will give the complete track of the lost person by once-click. On the built CUHK-PEDES-AUDIOS dataset, the system can achieve 82.46% rank-1 accuracy in real-time speed. Our code of CMPS is available at https://github.com/SheldongChen/Search-People-With-Audio.
Xiaodong Chen 0011, Wu Liu 0005, Xinchen Liu, Yongdong Zhang 0001, Tao Mei 0001
ACM Multimedia2
2020 PyAnomaly: A Pytorch-based Toolkit for Video Anomaly Detection
abstract
Video anomaly detection is an essential task in computer vision which attracts massive attention from academia and industry. The existing approaches are implemented in diverse deep learning frameworks and settings, making it difficult to reproduce the results published by the original authors. Undoubtedly, this phenomenon is detrimental to the development of Video Anomaly detection and community communication. In this paper, we present a PyTorch-based video anomaly detection toolbox, namely PyAnomaly that contains high modular and extensible components, comprehensive and impartial evaluation platforms, a friendly manageable system configuration, and the abundant engineering deployment functions. To make it easy-to-use and easy-to-extend, we implement the architecture by hooks and registers functionality. Remarkably, we have reproduced the comparable experimental results of six representative methods as those published by the original authors, and we will release these pre-trained models with more rich configurations. To our best knowledge, the PyAnomaly is the first open-source tool in video anomaly detection and is available at https://github.com/YuhaoCheng/PyAnomaly.
Yuhao Cheng, Wu Liu 0005, Pengrui Duan, Jingen Liu, Tao Mei 0001
ACM Multimedia2
2020 HUMA'20: 1st International Workshop on Human-Centric Multimedia Analysis
abstract
The First International Workshop on Human-Centric MultimediaAnalysis is concentrated on the tasks of human-centric analysis with multimedia and multimodal information. It is one of the fundamental and challenging problems of multimedia understanding. The human-centric multimedia analysis involves multiple tasks such as face detection and recognition, human body pattern analysis, person re-identification, human action detection, person tracking,human-object interaction, and so on. Today, multiple multimedia sensing technologies and large-scale computing infrastructures are producing at a rapid velocity a wide variety of big multi-modality data for human-centric analysis, which provides rich knowledge to help tackle these challenges. Researchers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as intelligent surveillance, retailing, fashion design, and services. Therefore, this workshop aims to provide a platform to bridge the gap between the communities of human analysis and multimedia.
Wu Liu 0005, Chuang Gan 0001, Jingkuan Song, Dingwen Zhang, Wenbing Huang 0001, John R. Smith
ACM Multimedia1
2020 Beyond the Parts: Learning Multi-view Cross-part Correlation for Vehicle Re-identification
abstract
Vehicle re-identification (Re-Id) is a challenging task due to the inter-class similarity, the intra-class difference, and the cross-view misalignment of vehicle parts. Although recent methods achieve great improvement by learning detailed features from keypoints or bounding boxes of parts, vehicle Re-Id is still far from being solved. Different from existing methods, we propose a Parsing-guided Cross-part Reasoning Network, named as PCRNet, for vehicle Re-Id. The PCRNet explores vehicle parsing to learn discriminative part-level features, model the correlation among vehicle parts, and achieve precise part alignment for vehicle Re-Id. To accurately segment vehicle parts, we first build a large-scale Multi-grained Vehicle Parsing (MVP) dataset from surveillance images. With the parsed parts, we extract regional features for each part and build a part-neighboring graph to explicitly model the correlation among parts. Then, the graph convolutional networks (GCNs) are adopted to propagate local information among parts, which can discover the most effective local features of varied viewpoints. Moreover, we propose a self-supervised part prediction loss to make the GCNs generate features of invisible parts from visible parts under different viewpoints. By this means, the same vehicle from different viewpoints can be matched with the well-aligned and robust feature representations. Through extensive experiments, our PCRNet significantly outperforms the state-of-the-art methods on three large-scale vehicle Re-Id datasets.
Xinchen Liu, Wu Liu 0005, Jinkai Zheng, Chenggang Yan 0001, Tao Mei 0001
ACM Multimedia2
2020 Multi-Features Fusion and Decomposition for Age-Invariant Face Recognition
abstract
Although the General Face Recognition (GFR) research achieves great success, Age-Invariant Face Recognition (AIFR) is still a challenging problem since facial appearance changing over time brings significant intra-class variations. The existing discriminative methods for the AIFR task mostly focus on decomposing the facial feature from a sigle image into age-related feature and age-independent feature for recognition, which suffer from the loss of facial identity information. To address this issue, in this work we propose a novel Multi-Features Fusion and Decomposition (MFFD) framework to learn more discriminative feature representations and alleviate the intra-class variations for AIFR. Specifically, we first sample multiple face images of different ages with the same identity as a face time series. Next, we combine feature decomposition with fusion based on the face time series to ensure that the final age-independent features effectively represent the identity information of the face and have stronger robustness against aging. Moreover, we also present two feature fusion methods and several different training strategies to explore the impact on the model. Extensive experiments on several cross-age datasets (CACD, CACD-VS) demonstrate the effectiveness of our proposed method. Besides, our method also shows comparable generalization performance on the well-known LFW dataset.
Lixuan Meng, Chenggang Yan 0001, Jian Yin 0003, Wu Liu 0005, Hongtao Xie 0001, Liang Li 0003
ACM Multimedia5
2020 Beyond the Attention: Distinguish the Discriminative and Confusable Features For Fine-grained Image Classification
abstract
Learning subtle discriminative features plays a significant role in fine-grained image classification. Existing methods usually extract the distinguishable parts through the attention module for classification. Although these learned distinguishable parts contain valuable features that are beneficial for classification, part of irrelevant features are also preserved, which may confuse the model to make a correct classification, especially for the fine-grained tasks due to their similarities. How to keep the discriminative features while removing confusable features from the distinguishable parts is an interesting yet changeling task. In this paper, we introduce a novel classification approach, named Logical-based Feature Extraction Model (LAFE for short) to address this issue. The main advantage of LAFE lies in the fact that it can explicitly add the significance of discriminative features and subtract the confusable features. Specifically, LAFE utilizes the region attention modules and channel attention modules to extract discriminative features and confusable features respectively. Based on this, two novel loss functions are designed to automatically induce attention over these features for fine-grained image classification. Our approach demonstrates its robustness, efficiency, and state-of-the-art performance on three benchmark datasets.
Xiruo Shi, Liutong Xu, Pengfei Wang 0009, Haifang Jian, Wu Liu 0005
ACM Multimedia6
2020 Black Re-ID: A Head-shoulder Descriptor for the Challenging Problem of Person Re-Identification
abstract
Person re-identification (Re-ID) aims at retrieving an input person image from a set of images captured by multiple cameras. Although recent Re-ID methods have made great success, most of them extract features in terms of the attributes of clothing (e.g., color, texture). However, it is common for people to wear black clothes or be captured by surveillance systems in low light illumination, in which cases the attributes of the clothing are severely missing. We call this problem the Black Re-ID problem. To solve this problem, rather than relying on the clothing information, we propose to exploit head-shoulder features to assist person Re-ID. The head-shoulder adaptive attention network (HAA) is proposed to learn the head-shoulder feature and an innovative ensemble method is designed to enhance the generalization of our model. Given the input person image, the ensemble method would focus on the head-shoulder feature by assigning a larger weight if the individual insides the image is in black clothing. Due to the lack of a suitable benchmark dataset for studying the Black Re-ID problem, we also contribute the first Black-reID dataset, which contains 1274 identities in training set. Extensive evaluations on the Black-reID, Market1501 and DukeMTMC-reID datasets show that our model achieves the best result compared with the state-of-the-art Re-ID methods on both Black and conventional Re-ID problems. Furthermore, our method is also proved to be effective in dealing with person Re-ID in similar clothing. Our code and dataset are avaliable on https://github.com/xbq1994/.
Boqiang Xu, Lingxiao He, Xingyu Liao, Wu Liu 0005, Zhenan Sun, Tao Mei 0001
ACM Multimedia4
2020 Hierarchical Gumbel Attention Network for Text-based Person Search
abstract
Text-based person search aims to retrieve the pedestrian images that best match a given textual description from gallery images. Previous methods utilize the soft-attention mechanism to infer the semantic alignments between the regions of image and the corresponding words in sentence. However, these methods may fuse the irrelevant multi-modality features together which cause matching redundancy problem. In this work, we propose a novel hierarchical Gumbel attention network for text-based person search via Gumbel top-k re-parameterization algorithm. Specifically, it adaptively selects the strong semantically relevant image regions and words/phrases from images and texts for precise alignment and similarity calculation. This hard selection strategy is able to fuse the strong-relevant multi-modality features for alleviating the problem of matching redundancy. Meanwhile, a Gumbel top-k re-parameterization algorithm is designed as a low-variance, unbiased gradient estimator to handle the discreteness problem of hard attention mechanism by an end-to-end manner. Moreover, a hierarchical adaptive matching strategy is employed by the model from three different granularities, i.e., word-level, phrase-level, and sentence-level, towards fine-grained matching. Extensive experimental results demonstrate the state-of-the-art performance. Compared the existed best method, we achieve the 8.24% Rank-1 and 7.6% mAP relative improvements in the text-to-image retrieval task, and 5.58% Rank-1 and 6.3% mAP relative improvements in the image-to-text retrieval task on CUHK-PEDES dataset, respectively.
Kecheng Zheng, Wu Liu 0005, Jiawei Liu 0001, Zhengjun Zha, Tao Mei 0001
ACM Multimedia2
2020 MetaSearch: Incremental Product Search via Deep Meta-Learning
abstract
With the advancement of image processing and computer vision technology, content-based product search is applied in a wide variety of common tasks, such as online shopping, automatic checkout systems, and intelligent logistics. Given a product image as a query, existing product search systems mainly perform the retrieval process using predefined databases with fixed product categories. However, real-world applications often require inserting new categories or updating existing products in the product database. When using existing product search methods, the image feature extraction models must be retrained and database indexes must be rebuilt to accommodate the updated data, and these operations incur high costs for data annotation and training time. To this end, we propose a few-shot incremental product search framework with meta-learning, which requires very few annotated images and has a reasonable training time. In particular, our framework contains a multipooling-based product feature extractor that learns a discriminative representation for each product, and we also design a meta-learning-based feature adapter to guarantee the robustness of the few-shot features. Furthermore, when expanding new categories in batches during a product search, we reconstruct the few-shot features by using an incremental weight combiner to accommodate the incremental search task. Through extensive experiments, we demonstrate that the proposed framework achieves excellent performance for new products while still guaranteeing the high search accuracy of the base categories after gradually expanding new product categories without forgetting.
Qi Wang 0079, Xinchen Liu, Wu Liu 0005, Anan Liu, Wenyin Liu, Tao Mei 0001
IEEE Trans. Image Process.3
2020 Listen, Look, and Find the One: Robust Person Search with Multimodality Index
abstract
Person search with one portrait, which attempts to search the targets in arbitrary scenes using one portrait image at a time, is an essential yet unexplored problem in the multimedia field. Existing approaches, which predominantly depend on the visual information of persons, cannot solve problems when there are variations in the person’s appearance caused by complex environments and changes in pose, makeup, and clothing. In contrast to existing methods, in this article, we propose an associative multimodality index for person search with face, body, and voice information. In the offline stage, an associative network is proposed to learn the relationships among face, body, and voice information. It can adaptively estimate the weights of each embedding to construct an appropriate representation. The multimodality index can be built by using these representations, which exploit the face and voice as long-term keys and the body appearance as a short-term connection. In the online stage, through the multimodality association in the index, we can retrieve all targets depending only on the facial features of the query portrait. Furthermore, to evaluate our multimodality search framework and facilitate related research, we construct the Cast Search in Movies with Voice (CSM-V) dataset, a large-scale benchmark that contains 127K annotated voices corresponding to tracklets from 192 movies. According to extensive experiments on the CSM-V dataset, the proposed multimodality person search framework outperforms the state-of-the-art methods.
Xiao Wang 0029, Wu Liu 0005, Jun Chen 0001, Xiaobo Wang 0001, Chenggang Yan 0001, Tao Mei 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2019 Structured Two-Stream Attention Network for Video Question Answering
abstract
To date, visual question answering (VQA) (i.e., image QA and video QA) is still a holy grail in vision and language understanding, especially for video QA. Compared with image QA that focuses primarily on understanding the associations between image region-level details and corresponding questions, video QA requires a model to jointly reason across both spatial and long-range temporal structures of a video as well as text to provide an accurate answer. In this paper, we specifically tackle the problem of video QA by proposing a Structured Two-stream Attention network, namely STA, to answer a free-form or open-ended natural language question about the content of a given video. First, we infer rich longrange temporal structures in videos using our structured segment component and encode text features. Then, our structured two-stream attention component simultaneously localizes important visual instance, reduces the influence of background video and focuses on the relevant text. Finally, the structured two-stream fusion component incorporates different segments of query and video aware context representation and infers the answers. Experiments on the large-scale video QA dataset TGIF-QA show that our proposed method significantly surpasses the best counterpart (i.e., with one representation for the video input) by 13.0%, 13.5%, 11.0% and 0.3 for Action, Trans., TrameQA and Count tasks. It also outperforms the best competitor (i.e., with two representations) on the Action, Trans., TrameQA tasks by 4.1%, 4.7%, and 5.1%.
Lianli Gao, Pengpeng Zeng, Jingkuan Song, Yuan-Fang Li, Wu Liu 0005, Tao Mei 0001, Heng Tao Shen
AAAI5
2019 Social Relation Recognition From Videos via Multi-Scale Spatial-Temporal Reasoning
abstract
Discovering social relations, e.g., kinship, friendship, etc., from visual contents can make machines better interpret the behaviors and emotions of human beings. Existing studies mainly focus on recognizing social relations from still images while neglecting another important media--video. On one hand, the actions and storylines in videos provide more important cues for social relation recognition. On the other hand, the key persons may appear at arbitrary spatial-temporal locations, even not in one same image from beginning to the end. To overcome these challenges, we propose a Multi-scale Spatial-Temporal Reasoning (MSTR) framework to recognize social relations from videos. For the spatial representation, we not only adopt a temporal segment network to learn global action and scene information, but also design a Triple Graphs model to capture visual relations between persons and objects. For the temporal domain, we propose a Pyramid Graph Convolutional Network to perform temporal reasoning with multi-scale receptive fields, which can obtain both long-term and short-term storylines in videos. By this means, MSTR can comprehensively explore the multi-scale actions and storylines in spatial-temporal dimensions for social relation reasoning in videos. Extensive experiments on a new large-scale Video Social Relation dataset demonstrate the effectiveness of the proposed framework.
Xinchen Liu, Wu Liu 0005, Jingwen Chen 0001, Lianli Gao, Chenggang Yan 0001, Tao Mei 0001
CVPR2
2019 Foreground-Aware Pyramid Reconstruction for Alignment-Free Occluded Person Re-Identification
abstract
Re-identifying a person across multiple disjoint camera views is important for intelligent video surveillance, smart retailing and many other applications. However, existing person re-identification methods are challenged by the ubiquitous occlusion over persons and suffer performance degradation. This paper proposes a novel occlusion-robust and alignment-free model for occluded person ReID and extends its application to realistic and crowded scenarios. The proposed model first leverages the fully convolution network (FCN) and pyramid pooling to extract spatial pyramid features. Then an alignment-free matching approach namely Foreground-aware Pyramid Reconstruction (FPR) is developed to accurately compute matching scores between occluded persons, regardless of their different scales and sizes. FPR uses the error from robust reconstruction over spatial pyramid features to measure similarities between two persons. More importantly, we design a occlusion-sensitive foreground probability generator that focuses more on clean human body parts to robustify the similarity computation with less contamination from occlusion. The FPR is easily embedded into any end-to-end person ReID models. The effectiveness of the proposed method is clearly demonstrated by the experimental results (Rank-1 accuracy) on three occluded person datasets: Partial REID (78.30%), Partial iLIDS (68.08%), Occluded REID (81.00%), and three benchmark person datasets: Market1501 (95.42%), DukeMTMC (88.64%), CUHK03 (76.08%).
Lingxiao He, Yinggang Wang, Wu Liu 0005, Zhenan Sun, Jiashi Feng
ICCV3
2019 Human Mesh Recovery From Monocular Images via a Skeleton-Disentangled Representation
abstract
We describe an end-to-end method for recovering 3D human body mesh from single images and monocular videos. Different from the existing methods try to obtain all the complex 3D pose, shape, and camera parameters from one coupling feature, we propose a skeleton-disentangling based framework, which divides this task into multi-level spatial and temporal granularity in a decoupling manner. In spatial, we propose an effective and pluggable “disentangling the skeleton from the details” (DSD) module. It reduces the complexity and decouples the skeleton, which lays a good foundation for temporal modeling. In temporal, the self-attention based temporal convolution network is proposed to efficiently exploit the short and long-term temporal cues. Furthermore, an unsupervised adversarial training strategy, temporal shuffles and order recovery, is designed to promote the learning of motion dynamics. The proposed method outperforms the state-of-the-art 3D human mesh recovery methods by 15.4% MPJPE and 23.8% PA-MPJPE on Human3.6M. State-of-the-art results are also achieved on the 3D pose in the wild (3DPW) dataset without any fine-tuning. Especially, ablation studies demonstrate that skeleton-disentangled representation is crucial for better temporal modeling and generalization.
Yu Sun 0030, Yun Ye 0001, Wu Liu 0005, Wenpeng Gao, Yili Fu 0001, Tao Mei 0001
ICCV3
2019 Multi-Granularity Reasoning for Social Relation Recognition From Images
abstract
Discovering social relations in images can make machines better interpret the behavior of human beings. However, automatically recognizing social relations in images is a challenging task due to the significant gap between the domains of visual content and social relation. Existing studies separately process various features such as faces expressions, body appearance, and contextual objects, thus they cannot comprehensively capture the multi-granularity semantics, such as scenes, regional cues of persons, and interactions among persons and objects. To bridge the domain gap, we propose a Multi-Granularity Reasoning framework for social relation recognition from images. The global knowledge and mid-level details are learned from the whole scene and the regions of persons and objects, respectively. Most importantly, we explore the fine-granularity pose keypoints of persons to discover the interactions among persons and objects. Specifically, the pose-guided Person-Object Graph and Person-Pose Graph are proposed to model the actions from persons to object and the interactions between paired persons, respectively. Based on the graphs, social relation reasoning is performed by graph convolutional networks. Finally, the global features and reasoned knowledge are integrated as a comprehensive representation for social relation recognition. Extensive experiments on two public datasets show the effectiveness of the proposed framework.
Xinchen Liu, Wu Liu 0005, Anfu Zhou, Huadong Ma, Tao Mei 0001
ICME3
2019 Deep Recurrent Quantization for Generating Sequential Binary Codes
abstract
Quantization has been an effective technology in ANN (approximate nearest neighbour) search due to its high accuracy and fast search speed. To meet the requirement of different applications, there is always a trade-off between retrieval accuracy and speed, reflected by variable code lengths. However, to encode the dataset into different code lengths, existing methods need to train several models, where each model can only produce a specific code length. This incurs a considerable training time cost, and largely reduces the flexibility of quantization methods to be deployed in real applications. To address this issue, we propose a Deep Recurrent Quantization (DRQ) architecture which can generate sequential binary codes. To the end, when the model is trained, a sequence of binary codes can be generated and the code length can be easily controlled by adjusting the number of recurrent iterations. A shared codebook and a scalar factor is designed to be the learnable weights in the deep recurrent quantization block, and the whole framework can be trained in an end-to-end manner. As far as we know, this is the first quantization method that can be trained once and generate sequential binary codes. Experimental results on the benchmark datasets show that our model achieves comparable or even better performance compared with the state-of-the-art for image retrieval. But it requires significantly less number of parameters and training times. Our code is published online: https://github.com/cfm-uestc/DRQ.
Jingkuan Song, Xiaosu Zhu, Lianli Gao, Xin-Shun Xu, Wu Liu 0005, Heng Tao Shen
IJCAI5
2019 Learnable Aggregating Net with Diversity Learning for Video Question Answering
abstract
Video visual question answering (V-VQA) remains challenging at the intersection of vision and language, where it requires joint comprehension of video and natural language question. Image-Question co-attention mechanism, which aims at generating a spatial map highlighting image regions relevant to answering the question and vice versa, has obtained impressive results. Despite the success, simply applying co-attention to video visual question answering results in unsatisfactory performance due to the complexity and temporal nature of videos. In this paper, we proposed a novel architecture, namely Learnable Aggregating Net with Diversity learning (LAD-Net), for V-VQA. In the proposed method, we address two central problems: 1) how to deploy co-attention to V-VQA task considering the complex and diverse content of videos; and 2) how to aggregate the frame-level features without destroying the feature distributions and temporal information. To solve these problems, our LAD-Net first extends single-path based co-attention mechanism to a multi-path pyramid co-attention structure with a novel diversity learning to explicitly encourage attention diversity. For video-level (or question-level) descriptor, instead of taking a simple temporal pooling (i.e., average pooling), we propose a new learnable aggregation method with a set of evidence gates. It automatically aggregates adaptively-weighted frame-level features (or word-level features) to extract rich video (or question) context semantic information by imitating Bags-of-Words (BoW) quantization. With evidence gates, it then further chooses the most related signals representing the evidence information to predict the answer.Extensive validations on the two challenging video visual question answering datasets TGIF-QA and TVQA show that LAD-Net achieves the state-of-the-art performance under various settings and metrics. Our proposed strategies are of particular importance for improving the performance of the baseline co-attention V-VQA.
Lianli Gao, Xuanhan Wang, Wu Liu 0005, Xing Xu 0001, Heng Tao Shen, Jingkuan Song
ACM Multimedia4
2019 Fine-grained Cross-media Representation Learning with Deep Quantization Attention Network
abstract
Cross-media search is useful for getting more comprehensive and richer information about social network hot topics or events. To solve the problems of feature heterogeneity and semantic gap of different media data, existing deep cross-media quantization technology provides an efficient and effective solution for cross-media common semantic representation learning. However, due to the fact that social network data often exhibits semantic sparsity, diversity, and contains a lot of noise, the performance of existing cross-media search methods often degrades. To address the above issue, this paper proposes a novel fine-grained cross-media representation learning model with deep quantization attention network for social network cross-media search (CMSL). First, we construct the image-word semantic correlation graph, and perform deep random walks on the graph to realize semantic expansion and semantic embedding learning, which can discover some potential semantic correlations between images and words. Then, in order to discover more fine-grained cross-media semantic correlations, a multi-scale fine-grained cross-media semantic correlation learning method that combines global and local saliency semantic similarity is proposed. Third, the fine-grained cross-media representation, cross-media semantic correlations and binary quantization code are jointly learned by a unified deep quantization attention network, which can preserve both inter-media correlations and intra-media similarities, by minimizing both cross-media correlation loss and binary quantization loss. Experimental results demonstrate that CMSL can generate high-quality cross-media common semantic representation, which yields state-of-the-art cross-media search performance on two benchmark datasets, NUS-WIDE and MIR-Flickr 25k.
Meiyu Liang, Junping Du 0001, Wu Liu 0005, Zhe Xue, Yue Geng, Cong-Xian Yang
ACM Multimedia3
2019 BraidNet: Braiding Semantics and Details for Accurate Human Parsing
abstract
This paper focuses on fine-grained human parsing in images. This is a very challenging task due to the diverse person appearance, semantic ambiguity of different body parts and clothing, and extremely small parsing targets. Although existing approaches can achieve significant improvement by pyramid feature learning, multi-level supervision, and joint learning with pose estimation, human parsing is still far from being solved. Different from existing approaches, we propose a Braiding Network, named as BraidNet, to learn complementary semantics and details for fine-grained human parsing. The BraidNet contains a two-stream braid-like architecture. The first stream is a semantic abstracting net with a deep yet narrow structure which can learn semantic knowledge by a hierarchy of fully convolution layers to overcome the challenges of diverse person appearance. To capture low-level details of small targets, the detail-preserving net is designed to exploit a shallow yet wide network without down-sampling, which can retain sufficient local structures for small objects. Moreover, we design a group of braiding modules across the two sub-nets, by which complementary information can be exchanged during end-to-end training. Besides, in the end of BraidNet, a Pairwise Hard Region Embedding strategy is propose to eliminate the semantic ambiguity of different body parts and clothing. Extensive experiments show that the proposed BraidNet achieves better performance than the state-of-the-art methods for fine-grained human parsing.
Xinchen Liu, Wu Liu 0005, Jingkuan Song, Tao Mei 0001
ACM Multimedia3
2019 POINet: Pose-Guided Ovonic Insight Network for Multi-Person Pose Tracking
abstract
Multi-person pose tracking aims to jointly estimate and track multi-person keypoints in the unconstrained videos. The most popular solution to this task follows the tracking-by-detection strategy that relies on human detection and data association. While human detection has been boosted by deep learning, existing works mainly exploit several separated stages with hand-crafted metrics to realize data association, leading to great uncertainty and feeble adaption in complex scenes. To handle these problems, we propose an end-to-end pose-guided ovonic insight network (POINet) for the data association in multi-person pose tracking, which jointly learns feature extraction, similarity estimation, and identity assignment. Specifically, we design a pose-guided representation network to integrate pose information into hierarchical convolutional features, generating a pose-aligned person representation for person, which helps handle partial occlusions. Moreover, we propose an ovonic insight network to adaptively encode the cross-frame identity transformation, which can cope with the tough tracking cases of person leaving and entering the scene. In general, the proposed POINet provides a new insight to realize multi-person pose tracking in an end-to-end fashion. Extensive experiments conducted on the PoseTrack benchmark demonstrate that our POINet outperforms the state-of-the-art methods.
Weijian Ruan, Wu Liu 0005, Qian Bao, Jun Chen 0001, Yuhao Cheng, Tao Mei 0001
ACM Multimedia2
2019 A common subgraph correspondence mining framework for map search services
Wu Liu 0005, Lingheng Zhu, Lingyang Chu, Huadong Ma
Multim. Tools Appl.1
2019 DELTA: A deep dual-stream network for multi-label image classification
Wan-Jin Yu, Zhen-Duo Chen 0001, Xin Luo 0006, Wu Liu 0005, Xin-Shun Xu
Pattern Recognit.4
2019 Attentive Spatial-Temporal Summary Networks for Feature Learning in Irregular Gait Recognition
abstract
Gait recognition is an attractive human recognition technology. However, existing gait recognition methods mainly focus on the regular gait cycles, which ignore the irregular situation. In real-world surveillance, human gait is almost irregular, which contains arbitrary dynamic characteristics (e.g., duration, speed, and phase) and varied viewpoints. In this paper, we propose the attentive spatial-temporal summary networks to learn salient spatial-temporal and view-independence features for irregular gait recognition. First of all, we design the gate mechanism with attentive spatial-temporal summary to extract the discriminative sequence-level features for representing the periodic motion cues of irregular gait sequences. The designed general attention and residual attention components can concentrate on the discriminative identity-related semantic regions from the spatial feature maps. The proposed attentive temporal summary component can automatically assign adaptive attention to enhance the discriminative gait timesteps and suppress the redundant ones. Furthermore, to improve the accuracy of cross-view gait recognition, we combine the Siamese structure and Null Foley-Sammon transform to obtain the view-invariant gait features from irregular gait sequences. Finally, we quantitatively evaluate the impact of the irregular gait and viewpoint interval between matching pairs on gait recognition accuracy. Experimental results show that our method achieves state-of-the-art performance in irregular gait recognition on the OULP and CASIA-B datasets.
Shuangqun Li, Wu Liu 0005, Huadong Ma
IEEE Trans. Multim.2
2019 Generalized zero-shot learning for action recognition with web-scale video data
Kun Liu 0016, Wu Liu 0005, Huadong Ma, Wenbing Huang 0001, Xiongxiong Dong
World Wide Web2
2018 T-C3D: Temporal Convolutional 3D Network for Real-Time Action Recognition
abstract
Video-based action recognition with deep neural networks has shown remarkable progress. However, most of the existing approaches are too computationally expensive due to the complex network architecture. To address these problems, we propose a new real-time action recognition architecture, called Temporal Convolutional 3D Network (T-C3D), which learns video action representations in a hierarchical multi-granularity manner. Specifically, we combine a residual 3D convolutional neural network which captures complementary information on the appearance of a single frame and the motion between consecutive frames with a new temporal encoding method to explore the temporal dynamics of the whole video. Thus heavy calculations are avoided when doing the inference, which enables the method to be capable of real-time processing. On two challenging benchmark datasets, UCF101 and HMDB51, our method is significantly better than state-of-the-art real-time methods by over 5.4% in terms of accuracy and 2 times faster in terms of inference speed (969 frames per second), demonstrating comparable recognition performance to the state-of-the-art methods. The source code for the complete system as well as the pre-trained models are publicly available at https://github.com/tc3d.
Kun Liu 0016, Wu Liu 0005, Chuang Gan 0001, Mingkui Tan, Huadong Ma
AAAI2
2018 Joint License Plate Super-Resolution and Recognition in One Multi-Task Gan Framework
abstract
License plate recognition (LPR) plays an important role in intelligent transport systems. The existed LPR systems are mostly based on hand-crafted methods for detection, segmentation, and recognition, which cannot accurately recognize the license plate in unconstrained surveillance environments. In this paper, we propose a Multi-Task Generative Adversarial Network (MTGAN) based LPR system, which combines the license plate super-resolution and recognition in one end-to-end framework. In the proposed MTGAN, we design a Fully Connected Network (FCN) as generative network (GN), which can combine knowledge from data distribution and domain prior knowledge of license plate to generate the spatial corresponding and high-resolution plate images in the synthesis pipeline. More important, a multi-task discriminative network is designed in MTGAN to combine the super-resolution and recognition in an adversarial manner to enhance each other. The experiments on the built real-world license plate dataset show that the proposed LPR system can generate high-resolution license plates as well as recognize them with higher accuracy than state-of-the-art LPR systems.
Wu Liu 0005, Huadong Ma
ICASSP2
2018 Beyond View Transformation: Cycle-Consistent Global and Partial Perception Gan for View-Invariant Gait Recognition
abstract
Cross-view gait recognition is a challenging problem when view-interval and pose variation are relatively large. In this paper, we propose Cycle-consistent Attentive Generative Adversarial Networks (CA-GAN) to map different views' gait images to view-consistent and photorealistic gait images for cross-view gait recognition. In CA-GAN, the generative network is composed of two branches, which simultaneously perceives human's global contexts and local body parts information respectively. Moreover, we design a novel Attentive Adversarial Network (AAN) to adaptively learn different weights for the discriminator's receptive fields with attention mechanism. Furthermore, as it is hard to collect the pose-aligned gait image pairs from different views for training CA-GAN’ we combine forward cycle-consistency loss and adver-sarial loss to learn the transformation relationship from source views to target view. The combined loss function can also preserve the discriminative gait structures of different identities at the training stage. Finally, we directly exploit the synthesized view-consistent gait images for cross-view gait recognition task. Experimental results on CASIA-B demonstrate that our method not only outperforms the state-of-the-art methods in cross-view gait recognition, but also presents compelling perceptual results even across the large view-interval.
Shuangqun Li, Wu Liu 0005, Huadong Ma, Shaopeng Zhu
ICME2
2018 LOCO: Local Context Based Faster R-CNN for Small Traffic Sign Detection
Peng Cheng 0002, Wu Liu 0005, Huadong Ma
MMM (1)2
2018 Multi-stream Fusion Model for Social Relation Recognition from Videos
Jinna Lv, Wu Liu 0005, Bin Wu 0001, Huadong Ma
MMM (1)2
2018 PROVID: Progressive and Multimodal Vehicle Reidentification for Large-Scale Urban Surveillance
abstract
Compared with person reidentification, which has attracted concentrated attention, vehicle reidentification is an important yet frontier problem in video surveillance and has been neglected by the multimedia and vision communities. Since most existing approaches mainly consider the general vehicle appearance for reidentification while overlooking the distinct vehicle identifier, such as the license plate number, they attain suboptimal performance. In this paper, we propose PROVID, a PROgressive Vehicle re-IDentification framework based on deep neural networks. In particular, our framework not only utilizes the multimodality data in large-scale video surveillance, such as visual features, license plates, camera locations, and contextual information, but also considers vehicle reidentification in two progressive procedures: coarse-to-fine search in the feature domain, and near-to-distant search in the physical space. Furthermore, to evaluate our progressive search framework and facilitate related research, we construct the VeRi dataset, which is the most comprehensive dataset from real-world surveillance videos. It not only provides large numbers of vehicles with varied labels and sufficient cross-camera recurrences but also contains license plate numbers and contextual information. Extensive experiments on the VeRi dataset demonstrate both the accuracy and efficiency of our progressive vehicle reidentification framework.
Xinchen Liu, Wu Liu 0005, Tao Mei 0001, Huadong Ma
IEEE Trans. Multim.2
2017 Weighted sequence loss based spatial-temporal deep learning framework for human body orientation estimation
abstract
Accurate human body orientation estimation (HBOE) can significantly promote the analysis of human behavior. However, conventional methods cannot holistically exploit the complementary nature of spatial and temporal information for H-BOE. Different from existing methods, we propose an end-to-end temporal-spatial deep learning framework to accurately estimate the human body orientation. In this framework, we firstly utilize the convolutional neural network to capture the spatial information for human orientation. Furthermore, the spatial-temporal information are fused in the recurrent neural networks (RNNs), which can automatically memorize a long-term temporal information of human orientation transformation. More important, to effectively adapt different moving speeds and diversity actions of people, we design a weighted sequence loss function, which can capture the significant orientation conversion to guide the RNN training. According to the comprehensive evaluations, the proposed method greatly outperforms the states-of-the-art methods. Although only utilizing the 2D information, it can perform better than the 3-D/RGB-D based approaches.
Peiye Liu, Wu Liu 0005, Huadong Ma
ICME2
2017 Hierarchical LSTM with Adjusted Temporal Attention for Video Captioning
abstract
Recent progress has been made in using attention based encoder-decoder framework for video captioning. However, most existing decoders apply the attention mechanism to every generated words including both visual words (e.g., “gun” and "shooting“) and non-visual words (e.g. "the“, "a”).However, these non-visual words can be easily predicted using natural language model without considering visual signals or attention.Imposing attention mechanism on non-visual words could mislead and decrease the overall performance of video captioning.To address this issue, we propose a hierarchical LSTM with adjusted temporal attention (hLSTMat) approach for video captioning. Specifically, the proposed framework utilizes the temporal attention for selecting specific frames to predict related words, while the adjusted temporal attention is for deciding whether to depend on the visual information or the language context information. Also, a hierarchical LSTMs is designed to simultaneously consider both low-level visual information and deep semantic information to support the video caption generation. To demonstrate the effectiveness of our proposed framework, we test our method on two prevalent datasets: MSVD and MSR-VTT, and experimental results show that our approach outperforms the state-of-the-art methods on both two datasets.
Jingkuan Song, Lianli Gao, Zhao Guo, Wu Liu 0005, Dongxiang Zhang, Heng Tao Shen
IJCAI4
2017 Beyond Human-level License Plate Super-resolution with Progressive Vehicle Search and Domain Priori GAN
abstract
In this paper, we address the challenging problem of vehicle license plate image super-resolution. Different from existing image super-resolution approaches only resorted to one single image, we propose to leverage complementary information from multiple images to recover the license plate numbers. To achieve this goal, we design a principled license plate images super-resolution framework which is composed of two components: progressive vehicle search and Domain Priori GAN (DP-GAN). Particularly, we design a null space based progressive vehicle search approach to retrieve the relevant images captured by different cameras given one vehicle with a low-resolution license plate. To handle the extremely varied license plate images caused by different sensors, times, depths, and viewpoints, we also propose a DP-GAN framework to generate multiple spatial correspondences and high-resolution plate images. In the generator network of DP-GAN, a license plate synthesis pipeline is exploited to generate the nearly canonical license plates. In the discriminator network, a spatial split layer is designed to simultaneously preserve the global and local manufacture standards of the license plate. Finally, a multiple images super-resolution GAN is exploited to combine all the synthetic license plates into one high-resolution image. Different from previous super-resolution criteria mainly focus on pixel-level detail recovery condition, we leverage the downstream tasks, i.e. license plate recognition and vehicle search as criteria. The results on a new collected real-world dataset demonstrate that the proposed method achieves the beyond human-level license plate super-resolution performance for automatic license plate recognition and vehicle search.
Wu Liu 0005, Xinchen Liu, Huadong Ma, Peng Cheng 0002
ACM Multimedia1
2017 Multi-feature Fusion for Predicting Social Media Popularity
abstract
This paper presents the method that underlies our submission to the popularity prediction task of Social Media Prediction Challenge 2017. The task is designed to predict the impact of sharing different posts for a publisher on social media. There are many factors that influence image popularity; these include not only the visual features of the image, but also the social features, such as user characteristics of its poster and even the upload time. In this project, we propose a fast and effective framework for popularity prediction. First, we investigate and extract visual and social features of images. For the visual feature, we introduce 1) global feature descriptors, such as Local Binary Pattern and Color Names, 2) local feature descriptors, such as Local Maximal Occurrence, and 3) deep features. For the social feature, we adopt users features (average views, group count, and member count), post features (title length, description length, and tag count), and time features (month, weekday, day, and hour). Furthermore, we fed a fusion of multi-feature to Linear Regression, Matrix Factorization based on Time and feature Cluster, and Support Vector Regression models respectively, and present comparative analysis of the prediction results. Finally, we choose the best model to predict the popularity scores of the test images. Experimental results demonstrate that our method can achieve 0.8581, 1.4062 and 0.8625 in terms of Spearman Ranking Correlation, Mean Absolute Error, and Mean Squared Error, respectively.
Jinna Lv, Wu Liu 0005, He Gong, Bin Wu 0001, Huadong Ma
ACM Multimedia2
2017 Multi-attribute Based Fire Detection in Diverse Surveillance Videos
Shuangqun Li, Wu Liu 0005, Huadong Ma, Huiyuan Fu
MMM (1)2
2017 Deep Learning Based Intelligent Basketball Arena with Energy Image
Wu Liu 0005, Jiangyu Liu, Xiaoyan Gu 0001, Kun Liu 0016, Xiaowei Dai, Huadong Ma
MMM (1)1
2017 Minimizing Resource Cost for Camera Stream Scheduling in Video Data Center
Yihong Gao, Huadong Ma, Wu Liu 0005
J. Comput. Sci. Technol.3
2017 Online multi-objective optimization for live video forwarding across video data centers
Wu Liu 0005, Yihong Gao, Huadong Ma, Shui Yu 0001, Jie Nie
J. Vis. Commun. Image Represent.1
2017 Multi-modal tag localization for mobile video search
Rui Zhang 0040, Sheng Tang, Wu Liu 0005, Yongdong Zhang 0001, Jintao Li 0001
Multim. Syst.3
2017 Deep learning based basketball video analysis for intelligent arena application
Wu Liu 0005, Chenggang Yan 0001, Jiangyu Liu, Huadong Ma
Multim. Tools Appl.1
2017 Special issue on intelligent urban computing with big data
Wu Liu 0005, Peng Cui 0001, Jukka K. Nurminen, Jingdong Wang 0001
Mach. Vis. Appl.1
2016 A Deep Learning-Based Approach to Progressive Vehicle Re-identification for Urban Surveillance
Xinchen Liu, Wu Liu 0005, Tao Mei 0001, Huadong Ma
ECCV (2)2
2016 Siamese neural network based gait recognition for human identification
abstract
As the remarkable characteristics of remote accessed, robust and security, gait recognition has gained significant attention in the biometrics based human identification task. However, the existed methods mainly employ the handcrafted gait features, which cannot well handle the indistinctive inter-class differences and large intra-class variations of human gait in real-world situation. In this paper, we have developed a Siamese neural network based gait recognition framework to automatically extract robust and discriminative gait features for human identification. Different from conventional deep neural network, the Siamese network can employ distance metric learning to drive the similarity metric to be small for pairs of gait from the same person, and large for pairs from different persons. In particular, to further learn effective model with limited training data, we composite the gait energy images instead of raw sequence of gaits. Consequently, the experiments on the world's largest gait database show our framework impressively outperforms state-of-the-arts.
Cheng Zhang 0014, Wu Liu 0005, Huadong Ma, Huiyuan Fu
ICASSP2
2016 Large-scale vehicle re-identification in urban surveillance videos
abstract
Vehicle, as a significant object class in urban surveillance, attracts massive focuses in computer vision field, such as detection, tracking, and classification. Among them, vehicle re-identification (Re-Id) is an important yet frontier topic, which not only faces the challenges of enormous intra-class and subtle inter-class differences of vehicles in multicameras, but also suffers from the complicated environments in urban surveillance scenarios. Besides, the existing vehicle related datasets all neglect the requirements of vehicle Re-Id: 1) massive vehicles captured in real-world traffic environment; and 2) applicable recurrence rate to give cross-camera vehicle search for vehicle Re-Id. To facilitate vehicle Re-Id research, we propose a large-scale benchmark dataset for vehicle Re-Id in the real-world urban surveillance scenario, named “VeRi”. It contains over 40,000 bounding boxes of 619 vehicles captured by 20 cameras in unconstrained traffic scene. Moreover, each vehicle is captured by 2~18 cameras in different viewpoints, illuminations, and resolutions to provide high recurrence rate for vehicle Re-Id. Finally, we evaluate six competitive vehicle Re-Id methods on VeRi and propose a baseline which combines the color, texture, and highlevel semantic information extracted by deep neural network.
Xinchen Liu, Wu Liu 0005, Huadong Ma, Huiyuan Fu
ICME2
2015 Multi-task deep visual-semantic embedding for video thumbnail selection
abstract
Given the tremendous growth of online videos, video thumbnail, as the common visualization form of video content, is becoming increasingly important to influence user's browsing and searching experience. However, conventional methods for video thumbnail selection often fail to produce satisfying results as they ignore the side semantic information (e.g., title, description, and query) associated with the video. As a result, the selected thumbnail cannot always represent video semantics and the click-through rate is adversely affected even when the retrieved videos are relevant. In this paper, we have developed a multi-task deep visual-semantic embedding model, which can automatically select query-dependent video thumbnails according to both visual and side information. Different from most existing methods, the proposed approach employs the deep visual-semantic embedding model to directly compute the similarity between the query and video thumbnails by mapping them into a common latent semantic space, where even unseen query-thumbnail pairs can be correctly matched. In particular, we train the embedding model by exploring the large-scale and freely accessible click-through video and image data, as well as employing a multi-task learning strategy to holistically exploit the query-thumbnail relevance from these two highly related datasets. Finally, a thumbnail is selected by fusing both the representative and query relevance scores. The evaluations on 1,000 query-thumbnail dataset labeled by 191 workers in Amazon Mechanical Turk have demonstrated the effectiveness of our proposed method.
Wu Liu 0005, Tao Mei 0001, Yongdong Zhang 0001, Cherry Che, Jiebo Luo 0001
CVPR1
2014 Instant Mobile Video Search With Layered Audio-Video Indexing and Progressive Transmission
abstract
The proliferation of mobile devices is producing a new wave of applications that enable users to sense their surroundings with smart phones. People are preferring mobile devices to search and browse video content on the move. In this paper, we have developed an innovative mobile video search system through which users can discover videos by simply pointing their phones at a screen to capture a very few seconds of what they are watching. Different than most existing mobile video search applications, the proposed system is aiming at instant and progressive video search by leveraging the light-weight computing capacity of mobile devices. In particular, the system is able to index large-scale video data using a new layered audio-video indexing approach in the cloud, as well as generate lightweight joint audio-video signatures with progressive transmission and perform progressive search on mobile devices. Furthermore, we showcase that the system can be applied to two novel applications—video entity search and video clip localization. The evaluations on the real-world mobile video query dataset show that our system significantly improves user’s search experience due to search accuracy, low retrieval latency, and very short recording duration.
Wu Liu 0005, Tao Mei 0001, Yongdong Zhang 0001
IEEE Trans. Multim.1
2013 Listen, look, and gotcha: instant video search with mobile phones by layered audio-video indexing
abstract
Mobile video is quickly becoming a mass consumer phenomenon. More and more people are using their smartphones to search and browse video content while on the move. In this paper, we have developed an innovative instant mobile video search system through which users can discover videos by simply pointing their phones at a screen to capture a very few seconds of what they are watching. The system is able to index large-scale video data using a new layered audio-video indexing approach in the cloud, as well as extract light-weight joint audio-video signatures in real time and perform progressive search on mobile devices. Unlike most existing mobile video search applications that simply send the original video query to the cloud, the proposed mobile system is one of the first attempts at instant and progressive video search leveraging the light-weight computing capacity of mobile devices. The system is characterized by four unique properties: 1) a joint audio-video signature to deal with the large aural and visual variances associated with the query video captured by the mobile phone, 2) layered audio-video indexing to holistically exploit the complementary nature of audio and video signals, 3) light-weight fingerprinting to comply with mobile processing capacity, and 4) a progressive query process to significantly reduce computational costs and improve the user experience---the search process can stop anytime once a confident result is achieved. We have collected 1,400 query videos captured by 25 mobile users from a dataset of 600 hours of video. The experiments show that our system outperforms state-of-the-art methods by achieving 90.79% precision when the query video is less than 10 seconds and 70.07% even when the query video is less than 5 seconds.
Wu Liu 0005, Tao Mei 0001, Yongdong Zhang 0001, Jintao Li 0001, Shipeng Li 0001
ACM Multimedia1
2013 LAVES: an instant mobile video search system based on layered audio-video indexing
abstract
This demonstration presents an innovative instant mobile video search system based on layered audio-video indexing, called "LAVES." Through the system, users can discover videos by simply pointing their phones at a screen to capture a very few seconds of what they are watching. Unlike most existing mobile video search applications which simply send the original video query to the cloud, the proposed mobile system is one of the first attempts towards instant and progressive video search leveraging the light-weight computing capacity of mobile devices. The system is able to index large-scale video data using the layered audio-video indexing technique on the cloud, as well as extract light-weight joint audio-video signatures in real time and perform bipartite-graph-based progressive search process on the devices. On a 600 hours video dataset, the system can outperform the state-of-the-arts by achieving 90.79% precision when the query video is less than 10 seconds.
Wu Liu 0005, Feibin Yang, Yongdong Zhang 0001, Qinghua Huang, Tao Mei 0001
ACM Multimedia1
2013 Accurate Estimation of Human Body Orientation From RGB-D Sensors
abstract
Accurate estimation of human body orientation can significantly enhance the analysis of human behavior, which is a fundamental task in the field of computer vision. However, existing orientation estimation methods cannot handle the various body poses and appearances. In this paper, we propose an innovative RGB-D-based orientation estimation method to address these challenges. By utilizing the RGB-D information, which can be real time acquired by RGB-D sensors, our method is robust to cluttered environment, illumination change and partial occlusions. Specifically, efficient static and motion cue extraction methods are proposed based on the RGB-D superpixels to reduce the noise of depth data. Since it is hard to discriminate all the 360 (°) orientation using static cues or motion cues independently, we propose to utilize a dynamic Bayesian network system (DBNS) to effectively employ the complementary nature of both static and motion cues. In order to verify our proposed method, we build a RGB-D-based human body orientation dataset that covers a wide diversity of poses and appearances. Our intensive experimental evaluations on this dataset demonstrate the effectiveness and efficiency of the proposed method.
Wu Liu 0005, Yongdong Zhang 0001, Sheng Tang, Jinhui Tang 0001, Richang Hong, Jintao Li 0001
IEEE Trans. Cybern.1
2012 RGB-D Based Multi-attribute People Search in Intelligent Visual Surveillance
Wu Liu 0005, Tian Xia 0002, Ji Wan, Yongdong Zhang 0001, Jintao Li 0001
MMM1
2012 Web Video Geolocation by Geotagged Social Resources
abstract
This paper considers the problem of web video geolocation: we hope to determine where on the Earth a web video was taken. By analyzing a 6.5-million geotagged web video dataset, we observe that there exist inherent geography intimacies between a video with its relevant videos (related videos and same-author videos). This social relationship supplies a direct and effective cue to locate the video to a particular region on the earth. Based on this observation, we propose an effective web video geolocation algorithm by propagating geotags among the web video social relationship graph. For the video that have no geotagged relevant videos, we aim to collect those geotagged relevant images that are content similar with the video (share some visual or textual information with the video) as the cue to infer the location of the video. The experiments have demonstrated the effectiveness of both methods, with the geolocation accuracy much better than state-of-the-art approaches. Finally, an online web video geolocation system: Video2Locatoin (V2L) is developed to provide public access to our algorithm.
Yicheng Song, Yongdong Zhang 0001, Juan Cao 0001, Tian Xia 0002, Wu Liu 0005, Jintao Li 0001
IEEE Trans. Multim.5