VLDB 2026 Research / reviewers in the wild / expert
Jian Zhao 0006
dblp:70/2932-6
· DBLP profile ↗
102ranked-venue papers
12as first author
80since 2021 · last 2026
0000-0002-3508-756XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 63 · 7 first-author · 48 since 2021Artificial intelligence and machine learning · 49 · 11 first-author · 37 since 2021Computer networks · 11 · 11 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language ModelsabstractMultimodal Large Language Models (MLLMs) are susceptible to the implicit reasoning risk, wherein innocuous unimodal inputs synergistically assemble into risky multimodal data that produce harmful outputs. We attribute this vulnerability to the difficulty of MLLMs maintaining safety alignment through long-chain reasoning.To address this issue, we introduce Safe-Semantics-but-Unsafe-Interpretation (SSUI), the first dataset featuring interpretable reasoning paths tailored for such a cross-modal challenge.A novel training framework, Safety-aware Reasoning Path Optimization (SRPO), is also designed based on the SSUI dataset to align the MLLM's internal reasoning process with human safety values. Experimental results show that our SRPO-trained models achieve state-of-the-art results on key safety benchmarks, including the proposed Reasoning Path Benchmark (RSBench), significantly outperforming both open-source and top-tier commercial MLLMs. Shujuan Liu, Jian Zhao 0006, Ziyan Shi, Yusheng Zhao, Yuchen Yuan, Chi Zhang 0012, Xuelong Li 0001 |
AAAI | 3 |
| 2026 | GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative ReasoningabstractRecent advancements in Large Language Models (LLMs) have shown that it is promising to utilize Process Reward Models (PRMs) as verifiers to enhance the performance of LLMs. However, current PRMs face three key challenges: (1) limited process supervision and generalization capabilities, (2) dependence on scalar value prediction without leveraging the generative abilities of LLMs, and (3) inability to scale the test-time compute of PRMs. In this work, we introduce GenPRM, a generative process reward model that performs explicit Chain-of-Thought (CoT) reasoning with code verification before providing judgment for each reasoning step. To obtain high-quality process supervision labels and rationale data, we propose Relative Progress Estimation (RPE) and a rationale synthesis framework that incorporates code verification. Experimental results on ProcessBench and several mathematical reasoning tasks show that GenPRM significantly outperforms prior PRMs with only 23K training data from MATH dataset. Through test-time scaling, a 1.5B GenPRM outperforms GPT-4o, and a 7B GenPRM surpasses Qwen2.5-Math-PRM-72B on ProcessBench. Additionally, GenPRM demonstrates strong abilities to serve as a critic model for policy model refinement. This work establishes a new paradigm for process supervision that bridges the gap between PRMs and critic models in LLMs. Jian Zhao 0006, Runze Liu 0002, Zhimu Zhou, Junqi Gao, Dong Li 0016, Jiafei Lyu, Zhouyi Qian, Biqing Qi, Xiu Li 0001, Bowen Zhou 0002 |
AAAI | 1 |
| 2026 | Dual Activation-Weight Sparsity: A Training-Free Framework for Efficient Large Language Model CompressionabstractLuoyang Sun, Guangyan Li, Cheng Deng, Haifeng Zhang, Jian Zhao, Yongqiang Tang, Wensheng Zhang, Jun Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Luoyang Sun, Guangyan Li, Cheng Deng 0001, Haifeng Zhang 0002, Jian Zhao 0006, Yongqiang Tang, Wensheng Zhang 0002, Jun Wang 0012 |
ACL (1) | 5 |
| 2026 | ForgeryMoE: Mixture of Experts for Image Forgery Detection under JPEG CompressionabstractAI-generated imagery threatens information integrity, making reliable forgery detection crucial. However, the pervasive use of JPEG compression throughout image sharing pipelines critically undermines the robustness and generalization capability of existing detectors. To address this, we propose ForgeryMoE, a new Mixture of Experts (MoE) framework designed for universal and robust image forgery detection. Our framework strategically integrates three complementary experts: a Frequency Domain Expert that analyzes wavelet-based artifacts, a Pre-trained model expert that leverages CLIP and DINOv2 for semantic inconsistencies, and a Visual Domain Expert that captures pixel-level features and reconstruction anomalies. A dual gating mechanism dynamically estimates expert reliability and adaptively fuses their decisions. Extensive experiments on the GenImage benchmark show our method achieves state-of-the-art performance in both pristine and JPEG compressed settings, demonstrating strong generalization across diverse generative models and compression qualities. Saihui Hou, Jian Zhao 0006, Zhaofeng He 0001 |
ICMR | 4 |
| 2026 | XFir: Accelerating New-Flow Setup on Host Servers of a Large Cloud NetworkabstractIn today's cloud networks, host servers widely deploy Data Processing Units (DPUs) as network accelerators under the "Sep-Path" paradigm. However, as server capabilities scale with increasing CPU cores and network bandwidth, the software slow path (executed on a DPU's CPU) has become a critical bottleneck for workloads with high new-flow rates. Meanwhile, new-flow setup logic on host servers must continuously evolve to meet diverse and changing customer demands, making flexibility a key requirement alongside performance. To address this gap, we present XFir, the first hardware-accelerated new-flow setup system for cloud host servers that delivers high CPS throughput while preserving sufficient flexibility. XFir leverages a next-generation DPU equipped with a Cloud Network co-Processor (CNP) to execute the host server's new-flow setup logic. XFir redesigns the host-server flow-setup datapath and table layout, optimizes LPM lookups, and introduces CPU-CNP collaboration mechanisms to further improve performance and reliability. Our evaluation shows that XFir achieves over 776K new-flow CPS on a single host server with 11.7μs slow-path latency. Compared to prior work (Fornax), XFir achieves 4.8x CPS and reduces latency by 69.2%. Moreover, XFir is cost-effective to deploy, requiring only a single DPU per host. Overall, XFir improves new-flow throughput while maintaining development flexibility at low financial cost. Shihan Lin, Shunqiao Jiang, Chao Pei, Jian Zhao 0006, Wenjun Wu 0001, Lijun Zhuang, Qingmin Liu, Heng Yu 0005, Yibo Huang 0005, Yifei Zhu 0001, Yunming Xiao, Ang Chen 0001, Linghe Kong, Congcong Miao |
SIGCOMM | 6 |
| 2026 | DistDPU: A Disaggregated DPU Architecture for High-Performance and Cost-Efficient AI CloudsabstractAI training and inference are driving cloud networks toward terabit-per-second (Tbps) bandwidth per server, challenging the scalability and efficiency of today's cloud network architectures. A prevalent design scales bandwidth by stacking monolithic Data Processing Units (DPUs), but this approach tightly couples control and data plane resources, leading to excessive cost, power consumption, and operational complexity. We identify a fundamental control-data plane divergence in AI clouds: while data plane bandwidth demand grows rapidly, control plane demand remains largely flat due to the dominance of elephant flows. As a result, monolithic DPUs become systematically over-provisioned when used as bandwidth scaling primitives. Lizhou Gao, Yuanyi Zhu, Chao Pei, Chuhao Chen 0001, Zijian Li 0003, Jian Zhao 0006, Dongbo Gu, Hongchen Ren, Jiyuan Chen, Yunpeng Guan, Jianye Yuan, Yibo Huang 0005, Yang Xu 0010 |
SIGCOMM | 9 |
| 2026 | Dorado: Scaling SmartNIC Session Tables on Commodity DDRs
Heng Yu 0005, Jiajun Liang, Baozeng Zhang, Guozhi Lin, Xinyi Zhang 0004, Jian Zhao 0006, Ziyue Zhai, Chao Pei, Jilong Wang 0001, Gaogang Xie, Ang Chen 0001, Congcong Miao |
SIGCOMM | 8 |
| 2026 | LMT-SDNN: A Lightweight Malicious Traffic Detection Method for the Internet of Things Based on Multiteacher DistillationabstractThe rapid proliferation of Internet of Things (IoT) devices, coupled with their inherent security vulnerabilities, has significantly expanded the attack surface, intensifying threats such as man-in-the-middle attacks, traffic hijacking, and distributed denial-of-service (DDoS) attacks, thereby posing serious risks to the security and reliability of the entire ecosystem. The network traffic associated with IoT devices is diverse and dynamic, often exhibiting complex structural features such as periodic fluctuations, varying packet sizes, and time-varying patterns that make detection challenging. Although deep learning has demonstrated strong capabilities in efficiently identifying complex and dynamic malicious traffic through powerful feature extraction and adaptive learning abilities, its high model complexity, substantial computational demands, and large parameter sizes hinder direct deployment on resource-constrained IoT devices. In order to tackle this issue, this paper proposes a malicious traffic detection framework for the Internet of Things (IoT) based on multi-teacher knowledge distillation. The proposed model, termed the Lightweight Multi-Teacher Spatiotemporal Distillation Neural Network (LMT-SDNN), employs two high-performance teacher models: Residual Inception and a One-Dimension Convolution Netural Network (1D-CNN) integrated with idirectional Long Short-Term Memory (BiLSTM), to effectively capture the complex structural features of network traffic. Furthermore, a novel Time-Related Window Loss (TRW) function is design to enhance the student’s ability to capture temporal features, thereby improving its overall performance. The effectiveness of LMT-SDNN is validated through comparisons with five baseline models on two publicly available datasets, ToN_IoT and BoT_IoT. Experimental results show that LMT-SDNN achieves a compression rate of over 99% in both model complexity and parameter count, while maintaining an accuracy exceeding 99%, indicating its strong potential for multiclass malicious traffic detection in IoT environments. Yunfang Liang, Chunhai Li, Chuan Zhang 0003, Liehuang Zhu, Jian Zhao 0006 |
IEEE Internet Things J. | 7 |
| 2026 | NuwaDynamics+: A Causality-Aware Generative Framework for Spatio-Temporal Representation LearningabstractSpatio-temporal (ST) prediction is crucial in earth sciences, including meteorological forecasting and urban computing, to name just a few. Access to ample high-quality data, combined with deep models adept at inference, is essential for attaining significant outcomes. Yet, data scarcity and the substantial costs of sensor deployment result in notable data imbalances. Overly specialized models that lack causal linkages further undermine the generalizability of inference techniques. To address these challenges, we first introduce a causal framework for ST predictions, named $\mathtt{NuwaDynamics}$NuwaDynamicsNuwaDynamics, aimed at pinpointing causal regions in data and providing models with the capability for causal reasoning in a dual-phase process. Initially, we employ upstream self-supervision to identify causally significant patches, equipping the model with generalizable insights and performing targeted interventions on non-essential patches to approximate potential testing distributions. This stage is known as the discovery phase. Progressing from discovery, we apply the insights to downstream tasks tailored to specific ST goals, enhancing the model's recognition of a wider potential data distribution and augmenting its causal perceptual abilities (referred to as the Update phase). Additionally, we address environmental controllability and high computational complexity by implementing channel multiplication and conditional generation methods. This process, termed $\mathtt{NuwaDynamics+}$NuwaDynamics+NuwaDynamics+, can further be interpreted as the front-door adjustment technique in the causality domain. Through comprehensive experiments across ten real-world or simulated ST benchmarks, we demonstrate that integrating the $\mathtt{NuwaDynamics+}$NuwaDynamics+NuwaDynamics+ concept substantially improves various model performance. $\mathtt{NuwaDynamics+}$NuwaDynamics+NuwaDynamics+ concept also significantly enhances the versatility across various dynamic ST tasks, such as extreme weather forecasting and long-temporal-step super-resolution predictions. Kun Wang 0056, Yifan Duan, Hao Wu 0083, Jian Zhao 0006, Kai Wang 0036, Zhengyang Zhou, Yuxuan Liang 0002, Xu Wang 0029, Yang Wang 0015, Yu Zheng 0004, Xuelong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | DrawMotion: Generating 3D Human Motions by Freehand Drawing
Tao Wang 0011, Lei Jin 0003, Qiaozhi He, Jiaming Chu, Yu Cheng 0009, Junliang Xing, Jian Zhao 0006, Shuicheng Yan, Li Wang 0039 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2026 | An improved HotStuff consensus algorithm based on a double reputation evaluate modelabstractIn order to solve the situation that the Leader node in the blockchain consensus algorithm HotStuff is randomly selected and the consensus success rate decreases significantly when facing the Byzantine node attack, this paper designs a consensus algorithm DR-HotStuff(Double Reputation evaluate- HotStuff) with double reputation evaluation model applicable to this complex environment. HotStuff), which maintains the advantages of low communication complexity of the HotStuff algorithm, combines with the dual reputation model to quickly screen out the Byzantine nodes in the consensus nodes, and assigns each node a corresponding identity based on different reputation values to reduce the probability that the Leader node is a malicious node, which further improves the stability and work efficiency of the blockchain. The experimental results show that the consensus algorithm has high throughput and low consensus delay, and can still eliminate most of the Byzantine nodes when the number of Byzantine nodes is greater than 30% and less than 50%, so that the efficiency of the blockchain can be restored to the optimal state through up to four rounds of consensus. Yunfang Liang, Guogang Zhao, Jian Zhao 0006, Dawen Sun |
Peer Peer Netw. Appl. | 5 |
| 2026 | Phys-EdiGAN: A privacy-preserving method for editing physiological signals in facial videos
Xiaoguang Tu, Zhiyi Niu, Juhang Yin, Zhaoxin Fan, Jian Zhao 0006 |
Pattern Recognit. | 9 |
| 2026 | MA-FSAR: Multimodal Adaptation of CLIP for few-shot action recognition
Jiazheng Xing, Jian Zhao 0006, Chao Xu 0023, Mengmeng Wang 0005, Guang Dai, Yong Liu 0007, Jingdong Wang 0001, Xuelong Li 0001 |
Pattern Recognit. | 2 |
| 2026 | Corrigendum to "MA-FSAR: Multimodal Adaptation of CLIP for few-shot action recognition" [Pattern Recognition 169 (2026) 111902]
Jiazheng Xing, Jian Zhao 0006, Chao Xu 0023, Mengmeng Wang 0005, Guang Dai, Yong Liu 0007, Jingdong Wang 0001, Xuelong Li 0001 |
Pattern Recognit. | 2 |
| 2026 | Corrigendum to "FaceChain-MMID: Generating highly identity-consistent realistic portraits via dividing & merging multi-modal representations" [Pattern Recognition 168 (2025) 111858]
Chao Xu 0023, Fei Wang 0032, Baigui Sun, Jian Zhao 0006 |
Pattern Recognit. | 5 |
| 2026 | Beyond similarity: Mutual information-guided retrieval for in-context learning in VQA
Zezhong Lv, Jian Zhao 0006, Yan Wang 0122, Yuchen Yuan, Yuchu Jiang, Wenqi Ren, Xuelong Li 0001 |
Pattern Recognit. | 3 |
| 2026 | CLAP: Cross-Layer Adaptive Pipelining Inference Scheduling for Resource-Efficient Edge-Cloud Vision SystemsabstractWith the rapid growth of video-based applications, edge-cloud collaboration has become a mainstream paradigm for large-scale visual inference. However, existing edge-cloud systems primarily emphasize task offloading and static resource allocation, often overlooking the dynamic and heterogeneous nature of real-world scenarios. The significant variability in scene complexity across tasks leads to inefficient system performance. In this article, we propose CLAP, a cross-layer adaptive pipelining inference scheduling framework for edge-cloud vision systems. First, CLAP introduces a lightweight multiscale scene-aware module that accurately characterizes the visual complexity of incoming tasks at different granularities with minimal overhead. Based on this complexity profile, we design an adaptive multi-stage pipeline scheduling strategy, which dynamically adjusts processing granularity and selectively activates stages across edge and cloud nodes. Furthermore, we formulate the resource allocation as a multi-agent decision-making problem and employ cross-layer reinforcement learning to optimize task distribution under complex objectives, efficiently balancing accuracy, delay, and energy consumption. Extensive evaluations on public datasets demonstrate that CLAP can improve the throughput by more than 2.1x compared to traditional cloud-only and edge-only solutions while meeting accuracy requirements. Compared to state-of-the-art edge-cloud methods, CLAP achieves a 3% improvement in inference accuracy while simultaneously reducing end-to-end resource overhead, delay, and energy consumption by over 35%, proving its effectiveness in dynamic, large-scale vision applications. Zheming Yang, Wen Ji 0003, Qi Guo 0009, Jian Zhao 0006, Xingzhou Zhang, Yangyu Zhang, Yang You 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2026 | SynSP++: General Pose Sequences Refinement via Synergy of Smoothness and Precision
Lei Jin 0003, Tao Wang 0011, Junliang Xing, Jian Zhao 0006, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | CCIN: Compositional Conflict Identification and Neutralization for Composed Image RetrievalabstractComposed Image Retrieval (CIR) is a multi-modal task that seeks to retrieve target images by harmonizing a reference image with a modified instruction. A key challenge in CIR lies in compositional conflicts between the reference image (e.g., blue, long sleeve) and the modified instruction (e.g., grey, short sleeve). Previous works attempt to mitigate such conflicts through feature-level manipulation, commonly employing learnable masks to obscure conflicting features within the reference image. However, the inherent complexity of feature spaces poses significant challenges in precise conflict neutralization, thereby leading to uncontrollable results. To this end, this paper proposes the Compositional Conflict Identification and Neutralization (CCIN) framework, which sequentially identifies and neutralizes compositional conflicts for effective CIR. Specifically, CCIN comprises two core modules: 1) Compositional Conflict Identification module, which utilizes LLM-based analysis to identify specific conflicting attributes, and 2) Compositional Conflict Neutralization module, which first generates a kept instruction to preserve non-conflicting attributes, then neutralizes conflicts under collaborative guidance of both the kept and modified instructions. Extensive experiments demonstrate the superiority of CCIN over the state-of-the-arts. Code repository: https://github.com/LikaiTian/CCIN. Likai Tian, Jian Zhao 0006, Zechao Hu 0003, Zhengwei Yang 0001, Hao Li 0093, Lei Jin 0003, Zheng Wang 0007, Xuelong Li 0001 |
CVPR | 2 |
| 2025 | MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training SmoothingabstractDeep visual odometry has demonstrated great advancements by learning-to-optimize technology. This approach heavily relies on the visual matching across frames. However, ambiguous matching in challenging scenarios leads to significant errors in geometric modeling and bundle adjustment optimization, which undermines the accuracy and robustness of pose estimation. To address this challenge, this paper proposes MambaVO, which conducts robust initialization, Mamba-based sequential matching refinement, and smoothed training to enhance the matching quality and improve the pose estimation. Specifically, the new frame is matched with the closest keyframe in the maintained Point-Frame Graph (PFG) via the semi-dense based Geometric Initialization Module (GIM). Then the initialized PFG is processed by a proposed Geometric Mamba Module (GMM), which exploits the matching features to refine the overall inter-frame matching. The refined PFG is finally processed by differentiable BA to optimize the poses and the map. To deal with the gradient variance, a Trending-Aware Penalty (TAP) is proposed to smooth training and enhance convergence and stability. A loop closure module is finally applied to enable MambaVO++. On public benchmarks, MambaVO and MambaVO++ demonstrate SOTA performance, while ensuring real-time running. Shuo Wang 0015, Yongcai Wang, Zhaoxin Fan, Jian Zhao 0006, Deying Li 0001 |
CVPR | 7 |
| 2025 | StickMotion: Generating 3D Human Motions by Drawing a StickmanabstractText-to-motion generation, which translates textual descriptions into human motions, has been challenging in accurately capturing detailed user-imagined motions from simple text inputs. This paper introduces StickMotion, an efficient diffusion-based network designed for multi-condition scenarios, which generates desired motions based on traditional text and our proposed stickman conditions for global and local control of these motions, respectively. We address the challenges introduced by the user-friendly stickman from three perspectives: 1) Data generation. We develop an algorithm to generate hand-drawn stickmen automatically across different dataset formats. 2) Multi-condition fusion. We propose a multi-condition module that integrates into the diffusion process and obtains outputs of all possible condition combinations, reducing computational complexity and enhancing StickMotion’s performance compared to conventional approaches with the self-attention module. 3) Dynamic supervision. We empower StickMotion to make minor adjustments to the stickman’s position within the output sequences, generating more natural movements through our proposed dynamic supervision strategy. Through quantitative experiments and user studies, sketching stickmen saves users about 51.5% of their time generating motions consistent with their imagination. Our codes, demos, and relevant data will be released in https:// github.com/InvertedForest/StickMotion. Tao Wang 0011, Qiaozhi He, Jiaming Chu, Ling Qian, Yu Cheng 0009, Junliang Xing, Jian Zhao 0006, Lei Jin 0003 |
CVPR | 8 |
| 2025 | JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV SystemsabstractUnmanned Aerial Vehicles (UAVs) are widely adopted across various fields, yet they raise significant privacy and safety concerns, demanding robust monitoring solutions. Existing anti-UAV methods primarily focus on position tracking but fail to capture UAV behavior and intent. To address this, we introduce a novel task—UAV Tracking and Intent Understanding (UTIU)—which aims to track UAVs while inferring and describing their motion states and intent for a more comprehensive monitoring approach. To tackle the task, we propose JTD-UAV, the first joint tracking, and intent description framework based on large language models. Our dual-branch architecture integrates UAV tracking with Visual Question Answering (VQA), allowing simultaneous localization and behavior description. To benchmark this task, we introduce the TDUAV dataset, the largest dataset for joint UAV tracking and intent understanding, featuring 1,328 challenging video sequences, over 163K annotated thermal frames, and 3K VQA pairs. Our benchmark demonstrates the effectiveness of JTD-UAV. Jian Zhao 0006, Zhaoxin Fan, Xin Zhang 0093, Yudian Zhang, Lei Jin 0003, Gang Wang 0031, Mengxi Jia, Xuelong Li 0001 |
CVPR | 2 |
| 2025 | MAGiC: An LLM-Powered Multi-Agent Framework for Unleashing Visual CreativityabstractHumans can complete high-quality creative work, such as drawing a picture or creating a video based on text, and editing images or videos according to textual requirements. In the earlier period of artificial intelligence, the “best of N” strategy was often utilized to leverage the creative capability of multiple visual creators, which was computationally inefficient and labor-intensive. With the emergence of large language models (LLMs), the LLM-based agent dynamically plans the invocation of tools to accomplish creative tasks. However, these agent systems struggle to achieve optimal tool planning and creative performance, especially complex creative tasks. Toward these issues, we propose MAGiC, a LLM-Powered Multi-Agent Framework for Visual Generation and Editing to unleash Visual Creativity. MAGiC addresses users’ creation requirements through the collaboration of four modules, i.e., task assignment, planning, execution, and evaluation, with each controlled by agents configured for different roles. Specifically, the task assignment module iteratively releases new tasks based on the user requirements and its completion progress. The Planning module configures the corresponding Planner for different types of tasks, and these Planners create detailed plans for the tasks they are responsible for. The Execution module iteratively executes the plans set by the Planners. The evaluation module assesses the result obtained by the execution module to prevent errors from affecting subsequent tasks. Finally, MAGiC illustrates excellent creativity on multiple tasks, especially complex ones, and the high scalability of MAGiC is an initial step in applying multi-agent systems to llm-based visual system. Shilong Wang 0002, Jian Zhao 0006, Yawen Cui, Chi Zhang 0012, Xuelong Li 0001 |
ECAI | 2 |
| 2025 | BF-HFD: Hidden Follower Detection Based on Behavioral Features
TongQing Zhu, Jian Zhao 0006 |
PRCV (5) | 5 |
| 2025 | URMP: The Uncertainty Relation Reasoning and Minimum Permutation Distance for Text-Based Group Re-Identification
BinQuan Tan, XiaoMing Guo, Tongqing Zhu, Jian Zhao 0006, LuMei Zhou |
PRCV (15) | 5 |
| 2025 | AsynFusion: Towards Asynchronous Latent Consistency Models for Decoupled Whole-Body Audio-Driven Avatars
Tianbao Zhang, Jian Zhao 0006, Yuer Li, Zhaoxin Fan, Wenjun Wu 0001, Xuelong Li 0001 |
PRCV (7) | 2 |
| 2025 | Fornax: A Hardware-Centric Session Management in Large Public Cloud NetworkabstractSmartNIC is increasingly utilized to accelerate cloud network components. The effectiveness and correctness of hardware acceleration heavily rely on its management mechanism. Unfortunately, traditional management mechanisms adopt software-centric architecture, which treats flow as the basic management unit and completely relies on one-way commands to manage the flow table, making it challenging to support various cloud network scenarios while managing extremely large tables. In this paper, we advocate for a radical new mechanism to shift the management paradigm from software-centric architecture to hardware-centric architecture, which adopts session as the basic management unit and designs two-way protocols to facilitate the management process. We propose and implement a first-of-its-kind system, called Fornax, a novel management architecture for large public cloud networks. At the core of Fornax is leveraging a session-empowered hardware engine to provide various management capabilities. Besides, Fornax utilizes a light-weight software manager to enhance system scalability, and hardware-driven management protocols to improve resource efficiency. Our testbed evaluations demonstrate that Fornax can reduce the software storage usage by 80% and CPU usage by 77% with little hardware resource overhead. Our large-scale production results show that Fornax can manage up to 16M session entries while significantly reducing the resource overhead by over 79%. Heng Yu 0005, Jian Zhao 0006, Guozhi Lin, Baozeng Zhang, Yunpeng Guan, Jiajun Liang, Chao Pei, Yachen Wang, Xin Jin 0008, Jilong Wang 0001, Congcong Miao |
SIGCOMM | 3 |
| 2025 | LEAF: Unveiling two sides of the same coin in semi-supervised facial expression recognition
Fan Zhang 0111, Zhi-Qi Cheng, Jian Zhao 0006, Xiaojiang Peng, Xuelong Li 0001 |
Comput. Vis. Image Underst. | 3 |
| 2025 | UniFace++: Revisiting a Unified Framework for Face Reenactment and Swapping via 3D Priors
Chao Xu 0023, Yijie Qian, Shaoting Zhu, Baigui Sun, Jian Zhao 0006, Yong Liu 0007, Xuelong Li 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | SurANet: Surrounding-Aware Network for concealed object detection via highly-efficient interactive contrastive learning strategy
Yuhan Kang, Qingpeng Li, Leyuan Fang, Jian Zhao 0006, Xuelong Li 0001 |
Neurocomputing | 4 |
| 2025 | FaceChain-MMID: Generating highly identity-consistent realistic portraits via dividing & merging multi-modal representations
Chao Xu 0023, Fei Wang 0032, Baigui Sun, Jian Zhao 0006 |
Pattern Recognit. | 5 |
| 2025 | CMoA: Contrastive Mixture of Adapters for Generalized Few-Shot Continual LearningabstractThe goal of Few-Shot Continual Learning (FSCL) is to incrementally learn novel tasks with limited labeled samples and preserve previous capabilities simultaneously. However, current FSCL works lack research on domain increment and domain generalization ability, which cannot cope with changes in the visual perception environment. In this paper, we set up a Generalized FSCL (GFSCL) protocol involving both class- and domain-incremental scenarios together with domain generalization assessment. Firstly, two benchmark datasets and protocols are newly arranged, and detailed baselines are provided for this unexplored configuration. Furthermore, we find that common continual learning methods have poor generalization ability on unseen domains and cannot better tackle catastrophic forgetting issue in cross-incremental tasks. Hence, we propose a rehearsal-free framework based on Vision Transformer (ViT) named Contrastive Mixture of Adapters (CMoA). It contains two non-conflicting parts: (1) By applying the fast-adaptation characteristic of adapter-embedded ViT, the mixture of Adapters (MoA) module is incorporated into ViT. For stability purpose, cosine similarity regularization and dynamic weighting are designed to make each adapter learn specific knowledge and concentrate on particular classes. (2) To further enhance domain generalization ability, we alleviate the intra-class variation by prototype-calibrated contrastive learning to improve domain-invariant representation learning. Finally, six evaluation indicators showing the overall performance and forgetting are compared by comprehensive experiments on two benchmark datasets to validate the efficacy of CMoA, and the results illustrate that CMoA can achieve comparative performance with rehearsal-based continual learning methods. Yawen Cui, Jian Zhao 0006, Zitong Yu, Rizhao Cai, Lei Jin 0003, Alex Chichung Kot, Li Liu 0002, Xuelong Li 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | SynSP: Synergy of Smoothness and Precision in Pose Sequences RefinementabstractPredicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily concentrated on striking a balance between the two objectives, i.e., smoothness and precision, while optimizing the predicted pose sequences. However, it has come to our attention that the tension between these two objectives can provide additional quality cues about the predicted pose sequences. These cues, in turn, are able to aid the network in optimizing lower-quality poses. To leverage this quality information, we propose a motion refinement network, termed SynSP, to achieve a Synergy of Smoothness and Precision in the sequence refinement tasks. Moreover, SynSP can also address multi-view poses of one person simultaneously, fixing inaccuracies in predicted poses through heightened attention to similar poses from other views, thereby amplifying the resultant quality cues and overall performance. Compared with previous methods, SynSP benefits from both pose quality and multi-view information with a much shorter input sequence length, achieving state-of-the-art results among four challenging datasets involving 2D, 3D, and SMPL pose representations in both single-view and multi-view scenes. Github code: https://github.com/InvertedForest/SynSP. Tao Wang 0011, Lei Jin 0003, Zheng Wang 0007, Jianshu Li, Liang Li 0003, Fang Zhao 0006, Yu Cheng 0009, Li Yuan 0007, Junliang Xing, Jian Zhao 0006 |
CVPR | 11 |
| 2024 | DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous DrivingabstractVision-centric autonomous driving has recently raised wide attention due to its lower cost. Pretraining is essential for extracting a universal representation. However, current vision-centric pretraining typically relies on either 2D or 3D pre-text tasks, overlooking the temporal characteristics of autonomous driving as a 4D scene understanding task. In this paper, we address this challenge by introducing a world model-based autonomous driving 4D representation learning framework, dubbed DriveWorld, which is capable of pretraining from multi-camera driving videos in a spatiotemporal fashion. Specifically, we propose a Memory State-Space Model for spatiotemporal modelling, which consists of a Dynamic Memory Bank module for learning temporal-aware latent dynamics to predict future changes and a Static Scene Propagation module for learning spatial-aware latent statics to offer comprehensive scene contexts. We additionally introduce a Task Prompt to decouple task-aware features for various downstream tasks. The experiments demonstrate that DriveWorld delivers promising results on various autonomous driving tasks. When pretrained with the OpenScene dataset, DriveWorld achieves a 7.5% increase in mAP for 3D object detection, a 3.0% increase in IoU for online mapping, a 5.0% increase in AMOTA for multi-object tracking, a 0.1m decrease in minADE for motionforecasting, a 3.0% increase in IoU for occupancy prediction, and a 0.34m reduction in average L2 error for planning. Dawei Zhao 0003, Liang Xiao 0007, Jian Zhao 0006, Xinli Xu, Lei Jin 0003, Jianshu Li, Yulan Guo, Junliang Xing, Liping Jing, Yiming Nie, Bin Dai 0001 |
CVPR | 4 |
| 2024 | ASQuery: A Query-based Model for Action SegmentationabstractFor the task of temporal action segmentation, existing works commonly treat it as a frame-wise classification problem. In this paper, we propose a straight but effective model namely ASQuery by learning central representation of each action category, which transforms the classification problem to the similarity calculation between category-specific queries and frame features. These central representations are dynamically generated through our Transformer decoder module, endowing them more flexible and comprehensive perception of the whole video. Moreover, we first introduce the boundary query for refining segmentation results, aiding to alleviating the troublesome over-segmentation problem. ASQuery demonstrates superior performance compared to state-of-the-art models, achieving improvements of 0.9% and 4.1% in the mean metrics on two public action segmentation datasets, i.e., Breakfast and Assembly101, respectively. The source codes are available at https://github.com/zlngan/ASQuery. Ziliang Gan, Lei Jin 0003, Zheng Wang 0007, Liang Li 0003, Zhecan Wang, Jianshu Li, Junliang Xing, Jian Zhao 0006 |
ICME | 10 |
| 2024 | Unified Single-Stage Transformer Network for Efficient RGB-T Tracking
Jianqiang Xia, Dian-xi Shi, Linna Song, Songchang Jin, Chenran Zhao, Yu Cheng 0009, Lei Jin 0003, Jianan Li 0001, Gang Wang 0031, Junliang Xing, Jian Zhao 0006 |
IJCAI | 14 |
| 2024 | Two in One Go: Single-stage Emotion Recognition with Decoupled Subject-context TransformerabstractEmotion recognition aims to discern the emotional state of subjects within an image, relying on subject-centric and contextual visual cues. Current approaches typically follow a two-stage pipeline: first localize subjects by off-the-shelf detectors, then perform emotion classification through the late fusion of subject and context features. However, the complicated paradigm suffers from disjoint training stages and limited fine-grained interaction between subject-context elements. To address the challenge, we present a single-stage emotion recognition approach, employing a Decoupled Subject-Context Transformer (DSCT), for simultaneous subject localization and emotion classification. Rather than compartmentalizing training stages, we jointly leverage box and emotion signals as supervision to enrich subject-centric feature learning. Furthermore, we introduce DSCT to facilitate interactions between fine-grained subject-context cues in a ''decouple-then-fuse'' manner. The decoupled query tokens-subject queries and context queries-gradually intertwine across layers within DSCT, during which spatial and semantic relations are exploited and aggregated. We evaluate our single-stage framework on two widely used context-aware emotion recognition datasets, CAER-S and EMOTIC. Our approach surpasses two-stage alternatives with fewer parameter numbers, achieving a 3.39% accuracy improvement and a 6.46% average precision gain on CAER-S and EMOTIC datasets, respectively. Code and models are available at: https://github.com/Sampson-Lee/DSCT. Xinpeng Li 0004, Teng Wang 0007, Jian Zhao 0006, Shuyi Mao, Jinbao Wang 0001, Feng Zheng 0001, Xiaojiang Peng, Xuelong Li 0001 |
ACM Multimedia | 3 |
| 2024 | Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion ModelabstractSpatio-temporal (ST) prediction has garnered a De facto attention in earth sciences, such as meteorological prediction, human mobility perception. However, the scarcity of data coupled with the high expenses involved in sensor deployment results in notable data imbalances. Furthermore, models that are excessively customized and devoid of causal connections further undermine the generalizability and interpretability. To this end, we establish a causal framework for ST predictions, termed CaPaint, which targets to identify causal regions in data and endow model with causal reasoning ability in a two-stage process. Going beyond this process, we utilize the back-door adjustment to specifically address the sub-regions identified as non-causal in the upstream phase. Specifically, we employ a novel image inpainting technique. By using a fine-tuned unconditional Diffusion Probabilistic Model (DDPM) as the generative prior, we in-fill the masks defined as environmental parts, offering the possibility of reliable extrapolation for potential data distributions. CaPaint overcomes the high complexity dilemma of optimal ST causal discovery models by reducing the data generation complexity from exponential to quasi-linear levels. Extensive experiments conducted on five real-world ST benchmarks demonstrate that integrating the CaPaint concept allows models to achieve improvements ranging from 4.3% to 77.3%. Moreover, compared to traditional mainstream ST augmenters, CaPaint underscores the potential of diffusion models in ST enhancement, offering a novel paradigm for this field. Our project is available at https://anonymous.4open.science/r/12345-DFCC. Yifan Duan, Jian Zhao 0006, pengcheng, Junyuan Mao, Hao Wu 0098, Jingyu Xu 0002, Shilong Wang 0002, Caoyuan Ma, Kai Wang 0036, Kun Wang 0056, Xuelong Li 0001 |
NeurIPS | 2 |
| 2024 | PLRUT: Pseudo Label and Re-detection Boosted Unsupervised Tracking of Unmanned Aerial Vehicle Objects
Jun Wang 0041, Huadong Dai, Bo Zhang 0007, Shan Qin, Jian Zhao 0006 |
PRCV (12) | 5 |
| 2024 | GOP: A Group Object Perception Framework for Optical Remote Sensing
Lei Jin 0003, Xuechao Zou, Jian Zhao 0006, Junliang Xing |
PRCV (12) | 4 |
| 2024 | SkatingVerse: A large-scale benchmark for comprehensive evaluation on human action understandingabstractAbstract Human action understanding (HAU) is a broad topic that involves specific tasks, such as action localisation, recognition, and assessment. However, most popular HAU datasets are bound to one task based on particular actions. Combining different but relevant HAU tasks to establish a unified action understanding system is challenging due to the disparate actions across datasets. A large‐scale and comprehensive benchmark, namely SkatingVerse is constructed for action recognition, segmentation, proposal, and assessment. SkatingVerse focus on fine‐grained sport action, hence figure skating is chosen as the task object, which eliminates the biases of the object, scene, and space that exist in most previous datasets. In addition, skating actions have inherent complexity and similarity, which is an enormous challenge for current algorithms. A total of 1687 official figure skating competition videos was collected with a total of 184.4 h, exceeding four times over other datasets with a similar topic. SkatingVerse enables to formulate a unified task to output fine‐grained human action classification and assessment results from a raw figure skating competition video. In addition, SkatingVerse can facilitate the study of HAU foundation model due to its large scale and abundant categories. Moreover, image modality is incorporated for human pose estimation task into SkatingVerse . Extensive experimental results show that (1) SkatingVerse significantly helps the training and evaluation of HAU methods, (2) the performance of existing HAU methods has much room to improve, and SkatingVerse helps to reduce such gaps, and (3) unifying relevant tasks in HAU through a uniform dataset can facilitate more practical applications. SkatingVerse will be publicly available to facilitate further studies on relevant problems. Ziliang Gan, Lei Jin 0003, Yu Cheng 0009, Yinglei Teng, Zun Li 0001, Yawen Li 0001, Wenhan Yang, Junliang Xing, Jian Zhao 0006 |
IET Comput. Vis. | 11 |
| 2024 | PRCL: Probabilistic Representation Contrastive Learning for Semi-Supervised Semantic Segmentation
Haoyu Xie 0002, Changqi Wang, Jian Zhao 0006, Yang Liu 0356, Jun Dan, Chong Fu 0001, Baigui Sun |
Int. J. Comput. Vis. | 3 |
| 2024 | Anti-UAV410: A Thermal Infrared Benchmark and Customized Scheme for Tracking Drones in the WildabstractThe perception of drones, also known as Unmanned Aerial Vehicles (UAVs), particularly in infrared videos, is crucial for effective anti-UAV tasks. However, existing datasets for UAV tracking have limitations in terms of target size and attribute distribution characteristics, which do not fully represent complex realistic scenes. To address this issue, we introduce a generalized infrared UAV tracking benchmark called Anti-UAV410. The benchmark comprises a total of 410 videos with over 438 K manually annotated bounding boxes. To tackle the challenges of UAV tracking in complex environments, we propose a novel method called Siamese drone tracker (SiamDT). SiamDT incorporates a dual-semantic feature extraction mechanism that explicitly models targets in dynamic background clutter, enabling effective tracking of small UAVs. The SiamDT method consists of three key steps: Dual-Semantic RPN Proposals (DS-RPN), Versatile R-CNN (VR-CNN), and Background Distractors Suppression. These steps are responsible for generating candidate proposals, refining prediction scores based on dual-semantic features, and enhancing the discriminative capacity of the trackers against dynamic background clutter, respectively. Extensive experiments conducted on the Anti-UAV410 dataset and three other large-scale benchmarks demonstrate the superior performance of the proposed SiamDT method compared to recent state-of-the-art trackers. Bo Huang 0012, Jianan Li 0001, Gang Wang 0031, Jian Zhao 0006, Tingfa Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Adversarial Attacks on Video Object Segmentation With Hard Region DiscoveryabstractVideo object segmentation has been applied to various computer vision tasks, such as video editing, autonomous driving, and human-robot interaction. However, the methods based on deep neural networks are vulnerable to adversarial examples, which are the inputs attacked by almost human-imperceptible perturbations, and the adversary (i.e., attacker) will fool the segmentation model to make incorrect pixel-level predictions. This will rise the security issues in highly-demanding tasks because small perturbations to the input video will result in potential attack risks. Though adversarial examples have been extensively used for classification, it is rarely studied in video object segmentation. Existing related methods in computer vision either require prior knowledge of categories or cannot be directly applied due to the special design for certain tasks, failing to consider the pixel-wise region attack. Hence, this work develops an object-agnostic adversary that has adversarial impacts on VOS by first-frame attacking via hard region discovery. Particularly, the gradients from the segmentation model are exploited to discover the easily confused region, in which it is difficult to identify the pixel-wise objects from the background in a frame. This provides a hardness map that helps to generate perturbations with a stronger adversarial power for attacking the first frame. Empirical studies on three benchmarks indicate that our attacker significantly degrades the performance of several state-of-the-art video object segmentation models. Ping Li 0006, Li Yuan 0007, Jian Zhao 0006, Xianghua Xu, Xiaoqin Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | UniParser: Multi-Human Parsing With Unified Correlation Representation LearningabstractMulti-human parsing is an image segmentation task necessitating both instance-level and fine-grained category-level information. However, prior research has typically processed these two types of information through distinct branch types and output formats, leading to inefficient and redundant frameworks. This paper introduces UniParser, which integrates instance-level and category-level representations in three key aspects: 1) we propose a unified correlation representation learning approach, allowing our network to learn instance and category features within the cosine space; 2) we unify the form of outputs of each modules as pixel-level results while supervising instance and category features using a homogeneous label accompanied by an auxiliary loss; and 3) we design a joint optimization procedure to fuse instance and category representations. By unifying instance-level and category-level output, UniParser circumvents manually designed post-processing techniques and surpasses state-of-the-art methods, achieving 49.3% AP on MHPv2.0 and 60.4% AP on CIHP. We have released our source code, pretrained models, and demos to facilitate future studies on https://github.com/cjm-sfw/Uniparser. Jiaming Chu, Lei Jin 0003, Yinglei Teng, Jianshu Li, Yunchao Wei, Zheng Wang 0007, Junliang Xing, Shuicheng Yan, Jian Zhao 0006 |
IEEE Trans. Image Process. | 9 |
| 2024 | Rethinking the Person Localization for Single-Stage Multi-Person Pose EstimationabstractSingle-stage models for multi-person pose estimation have garnered significant attention due to their streamlined approach in generating person position localization and body structure perception in a single pass. These two parts, however, are processed individually by existing methods, leading to suboptimal results, e.g., candidates with high confidences for person localization while poor structure estimations. To this end, we propose a simple yet effective approach, namely Structure-guided Person Localization (SPL), jointly leveraging the advantages of the two aspects to solve the multi-person pose estimation problem, with two complementary novelties. First, we propose to incorporate body structure perception to guide person position localization, consequently, we introduce the Structure-guided Center Learning (SCL) to unify the quality of the body structure perception in the displacement map with the confidence of the person existence in the center map, thus achieving more accurate keypoint position localization results even with extreme poses. Second, to facilitate the end-to-end training of SPL, we propose the efficient Agency-based Scale-adaptive Learning (ASL). Specifically, we predict an agency map of the same size as the center map, which focuses on the foreground area and can adaptively adjust the scale size for each central area with the body structure perception confidence. Comprehensive experiments on challenging benchmarks including COCO and CrowdPose clearly verify the superiority of our framework, which achieves new state-of-the-art single-stage multi-person pose estimation results. Specifically, SPL obtains 72.1 AP scores and 69.5 AP scores in COCO test-dev2017 and CrowdPose test set, respectively. Lei Jin 0003, Xuecheng Nie, Wendong Wang 0003, Yandong Guo, Shuicheng Yan, Jian Zhao 0006 |
IEEE Trans. Multim. | 7 |
| 2024 | Review and Analysis of RGBT Single Object Tracking Methods: A Fusion PerspectiveabstractVisual tracking is a fundamental task in computer vision with significant practical applications in various domains, including surveillance, security, robotics, and human-computer interaction. However, it may face limitations in visible light data, such as low-light environments, occlusion, and camouflage, which can significantly reduce its accuracy. To cope with these challenges, researchers have explored the potential of combining the visible and infrared modalities to improve tracking performance. By leveraging the complementary strengths of visible and infrared data, RGB-infrared fusion tracking has emerged as a promising approach to address these limitations and improve tracking accuracy in challenging scenarios. In this article, we present a review on RGB-infrared fusion tracking. Specifically, we categorize existing RGBT tracking methods into four categories based on their underlying architectures, feature representations, and fusion strategies, namely feature decoupling based method, feature selecting based method, collaborative graph tracking method, and traditional fusion method. Furthermore, we provide a critical analysis of their strengths, limitations, representative methods, and future research directions. To further demonstrate the advantages and disadvantages of these methods, we present a review of publicly available RGBT tracking datasets and analyze the main results on public datasets. Moreover, we discuss some limitations in RGBT tracking at present and provide some opportunities and future directions for RGBT visual tracking, such as dataset diversity, unsupervised and weakly supervised applications. In conclusion, our survey aims to serve as a useful resource for researchers and practitioners interested in the emerging field of RGBT tracking, and to promote further progress and innovation in this area. Jun Wang 0041, Shengjie Li 0003, Lei Jin 0003, Hao Wu 0098, Jian Zhao 0006, Bo Zhang 0007 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | MSINet: Twins Contrastive Search of Multi-Scale Interaction for Object ReIDabstractNeural Architecture Search (NAS) has been increasingly appealing to the society of object Re-Identification (ReID), for that task-specific architectures significantly improve the retrieval performance. Previous works explore new optimizing targets and search spaces for NAS ReID, yet they neglect the difference of training schemes between image classification and ReID. In this work, we propose a novel Twins Contrastive Mechanism (TCM) to provide more appropriate supervision for ReID architecture search. TCM reduces the category overlaps between the training and validation data, and assists NAS in simulating real-world ReID training schemes. We then design a Multi-Scale Interaction (MSI) search space to search for rational interaction operations between multi-scale features. In addition, we introduce a Spatial Alignment Module (SAM) to further enhance the attention consistency confronted with images from different sources. Under the proposed NAS scheme, a specific architecture is automatically searched, named as MSINet. Extensive experiments demonstrate that our method surpasses state-of-the-art ReID methods on both indomain and cross-domain scenarios. Source code available in https://github.com/vimar-gu/MSINet. Jianyang Gu, Kai Wang 0036, Hao Luo 0004, Chen Chen 0114, Wei Jiang 0009, Yuqiang Fang, Shanghang Zhang, Yang You 0001, Jian Zhao 0006 |
CVPR | 9 |
| 2023 | 3D Implicit Transporter for Temporally Consistent Keypoint DiscoveryabstractKeypoint-based representation has proven advantageous in various visual and robotic tasks. However, the existing 2D and 3D methods for detecting keypoints mainly rely on geometric consistency to achieve spatial alignment, neglecting temporal consistency. To address this issue, the Transporter method was introduced for 2D data, which reconstructs the target frame from the source frame to incorporate both spatial and temporal information. However, the direct application of the Transporter to 3D point clouds is infeasible due to their structural differences from 2D images. Thus, we propose the first 3D version of the Transporter, which leverages hybrid 3D representation, cross attention, and implicit reconstruction. We apply this new learning system on 3D articulated objects and non-rigid animals (humans and rodents) and show that learned keypoints are spatio-temporally consistent. Additionally, we propose a closed-loop control strategy that utilizes the learned keypoints for 3D object manipulation and demonstrate its superior performance. Codes are available at https://github.com/zhongcl-thu/3D-Implicit-Transporter. Chengliang Zhong, Yuhang Zheng 0004, Yupeng Zheng, Hao Zhao 0002, Li Yi 0001, Xiaodong Mu, Ling Wang 0001, Pengfei Li 0007, Guyue Zhou, Chao Yang 0026, Jian Zhao 0006 |
ICCV | 12 |
| 2023 | Modality Meets Long-Term Tracker: A Siamese Dual Fusion Framework for Tracking UAVabstractTracking an Unmanned Aerial Vehicle (UAV) to obtain its locations and trajectory is a crucial task to avoid the unlawful use of UAVs. However, most existing UAV tracking methods fail when facing cluster environments, out-of-view, and occlusions because of their insufficient representation of global context information capacity. To mitigate these issues, we propose a new tracker, namely SiamFusion, to innovate a dual fusion procedure that leverages the advantages in both the feature and decision levels. In particular, we propose a novel feature fusion module named Modality-Fusion to utilize multi-modal information, enhancing the perception of the target. From the decision level, we further develop a local-global converter based on a multi-modal fusion decision-making mechanism to reduce the accumulation during tracking, which significantly increases the robustness of the tracking process. Extensive experiments demonstrate the superiority of the proposed SiamFusion, which achieves the best performance on Anti-UAV in terms of accuracy and speed. In particular, we exceed the state-of-the-art tracking algorithm in the tracking accuracy by 4.2% at a similar frame rate. Our source codes, pre-trained models, and online demos will be released upon acceptance. Lei Jin 0003, Shengjie Li 0003, Jianqiang Xia, Jun Wang 0041, Zun Li 0001, Wenhan Yang, Pengfei Zhang 0016, Jian Zhao 0006, Bo Zhang 0007 |
ICIP | 10 |
| 2023 | Single-Stage Multi-human Parsing via Point Sets and Center-based OffsetsabstractThis work studies the multi-human parsing problem. Existing methods, either following top-down or bottom-up two-stage paradigms, usually involve expensive computational costs. We instead present a high-performance Single-stage Multi-human Parsing (SMP) deep architecture that decouples the multi-human parsing problem into two fine-grained sub-problems,i.e., locating the human body and parts. SMP leverages the point features in the barycenter positions to obtain their segmentation and then generates a series of offsets from the barycenter of the human body to the barycenters of parts, thus performing human body and parts matching without the grouping process. Within the SMP architecture, we propose a Refined Feature Retain module to extract the global feature of instances through generated mask attention and a Mask of Interest Reclassify module as a trainable plug-in module to refine the classification results with the predicted segmentation. Extensive experiments on the MHPv2.0 dataset demonstrate the best effectiveness and efficiency of the proposed method, surpassing the state-of-the-art method by 2.1% in AP50p, 1.0% in APvolpsup>, and 1.2% in PCP50. Moreover, SMP also achieves superior performance in DensePose-COCO, verifying generalization of the model. In particular, the proposed method requires fewer training epochs and a less complex model architecture. Our codes are released in https://github.com/cjm-sfw/SMP. Jiaming Chu, Lei Jin 0003, Xiaojin Fan, Yinglei Teng, Yunchao Wei, Yuqiang Fang, Junliang Xing, Jian Zhao 0006 |
ACM Multimedia | 8 |
| 2023 | DecenterNet: Bottom-Up Human Pose Estimation Via Decentralized Pose RepresentationabstractMulti-person pose estimation in crowded scenes remains a very challenging task. This paper finds that most previous methods fail to estimate or group visible keypoints in crowded scenes rather than reasoning invisible keypoints. We thus categorize the crowded scenes into entanglement and occlusion based on the visibility of human parts and observe that entanglement is a significant problem in crowded scenes. With this observation, we propose DecenterNet, an end-to-end deep architecture to perform robust and efficient pose estimation in crowded scenes. Within DecenterNet, we introduce a decentralized pose representation that uses all visible keypoints as the root points to represent human poses, which is more robust in the entanglement area. We also propose a decoupled pose assessment mechanism, which introduces a location map to adaptively select optimal poses in the offset map. In addition, we have constructed a new dataset named SkatingPose, containing more entangled scenes. The proposed DecenterNet surpasses the best method on SkatingPose by 1.8 AP. Furthermore, DecenterNet obtains 71.2 AP and 71.4 AP on the COCO and CrowdPose datasets, respectively, demonstrating the superiority of our method. We will release our source code, trained models, and dataset to facilitate further studies in this research direction. Our code and dataset are available in https://github.com/InvertedForest/DecenterNet. Tao Wang 0011, Lei Jin 0003, Xiaojin Fan, Yu Cheng 0009, Yinglei Teng, Junliang Xing, Jian Zhao 0006 |
ACM Multimedia | 8 |
| 2023 | Uncovering the Unseen: Discover Hidden Intentions by Micro-Behavior Graph ReasoningabstractThis paper introduces a new and challenging Hidden Intention Discovery (HID) task. Unlike existing intention recognition tasks, which are based on obvious visual representations to identify common intentions for normal behavior, HID focuses on discovering hidden intentions when humans try to hide their intentions for abnormal behavior. HID presents a unique challenge in that hidden intentions lack the obvious visual representations to distinguish them from normal intentions. Fortunately, from a sociological and psychological perspective, we find that the difference between hidden and normal intentions can be reasoned from multiple micro-behaviors, such as gaze, attention, and facial expressions. Therefore, we first discover the relationship between micro-behavior and hidden intentions and use graph structure to reason about hidden intentions. To facilitate research in the field of HID, we also constructed a seminal dataset containing a hidden intention annotation of a typical theft scenario for HID. Extensive experiments show that the proposed network improves performance on the HID task by 9.9% over the state-of-the-art method SBP. Zhuo Zhou, Wenxuan Liu 0008, Danni Xu, Zheng Wang 0007, Jian Zhao 0006 |
ACM Multimedia | 5 |
| 2023 | Joint coupled representation and homogeneous reconstruction for multi-resolution small sample face recognition
Xiaojin Fan, Mengmeng Liao, Jingfeng Xue, Hao Wu 0098, Lei Jin 0003, Jian Zhao 0006, Liehuang Zhu |
Neurocomputing | 6 |
| 2023 | TransZero++: Cross Attribute-Guided Transformer for Zero-Shot LearningabstractZero-shot learning (ZSL) tackles the novel class recognition problem by transferring semantic knowledge from seen classes to unseen ones. Semantic knowledge is typically represented by attribute descriptions shared between different classes, which act as strong priors for localizing object attributes that represent discriminative region features, enabling significant and sufficient visual-semantic interaction for advancing ZSL. Existing attention-based models have struggled to learn inferior region features in a single image by solely using unidirectional attention, which ignore the transferable and discriminative attribute localization of visual features for representing the key semantic knowledge for effective knowledge transfer in ZSL. In this paper, we propose a cross attribute-guided Transformer network, termed TransZero++, to refine visual features and learn accurate attribute localization for key semantic knowledge representations in ZSL. Specifically, TransZero++ employs an attribute → visual Transformer sub-net (AVT) and a visual → attribute Transformer sub-net (VAT) to learn attribute-based visual features and visual-based attribute features, respectively. By further introducing feature-level and prediction-level semantical collaborative losses, the two attribute-guided transformers teach each other to learn semantic-augmented visual embeddings for key semantic knowledge representations via semantical collaborative learning. Finally, the semantic-augmented visual embeddings learned by AVT and VAT are fused to conduct desirable visual-semantic interaction cooperated with class semantic vectors for ZSL classification. Extensive experiments show that TransZero++ achieves the new state-of-the-art results on three golden ZSL benchmarks and on the large-scale ImageNet dataset. The project website is available at: https://shiming-chen.github.io/TransZero-pp/TransZero-pp.html. Shiming Chen 0002, Ziming Hong, Wenjin Hou, Guosen Xie, Yibing Song, Jian Zhao 0006, Xinge You, Shuicheng Yan, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | B-Tor: Anonymous communication system based on consortium blockchain
Liehuang Zhu, Feng Gao 0019, Jian Zhao 0006 |
Peer Peer Netw. Appl. | 6 |
| 2023 | Rethinking Sampling Strategies for Unsupervised Person Re-IdentificationabstractUnsupervised person re-identification (re-ID) remains a challenging task. While extensive research has focused on the framework design and loss function, this paper shows that sampling strategy plays an equally important role. We analyze the reasons for the performance differences between various sampling strategies under the same framework and loss function. We suggest that deteriorated over-fitting is an important factor causing poor performance, and enhancing statistical stability can rectify this problem. Inspired by that, a simple yet effective approach is proposed, termed group sampling, which gathers samples from the same class into groups. The model is thereby trained using normalized group samples, which helps alleviate the negative impact of individual samples. Group sampling updates the pipeline of pseudo-label generation by guaranteeing that samples are more efficiently classified into the correct classes. It regulates the representation learning process, enhancing statistical stability for feature representation in a progressive fashion. Extensive experiments on Market-1501, DukeMTMC-reID and MSMT17 show that group sampling achieves performance comparable to state-of-the-art methods and outperforms the current techniques under purely camera-agnostic settings. Code has been available at https://github.com/ucas-vg/GroupSampling. Xumeng Han, Xuehui Yu, Guorong Li, Jian Zhao 0006, Gang Pan 0002, Qixiang Ye, Jianbin Jiao, Zhenjun Han |
IEEE Trans. Image Process. | 4 |
| 2023 | 3D-Guided Frontal Face Generation for Pose-Invariant RecognitionabstractAlthough deep learning techniques have achieved extraordinary accuracy in recognizing human faces, the pose variances of images captured in real-world scenarios still hinder reliable model appliance. To mitigate this gap, we propose to recognize faces via generation frontal face images with a 3D -Guided Deep P ose- I nvariant Face Recognition M odel (3D-PIM) consisted of a simulator and a refiner module. The simulator employs a 3D Morphable Model (3D MM) to fit the shape and appearance features and recover primary frontal images with less training data. The refiner further enhances the image realism on both global facial structure and local details with adversarial training, while keeping the discriminative identity information consistent with original images. An Adaptive Weighting (AW) metric is then adopted to leverage the complimentary information from recovered frontal faces and original profile faces and to obtain credible similarity scores for recognition. Extended experiments verify the superiority of the proposed “recognition via generation” framework over state-of-the-art. Hao Wu 0098, Jianyang Gu, Xiaojin Fan, He Li 0034, Lidong Xie, Jian Zhao 0006 |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2023 | Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV TrackingabstractUnmanned Aerial Vehicles (UAV) have many applications in both commerce and recreation. However, irresponsibly operated UAVs will pose a threat to public safety. Therefore, developing our understanding of UAVs and their uses is of particular interest. This paper considers tracking UAVs, which provide multifaceted information around location, paths and trajectories. To facilitate research on this topic, we introduce a new benchmark, herein referred to as Anti-UAV, which provides a novel direction for UAV tracking with more than 300 video pairs containing over 580 k manually annotated bounding boxes. Addressing anti-UAV research challenges could help to design anti-UAV systems, which in turn may improve surveillance. Accordingly, we have proposed a simple yet effective approach, called dual-flow semantic consistency (DFSC) is proposed for UAV tracking. Modulated by the semantic flow across video sequences, tracker learns more robust class-level semantic information and obtains more discriminative instance-level features. Experiments highlight significant performance gain with the proposed approach over state-of-the-art trackers and the challenging aspects of Anti-UAV. The Anti-UAV benchmark and the code for the proposed approach have been made publicly available athttps://github.com/ucas-vg/Anti-UAVandhttps://github.com/ZhaoJ9014/Anti-UAV. Kuiran Wang, Xiaoke Peng, Xuehui Yu, Qiang Wang 0051, Junliang Xing, Guorong Li, Guodong Guo, Qixiang Ye, Jianbin Jiao, Jian Zhao 0006, Zhenjun Han |
IEEE Trans. Multim. | 11 |
| 2023 | Grouping by Center: Predicting Centripetal Offsets for the Bottom-up Human Pose EstimationabstractWe introduce Grouping by Center, a novel grouping approach for the bottom-up human pose estimation, which detects human joint first and then does grouping. The grouping strategy is the critical factor for the bottom-up pose estimation. To increase the conciseness and accuracy, we propose to use the center of the body as a grouping clue. More concretely, we predict the offsets from the keypoints to the body centers. Keypoints with aligned shifted results will be grouped as one person. However, the multi-scale variance of people can affect the prediction of the grouping clue, which has been neglected in previous research. To resolve the scale variance of the offset, we put forward a Multi-scale Translation Layer and an iterative refinement. Furthermore, we scheme a greedy grouping strategy with a dynamic threshold due to the various scales of instances. Through a comprehensive comparison, our framework is validated to be effective and practical. We also lay out the state-of-the-art performance revolving the bottom-up multi-person pose estimation on the MS-COCO dataset and the CrowdPose dataset. Lei Jin 0003, Xuecheng Nie, Luoqi Liu, Yandong Guo, Jian Zhao 0006 |
IEEE Trans. Multim. | 6 |
| 2023 | Toward High-quality Face-Mask Occluded RestorationabstractFace-mask occluded restoration aims at restoring the masked region of a human face, which has attracted increasing attention in the context of the COVID-19 pandemic. One major challenge of this task is the large visual variance of masks in the real world. To solve it we first construct a large-scale Face-mask Occluded Restoration (FMOR) dataset, which contains 5,500 unmasked images and 5,500 face-mask occluded images with various illuminations, and involves 1,100 subjects of different races, face orientations, and mask types. Moreover, we propose a Face-Mask Occluded Detection and Restoration (FMODR) framework, which can detect face-mask regions with large visual variations and restore them to realistic human faces. In particular, our FMODR contains a self-adaptive contextual attention module specifically designed for this task, which is able to exploit the contextual information and correlations of adjacent pixels for achieving high realism of the restored faces, which are however often neglected in existing contextual attention models. Our framework achieves state-of-the-art results of face restoration on three datasets, including CelebA, AR, and our FMOR datasets. Moreover, experimental results on AR and FMOR datasets demonstrate that our framework can significantly improve masked face recognition and verification performance. Feihong Lu, Qiliang Deng, Jian Zhao 0006, Kaipeng Zhang, Hong Han 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | MSDN: Mutually Semantic Distillation Network for Zero-Shot LearningabstractThe key challenge of zero-shot learning (ZSL) is how to infer the latent semantic knowledge between visual and attribute features on seen classes, and thus achieving a desirable knowledge transfer to unseen classes. Prior works either simply align the global features of an image with its associated class semantic vector or utilize unidirectional attention to learn the limited latent semantic representations, which could not effectively discover the intrinsic semantic knowledge (e.g., attribute semantics) between visual and attribute features. To solve the above dilemma, we propose a Mutually Semantic Distillation Network (MSDN), which progressively distills the intrinsic semantic representations between visual and attribute features for ZSL. MSDN incorporates an attribute→visual attention sub-net that learns attribute-based visual features, and a visual→attribute attention sub-net that learns visual-based attribute features. By further introducing a semantic distillation loss, the two mutual attention sub-nets are capable of learning collaboratively and teaching each other throughout the training process. The proposed MSDN yields significant improvements over the strong baselines, leading to new state-of-the-art performances on three popular challenging benchmarks. Our codes have been available at: https://github.com/shiming-chen/MSDN. Shiming Chen 0002, Ziming Hong, Guosen Xie, Wenhan Yang, Qinmu Peng, Kai Wang 0036, Jian Zhao 0006, Xinge You |
CVPR | 7 |
| 2022 | Single-Stage is Enough: Multi-Person Absolute 3D Pose EstimationabstractThe existing multi-person absolute 3D pose estimation methods are mainly based on two-stage paradigm, i.e., top-down or bottom-up, leading to redundant pipelines with high computation cost. We argue that it is more desirable to simplify such two-stage paradigm to a single-stage one to promote both efficiency and performance. To this end, we present an efficient single-stage solution, Decoupled Regression Model (DRM), with three distinct novelties. First, DRM introduces a new decoupled representation for 3D pose, which expresses the 2D pose in image plane and depth information of each 3D human instance via 2D center point (center of visible keypoints) and root point (denoted as pelvis), respectively. Second, to learn better feature representation for the human depth regression, DRM introduces a 2D Pose-guided Depth Query Module (PDQM) to extract the features in 2D pose regression branch, enabling the depth regression branch to perceive the scale information of instances. Third, DRM leverages a Decoupled Absolute Pose Loss (DAPL) to facilitate the absolute root depth and root-relative depth estimation, thus improving the accuracy of absolute 3D pose. Comprehensive experiments on challenging benchmarks including MuPoTS-3D and Panoptic clearly verify the superiority of our framework, which outperforms the state-of-the-art bottom-up absolute 3D pose estimation methods. Lei Jin 0003, Yabo Xiao, Yandong Guo, Xuecheng Nie, Jian Zhao 0006 |
CVPR | 7 |
| 2022 | Point-to-Box Network for Accurate Object Detection via Single Point Supervision
Pengfei Chen 0004, Xuehui Yu, Xumeng Han, Najmul Hassan, Kai Wang 0058, Jiachen Li 0003, Jian Zhao 0006, Humphrey Shi, Zhenjun Han, Qixiang Ye |
ECCV (9) | 7 |
| 2022 | Semantic Compression Embedding for Generative Zero-Shot LearningabstractGenerative methods have been successfully applied in zero-shot learning (ZSL) by learning an implicit mapping to alleviate the visual-semantic domain gaps and synthesizing unseen samples to handle the data imbalance between seen and unseen classes. However, existing generative methods simply use visual features extracted by the pre-trained CNN backbone. These visual features lack attribute-level semantic information. Consequently, seen classes are indistinguishable, and the knowledge transfer from seen to unseen classes is limited. To tackle this issue, we propose a novel Semantic Compression Embedding Guided Generation (SC-EGG) model, which cascades a semantic compression embedding network (SCEN) and an embedding guided generative network (EGGN). The SCEN extracts a group of attribute-level local features for each sample and further compresses them into the new low-dimension visual feature. Thus, a dense-semantic visual space is obtained. The EGGN learns a mapping from the class-level semantic space to the dense-semantic visual space, thus improving the discriminability of the synthesized dense-semantic unseen visual features. Extensive experiments on three benchmark datasets, i.e., CUB, SUN and AWA2, demonstrate the significant performance gains of SC-EGG over current state-of-the-art methods and its baselines. Ziming Hong, Shiming Chen 0002, Guosen Xie, Wenhan Yang, Jian Zhao 0006, Yuanjie Shao, Qinmu Peng, Xinge You |
IJCAI | 5 |
| 2022 | Waveform level adversarial example generation for joint attacks against both automatic speaker verification and spoofing countermeasures
Xiongwei Zhang, Wei Liu 0005, Xia Zou, Meng Sun 0001, Jian Zhao 0006 |
Eng. Appl. Artif. Intell. | 6 |
| 2022 | Face.evoLVe: A cross-platform library for high-performance face analytics
Qingzhong Wang, Pengfei Zhang 0016, Haoyi Xiong, Jian Zhao 0006 |
Neurocomputing | 4 |
| 2022 | Towards Age-Invariant Face RecognitionabstractDespite the remarkable progress in face recognition related technologies, reliably recognizing faces across ages remains a big challenge. The appearance of a human face changes substantially over time, resulting in significant intra-class variations. As opposed to current techniques for age-invariant face recognition, which either directly extract age-invariant features for recognition, or first synthesize a face that matches target age before feature extraction, we argue that it is more desirable to perform both tasks jointly so that they can leverage each other. To this end, we propose a deep Age-Invariant Model (AIM) for face recognition in the wild with three distinct novelties. First, AIM presents a novel unified deep architecture jointly performing cross-age face synthesis and recognition in a mutual boosting way. Second, AIM achieves continuous face rejuvenation/aging with remarkable photorealistic and identity-preserving properties, avoiding the requirement of paired data and the true age of testing samples. Third, effective and novel training strategies are developed for end-to-end learning of the whole deep architecture, which generates powerful age-invariant face representations explicitly disentangled from the age variation. Moreover, we construct a new large-scale Cross-Age Face Recognition (CAFR) benchmark dataset to facilitate existing efforts and push the frontiers of age-invariant face recognition research. Extensive experiments on both our CAFR dataset and several other cross-age datasets (MORPH, CACD, and FG-NET) demonstrate the superiority of the proposed AIM model over the state-of-the-arts. Benchmarking our model on the popular unconstrained face recognition datasets YTF and IJB-C additionally verifies its promising generalization ability in recognizing faces in the wild. Jian Zhao 0006, Shuicheng Yan, Jiashi Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Dense Attentive Feature Enhancement for Salient Object DetectionabstractAttention mechanisms have been proven highly effective for salient object detection. Most previous works utilize attention as a self-gated module to reweigh the feature maps at different levels independently. However, they are limited to certain-level guidance and could not satisfy the need of both accurately detecting intact objects and maintaining their detailed boundaries. In this paper, we build dense attention upon features from multiple levels simultaneously and propose a novel Dense Attentive Feature Enhancement (DAFE) module for efficient feature enhancement in saliency detection. DAFE stacks several attentional units and densely connects attentive feature output from current unit to its all subsequent units. This allows feature maps at deep units to absorb attentive information from shallow units, thus more discriminative information can be efficiently selected at the final output. Note that DAFE is plug and play, which can be effortlessly inserted into any saliency or video saliency models for their performance improvements. We further instantiate a highly effective Dense Attentive Feature Enhancement Network (DAFE-Net) for accurate salient object detection. DAFE-Net constructs DAFE over the aggregation feature that contains both semantics and saliency details, the entire salient objects and their boundaries can be well retained through dense attentions. Extensive experiments demonstrate that the proposed DAFE module is highly effective, and the DAFE-Net performs favorably compared with state-of-the-art approaches. Zun Li 0001, Congyan Lang, Liqian Liang, Jian Zhao 0006, Songhe Feng, Qibin Hou, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Joint Face Image Restoration and Frontalization for RecognitionabstractIn real-world scenarios, many factors may harm face recognition performance,e.g., large pose, bad illumination, low resolution, blur and noise. To address these challenges, previous efforts usually first restore the low-quality faces to high-quality ones and then perform face recognition. However, most of these methods are stage-wise, which is sub-optimal and deviates from the reality. In this paper, we address all these challenges jointly for unconstrained face recognition. We propose anMulti-DegradationFaceRestoration (MDFR) model to restore frontalized high-quality faces from the given low-quality ones under arbitrary facial poses, with three distinct novelties. First, MDFR is a well-designed encoder-decoder architecture which extracts feature representation from an input face image with arbitrary low-quality factors and restores it to a high-quality counterpart. Second, MDFR introduces a pose residual learning strategy along with a 3D-basedPoseNormalizationModule (PNM), which can perceive the pose gap between the input initial pose and its real-frontal pose to guide the face frontalization. Finally, MDFR can generate frontalized high-quality face images by a single unified network, showing a strong capability of preserving face identity. Qualitative and quantitative experiments on both controlled and in-the-wild benchmarks demonstrate the superiority of MDFR over state-of-the-art methods on both face frontalization and face restoration. Xiaoguang Tu, Jian Zhao 0006, Wenjie Ai, Guodong Guo, Zhifeng Li 0001, Wei Liu 0005, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Image-to-Video Generation via 3D Facial DynamicsabstractWe present a versatile model, FaceAnime, for various video generation tasks from still images. Video generation from a single face image is an interesting problem and usually tackled by utilizing Generative Adversarial Networks (GANs) to integrate information from the input face image and a sequence of sparse facial landmarks. However, the generated face images usually suffer from quality loss, image distortion, identity change, and expression mismatching due to the weak representation capacity of the facial landmarks. In this paper, we propose to “imagine” a face video from a single face image according to the reconstructed 3D face dynamics, aiming to generate a realistic and identity-preserving face video, with precisely predicted pose and facial expression. The 3D dynamics reveal changes of the facial expression and motion, and can serve as a strong prior knowledge for guiding highly realistic face video generation. In particular, we explore face video prediction and exploit a well-designed 3D dynamic prediction network to predict a 3D dynamic sequence for a single face image. The 3D dynamics are then further rendered by the sparse texture mapping algorithm to recover structural details and sparse textures for generating face frames. Our model is versatile for various AR/VR and entertainment applications, such as face video retargeting and face video prediction. Superior experimental results have well demonstrated its effectiveness in generating high-fidelity, identity-preserving, and visually pleasant face video clips from a single source face image. Xiaoguang Tu, Yingtian Zou, Jian Zhao 0006, Wenjie Ai, Jian Dong 0011, Yuan Yao 0011, Zhikang Wang, Guodong Guo, Zhifeng Li 0001, Wei Liu 0005, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Robust Video-Based Person Re-Identification by Hierarchical MiningabstractVideo-based person re-identification (Re-ID) aims at retrieving the person through the video sequences across non-overlapping cameras. Some characteristics of pedestrians are not consecutive across frames due to the variations of viewpoints, postures, and occlusions over time. However, existing methods ignore such data peculiarity and the networks tend to only learn those salient consecutive characteristics among frames in video sequences. As a result, the learned representations fail to cover all the characteristics of pedestrians, thus lacking integrity and discrimination. To tackle this problem, we present a novel deep architecture termed Hierarchical Mining Network (HMN), which mines as many pedestrians’ characteristics by referring to the temporal and intra-class knowledge. It consists of a novel Attentive Temporal Module (ATM) and a Dynamic Supervising Branch (DSB), with a Balancing Triplet Loss (BTL) assisting the training. The proposed ATM, with pedestrian perceiving capacity, is capable of evaluating each activation of features through temporal analysis, so that the temporally scattered characteristics of pedestrians can be better aggregated and the contaminated ones can be eliminated. Then, the DSB along with the BTL further enhances the integrity of representations by multiple supervision. Specifically, the DSB perceives the diversities of intra-class samples in each mini-batch and generates targeted supervising signals for them, in which process the BTL guarantees the signals with smaller intra-class variations and larger inter-class variations. Comprehensive experiments on two video-based datasets, i.e., MARS, and DukeMTMC-VideoReID, demonstrate the contribution of each component and the superiority of the proposed HMN over the state-of-the-arts. Benchmarking our model on three popular image-based datasets, i.e., Market1501, DukeMTMC-Reid, and MSMT17 additionally verifies the promising generalizability of the proposed DSB and BTL. Zhikang Wang, Lihuo He, Xiaoguang Tu, Jian Zhao 0006, Xinbo Gao 0001, Shengmei Shen, Jiashi Feng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Diverse Complementary Part Mining for Weakly Supervised Object LocalizationabstractWeakly Supervised Object Localization (WSOL) aims to localize objects with only image-level labels, which has better scalability and practicability than fully supervised methods in the actual deployment. However, a common limitation for available techniques based on classification networks is that they only highlight the most discriminative part of the object, not the entire object. To alleviate this problem, we propose a novel end-to-end part discovery model (PDM) to learn multiple discriminative object parts in a unified network for accurate object localization and classification. The proposed PDM enjoys several merits. First, to the best of our knowledge, it is the first work to directly model diverse and robust object parts by exploiting part diversity, compactness, and importance jointly for WSOL. Second, three effective mechanisms including diversity, compactness, and importance learning mechanisms are designed to learn robust object parts. Therefore, our model can exploit complementary spatial information and local details from the learned object parts, which help to produce precise bounding boxes and discriminate different object categories. Extensive experiments on two standard benchmarks demonstrate that our PDM performs favorably against state-of-the-art WSOL approaches. Tianzhu Zhang 0001, Wenfei Yang, Jian Zhao 0006, Yongdong Zhang 0001, Feng Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2022 | GrOD: Deep Learning with Gradients Orthogonal Decomposition for Knowledge Transfer, Distillation, and Adversarial TrainingabstractRegularization that incorporates the linear combination of empirical loss and explicit regularization terms as the loss function has been frequently used for many machine learning tasks. The explicit regularization term is designed in different types, depending on its applications. While regularized learning often boost the performance with higher accuracy and faster convergence, the regularization would sometimes hurt the empirical loss minimization and lead to poor performance. To deal with such issues in this work, we propose a novel strategy, namely Gr adients O rthogonal D ecomposition ( GrOD ), that improves the training procedure of regularized deep learning. Instead of linearly combining gradients of the two terms, GrOD re-estimates a new direction for iteration that does not hurt the empirical loss minimization while preserving the regularization affects, through orthogonal decomposition. We have performed extensive experiments to use GrOD improving the commonly used algorithms of transfer learning [ 2 ], knowledge distillation [ 3 ], and adversarial learning [ 4 ]. The experiment results based on large datasets, including Caltech 256 [ 5 ], MIT indoor 67 [ 6 ], CIFAR-10 [ 7 ], and ImageNet [ 8 ], show significant improvement made by GrOD for all three algorithms in all cases. Haoyi Xiong, Ruosi Wan, Jian Zhao 0006, Xingjian Li 0002, Zhanxing Zhu, Jun Huan |
ACM Trans. Knowl. Discov. Data | 3 |
| 2022 | Seeing Crucial Parts: Vehicle Model Verification via a Discriminative Representation ModelabstractWidely used surveillance cameras have promoted large amounts of street scene data, which contains one important but long-neglected object: the vehicle. Here we focus on the challenging problem of vehicle model verification. Most previous works usually employ global features (e.g., fully connected features) to further perform vehicle-level deep metric learning (e.g., triplet-based network). However, we argue that it is noteworthy to investigate the distinctiveness of local features and consider vehicle-part-level metric learning by reducing the intra-class variance as much as possible. In this article, we introduce a simple yet powerful deep model—the enforced intra-class alignment network (EIA-Net)—which can learn a more discriminative image representation by localizing key vehicle parts and jointly incorporating two distance metrics: vehicle-level embedding and vehicle-part-sensitive embedding. For learning features, we propose an effective feature extraction module that is composed of two components: the regional proposal network (RPN)-based network and part-based CNN. The RPN is used to define key vehicle regions and aggregate local features on these regions, whereas part-based CNN offers supplementary global features for the RPN-based network. The fusion features learned by feature extraction module are cast into the deep metric learning module. Especially, we derived an enforced intra-class alignment loss by re-utilizing key vehicle part information to enhance reducing intra-class variance. Furthermore, we modify the coupled cluster loss to model the vehicle-level embedding by enlarging the inter-class variance while shortening intra-class variance. Extensive experiments over benchmark datasets VehicleID and CompCars have shown that the proposed EIA-Net significantly outperforms the state-of-the-art approaches for vehicle model verification. Furthermore, we also conduct comprehensive experiments on vehicle re-identification datasets (i.e., VehicleID and VeRi776) to validate the generalization ability effectiveness of our proposed method. Liqian Liang, Congyan Lang, Zun Li 0001, Jian Zhao 0006, Tao Wang 0011, Songhe Feng |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Multi-caption Text-to-Face Synthesis: Dataset and AlgorithmabstractText-to-Face synthesis with multiple captions is still an important yet less addressed problem because of the lack of effective algorithms and large-scale datasets. We accordingly propose a Semantic Embedding and Attention (SEA-T2F) network that allows multiple captions as input to generate highly semantically related face images. With a novel Sentence Features Injection Module, SEA-T2F can integrate any number of captions into the network. In addition, an attention mechanism named Attention for Multiple Captions is proposed to fuse multiple word features and synthesize fine-grained details. Considering text-to-face generation is an ill-posed problem, we also introduce an attribute loss to guide the network to generate sentence-related attributes. Existing datasets for text-to-face are either too small or roughly generated according to attribute labels, which is not enough to train deep learning based methods to synthesize natural face images. Therefore, we build a large-scale dataset named CelebAText-HQ, in which each image is manually annotated with 10 captions. Extensive experiments demonstrate the effectiveness of our algorithm. Jianxin Sun 0003, Qi Li 0005, Weining Wang 0001, Jian Zhao 0006, Zhenan Sun |
ACM Multimedia | 4 |
| 2021 | Effective Fusion Factor in FPN for Tiny Object DetectionabstractFPN-based detectors have made significant progress in general object detection, e.g., MS COCO and PASCAL VOC. However, these detectors fail in certain application scenarios, e.g., tiny object detection. In this paper, we argue that the top-down connections between adjacent layers in FPN bring two-side influences for tiny object detection, not only positive. We propose a novel concept, fusion factor, to control information that deep layers deliver to shallow layers, for adapting FPN to tiny object detection. After series of experiments and analysis, we explore how to estimate an effective value of fusion factor for a particular dataset by a statistical method. The estimation is dependent on the number of objects distributed in each layer. Comprehensive experiments are conducted on tiny object detection datasets, e.g., TinyPerson and Tiny CityPersons. Our results show that when configuring FPN with a proper fusion factor, the network is able to achieve significant performance gains over the baseline on tiny object detection datasets. Codes and models will be released. Yuqi Gong, Xuehui Yu, Yao Ding 0006, Xiaoke Peng, Jian Zhao 0006, Zhenjun Han |
WACV | 5 |
| 2021 | Fine-Grained Facial Expression Recognition in the WildabstractOver the past decades, researches on facial expression recognition have been restricted within six basic expressions (anger, fear, disgust, happiness, sadness and surprise). However, these six words can not fully describe the richness and diversity of human beings' emotions. To enhance the recognitive capabilities for computers, in this paper, we focus on fine-grained facial expression recognition in the wild and build a brand new benchmark FG-Emotions to push the research frontiers on this topic, which extends the original six classes to more elaborate thirty-three classes. Our FG-Emotions contains 10,371 images and 1,491 video clips annotated with corresponding fine-grained facial expression categories and landmarks. FG-Emotions also provides several features (e.g., LBP features and dense trajectories features) to facilitate related research. Moreover, on top of FG-Emotions, we propose a new end-to-end Multi-Scale Action Unit (AU)-based Network (MSAU-Net) for facial expression recognition with image which learns a more powerful facial representation by directly focusing on locating facial action units and utilizing “zoom in” operation to aggregate distinctive local features. As for recognition with video, we further extend the MSAU-Net to a two-stream model (TMSAU-Net) by adding a module with attention mechanism and a temporal stream branch to jointly learn spatial and temporal features. (T)MSAU-Net consistently outperforms existing state-of-the-art solutions on our FG-Emotions and several other datasets, and serves as a strong baseline to drive the future research towards fine-grained facial expression recognition in the wild. Liqian Liang, Congyan Lang, Yidong Li, Songhe Feng, Jian Zhao 0006 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2021 | 3D Face Reconstruction From A Single Image Assisted by 2D Face Images in the Wildabstract3D face reconstruction from a single image is an important task in many multimedia applications. Recent works typically learn a CNN-based 3D face model that regresses coefficients of a 3D Morphable Model (3DMM) from 2D images to perform 3D face reconstruction. However, the shortage of training data with 3D annotations considerably limits performance of these methods. To alleviate this issue, we propose a novel 2D-Assisted Learning (2DAL) method that can effectively use “in the wild” 2D face images with noisy landmark information to substantially improve 3D face model learning. Specifically, taking the sparse 2D facial landmark heatmaps as additional information, 2DAL introduces four novel self-supervision schemes that view the 2D landmark and 3D landmark prediction as a self-mapping process, including the landmark self-prediction consistency for 2D and 3D faces respectively, cycle-consistency over the 2D landmark prediction and self-critic over the predicted 3DMM coefficients based on landmark prediction. Using these four self-supervision schemes, 2DAL significantly relieves the demands for the the conventional paired 2D-to-3D annotations and gives much higher-quality 3D face models without requiring any additional 3D annotations. Experiments on AFLW2000-3D, AFLW-LFPA and Florence benchmarks show that our method outperforms state-of-the-arts for both 3D face reconstruction and dense face alignment by a large margin. Xiaoguang Tu, Jian Zhao 0006, Mei Xie, Zihang Jiang, Akshaya Balamurugan, Yao Luo, Yang Zhao 0003, Lingxiao He, Zheng Ma 0005, Jiashi Feng |
IEEE Trans. Multim. | 2 |
| 2021 | Multi-human Parsing with a Graph-based Generative Adversarial ModelabstractHuman parsing is an important task in human-centric image understanding in computer vision and multimedia systems. However, most existing works on human parsing mainly tackle the single-person scenario, which deviates from real-world applications where multiple persons are present simultaneously with interaction and occlusion. To address such a challenging multi-human parsing problem, we introduce a novel multi-human parsing model named MH-Parser, which uses a graph-based generative adversarial model to address the challenges of close-person interaction and occlusion in multi-human parsing. To validate the effectiveness of the new model, we collect a new dataset named Multi-Human Parsing (MHP), which contains multiple persons with intensive person interaction and entanglement. Experiments on the new MHP dataset and existing datasets demonstrate that the proposed method is effective in addressing the multi-human parsing problem compared with existing solutions in the literature. Jianshu Li, Jian Zhao 0006, Congyan Lang, Yidong Li, Yunchao Wei, Guodong Guo, Terence Sim, Shuicheng Yan, Jiashi Feng |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Unsupervised Domain Adaptation with Noise Resistible Mutual-Training for Person Re-identification
Fang Zhao 0006, Shengcai Liao, Guosen Xie, Jian Zhao 0006, Kaihao Zhang, Ling Shao 0001 |
ECCV (11) | 4 |
| 2020 | Learning to Detect Head Movement in Unconstrained Remote Gaze Estimation in the WildabstractUnconstrained remote gaze estimation remains challenging mostly due to its vulnerability to the large variability in head-pose. Prior solutions struggle to maintain reliable accuracy in unconstrained remote gaze tracking. Among them, appearance-based solutions demonstrate tremendous potential in improving gaze accuracy. However, existing works still suffer from head movement and are not robust enough to handle real-world scenarios. Especially most of them study gaze estimation under controlled scenarios where the collected datasets often cover limited ranges of both head-pose and gaze which introduces further bias. In this paper, we propose novel end-to-end appearance-based gaze estimation methods that could more robustly incorporate different levels of head-pose representations into gaze estimation. Our method could generalize to real-world scenarios with low image quality, different lightings and scenarios where direct head-pose information is not available. To better demonstrate the advantage of our methods, we further propose a new benchmark dataset with the most rich distribution of head-gaze combination reflecting real-world scenarios. Extensive evaluations on several public datasets and our own dataset demonstrate that our method consistently outperforms the state-of-the-art by a significant margin. Zhecan Wang, Jian Zhao 0006, Cheng Lu 0006, Han Huang 0005, Fan Yang 0035, Lianji Li, Yandong Guo |
WACV | 2 |
| 2020 | Fine-Grained Multi-human Parsing
Jian Zhao 0006, Jianshu Li, Hengzhu Liu, Shuicheng Yan, Jiashi Feng |
Int. J. Comput. Vis. | 1 |
| 2020 | Recognizing Profile Faces by Imagining Frontal View
Jian Zhao 0006, Junliang Xing, Shuicheng Yan, Jiashi Feng |
Int. J. Comput. Vis. | 1 |
| 2020 | Learning Generalizable and Identity-Discriminative Representations for Face Anti-SpoofingabstractFace anti-spoofing aims to detect presentation attack to face recognition--based authentication systems. It has drawn growing attention due to the high security demand. The widely adopted CNN-based methods usually well recognize the spoofing faces when training and testing spoofing samples display similar patterns, but their performance would drop drastically on testing spoofing faces of novel patterns or unseen scenes, leading to poor generalization performance. Furthermore, almost all current methods treat face anti-spoofing as a prior step to face recognition, which prolongs the response time and makes face authentication inefficient. In this article, we try to boost the generalizability and applicability of face anti-spoofing methods by designing a new generalizable face authentication CNN (GFA-CNN) model with three novelties. First, GFA-CNN introduces a simple yet effective total pairwise confusion loss for CNN training that properly balances contributions of all spoofing patterns for recognizing the spoofing faces. Second, it incorporate a fast domain adaptation component to alleviate negative effects brought by domain variation. Third, it deploys filter diversification learning to make the learned representations more adaptable to new scenes. In addition, the proposed GFA-CNN works in a multi-task manner—it performs face anti-spoofing and face recognition simultaneously. Experimental results on five popular face anti-spoofing and face recognition benchmarks show that GFA-CNN outperforms previous face anti-spoofing methods on cross-test protocols significantly and also well preserves the identity information of input face images. Xiaoguang Tu, Zheng Ma 0005, Jian Zhao 0006, Guodong Du 0004, Mei Xie, Jiashi Feng |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2019 | Look across Elapse: Disentangled Representation Learning and Photorealistic Cross-Age Face Synthesis for Age-Invariant Face RecognitionabstractDespite the remarkable progress in face recognition related technologies, reliably recognizing faces across ages still remains a big challenge. The appearance of a human face changes substantially over time, resulting in significant intraclass variations. As opposed to current techniques for ageinvariant face recognition, which either directly extract ageinvariant features for recognition, or first synthesize a face that matches target age before feature extraction, we argue that it is more desirable to perform both tasks jointly so that they can leverage each other. To this end, we propose a deep Age-Invariant Model (AIM) for face recognition in the wild with three distinct novelties. First, AIM presents a novel unified deep architecture jointly performing cross-age face synthesis and recognition in a mutual boosting way. Second, AIM achieves continuous face rejuvenation/aging with remarkable photorealistic and identity-preserving properties, avoiding the requirement of paired data and the true age of testing samples. Third, we develop effective and novel training strategies for end-to-end learning the whole deep architecture, which generates powerful age-invariant face representations explicitly disentangled from the age variation. Extensive experiments on several cross-age datasets (MORPH, CACD and FG-NET) demonstrate the superiority of the proposed AIM model over the state-of-the-arts. Benchmarking our model on one of the most popular unconstrained face recognition datasets IJB-C additionally verifies the promising generalizability of AIM in recognizing faces in the wild. Jian Zhao 0006, Yu Cheng 0009, Yang Yang 0002, Fang Zhao 0006, Jianshu Li, Hengzhu Liu, Shuicheng Yan, Jiashi Feng |
AAAI | 1 |
| 2019 | Multi-Prototype Networks for Unconstrained Set-based Face RecognitionabstractIn this paper, we address the challenging unconstrained set-based face recognition problem where each subject face is instantiated by a set of media (images and videos) instead of a single image. Naively aggregating information from all the media within a set would suffer from the large intra-set variance caused by heterogeneous factors (e.g., varying media modalities, poses and illumination) and fail to learn discriminative face representations. A novel Multi-Prototype Network (MP- Net) model is thus proposed to learn multiple prototype face representations adaptively from the media sets. Each learned prototype is representative for the subject face under certain condition in terms of pose, illumination and media modality. Instead of handcrafting the set partition for prototype learn- ing, MPNet introduces a Dense SubGraph (DSG) learning sub-net that implicitly untangles inconsistent media and learns a number of representative prototypes. Qualitative and quantitative experiments clearly demonstrate the superiority of the proposed model over state-of-the-arts. Jian Zhao 0006, Jianshu Li, Xiaoguang Tu, Fang Zhao 0006, Yuan Xin, Junliang Xing, Hengzhu Liu, Shuicheng Yan, Jiashi Feng |
IJCAI | 1 |
| 2019 | Task Relation NetworksabstractMulti-task learning is popular in machine learning and computer vision. In multitask learning, properly modeling task relations is important for boosting the performance of jointly learned tasks. Task covariance modeling has been successfully used to model the relations of tasks but is limited to homogeneous multi-task learning. In this paper, we propose a feature based task relation modeling approach, suitable for both homogeneous and heterogeneous multi-task learning. First, we propose a new metric to quantify the relations between tasks. Based on the quantitative metric, we then develop the task relation layer, which can be combined with any deep learning architecture to form task relation networks to fully exploit the relations of different tasks in an online fashion. Benefiting from the task relation layer, the task relation networks can better leverage the mutual information from the data. We demonstrate our proposed task relation networks are effective in improving the performance in both homogeneous and heterogeneous multi-task learning settings through extensive experiments on computer vision tasks. Jianshu Li, Pan Zhou 0002, Yunpeng Chen, Jian Zhao 0006, Sujoy Roy, Shuicheng Yan, Jiashi Feng, Terence Sim |
WACV | 4 |
| 2019 | 3D-Aided Dual-Agent GANs for Unconstrained Face RecognitionabstractSynthesizing realistic profile faces is beneficial for more efficiently training deep pose-invariant models for large-scale unconstrained face recognition, by augmenting the number of samples with extreme poses and avoiding costly annotation work. However, learning from synthetic faces may not achieve the desired performance due to the discrepancy betwedistributions of the synthetic and real face images. To narrow this gap, we propose a Dual-Agent Generative Adversarial Network (DA-GAN) model, which can improve the realism of a face simulator's output using unlabeled real faces while preserving the identity information during the realism refinement. The dual agents are specially designed for distinguishing real versus fake and identities simultaneously. In particular, we employ an off-the-shelf 3D face model as a simulator to generate profile face images with varying poses. DA-GAN leverages a fully convolutional network as the generator to generate high-resolution images and an auto-encoder as the discriminator with the dual agents. Besides the novel architecture, we make several key modifications to the standard GAN to preserve pose, texture as well as identity, and stabilize the training process: (i) a pose perception loss; (ii) an identity perception loss; (iii) an adversarial loss with a boundary equilibrium regularization term. Experimental results show that DA-GAN not only achieves outstanding perceptual results but also significantly outperforms state-of-the-arts on the large-scale and challenging NIST IJB-A and CFP unconstrained face recognition benchmarks. In addition, the proposed DA-GAN is also a promising new approach for solving generic transfer learning problems more effectively. DA-GAN is the foundation of our winning entry to the NIST IJB-A face recognition competition in which we secured the $1^{st}$ places on the tracks of verification and identification. Jian Zhao 0006, Jianshu Li, Junliang Xing, Shuicheng Yan, Jiashi Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Weakly Supervised Phrase Localization With Multi-Scale Anchored Transformer NetworkabstractIn this paper, we propose a novel weakly supervised model, Multi-scale Anchored Transformer Network (MATN), to accurately localize free-form textual phrases with only image-level supervision. The proposed MATN takes region proposals as localization anchors, and learns a multiscale correspondence network to continuously search for phrase regions referring to the anchors. In this way, MATN can exploit useful cues from these anchors to reliably reason about locations of the regions described by the phrases given only image-level supervision. Through differentiable sampling on image spatial feature maps, MATN introduces a novel training objective to simultaneously minimize a contrastive reconstruction loss between different phrases from a single image and a set of triplet losses among multiple images with similar phrases. Superior to existing region proposal based methods, MATN searches for the optimal bounding box over the entire feature map instead of selecting a sub-optimal one from discrete region proposals. We evaluate MATN on the Flickr30K Entities and ReferItGame datasets. The experimental results show that MATN significantly outperforms the state-of-the-art methods. Fang Zhao 0006, Jianshu Li, Jian Zhao 0006, Jiashi Feng |
CVPR | 3 |
| 2018 | Towards Pose Invariant Face Recognition in the WildabstractPose variation is one key challenge in face recognition. As opposed to current techniques for pose invariant face recognition, which either directly extract pose invariant features for recognition, or first normalize profile face images to frontal pose before feature extraction, we argue that it is more desirable to perform both tasks jointly to allow them to benefit from each other. To this end, we propose a Pose Invariant Model (PIM) for face recognition in the wild, with three distinct novelties. First, PIM is a novel and unified deep architecture, containing a Face Frontalization sub-Net (FFN) and a Discriminative Learning sub-Net (DLN), which are jointly learned from end to end. Second, FFN is a well-designed dual-path Generative Adversarial Network (GAN) which simultaneously perceives global structures and local details, incorporated with an unsupervised cross-domain adversarial training and a "learning to learn" strategy for high-fidelity and identity-preserving frontal view synthesis. Third, DLN is a generic Convolutional Neural Network (CNN) for face recognition with our enforced cross-entropy optimization strategy for learning discriminative yet generalized feature representation. Qualitative and quantitative experiments on both controlled and in-the-wild benchmarks demonstrate the superiority of the proposed model over the state-of-the-arts. Jian Zhao 0006, Yu Cheng 0009, Yan Xu 0009, Jianshu Li, Fang Zhao 0006, Jayashree Karlekar, Sugiri Pranata, Shengmei Shen, Junliang Xing, Shuicheng Yan, Jiashi Feng |
CVPR | 1 |
| 2018 | Dynamic Conditional Networks for Few-Shot Learning
Fang Zhao 0006, Jian Zhao 0006, Shuicheng Yan, Jiashi Feng |
ECCV (15) | 2 |
| 2018 | 3D-Aided Deep Pose-Invariant Face RecognitionabstractLearning from synthetic faces, though perhaps appealing for high data efficiency, may not bring satisfactory performance due to the distribution discrepancy of the synthetic and real face images. To mitigate this gap, we propose a 3D-Aided Deep Pose-Invariant Face Recognition Model (3D-PIM), which automatically recovers realistic frontal faces from arbitrary poses through a 3D face model in a novel way. Specifically, 3D-PIM incorporates a simulator with the aid of a 3D Morphable Model (3D MM) to obtain shape and appearance prior for accelerating face normalization learning, requiring less training data. It further leverages a global-local Generative Adversarial Network (GAN) with multiple critical improvements as a refiner to enhance the realism of both global structures and local details of the face simulator’s output using unlabelled real data only, while preserving the identity information. Qualitative and quantitative experiments on both controlled and in-the-wild benchmarks clearly demonstrate superiority of the proposed model over state-of-the-arts. Jian Zhao 0006, Yu Cheng 0009, Jianshu Li, Yan Xu 0009, Jayashree Karlekar, Sugiri Pranata, Shengmei Shen, Junliang Xing, Shuicheng Yan, Jiashi Feng |
IJCAI | 1 |
| 2018 | Multi-Human Parsing MachinesabstractHuman parsing is an important task in human-centric analysis. Despite the remarkable progress in single-human parsing, the more realistic case of multi-human parsing remains challenging in terms of the data and the model. Compared with the considerable number of available single-human parsing datasets, the datasets for multi-human parsing are very limited in number mainly due to the huge annotation effort required. Besides the data challenge to multi-human parsing, the persons in real-world scenarios are often entangled with each other due to close interaction and body occlusion, making it difficult to distinguish body parts from different person instances. In this paper we propose the Multi-Human Parsing Machines (MHPM) system, which contains an MHP Montage model and an MHP Solver, to address both challenges in multi-human parsing. Specifically, the MHP Montage model in MHPM generates realistic images with multiple persons together with the parsing labels. It intelligently composes single persons onto background scene images while maintaining the structural information between persons and the scene. The generated images can be used to train better multi-human parsing algorithms. On the other hand, the MHP Solver in MHPM solves the bottleneck of distinguishing multiple entangled persons with close interaction. It employs a Group-Individual Push and Pull (GIPP) loss function, which can effectively separate persons with close interaction. We experimentally show that the proposed MHPM can achieve state-of-the-art performance on the multi-human parsing benchmark and the person individualization benchmark, which distinguishes closely entangled person instances. Jianshu Li, Jian Zhao 0006, Yunpeng Chen, Sujoy Roy, Shuicheng Yan, Jiashi Feng, Terence Sim |
ACM Multimedia | 2 |
| 2018 | Understanding Humans in Crowded Scenes: Deep Nested Adversarial Learning and A New Benchmark for Multi-Human ParsingabstractDespite the noticeable progress in perceptual tasks like detection, instance segmentation and human parsing, computers still perform unsatisfactorily on visually understanding humans in crowded scenes, such as group behavior analysis, person re-identification and autonomous driving, etc. To this end, models need to comprehensively perceive the semantic information and the differences between instances in a multi-human image, which is recently defined as the multi-human parsing task. In this paper, we present a new large-scale database "Multi-Human Parsing (MHP)" for algorithm development and evaluation, and advances the state-of-the-art in understanding humans in crowded scenes. MHP contains 25,403 elaborately annotated images with 58 fine-grained semantic category labels, involving 2-26 persons per image and captured in real-world scenes from various viewpoints, poses, occlusion, interactions and background. We further propose a novel deep Nested Adversarial Network (NAN) model for multi-human parsing. NAN consists of three Generative Adversarial Network (GAN)-like sub-nets, respectively performing semantic saliency prediction, instance-agnostic parsing and instance-aware clustering. These sub-nets form a nested structure and are carefully designed to learn jointly in an end-to-end way. NAN consistently outperforms existing state-of-the-art solutions on our MHP and several other datasets, and serves as a strong baseline to drive the future research for multi-human parsing. Jian Zhao 0006, Jianshu Li, Yu Cheng 0009, Terence Sim, Shuicheng Yan, Jiashi Feng |
ACM Multimedia | 1 |
| 2018 | Robust LSTM-Autoencoders for Face De-Occlusion in the WildabstractFace recognition techniques have been developed significantly in recent years. However, recognizing faces with partial occlusion is still challenging for existing face recognizers, which is heavily desired in real-world applications concerning surveillance and security. Although much research effort has been devoted to developing face de-occlusion methods, most of them can only work well under constrained conditions, such as all of faces are from a pre-defined closed set of subjects. In this paper, we propose a robust LSTM-Autoencoders (RLA) model to effectively restore partially occluded faces even in the wild. The RLA model consists of two LSTM components, which aims at occlusion-robust face encoding and recurrent occlusion removal respectively. The first one, named multi-scale spatial LSTM encoder, reads facial patches of various scales sequentially to output a latent representation, and occlusion-robustness is achieved owing to the fact that the influence of occlusion is only upon some of the patches. Receiving the representation learned by the encoder, the LSTM decoder with a dual channel architecture reconstructs the overall face and detects occlusion simultaneously, and by feat of LSTM, the decoder breaks down the task of face de-occlusion into restoring the occluded part step by step. Moreover, to minimize identify information loss and guarantee face recognition accuracy over recovered faces, we introduce an identity-preserving adversarial training scheme to further improve RLA. Extensive experiments on both synthetic and real data sets of faces with occlusion clearly demonstrate the effectiveness of our proposed RLA in removing different types of facial occlusion at various locations. The proposed method also provides significantly larger performance gain than other de-occlusion methods in promoting recognition performance over partially-occluded faces. Fang Zhao 0006, Jiashi Feng, Jian Zhao 0006, Wenhan Yang, Shuicheng Yan |
IEEE Trans. Image Process. | 3 |
| 2017 | Marginalized CNN: Learning Deep Invariant Representations
Jian Zhao 0006, Jianshu Li, Fang Zhao 0006, Xuecheng Nie, Yunpeng Chen, Shuicheng Yan, Jiashi Feng |
BMVC | 1 |
| 2017 | Integrated Face Analytics Networks through Cross-Dataset Hybrid TrainingabstractFace analytics benefits many multimedia applications. It consists of a number of tasks, such as facial emotion recognition and face parsing, and most existing approaches generally treat these tasks independently, which limits their deployment in real scenarios. In this paper we propose an integrated Face Analytics Network (iFAN), which is able to perform multiple tasks jointly for face analytics with a novel carefully designed network architecture to fully facilitate the informative interaction among different tasks. The proposed integrated network explicitly models the interactions between tasks so that the correlations between tasks can be fully exploited for performance boost. In addition, to solve the bottleneck of the absence of datasets with comprehensive training data for various tasks, we propose a novel cross-dataset hybrid training strategy. It allows "plug-in and play'' of multiple datasets annotated for different tasks without the requirement of a fully labeled common dataset for all the tasks. We experimentally show that the proposed iFAN achieves state-of-the-art performance on multiple face analytics tasks using a single integrated model. Specifically, iFAN achieves an overall F-score of 91.15% on the Helen dataset for face parsing, a normalized mean error of 5.81% on the MTFL dataset for facial landmark localization and an accuracy of 45.73% on the BNU dataset for emotion recognition with a single model. Jianshu Li, Shengtao Xiao, Fang Zhao 0006, Jian Zhao 0006, Jianan Li 0001, Jiashi Feng, Shuicheng Yan, Terence Sim |
ACM Multimedia | 4 |
| 2017 | Dual-Agent GANs for Photorealistic and Identity Preserving Profile Face SynthesisabstractSynthesizing realistic profile faces is promising for more efficiently training deep pose-invariant models for large-scale unconstrained face recognition, by populating samples with extreme poses and avoiding tedious annotations. However, learning from synthetic faces may not achieve the desired performance due to the discrepancy between distributions of the synthetic and real face images. To narrow this gap, we propose a Dual-Agent Generative Adversarial Network (DA-GAN) model, which can improve the realism of a face simulator's output using unlabeled real faces, while preserving the identity information during the realism refinement. The dual agents are specifically designed for distinguishing real v.s. fake and identities simultaneously. In particular, we employ an off-the-shelf 3D face model as a simulator to generate profile face images with varying poses. DA-GAN leverages a fully convolutional network as the generator to generate high-resolution images and an auto-encoder as the discriminator with the dual agents. Besides the novel architecture, we make several key modifications to the standard GAN to preserve pose and texture, preserve identity and stabilize training process: (i) a pose perception loss; (ii) an identity perception loss; (iii) an adversarial loss with a boundary equilibrium regularization term. Experimental results show that DA-GAN not only presents compelling perceptual results but also significantly outperforms state-of-the-arts on the large-scale and challenging NIST IJB-A unconstrained face recognition benchmark. In addition, the proposed DA-GAN is also promising as a new approach for solving generic transfer learning problems more effectively. Jian Zhao 0006, Jayashree Karlekar, Jianshu Li, Fang Zhao 0006, Zhecan Wang, Sugiri Pranata, Shengmei Shen, Shuicheng Yan, Jiashi Feng |
NIPS | 1 |
| 2016 | Global localization in 3D maps for structured environmentabstractThis paper presents a global localization method for mobile robots based on the geometric information of structured indoor environments. With a global/local point cloud and projection map, lines are extracted from the projection maps using Hough transform. According to the directions of the obtained lines, the orientations of projection maps and point clouds are normalized. Next, the template matching algorithm is applied to the normalized global and local projection maps. Once coarse localization is completed, final accurate localization is achieved using the Iterative Closest Points (ICP) algorithm. Experimental results on several point clouds show that the proposed method can achieve high localization accuracy in real-time. The proposed method can be used for other global localization applications in structured environments. Yanxin Ma, Yulan Guo, Min Lu 0001, Jian Zhao 0006, Jun Zhang 0044 |
IGARSS | 4 |
| 2016 | Robust Face Recognition with Deep Multi-View Representation LearningabstractThis paper describes our proposed method targeting at the MSR Image Recognition Challenge MS-Celeb-1M. The challenge is to recognize one million celebrities from their face images captured in the real world. The challenge provides a large scale dataset crawled from the Web, which contains a large number of celebrities with many images for each subject. Given a new testing image, the challenge requires an identify for the image and the corresponding confidence score. To complete the challenge, we propose a two-stage approach consisting of data cleaning and multi-view deep representation learning. The data cleaning can effectively reduce the noise level of training data and thus improves the performance of deep learning based face recognition models. The multi-view representation learning enables the learned face representations to be more specific and discriminative. Thus the difficulties of recognizing faces out of a huge number of subjects are substantially relieved. Our proposed method achieves a coverage of 46.1% at 95% precision on the random set and a coverage of 33.0% at 95% precision on the hard set of this challenge. Jianshu Li, Jian Zhao 0006, Fang Zhao 0006, Hao Liu 0003, Jing Li 0050, Shengmei Shen, Jiashi Feng, Terence Sim |
ACM Multimedia | 2 |
| 2016 | Accelerated Coherent Point Drift for Automatic Three-Dimensional Point Cloud RegistrationabstractFully automatic 3-D point cloud registration is a highly challenging task in light detection and ranging (LiDAR) remote sensing. The coherent point drift (CPD) algorithm provides an appropriate solution for point cloud registration because of its high accuracy. However, real application of the traditional CPD algorithm is limited due to its demanding computational complexity. In this letter, we present a novel accelerated CPD (ACPD) algorithm for fast, accurate, and automatic registration of 3-D point clouds. First, a global squared iterative expectation–maximization (gSQUAREM) technique is integrated to the ACPD algorithm. Then, the dual-tree improved fast Gauss transform method is used to further accelerate the Gaussian summation process during the correspondence probability matrix calculation. Experimental results on two real data sets show that the proposed algorithm can perform fast and accurate registration on LiDAR point clouds. Min Lu 0001, Jian Zhao 0006, Yulan Guo, Yanxin Ma |
IEEE Geosci. Remote. Sens. Lett. | 2 |