Junliang Xing

dblp:43/7659 · DBLP profile ↗
← Back
190ranked-venue papers
10as first author
88since 2021 · last 2026
0000-0001-6801-0510ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 125 · 8 first-author · 45 since 2021Artificial intelligence and machine learning · 123 · 6 first-author · 58 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Computer networks · 4 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
abstract
While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional feature fusion approaches often fail to enable effective cross-modal interactions, thereby leading to suboptimal generation outcomes. To address this challenge, we introduce MDiTFace—a customized diffusion transformer framework that employs a unified tokenization strategy to process semantic mask and text inputs, eliminating discrepancies between heterogeneous modality representations. The framework facilitates comprehensive multimodal feature interaction through stacked, newly designed multivariate transformer blocks that process all conditions synchronously. Additionally, we design a novel decoupled attention mechanism by dissociating implicit dependencies between mask tokens and temporal embeddings. This mechanism segregates internal computations into dynamic and static pathways, enabling caching and reuse of features computed in static pathways after initial calculation, thereby reducing additional computational overhead introduced by mask condition by over 94% while maintaining performance. Extensive experiments demonstrate that MDiTFace significantly outperforms other competing methods in terms of both facial fidelity and conditional consistency.
Yushe Cao, Dian-xi Shi, Xuechao Zou, Haikuo Peng, Chun Yu, Junliang Xing
AAAI8
2026 MARPO: A Reflective Policy Optimization for Multi-Agent Reinforcement Learning
abstract
We propose Multi-Agent Reflective Policy Optimization MARPO to alleviate the issue of sample inefficiency in multi-agent reinforcement learning. MARPO consists of two key components: a reflection mechanism that leverages subsequent trajectories to enhance sample efficiency, and an asymmetric clipping mechanism that is derived from the KL divergence and dynamically adjusts the clipping range to improve training stability. We evaluate MARPO in classic multi-agent environments, where it consistently outperforms other methods.
Cuiling Wu, Yaozhong Gan, Junliang Xing
AAAI3
2026 Deep (Predictive) Discounted Counterfactual Regret Minimization
abstract
Counterfactual regret minimization (CFR) is a family of algorithms for effectively solving imperfect-information games. To enhance CFR's applicability in large games, researchers use neural networks to approximate its behavior. However, existing methods are mainly based on vanilla CFR and struggle to effectively integrate more advanced CFR variants. In this work, we propose an efficient model-free neural CFR algorithm, overcoming the limitations of existing methods in approximating advanced CFR variants. At each iteration, it collects variance-reduced sampled advantages based on a value network, fits cumulative advantages by bootstrapping, and applies discounting and clipping operations to simulate the update mechanisms of advanced CFR variants. Experimental results show that, compared with model-free neural algorithms, it exhibits faster convergence in typical imperfect-information games and demonstrates stronger adversarial performance in a large poker game.
Hang Xu 0006, Kai Li 0022, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
AAAI5
2026 Dual-Pathway Diffusion for Hand Correction in Synthetic Portraits: Global Context Aware and Local Structure Refinement
abstract
Despite significant advancements in diffusion models for generating high-quality portrait images, the problem of malformed hands remains a critical and unresolved challenge. Existing methods primarily focus on structural restoration, often neglecting the semantic coherence between the corrected hand region and the overall image. To address this limitation, we propose a novel dual-pathway diffusion malformed hand correction method, named DD-MHC, which integrates a Global Context Aware Module (GCAM) and a Local Structure Refinement Module (LSRM) through a dual-path architecture. Under the synergy of cross-attention and spatial attention, these modules can effectively fuse global contextual features and local guiding cues, enabling precise restoration of the hand region. Comprehensive experiments demonstrate that DD-MHC significantly outperforms existing competing methods, particularly in enhancing the semantic consistency between the corrected hand region and the overall image. Additionally, to address the challenge of data scarcity in the research on correcting malformed hands within unconstrained scene portraits, we construct a brand-new General-Scene Portrait Dataset (GSPD), providing a standardized and reproducible data platform for subsequent related research.
Yushe Cao, Luoxi Jing, Yuanze Wang, Dian-xi Shi, Chun Yu, Junliang Xing
ICMR6
2026 Hierarchical fusion of local and global visual features with mixture-of-experts for remote sensing image scene classification
Yuanhao Tang, Xuechao Zou, Zhengpei Hu, Junliang Xing, Jianqiang Huang 0002
Neurocomputing4
2026 DrawMotion: Generating 3D Human Motions by Freehand Drawing
Tao Wang 0011, Lei Jin 0003, Qiaozhi He, Jiaming Chu, Yu Cheng 0009, Junliang Xing, Jian Zhao 0006, Shuicheng Yan, Li Wang 0039
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 SynSP++: General Pose Sequences Refinement via Synergy of Smoothness and Precision
Lei Jin 0003, Tao Wang 0011, Junliang Xing, Jian Zhao 0006, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Position-Aware Guided Point Cloud Completion with CLIP Model
abstract
Point cloud completion aims to recover partial geometric and topological shapes caused by equipment defects or limited viewpoints. Current methods either solely rely on the 3D coordinates of the point cloud to complete it or incorporate additional images with well-calibrated intrinsic parameters to guide the geometric estimation of the missing parts. Although these methods have achieved excellent performance by directly predicting the location of complete points, the extracted features lack fine-grained information regarding the location of the missing area. To address this issue, we propose a rapid and efficient method to expand an unimodal framework into a multimodal framework. This approach incorporates a position-aware module designed to enhance the spatial information of the missing parts through a weighted map learning mechanism. In addition, we establish a Point-Text-Image triplet corpus PCI-TI and MVP-TI based on the existing unimodal point cloud completion dataset and use the pre-trained vision-language model CLIP to provide richer detail information for 3D shapes, thereby enhancing performance. Extensive quantitative and qualitative experiments demonstrate that our method outperforms state-of-the-art point cloud completion methods.
Feng Zhou 0007, Ju Dai, Lei Li 0050, Junliang Xing
AAAI6
2025 An Open-Ended Learning Framework for Opponent Modeling
abstract
Opponent Modeling (OM) aims to enhance decision-making by modeling other agents in multi-agent environments. Existing works typically learn opponent models against a pre-designated fixed set of opponents during training. However, this will cause poor generalization when facing unknown opponents during testing, as previously unseen opponents can exhibit out-of-distribution (OOD) behaviors that the learned opponent models cannot handle. To tackle this problem, we introduce a novel Open-Ended Opponent Modeling (OEOM) framework, which continuously generates opponents with diverse strengths and styles to reduce the possibility of OOD situations occurring during testing. Founded on population-based training and information-theoretic trajectory space diversity regularization, OEOM generates a dynamic set of opponents. This set is then fed to any OM approaches to train a potentially generalizable opponent model. Upon this, we further propose a simple yet effective OM approach that naturally fits within the OEOM framework. This approach is based on in-context reinforcement learning and learns a Transformer that dynamically recognizes and responds to opponents based on their trajectories. Extensive experiments in cooperative, competitive, and mixed environments demonstrate that OEOM is an approach-agnostic framework that improves generalizability compared to training against a fixed set of opponents, regardless of OM approaches or testing opponent settings. The results also indicate that our proposed approach generally outperforms existing OM baselines.
Yuheng Jing, Kai Li 0022, Bingyun Liu, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
AAAI6
2025 Multi-View 3D Human Pose Estimation with Weakly Synchronized Images
abstract
Multi-view 3D human pose estimation (MHPE) is an important research task in computer vision. To maintain consistency during the data collection, hardware synchronization devices are commonly used to connect cameras, ensuring that images from different views are captured simultaneously. However, synchronizing with extra devices has two apparent limitations: the hardware is i) usually expensive and ii) less flexible for deployment in outdoor open scenarios. Suppose the model can improve its tolerance for the time differences in multi-view image capture. In that case, the difficulty and cost of deployment will be greatly reduced, and MHPE will become more widespread. In this paper, we try to answer how to build a model that performs pose estimation directly using ''weakly synchronized images" from multiple views, where the captured images shift from each other within a frame. To this end, we introduce a new multi-view 3D human pose estimation task given weakly synchronized image inputs. Apart from existing well-synchronized datasets, we present the first weakly synchronized dataset comprising 800k images. Thereon, we propose SyncDiffPose, a novel model based on the diffusion method for pose estimation to denoise the error in such data. By combining simple synchronization strategies, e.g., the timer method, our approach can perform pose estimation without hardware calibration.
Ruiwen Gu, Junliang Xing, Xinchun Yu, Xiao-Ping Zhang 0002
AAAI4
2025 StickMotion: Generating 3D Human Motions by Drawing a Stickman
abstract
Text-to-motion generation, which translates textual descriptions into human motions, has been challenging in accurately capturing detailed user-imagined motions from simple text inputs. This paper introduces StickMotion, an efficient diffusion-based network designed for multi-condition scenarios, which generates desired motions based on traditional text and our proposed stickman conditions for global and local control of these motions, respectively. We address the challenges introduced by the user-friendly stickman from three perspectives: 1) Data generation. We develop an algorithm to generate hand-drawn stickmen automatically across different dataset formats. 2) Multi-condition fusion. We propose a multi-condition module that integrates into the diffusion process and obtains outputs of all possible condition combinations, reducing computational complexity and enhancing StickMotion’s performance compared to conventional approaches with the self-attention module. 3) Dynamic supervision. We empower StickMotion to make minor adjustments to the stickman’s position within the output sequences, generating more natural movements through our proposed dynamic supervision strategy. Through quantitative experiments and user studies, sketching stickmen saves users about 51.5% of their time generating motions consistent with their imagination. Our codes, demos, and relevant data will be released in https:// github.com/InvertedForest/StickMotion.
Tao Wang 0011, Qiaozhi He, Jiaming Chu, Ling Qian, Yu Cheng 0009, Junliang Xing, Jian Zhao 0006, Lei Jin 0003
CVPR7
2025 UV-Mamba: A DCN-Enhanced State Space Model for Urban Village Boundary Identification in High-Resolution Remote Sensing Images
abstract
Due to the diverse geographical environments, intricate landscapes, and high-density settlements, the automatic identification of urban village boundaries using remote sensing images remains a highly challenging task. This paper proposes a novel and efficient neural network model called UV-Mamba for accurate boundary detection in high-resolution remote sensing images. UV-Mamba mitigates the memory loss problem in lengthy sequence modeling, which arises in state space models (SSM) with increasing image size, by incorporating deformable convolutions (DCN). Its architecture utilizes an encoder-decoder framework and includes an encoder with four deformable state space augmentation (DSSA) blocks for efficient multi-level semantic extraction and a decoder to integrate the extracted semantic information. We conducted experiments on two large datasets showing that UV-Mamba achieves state-of-the-art performance. Specifically, our model achieves 73.3% and 78.1% IoU on the Beijing and Xi’an datasets, respectively, representing improvements of 1.2% and 3.4% IoU over the previous best model while also being 6× faster in inference speed and 40× smaller in parameter count. Source code and pre-trained models are available at https://github.com/Devin-Egber/UV-Mamba.
Lulin Li, Xuechao Zou, Junliang Xing, Pin Tao
ICASSP4
2025 GTR: Guided Thought Reinforcement Prevents Thought Collapse in RL-Based VLM Agent Training
abstract
Reinforcement learning with verifiable outcome rewards (RLVR) has effectively scaled up chain-of-thought (CoT) reasoning in large language models (LLMs). Yet, its efficacy in training vision-language model (VLM) agents for goal-directed action reasoning in visual environments is less established. This work investigates this problem through extensive experiments on complex card games, such as 24 points, and embodied tasks from ALFWorld. We find that when rewards are based solely on action outcomes, RL fails to incentivize CoT reasoning in VLMs, instead leading to a phenomenon we termed thought collapse, characterized by a rapid loss of diversity in the agent's thoughts, state-irrelevant and incomplete reasoning, and subsequent invalid actions, resulting in negative rewards. To counteract thought collapse, we highlight the necessity of process guidance and propose an automated corrector that evaluates and refines the agent's reasoning at each RL step. This simple and scalable GTR (Guided Thought Reinforcement) framework trains reasoning and action simultaneously without the need for dense, per-step human labeling. Our experiments demonstrate that GTR significantly enhances the performance and generalization of the LLaVA-7b model across various visual environments, achieving 3-5 times higher task success rates compared to SoTA models with notably smaller model sizes.
Junliang Xing, Yuanchun Shi, Zongqing Lu 0002, Deheng Ye
ICCV3
2025 Entropy-Adaptive Diffusion Policy Optimization with Dynamic Step Alignment
Renye Yan, Jikang Cheng, Yaozhong Gan, Shikun Sun, Yunfan Yang, Ling Liang 0003, Jinlong Lin, Yeshuang Zhu, Jie Zhou 0001, Junliang Xing, Yimao Cai, Ru Huang 0001
ICCV12
2025 Dynamic Dictionary Learning for Remote Sensing Image Segmentation
Xuechao Zou, Kai Li 0023, Pin Tao, Junliang Xing, Congyan Lang
ICCV7
2025 BodyGen: Advancing Towards Efficient Embodiment Co-Design
abstract
Embodiment co-design aims to optimize a robot's morphology and control policy simultaneously. While prior work has demonstrated its potential for generating environment-adaptive robots, this field still faces persistent challenges in optimization efficiency due to the (i) combinatorial nature of morphological search spaces and (ii) intricate dependencies between morphology and control. We prove that the ineffective morphology representation and unbalanced reward signals between the design and control stages are key obstacles to efficiency. To advance towards efficient embodiment co-design, we propose **BodyGen**, which utilizes (1) topology-aware self-attention for both design and control, enabling efficient morphology representation with lightweight model sizes; (2) a temporal credit assignment mechanism that ensures balanced reward signals for optimization. With our findings, BodyGen achieves an average **60.03%** performance improvement against state-of-the-art baselines. We provide codes and more results on the website: https://genesisorigin.github.io.
Haofei Lu, Junliang Xing, Jianshu Li, Yuanchun Shi
ICLR3
2025 Minimal Impact ControlNet: Advancing Multi-ControlNet Integration
abstract
With the advancement of diffusion models, there is a growing demand for high-quality, controllable image generation, particularly through methods that utilize one or multiple control signals based on ControlNet. However, in current ControlNet training, each control is designed to influence all areas of an image, which can lead to conflicts when different control signals are expected to manage different parts of the image in practical applications. This issue is especially pronounced with edge-type control conditions, where regions lacking boundary information often represent low-frequency signals, referred to as silent control signals. When combining multiple ControlNets, these silent control signals can suppress the generation of textures in related areas, resulting in suboptimal outcomes. To address this problem, we propose Minimal Impact ControlNet. Our approach mitigates conflicts through three key strategies: constructing a balanced dataset, combining and injecting feature signals in a balanced manner, and addressing the asymmetry in the score function’s Jacobian matrix induced by ControlNet. These improvements enhance the compatibility of control signals, allowing for freer and more harmonious generation in areas with silent control signals.
Shikun Sun, Zixuan Wang 0026, Xubin Li, Tiezheng Ge, Zijie Ye, Xiaoyu Qin 0001, Junliang Xing, Bo Zheng 0007, Jia Jia 0001
ICLR8
2025 Diverse Policies Recovering via Pointwise Mutual Information Weighted Imitation Learning
abstract
Recovering a spectrum of diverse policies from a set of expert trajectories is an important research topic in imitation learning. After determining a latent style for a trajectory, previous diverse polices recovering methods usually employ a vanilla behavioral cloning learning objective conditioned on the latent style, treating each state-action pair in the trajectory with equal importance. Based on an observation that in many scenarios, behavioral styles are often highly relevant with only a subset of state-action pairs, this paper presents a new principled method in diverse polices recovering. In particular, after inferring or assigning a latent style for a trajectory, we enhance the vanilla behavioral cloning by incorporating a weighting mechanism based on pointwise mutual information. This additional weighting reflects the significance of each state-action pair's contribution to learning the style, thus allowing our method to focus on state-action pairs most representative of that style. We provide theoretical justifications for our new objective, and extensive empirical evaluations confirm the effectiveness of our method in recovering diverse polices from expert data.
Jian Yao 0008, Weiming Liu 0004, Hanmin Qin, Hansheng Kong, Kirk Tang, Jiechao Xiong, Chao Yu 0004, Kai Li 0022, Junliang Xing, Hongwu Chen, Juchao Zhuo, Qiang Fu 0016, Haobo Fu
ICLR11
2025 Goal-Oriented Skill Abstraction for Offline Multi-Task Reinforcement Learning
abstract
Offline multi-task reinforcement learning aims to learn a unified policy capable of solving multiple tasks using only pre-collected task-mixed datasets, without requiring any online interaction with the environment. However, it faces significant challenges in effectively sharing knowledge across tasks. Inspired by the efficient knowledge abstraction observed in human learning, we propose Goal-Oriented Skill Abstraction (GO-Skill), a novel approach designed to extract and utilize reusable skills to enhance knowledge transfer and task performance. Our approach uncovers reusable skills through a goal-oriented skill extraction process and leverages vector quantization to construct a discrete skill library. To mitigate class imbalances between broadly applicable and task-specific skills, we introduce a skill enhancement phase to refine the extracted skills. Furthermore, we integrate these skills using hierarchical policy learning, enabling the construction of a high-level policy that dynamically orchestrates discrete skills to accomplish specific tasks. Extensive experiments on diverse robotic manipulation tasks within the MetaWorld benchmark demonstrate the effectiveness and versatility of GO-Skill.
Jinmin He, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICML6
2025 Offline Opponent Modeling with Truncated Q-driven Instant Policy Refinement
abstract
Offline Opponent Modeling (OOM) aims to learn an adaptive autonomous agent policy that dynamically adapts to opponents using an offline dataset from multi-agent games. Previous work assumes that the dataset is optimal. However, this assumption is difficult to satisfy in the real world. When the dataset is suboptimal, existing approaches struggle to work. To tackle this issue, we propose a simple and general algorithmic improvement framework, Truncated Q-driven Instant Policy Refinement (TIPR), to handle the suboptimality of OOM algorithms induced by datasets. The TIPR framework is plug-and-play in nature. Compared to original OOM algorithms, it requires only two extra steps: (1) Learn a horizon-truncated in-context action-value function, namely Truncated Q, using the offline dataset. The Truncated Q estimates the expected return within a fixed, truncated horizon and is conditioned on opponent information. (2) Use the learned Truncated Q to instantly decide whether to perform policy refinement and to generate policy after refinement during testing. Theoretically, we analyze the rationale of Truncated Q from the perspective of No Maximization Bias probability. Empirically, we conduct extensive comparison and ablation experiments in four representative competitive environments. TIPR effectively improves various OOM algorithms pretrained with suboptimal datasets.
Yuheng Jing, Kai Li 0022, Bingyun Liu, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICML7
2025 Habitizing Diffusion Planning for Efficient and Effective Decision Making
abstract
Diffusion models have shown great promise in decision-making, also known as diffusion planning. However, the slow inference speeds limit their potential for broader real-world applications. Here, we introduce Habi, a general framework that transforms powerful but slow diffusion planning models into fast decision-making models, which mimics the cognitive process in the brain that costly goal-directed behavior gradually transitions to efficient habitual behavior with repetitive practice. Even using a laptop CPU, the habitized model can achieve an average 800+ Hz decision-making frequency (faster than previous diffusion planners by orders of magnitude) on standard offline reinforcement learning benchmarks D4RL, while maintaining comparable or even higher performance compared to its corresponding diffusion planner. Our work proposes a fresh perspective of leveraging powerful diffusion models for real-world decision-making tasks. We also provide robust evaluations and analysis, offering insights from both biological and engineering perspectives for efficient and effective decision-making.
Haofei Lu, Yifei Shen 0004, Dongsheng Li 0002, Junliang Xing
ICML4
2025 Pattern Extraction Learning for Cooperative Multi-Agent Reinforcement Learning
abstract
Multi-agent reinforcement learning (MARL) has demonstrated its superiority in addressing complex decision-making tasks involving multiple agents. However, the intricate and dynamic interactions among agents make this problem exceptionally challenging. Existing MARL methods often simplify the problem by implicitly decomposing shared rewards into individual utilities, neglecting the underlying interconnections between relevant entities. To overcome these limitations, we propose a novel framework of Multi-Agent Pattern Extraction (MAPE), which captures cooperation patterns from both agent-level and global perspectives to enhance decision-making and collaboration efficiency. Specifically, MAPE introduces two key modules: the Agent Pattern Extractor (APE) and the Global Pattern Extractor (GPE) that focus on specific interactions between the entities of interests from individual and global perspectives, respectively. The APE module focuses on computing each agent’s attention to other entities across different interaction patterns, providing this information to the GPE module. The GPE module then integrates the agent-specific pattern information and state features to identify the overall interaction pattern of the multi-agent system. By filtering out irrelevant interactions between unrelated entities and highlighting meaningful relationships, MAPE fosters more focused cooperation and facilitates more efficient learning. Extensive experiments on the StarCraft II micromanagement benchmark showcase the effectiveness of MAPE in improving efficiency in complex multi-agent environments.
Yifan Zang 0001, Jinmin He, Kai Li 0022, Junliang Xing, Jian Cheng 0001
IJCNN4
2025 Dual-Flow: Transferable Multi-Target, Instance-Agnostic Attacks via In-the-wild Cascading Flow Optimization
Shikun Sun, Jianshu Li, Junliang Xing
NeurIPS6
2025 Asymmetric interaction preference induces cooperation in human-agent hybrid game
Danyang Jia, Xiangfeng Dai, Junliang Xing, Pin Tao, Yuanchun Shi, Zhen Wang 0004
Sci. China Inf. Sci.3
2025 MT-Agent: Constructing a GUI Agent via Modality Enhancement and Text-Guided Fusion
abstract
Graphical User Interfaces (GUIs) play a crucial role in facilitating user-computer interactions, making them an essential focus of research. However, current automated GUI agents face significant challenges in effectively associating task implementations with specific visual elements, and the resolution constraints of Vision-Language Models (VLMs) also limit the richness of visual information. To this end, we propose a novel multi-modal agent named MT-Agent, which enhances both textual and visual input modalities to enable the model to perceive visual elements in GUIs more effectively. Specifically, Textual Modality Enhancement improves the semantic richness of input text by capturing task-specific details via an external VLM, while Visual Modality Enhancement incorporates fine-grained visual details to better represent critical GUI elements. In addition, we introduce an innovative text-guided directional feature fusion mechanism, which leverages enriched text features to guide the integration with visual information. In experiments, MT-Agent demonstrated exceptional performance on AITZ dataset, achieving an action type prediction accuracy of 84.80% and a step prediction accuracy of 58.07%, surpassing previous state-of-the-art models. Furthermore, on the GUI Odyssey benchmark, MT-Agent achieves performance comparable to previous state-of-the-art models while using only about 1/20 of their trainable parameters. Our codes, demos, and relevant data will be released to facilitate further research and validation within the scientific community.
Jinhan Dong, Lei Jin 0003, Zhihong Zhang 0006, Runqing Zhang, Liqiang Xu, Junliang Xing
IEEE Internet Things J.7
2025 Leveraging Privileged Information for Partially Observable Reinforcement Learning
abstract
Reinforcement learning has achieved remarkable success across diverse scenarios. However, learning optimal policies within partially observable games remains a formidable challenge. Crucial privileged information in states is often shrouded during gameplay, yet ideally, it should be accessible and exploitable during training. Previous studies have concentrated on formulating policies based wholly on partial observations or oracle states. Nevertheless, these approaches often face hindrances in attaining effective generalization. To surmount this challenge, we propose the actor–cross-critic (ACC) learning framework, integrating both partial observations and oracle states. ACC achieves this by coordinating two critics and invoking a maximization operation mechanism to switch between them dynamically. This approach encourages the selection of the higher values when computing advantages within the actor–critic framework, thereby accelerating learning and mitigating bias under partial observability. Some theoretical analyses show that ACC exhibits better learning ability toward optimal policies than actor–critic learning using the oracle states. We highlight its superior performance through comprehensive evaluations in decision-making tasks, such asQuestBall,Minigrid, andAtari, and the challenging card gameDouDizhu.
Jinqiu Li, Enmin Zhao, Junliang Xing, Shiming Xiang
IEEE Trans. Games4
2025 Adapting Vision Foundation Models for Robust Cloud Segmentation in Remote Sensing Images
abstract
Cloud segmentation is a critical challenge in remote sensing image interpretation, as its accuracy directly impacts the effectiveness of subsequent data processing and analysis. Recently, vision foundation models (VFM) have demonstrated powerful generalization capabilities across various visual tasks. In this paper, we present a parameter-efficient adaptive approach, termed Cloud-Adapter, designed to enhance the accuracy and robustness of cloud segmentation. Our method leverages a VFM pretrained on general domain data, which remains frozen, eliminating the need for additional training. Cloud-Adapter incorporates a lightweight spatial perception module that initially utilizes a convolutional neural network (ConvNet) to extract dense spatial representations. These multi-scale features are then aggregated and serve as contextual inputs to an adapting module, which modulates the frozen transformer layers within the VFM. Experimental results demonstrate that the Cloud-Adapter approach, utilizing only 0.6% of the trainable parameters of the frozen backbone, achieves substantial performance gains. Cloud-Adapter consistently achieves state-of-the-art performance across various cloud segmentation datasets from multiple satellite sources, sensor series, data processing levels, land cover scenarios, and annotation granularities. Code and model checkpoints are available at https://xavierjiezou.github.io/Cloud-Adapter/.
Xuechao Zou, Kai Li 0023, Junliang Xing, Lei Jin 0003, Congyan Lang, Pin Tao
IEEE Trans. Geosci. Remote. Sens.5
2025 UBG: An Unreal BattleGround Benchmark With Object-Aware Hierarchical Proximal Policy Optimization
abstract
The deep reinforcement learning (DRL) has made significant progress in various simulation environments. However, applying DRL methods to real-world scenarios poses certain challenges due to limitations in visual fidelity, scene complexity, and task diversity within existing environments. To address limitations and explore the potential ability of DRL, we developed a 3-D open-world first-person shooter (FPS) game called Unreal BattleGround (UBG) using the unreal engine (UE). UBG provides a realistic 3-D environment with variable complexity, random scenes, diverse tasks, and multiple scene interaction methods. This benchmark involves far more complex state-action spaces than classic pseudo-3-D FPS games (e.g., ViZDoom), making it challenging for DRL to learn human-level decision sequences. Then, we propose the object-aware hierarchically proximal policy optimization (OaH-PPO) method in the UBG. It involves a two-level hierarchy, where the high-level controller is tasked with learning option control, and the low-level workers focus on mastering subtasks. To boost the learning of subtasks, we propose three modules: an object-aware module for extracting depth detection information from the environment, potential-based intrinsic reward shaping for efficient exploration, and annealing imitation learning (IL) to guide the initialization. Experimental results have demonstrated the broad applicability of the UBG and the effectiveness of the OaH-PPO. We will release the code of the UBG and OaH-PPO after publication.
Longyu Niu, Baihui Li, Xingjian Fan, Jun Li 0033, Junliang Xing, Jun Wan 0001, Zhen Lei 0001
IEEE Trans. Neural Networks Learn. Syst.6
2024 Not All Tasks Are Equally Difficult: Multi-Task Deep Reinforcement Learning with Dynamic Depth Routing
abstract
Multi-task reinforcement learning endeavors to accomplish a set of different tasks with a single policy. To enhance data efficiency by sharing parameters across multiple tasks, a common practice segments the network into distinct modules and trains a routing network to recombine these modules into task-specific policies. However, existing routing approaches employ a fixed number of modules for all tasks, neglecting that tasks with varying difficulties commonly require varying amounts of knowledge. This work presents a Dynamic Depth Routing (D2R) framework, which learns strategic skipping of certain intermediate modules, thereby flexibly choosing different numbers of modules for each task. Under this framework, we further introduce a ResRouting method to address the issue of disparate routing paths between behavior and target policies during off-policy training. In addition, we design an automatic route-balancing mechanism to encourage continued routing exploration for unmastered tasks without disturbing the routing of mastered ones. We conduct extensive experiments on various robotics manipulation tasks in the Meta-World benchmark, where D2R achieves state-of-the-art performance with significantly improved learning efficiency.
Jinmin He, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
AAAI6
2024 Advancing Video Synchronization with Fractional Frame Analysis: Introducing a Novel Dataset and Model
abstract
Multiple views play a vital role in 3D pose estimation tasks. Ideally, multi-view 3D pose estimation tasks should directly utilize naturally collected videos for pose estimation. However, due to the constraints of video synchronization, existing methods often use expensive hardware devices to synchronize the initiation of cameras, which restricts most 3D pose collection scenarios to indoor settings. Some recent works learn deep neural networks to align desynchronized datasets derived from synchronized cameras and can only produce frame-level accuracy. For fractional frame video synchronization, this work proposes an Inter-Frame and Intra-Frame Desynchronized Dataset (IFID), which labels fractional time intervals between two video clips. IFID is the first dataset that annotates inter-frame and intra-frame intervals, with a total of 382,500 video clips annotated, making it the largest dataset to date. We also develop a novel model based on the Transformer architecture, named InSynFormer, for synchronizing inter-frame and intra-frame. Extensive experimental evaluations demonstrate its promising performance. The dataset and source code of the model are available at https://github.com/yuxuan-cser/InSynFormer.
Haizhou Ai, Junliang Xing, Xuri Li, Pin Tao
AAAI3
2024 SynSP: Synergy of Smoothness and Precision in Pose Sequences Refinement
abstract
Predicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily concentrated on striking a balance between the two objectives, i.e., smoothness and precision, while optimizing the predicted pose sequences. However, it has come to our attention that the tension between these two objectives can provide additional quality cues about the predicted pose sequences. These cues, in turn, are able to aid the network in optimizing lower-quality poses. To leverage this quality information, we propose a motion refinement network, termed SynSP, to achieve a Synergy of Smoothness and Precision in the sequence refinement tasks. Moreover, SynSP can also address multi-view poses of one person simultaneously, fixing inaccuracies in predicted poses through heightened attention to similar poses from other views, thereby amplifying the resultant quality cues and overall performance. Compared with previous methods, SynSP benefits from both pose quality and multi-view information with a much shorter input sequence length, achieving state-of-the-art results among four challenging datasets involving 2D, 3D, and SMPL pose representations in both single-view and multi-view scenes. Github code: https://github.com/InvertedForest/SynSP.
Tao Wang 0011, Lei Jin 0003, Zheng Wang 0007, Jianshu Li, Liang Li 0003, Fang Zhao 0006, Yu Cheng 0009, Li Yuan 0007, Junliang Xing, Jian Zhao 0006
CVPR10
2024 DriveWorld: 4D Pre-Trained Scene Understanding via World Models for Autonomous Driving
abstract
Vision-centric autonomous driving has recently raised wide attention due to its lower cost. Pretraining is essential for extracting a universal representation. However, current vision-centric pretraining typically relies on either 2D or 3D pre-text tasks, overlooking the temporal characteristics of autonomous driving as a 4D scene understanding task. In this paper, we address this challenge by introducing a world model-based autonomous driving 4D representation learning framework, dubbed DriveWorld, which is capable of pretraining from multi-camera driving videos in a spatiotemporal fashion. Specifically, we propose a Memory State-Space Model for spatiotemporal modelling, which consists of a Dynamic Memory Bank module for learning temporal-aware latent dynamics to predict future changes and a Static Scene Propagation module for learning spatial-aware latent statics to offer comprehensive scene contexts. We additionally introduce a Task Prompt to decouple task-aware features for various downstream tasks. The experiments demonstrate that DriveWorld delivers promising results on various autonomous driving tasks. When pretrained with the OpenScene dataset, DriveWorld achieves a 7.5% increase in mAP for 3D object detection, a 3.0% increase in IoU for online mapping, a 5.0% increase in AMOTA for multi-object tracking, a 0.1m decrease in minADE for motionforecasting, a 3.0% increase in IoU for occupancy prediction, and a 0.34m reduction in average L2 error for planning.
Dawei Zhao 0003, Liang Xiao 0007, Jian Zhao 0006, Xinli Xu, Lei Jin 0003, Jianshu Li, Yulan Guo, Junliang Xing, Liping Jing, Yiming Nie, Bin Dai 0001
CVPR10
2024 LEFormer: A Hybrid CNN-Transformer Architecture for Accurate Lake Extraction from Remote Sensing Imagery
abstract
Lake extraction from remote sensing images is challenging due to the complex lake shapes and inherent data noises. Existing methods suffer from blurred segmentation boundaries and poor foreground modeling. This paper proposes a hybrid CNN-Transformer architecture, called LEFormer, for accurate lake extraction. LEFormer contains three main modules: CNN encoder, Transformer encoder, and cross-encoder fusion. The CNN encoder effectively recovers local spatial information and improves fine-scale details. Simultaneously, the Transformer encoder captures long-range dependencies between sequences of any length, allowing them to obtain global features and context information. The cross-encoder fusion module integrates the local and global features to improve mask prediction. Experimental results show that LEFormer consistently achieves state-of-the-art performance and efficiency on the Surface Water and the Qinghai-Tibet Plateau Lake datasets. Specifically, LEFormer achieves 90.86% and 97.42% mIoU on two datasets with a parameter count of 3.61M, respectively, while being 20× minor than the previous best lake extraction method. The source code is available at https://github.com/BastianChen/LEFormer.
Xuechao Zou, Yu Zhang 0165, Jiayu Li 0008, Kai Li 0023, Junliang Xing, Pin Tao
ICASSP6
2024 ARFA: An Asymmetric Receptive Field Autoencoder Model for Spatiotemporal Prediction
abstract
Spatiotemporal prediction aims to generate future sequences by paradigms learned from historical contexts. It is essential in numerous domains, such as traffic flow prediction and weather forecasting. Recently, research in this field has been predominantly driven by deep neural networks based on autoencoder architectures. However, existing methods commonly adopt autoencoder architectures with identical receptive field sizes. To address this issue, we propose an Asymmetric Receptive Field Autoencoder (ARFA) model, which introduces corresponding sizes of receptive field modules tailored to the distinct functionalities of the encoder and decoder. In the encoder, we present a large kernel module for global spatiotemporal feature extraction. In the decoder, we develop a small kernel module for local spatiotemporal information reconstruction. Experimental results demonstrate that ARFA consistently achieves state-of-the-art performance on popular datasets. Additionally, we construct the RainBench, a large-scale radar echo dataset for precipitation prediction, to address the scarcity of meteorological data in the domain.
Xuechao Zou, Xiaoying Wang 0002, Jianqiang Huang 0002, Junliang Xing
ICASSP6
2024 Towards Offline Opponent Modeling with In-context Learning
abstract
Opponent modeling aims at learning the opponent's behaviors, goals, or beliefs to reduce the uncertainty of the competitive environment and assist decision-making. Existing work has mostly focused on learning opponent models online, which is impractical and inefficient in practical scenarios. To this end, we formalize an Offline Opponent Modeling (OOM) problem with the objective of utilizing pre-collected offline datasets to learn opponent models that characterize the opponent from the viewpoint of the controlled agent, which aids in adapting to the unknown fixed policies of the opponent. Drawing on the promises of the Transformers for decision-making, we introduce a general approach, Transformer Against Opponent (TAO), for OOM. Essentially, TAO tackles the problem by harnessing the full potential of the supervised pre-trained Transformers' in-context learning capabilities. The foundation of TAO lies in three stages: an innovative offline policy embedding learning stage, an offline opponent-aware response policy training stage, and a deployment stage for opponent adaptation with in-context learning. Theoretical analysis establishes TAO's equivalence to Bayesian posterior sampling in opponent modeling and guarantees TAO's convergence in opponent policy recognition. Extensive experiments and ablation studies on competitive environments with sparse and dense rewards demonstrate the impressive performance of TAO. Our approach manifests remarkable prowess for fast adaptation, especially in the face of unseen opponent policies, confirming its in-context learning potency.
Yuheng Jing, Kai Li 0022, Bingyun Liu, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICLR7
2024 Inner Classifier-Free Guidance and Its Taylor Expansion for Diffusion Models
abstract
Classifier-free guidance (CFG) is a pivotal technique for balancing the diversity and fidelity of samples in conditional diffusion models. This approach involves utilizing a single model to jointly optimize the conditional score predictor and unconditional score predictor, eliminating the need for additional classifiers. It delivers impressive results and can be employed for continuous and discrete condition representations. However, when the condition is continuous, it prompts the question of whether the trade-off can be further enhanced. Our proposed inner classifier-free guidance (ICFG) provides an alternative perspective on the CFG method when the condition has a specific structure, demonstrating that CFG represents a first-order case of ICFG. Additionally, we offer a second-order implementation, highlighting that even without altering the training policy, our second-order approach can introduce new valuable information and achieve an improved balance between fidelity and diversity for Stable Diffusion.
Shikun Sun, Longhui Wei, Zhicai Wang, Zixuan Wang 0026, Junliang Xing, Jia Jia 0001, Qi Tian 0001
ICLR5
2024 PAE: Reinforcement Learning from External Knowledge for Efficient Exploration
abstract
Human intelligence is adept at absorbing valuable insights from external knowledge. This capability is equally crucial for artificial intelligence. In contrast, classical reinforcement learning agents lack such capabilities and often resort to extensive trial and error to explore the environment. This paper introduces $\textbf{PAE}$: $\textbf{P}$lanner-$\textbf{A}$ctor-$\textbf{E}$valuator, a novel framework for teaching agents to $\textit{learn to absorb external knowledge}$. PAE integrates the Planner's knowledge-state alignment mechanism, the Actor's mutual information skill control, and the Evaluator's adaptive intrinsic exploration reward to achieve 1) effective cross-modal information fusion, 2) enhanced linkage between knowledge and state, and 3) hierarchical mastery of complex tasks. Comprehensive experiments across 11 challenging tasks from the BabyAI and MiniHack environment suites demonstrate PAE's superior exploration efficiency with good interpretability.
Haofei Lu, Junliang Xing, Renye Yan, Yaozhong Gan, Yuanchun Shi
ICLR3
2024 Dynamic Discounted Counterfactual Regret Minimization
abstract
Counterfactual regret minimization (CFR) is a family of iterative algorithms showing promising results in solving imperfect-information games. Recent novel CFR variants (e.g., CFR+, DCFR) have significantly improved the convergence rate of the vanilla CFR. The key to these CFR variants’ performance is weighting each iteration non-uniformly, i.e., discounting earlier iterations. However, these algorithms use a fixed, manually-specified scheme to weight each iteration, which enormously limits their potential. In this work, we propose Dynamic Discounted CFR (DDCFR), the first equilibrium-finding framework that discounts prior iterations using a dynamic, automatically-learned scheme. We formalize CFR’s iteration process as a carefully designed Markov decision process and transform the discounting scheme learning problem into a policy optimization problem within it. The learned discounting scheme dynamically weights each iteration on the fly using information available at runtime. Experimental results across multiple games demonstrate that DDCFR’s dynamic discounting scheme has a strong generalization ability and leads to faster convergence with improved performance. The code is available at https://github.com/rpSebastian/DDCFR.
Hang Xu 0006, Kai Li 0022, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
ICLR5
2024 High-Fidelity Lake Extraction Via Two-Stage Prompt Enhancement: Establishing A Novel Baseline and Benchmark
abstract
Lake extraction from remote sensing imagery is a complex challenge due to the varied lake shapes and data noise. Current methods rely on multispectral image datasets, making it challenging to learn lake features accurately from pixel arrangements. This, in turn, affects model learning and the creation of accurate segmentation masks. This paper introduces a prompt-based dataset construction approach that provides approximate lake locations using point, box, and mask prompts. We also propose a two-stage prompt enhancement framework, LEPrompter, with prompt-based and prompt-free stages during training. The prompt-based stage employs a prompt encoder to extract prior information, integrating prompt tokens and image embedding through self- and cross-attention in the prompt decoder. Prompts are deactivated to ensure independence during inference, enabling automated lake extraction without introducing additional parameters and GFlops. Extensive experiments showcase performance improvements of our proposed approach compared to the previous state- of-the-art method. The source code is available at https://github.com/BastianChen/LEPrompter.
Xuechao Zou, Yu Zhang 0165, Junliang Xing, Pin Tao
ICME5
2024 ASQuery: A Query-based Model for Action Segmentation
abstract
For the task of temporal action segmentation, existing works commonly treat it as a frame-wise classification problem. In this paper, we propose a straight but effective model namely ASQuery by learning central representation of each action category, which transforms the classification problem to the similarity calculation between category-specific queries and frame features. These central representations are dynamically generated through our Transformer decoder module, endowing them more flexible and comprehensive perception of the whole video. Moreover, we first introduce the boundary query for refining segmentation results, aiding to alleviating the troublesome over-segmentation problem. ASQuery demonstrates superior performance compared to state-of-the-art models, achieving improvements of 0.9% and 4.1% in the mean metrics on two public action segmentation datasets, i.e., Breakfast and Assembly101, respectively. The source codes are available at https://github.com/zlngan/ASQuery.
Ziliang Gan, Lei Jin 0003, Zheng Wang 0007, Liang Li 0003, Zhecan Wang, Jianshu Li, Junliang Xing, Jian Zhao 0006
ICME9
2024 Deviation Wing Loss for High-Performance 2D Pose Estimation
abstract
Heatmap regression using deep neural networks has become the dominant approach in 2D pose estimation. Nonetheless, the intrinsic variability in the flexibility of distinct keypoints engenders discernible momentum deviation among them, leading to training biases. Moreover, previous methods indiscriminately treat all pixels in a heatmap, further exacerbating biases. Regrettably, conventional loss functions, including the Mean Squared Error (MSE) loss, fail to rectify this issue adequately. Consequently, the need arises to recalibrate weights to concentrate the loss’s impact on specific regions. To this end, we introduce a novel Deviation Wing (DW) loss function for high-performance 2D pose estimation, incorporating two improvement aspects. Firstly, we use Gaussian Momentum Deviation encoding craft deviation maps, leveraging momentum deviation as a source of prior knowledge to enhance supervision over keypoints characterized by substantial momentum deviation. Subsequently, we harness the Cosine Wing function to amplify the loss concerning minor errors residing within the keypoint region and supervise errors across diverse scales. Our comprehensive empirical exploration spans multiple datasets encompassing 2D human pose and hand pose estimation. The experimental results demonstrate the efficacy of our proposed loss function in enhancing heatmap regression performances.
Junliang Xing, Xinchun Yu, Xiao-Ping Zhang 0002
ICME2
2024 A Parallel Attention Network For Cattle Face Recognition
abstract
Cattle face recognition holds paramount significance in domains such as animal husbandry and behavioral research. Despite significant progress in confined environments, applying these accomplishments in wild settings remains challenging. Thus, we create the first large-scale cattle face recognition dataset, ICRWE, for wild environments. It encompasses 483 cattle and 9,816 high-resolution image samples. Each sample undergoes annotation for face features, light conditions, and face orientation. Furthermore, we introduce a novel parallel attention network, PANet. Comprising several cascaded Transformer modules, each module incorporates two parallel Position Attention Modules (PAM) and Feature Mapping Modules (FMM). PAM focuses on local and global features at each image position through parallel channel attention, and FMM captures intricate feature patterns through non-linear mappings. Experimental results indicate that PANet achieves a recognition accuracy of 88.03% on the ICRWE dataset, establishing itself as the current state-of-the-art approach. The source code is available at https://github.com/1jy-0124/PANet
Jiayu Li 0008, Xuechao Zou, Junliang Xing, Pin Tao
ICME5
2024 Reflective Policy Optimization
abstract
On-policy reinforcement learning methods, like Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), often demand extensive data per update, leading to sample inefficiency. This paper introduces Reflective Policy Optimization (RPO), a novel on-policy extension that amalgamates past and future state-action information for policy optimization. This approach empowers the agent for introspection, allowing modifications to its actions within the current state. Theoretical analysis confirms that policy performance is monotonically improved and contracts the solution space, consequently expediting the convergence procedure. Empirical results demonstrate RPO's feasibility and efficacy in two reinforcement learning benchmarks, culminating in superior sample efficiency. The source code of this work is available at https://github.com/Edgargan/RPO.
Yaozhong Gan, Renye Yan, Junliang Xing
ICML4
2024 Unified Single-Stage Transformer Network for Efficient RGB-T Tracking
Jianqiang Xia, Dian-xi Shi, Linna Song, Songchang Jin, Chenran Zhao, Yu Cheng 0009, Lei Jin 0003, Jianan Li 0001, Gang Wang 0031, Junliang Xing, Jian Zhao 0006
IJCAI13
2024 Minimizing Weighted Counterfactual Regret with Optimistic Online Mirror Descent
Hang Xu 0006, Kai Li 0022, Bingyun Liu, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
IJCAI6
2024 Absorb What You Need: Accelerating Exploration via Valuable Knowledge Extraction
abstract
Leveraging external knowledge and extracting valuable insights are efficient human practices when handling various tasks. In contrast, current artificial intelligence lacks this capability. Recent research aims to teach Reinforcement Learning (RL) agents to incorporate external knowledge in the form of natural language to accelerate exploration. A common assumption in many of these approaches is that all introduced external knowledge is inherently valuable. To eliminate this assumption, we introduce the Knowledge Extraction Exploration Framework (KEEF). KEEF comprises two key components: 1) a knowledge extractor, designed to filter useful external knowledge based on three dimensions, task relevance, environment relevance, and achievement difficulty, through a prediction network and a policy network; and 2) a policy executor, which is a knowledge-conditioned network facilitating joint reasoning between the useful knowledge extracted by the knowledge extractor and the current state of the environment. In eight challenging sparse reward BabyAI environments, KEEF has consistently demonstrated superior sampling efficiency compared to knowledge-based and traditional RL methods.
Renye Yan, Pin Tao, Junliang Xing
IJCNN5
2024 Translating Motion to Notation: Hand Labanotation for Intuitive and Comprehensive Hand Movement Documentation
Wenrui Yang, Xinchun Yu, Junliang Xing, Xiao-Ping Zhang 0002
ACM Multimedia4
2024 Efficient Multi-task Reinforcement Learning with Cross-Task Policy Guidance
abstract
Multi-task reinforcement learning endeavors to efficiently leverage shared information across various tasks, facilitating the simultaneous learning of multiple tasks. Existing approaches primarily focus on parameter sharing with carefully designed network structures or tailored optimization procedures. However, they overlook a direct and complementary way to exploit cross-task similarities: the control policies of tasks already proficient in some skills can provide explicit guidance for unmastered tasks to accelerate skills acquisition. To this end, we present a novel framework called Cross-Task Policy Guidance (CTPG), which trains a guide policy for each task to select the behavior policy interacting with the environment from all tasks' control policies, generating better training trajectories. In addition, we propose two gating mechanisms to improve the learning efficiency of CTPG: one gate filters out control policies that are not beneficial for guidance, while the other gate blocks tasks that do not necessitate guidance. CTPG is a general framework adaptable to existing parameter sharing approaches. Empirical evaluations demonstrate that incorporating CTPG with these approaches significantly enhances performance in manipulation and locomotion benchmarks.
Jinmin He, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
NeurIPS6
2024 Opponent Modeling with In-context Search
abstract
Opponent modeling is a longstanding research topic aimed at enhancing decision-making by modeling information about opponents in multi-agent environments. However, existing approaches often face challenges such as having difficulty generalizing to unknown opponent policies and conducting unstable performance. To tackle these challenges, we propose a novel approach based on in-context learning and decision-time search named Opponent Modeling with In-context Search (OMIS). OMIS leverages in-context learning-based pretraining to train a Transformer model for decision-making. It consists of three in-context components: an actor learning best responses to opponent policies, an opponent imitator mimicking opponent actions, and a critic estimating state values. When testing in an environment that features unknown non-stationary opponent agents, OMIS uses pretrained in-context components for decision-time search to refine the actor's policy. Theoretically, we prove that under reasonable assumptions, OMIS without search converges in opponent policy recognition and has good generalization properties; with search, OMIS provides improvement guarantees, exhibiting performance stability. Empirically, in competitive, cooperative, and mixed environments, OMIS demonstrates more effective and stable adaptation to opponents than other approaches. See our project website at https://sites.google.com/view/nips2024-omis.
Yuheng Jing, Bingyun Liu, Kai Li 0022, Yifan Zang 0001, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
NeurIPS7
2024 Dual Critic Reinforcement Learning under Partial Observability
abstract
Partial observability in environments poses significant challenges that impede the formation of effective policies in reinforcement learning. Prior research has shown that borrowing the complete state information can enhance sample efficiency. This strategy, however, frequently encounters unstable learning with high variance in practical applications due to the over-reliance on complete information. This paper introduces DCRL, a Dual Critic Reinforcement Learning framework designed to adaptively harness full-state information during training to reduce variance for optimized online performance. In particular, DCRL incorporates two distinct critics: an oracle critic with access to complete state information and a standard critic functioning within the partially observable context. It innovates a synergistic strategy to meld the strengths of the oracle critic for efficiency improvement and the standard critic for variance reduction, featuring a novel mechanism for seamless transition and weighting between them. We theoretically prove that DCRL mitigates the learning variance while maintaining unbiasedness. Extensive experimental analyses across the Box2D and Box3D environments have verified DCRL's superior performance. The source code is available in the supplementary.
Jinqiu Li, Enmin Zhao, Junliang Xing, Shiming Xiang
NeurIPS4
2024 GOP: A Group Object Perception Framework for Optical Remote Sensing
Lei Jin 0003, Xuechao Zou, Jian Zhao 0006, Junliang Xing
PRCV (12)5
2024 Automatically designing counterfactual regret minimization algorithms for solving imperfect-information games
Kai Li 0022, Hang Xu 0006, Haobo Fu, Qiang Fu 0016, Junliang Xing
Artif. Intell.5
2024 SkatingVerse: A large-scale benchmark for comprehensive evaluation on human action understanding
abstract
Abstract Human action understanding (HAU) is a broad topic that involves specific tasks, such as action localisation, recognition, and assessment. However, most popular HAU datasets are bound to one task based on particular actions. Combining different but relevant HAU tasks to establish a unified action understanding system is challenging due to the disparate actions across datasets. A large‐scale and comprehensive benchmark, namely SkatingVerse is constructed for action recognition, segmentation, proposal, and assessment. SkatingVerse focus on fine‐grained sport action, hence figure skating is chosen as the task object, which eliminates the biases of the object, scene, and space that exist in most previous datasets. In addition, skating actions have inherent complexity and similarity, which is an enormous challenge for current algorithms. A total of 1687 official figure skating competition videos was collected with a total of 184.4 h, exceeding four times over other datasets with a similar topic. SkatingVerse enables to formulate a unified task to output fine‐grained human action classification and assessment results from a raw figure skating competition video. In addition, SkatingVerse can facilitate the study of HAU foundation model due to its large scale and abundant categories. Moreover, image modality is incorporated for human pose estimation task into SkatingVerse . Extensive experimental results show that (1) SkatingVerse significantly helps the training and evaluation of HAU methods, (2) the performance of existing HAU methods has much room to improve, and SkatingVerse helps to reduce such gaps, and (3) unifying relevant tasks in HAU through a uniform dataset can facilitate more practical applications. SkatingVerse will be publicly available to facilitate further studies on relevant problems.
Ziliang Gan, Lei Jin 0003, Yu Cheng 0009, Yinglei Teng, Zun Li 0001, Yawen Li 0001, Wenhan Yang, Junliang Xing, Jian Zhao 0006
IET Comput. Vis.10
2024 Freedom of choice disrupts cyclic dominance but maintains cooperation in voluntary prisoner's dilemma game
Danyang Jia, Chen Shen 0006, Xiangfeng Dai, Xinyu Wang 0022, Junliang Xing, Pin Tao, Yuanchun Shi, Zhen Wang 0004
Knowl. Based Syst.5
2024 DiffCR: A Fast Conditional Diffusion Framework for Cloud Removal From Optical Satellite Images
abstract
Optical satellite images are a critical data source; however, cloud cover often compromises their quality, hindering image applications and analysis. Consequently, effectively removing clouds from optical satellite images has emerged as a prominent research direction. Recent advances in deep learning-based cloud removal methods have been significant, but image generation quality still needs improvement. Diffusion models have demonstrated remarkable success in diverse image-generation tasks, showcasing their potential in addressing this challenge. This paper presents a novel framework called DiffCR, which leverages conditional guided diffusion with deep convolutional networks for high-performance cloud removal for optical satellite imagery. Specifically, we introduce a decoupled encoder for conditional image feature extraction, providing a robust color representation to ensure the close similarity of appearance information between the conditional input and the synthesized output. Moreover, we propose a novel and efficient time and condition fusion block within the cloud removal model to accurately simulate the correspondence between the appearance in the conditional image and the target image at a low computational cost. Extensive experimental evaluations on three commonly used benchmark datasets demonstrate that DiffCR consistently achieves state-of-the-art performance on all metrics, with parameter and computational complexities amounting to only 5.1% and 5.4%, respectively, of those previous best methods. The source code, pre-trained models, and all the experimental results will be publicly available at https://github.com/XavierJiezou/DiffCR upon the paper’s acceptance of this work.
Xuechao Zou, Kai Li 0023, Junliang Xing, Yu Zhang 0165, Lei Jin 0003, Pin Tao
IEEE Trans. Geosci. Remote. Sens.3
2024 UniParser: Multi-Human Parsing With Unified Correlation Representation Learning
abstract
Multi-human parsing is an image segmentation task necessitating both instance-level and fine-grained category-level information. However, prior research has typically processed these two types of information through distinct branch types and output formats, leading to inefficient and redundant frameworks. This paper introduces UniParser, which integrates instance-level and category-level representations in three key aspects: 1) we propose a unified correlation representation learning approach, allowing our network to learn instance and category features within the cosine space; 2) we unify the form of outputs of each modules as pixel-level results while supervising instance and category features using a homogeneous label accompanied by an auxiliary loss; and 3) we design a joint optimization procedure to fuse instance and category representations. By unifying instance-level and category-level output, UniParser circumvents manually designed post-processing techniques and surpasses state-of-the-art methods, achieving 49.3% AP on MHPv2.0 and 60.4% AP on CIHP. We have released our source code, pretrained models, and demos to facilitate future studies on https://github.com/cjm-sfw/Uniparser.
Jiaming Chu, Lei Jin 0003, Yinglei Teng, Jianshu Li, Yunchao Wei, Zheng Wang 0007, Junliang Xing, Shuicheng Yan, Jian Zhao 0006
IEEE Trans. Image Process.7
2024 OpenHoldem: A Benchmark for Large-Scale Imperfect-Information Game Research
abstract
Owing to the unremitting efforts from a few institutes, researchers have recently made significant progress in designing superhuman artificial intelligence (AI) in no-limit Texas hold'em (NLTH), the primary testbed for large-scale imperfect-information game research. However, it remains challenging for new researchers to study this problem since there are no standard benchmarks for comparing with existing methods, which hinders further developments in this research area. This work presents OpenHoldem, an integrated benchmark for large-scale imperfect-information game research using NLTH. OpenHoldem makes three main contributions to this research direction: 1) a standardized evaluation protocol for thoroughly evaluating different NLTH AIs; 2) four publicly available strong baselines for NLTH AI; and 3) an online testing platform with easy-to-use APIs for public NLTH AI evaluation. We will publicly release OpenHoldem and hope it facilitates further studies on the unsolved theoretical and computational issues in this area and cultivates crucial research problems like opponent modeling and human-computer interactive learning.
Kai Li 0022, Hang Xu 0006, Enmin Zhao, Junliang Xing
IEEE Trans. Neural Networks Learn. Syst.5
2023 PMAA: A Progressive Multi-Scale Attention Autoencoder Model for High-Performance Cloud Removal from Multi-Temporal Satellite Imagery
abstract
Satellite imagery analysis plays a pivotal role in remote sensing; however, information loss due to cloud cover significantly impedes its application. Although existing deep cloud removal models have achieved notable outcomes, they scarcely consider contextual information. This study introduces a high-performance cloud removal architecture, termed Progressive Multi-scale Attention Autoencoder (PMAA), which concurrently harnesses global and local information to construct robust contextual dependencies using a novel Multi-scale Attention Module (MAM) and a novel Local Interaction Module (LIM). PMAA establishes long-range dependencies of multi-scale features using MAM and modulates the reconstruction of fine-grained details utilizing LIM, enabling simultaneous representation of fine- and coarse-grained features at the same level. With the help of diverse and multi-scale features, PMAA consistently outperforms the previous state-of-the-art model CTGAN on two benchmark datasets. Moreover, PMAA boasts considerable efficiency advantages, with only 0.5% and 14.6% of the parameters and computational complexity of CTGAN, respectively. These comprehensive results underscore PMAA’s potential as a lightweight cloud removal network suitable for deployment on edge devices to accomplish large-scale cloud removal tasks. Our source code and pre-trained models are available at https://github.com/XavierJiezou/PMAA.
Xuechao Zou, Kai Li 0023, Junliang Xing, Pin Tao, Yachao Cui
ECAI3
2023 Shuffled Autoregression for Motion Interpolation
abstract
This work aims to provide a deep-learning solution for the motion interpolation task. Previous studies solve it with geometric weight functions. Some other works propose neural networks for different problem settings with consecutive pose sequences as input. However, motion interpolation is a more complex problem that takes isolated poses (e.g., only one start pose and one end pose) as input. When applied to motion interpolation, these deep learning methods have limited performance since they do not leverage the flexible dependencies between interpolation frames as the original geometric formulas do. To realize this interpolation characteristic, we propose a novel framework, referred to as Shuffled AutoRegression, which expands the autoregression to generate in arbitrary (shuffled) order and models any inter-frame dependencies as a directed acyclic graph. We further propose an approach to constructing a particular kind of dependency graph, with three stages assembled into an end-to-end spatial-temporal motion Transformer. Experimental results on one of the current largest datasets show that our model generates vivid and coherent motions from only one start frame to one end frame and outperforms competing methods by a large margin. The proposed model is also extensible to multiple keyframes’ motion interpolation tasks and other areas’ interpolation.
Shuo Huang 0005, Jia Jia 0001, Zongxin Yang, Wei Wang 0010, Haozhe Wu, Yi Yang 0001, Junliang Xing
ICASSP7
2023 MSNet: A Deep Architecture Using Multi-Sentiment Semantics for Sentiment-Aware Image Style Transfer
abstract
Sentiment plays an essential role in people’s perception of images. To incorporate the sentiment information into the image style transfer task for better sentiment-aware performance, we introduce a new task named sentiment-aware image style transfer. To solve this problem, we first introduce a novel Multi-Sentiment Semantics Space (MSS-Space) to capture the non-deterministic and complicated nature of sentiment semantics. With the MSS-Space, we establish tight associations between the visual attributes of images and the multi-sentiment semantics by minimizing their distance in MSS-Space and then propose the Multi-Sentiment Style Transfer Net (MSNet). Experiments demonstrate that, compared with three competing models, our proposed MSNet generates more explicit images and better preserves the integrity of salient objects, local details, and multi-sentiment. In particular, our model outperforms the state-of-the-art by +28.72% in terms of the top-3 accuracy on average.
Shikun Sun, Jia Jia 0001, Haozhe Wu, Zijie Ye, Junliang Xing
ICASSP5
2023 Salient Co-Speech Gesture Synthesizing with Discrete Motion Representation
abstract
Synthesizing co-speech gestures is challenging because the mapping from speech to gesticulation is inherently non-deterministic. When giving talks, people conduct not only gentle and rhythmic motions but also abrupt and salient gesticulations. Most previous research efforts, however, ignore this nature of co-speech gestures and synthesize deterministic results, producing over-smoothed movements with limited expressiveness. To address this issue, we propose a new co-speech gesture generation approach that produces high-quality salient gesticulations. Specifically, we build a discrete motion representation (DMR) space to bridge the speech-gesture mapping and the gesture generation stages. The incorporation of DMR enables random sampling in motion space and avoids the over-smooth problem in speech-gesture mapping. Based on DMR, we devise a novel multi-modal co-speech gesture synthesis model with temporal attention (MCGT). MCGT explicitly models DMR’s categorical distribution conditioned on the speech context, which captures complex context patterns and produces more salient gesticulations in sync with the context. In addition, we construct a new benchmark for evaluating salient motion quality in co-speech gestures, containing a large-scale co-speech gesture dataset with salient gesticulations. We also introduce a new metric, referred to as salient motion similarity, to evaluate the salient motion quality. Experiments demonstrate superior results from our approach over several competing baselines.
Zijie Ye, Jia Jia 0001, Haozhe Wu, Shuo Huang 0005, Shikun Sun, Junliang Xing
ICASSP6
2023 SDDM: Score-Decomposed Diffusion Models on Manifolds for Unpaired Image-to-Image Translation
abstract
Recent score-based diffusion models (SBDMs) show promising results in unpaired image-to-image translation (I2I). However, existing methods, either energy-based or statistically-based, provide no explicit form of the interfered intermediate generative distributions. This work presents a new score-decomposed diffusion model (SDDM) on manifolds to explicitly optimize the tangled distributions during image generation. SDDM derives manifolds to make the distributions of adjacent time steps separable and decompose the score function or energy guidance into an image "denoising" part and a content "refinement" part. To refine the image in the same noise level, we equalize the refinement parts of the score function and energy guidance, which permits multi-objective optimization on the manifold. We also leverage the block adaptive instance normalization module to construct manifolds with lower dimensions but still concentrated with the perturbed reference image. SDDM outperforms existing SBDM-based methods with much fewer diffusion steps on several I2I benchmarks.
Shikun Sun, Longhui Wei, Junliang Xing, Jia Jia 0001, Qi Tian 0001
ICML3
2023 Mnemonic Dictionary Learning for Intrinsic Motivation in Reinforcement Learning
abstract
Reinforcement learning for hard-exploration tasks remains challenging due to the long-term dependence and sparse-and-delay rewards in complex environments. In these challenging tasks, intrinsic motivation has become a dominant paradigm to enable the agent to explore the environment when no external reward feedback is available. In this work, inspired by studies from the human memory mechanism, we present a mnemonic dictionary learning (MDL) model for intrinsic motivation in reinforcement learning. The MDL model leverages sparse dictionary learning to incremental abstract the exploration histories into a compact memory-like dictionary, providing an excellent intrinsic motivation model. This mnemonic dictionary model not only drives the agent to explore novel stats in the environments indicated by the memory reconstruction error but also helps the agent to remember the key states and structure of the environments using its learned bases and reconstruction coefficients. The proposed MDL model can serve as a generative module for existing exploration methods. Extensive experimental results on typical sparse-reward tasks demonstrate its effectiveness and applicability over several competing algorithms. We will release the source code and trained models to facilitate further studies in this research direction.
Renye Yan, Yuan Zhan, Pin Tao, Zongwei Wang 0001, Yimao Cai, Junliang Xing
IJCNN7
2023 TLMIX: Twin Leader Mixing Network for Cooperative Multi-Agent Reinforcement Learning
abstract
Recent methods of cooperative multiagent rein-forcement learning built upon the individual global max value decomposition principle show promising results using variants of deep mixing networks, and credit assignment plays a crucial role in it. However, each agent in a multiagent system requires not only credit assignment but also credit feedback which tells each agent how many rewards it should obtain to maximize expected cumulative global rewards. In this work, we propose TLMIX, a novel Twin Leader Mixing Network for multiagent cooperation reinforcement learning while maintaining the centralized training and decentralized execution paradigm. TLMIX introduces a leader network to address the credit feedback issue by utilizing global information to provide reasonable objectives for agent networks. TLMIX also introduces a twin mixing network to find a more accurate target function from the Q-value functions, which avoids the rapid increase in parameter scale caused by introducing individual agents' twin networks and effectively mitigates the accumulation of high overestimation errors caused by temporal difference updates. Extensive results on SMAC experimental scenarios and the Predator-Prey environment demonstrate that TLMIX significantly outperforms comparable benchmark algorithms on convergence speed and performance.
Yu Zhang 0165, Pengyu Gao, Yusheng Jiang, Junliang Xing, Pin Tao
IJCNN5
2023 Pseudo Value Network Distillation for High-Performance Exploration
abstract
Solving hard exploration tasks with sparse rewards is notoriously challenging in reinforcement learning (RL), which needs to address two key issues simultaneously: exploiting past successful experiences and exploring the unknown environment. Many prior works take expert demonstrations as successful experiences and learn to imitate them directly. However, these demonstrations are often not available in practice. Recently, curiosity-driven RL methods provide intrinsic rewards, encouraging the agent to explore states with high novelty. Nonetheless, they lack a mechanism for leveraging past good experiences effectively. This work presents a Pseudo Value Network Distillation (PVND) framework to balance the RL agent's exploitative and exploratory behaviors effectively and automatically. In particular, PVND learns to set high exploitation bonuses to the critical states in rewarded trajectories from past experiences and high exploration bonuses to the novel states that agents rarely visit during exploration. We theoretically demonstrate that PVND gives larger positive intrinsic rewards to more critical states. Furthermore, PVND automatically finds meaningful and critical hierarchical sub-tasks for agents to accomplish the final goal progressively. Competitive results in several hard exploration sparse reward problems have verified its effectiveness and efficiency.
Enmin Zhao, Junliang Xing, Kai Li 0022, Yongxin Kang, Pin Tao
IJCNN2
2023 Single-Stage Multi-human Parsing via Point Sets and Center-based Offsets
abstract
This work studies the multi-human parsing problem. Existing methods, either following top-down or bottom-up two-stage paradigms, usually involve expensive computational costs. We instead present a high-performance Single-stage Multi-human Parsing (SMP) deep architecture that decouples the multi-human parsing problem into two fine-grained sub-problems,i.e., locating the human body and parts. SMP leverages the point features in the barycenter positions to obtain their segmentation and then generates a series of offsets from the barycenter of the human body to the barycenters of parts, thus performing human body and parts matching without the grouping process. Within the SMP architecture, we propose a Refined Feature Retain module to extract the global feature of instances through generated mask attention and a Mask of Interest Reclassify module as a trainable plug-in module to refine the classification results with the predicted segmentation. Extensive experiments on the MHPv2.0 dataset demonstrate the best effectiveness and efficiency of the proposed method, surpassing the state-of-the-art method by 2.1% in AP50p, 1.0% in APvolpsup>, and 1.2% in PCP50. Moreover, SMP also achieves superior performance in DensePose-COCO, verifying generalization of the model. In particular, the proposed method requires fewer training epochs and a less complex model architecture. Our codes are released in https://github.com/cjm-sfw/SMP.
Jiaming Chu, Lei Jin 0003, Xiaojin Fan, Yinglei Teng, Yunchao Wei, Yuqiang Fang, Junliang Xing, Jian Zhao 0006
ACM Multimedia7
2023 DecenterNet: Bottom-Up Human Pose Estimation Via Decentralized Pose Representation
abstract
Multi-person pose estimation in crowded scenes remains a very challenging task. This paper finds that most previous methods fail to estimate or group visible keypoints in crowded scenes rather than reasoning invisible keypoints. We thus categorize the crowded scenes into entanglement and occlusion based on the visibility of human parts and observe that entanglement is a significant problem in crowded scenes. With this observation, we propose DecenterNet, an end-to-end deep architecture to perform robust and efficient pose estimation in crowded scenes. Within DecenterNet, we introduce a decentralized pose representation that uses all visible keypoints as the root points to represent human poses, which is more robust in the entanglement area. We also propose a decoupled pose assessment mechanism, which introduces a location map to adaptively select optimal poses in the offset map. In addition, we have constructed a new dataset named SkatingPose, containing more entangled scenes. The proposed DecenterNet surpasses the best method on SkatingPose by 1.8 AP. Furthermore, DecenterNet obtains 71.2 AP and 71.4 AP on the COCO and CrowdPose datasets, respectively, demonstrating the superiority of our method. We will release our source code, trained models, and dataset to facilitate further studies in this research direction. Our code and dataset are available in https://github.com/InvertedForest/DecenterNet.
Tao Wang 0011, Lei Jin 0003, Xiaojin Fan, Yu Cheng 0009, Yinglei Teng, Junliang Xing, Jian Zhao 0006
ACM Multimedia7
2023 Versatile Face Animator: Driving Arbitrary 3D Facial Avatar in RGBD Space
abstract
Creating realistic 3D facial animation is crucial for various applications in the movie production and gaming industry, especially with the burgeoning demand in the metaverse. However, prevalent methods such as blendshape-based approaches and facial rigging techniques are time-consuming, labor-intensive, and lack standardized configurations, making facial animation production challenging and costly. In this paper, we propose a novel self-supervised framework, Versatile Face Animator, which combines facial motion capture with motion retargeting in an end-to-end manner, eliminating the need for blendshapes or rigs. Our method has the following two main characteristics: 1) we propose an RGBD animation module to learn facial motion from raw RGBD videos by hierarchical motion dictionaries and animate RGBD images rendered from 3D facial mesh coarse-to-fine, enabling facial animation on arbitrary 3D characters regardless of their topology, textures, blendshapes, and rigs; and 2) we introduce a mesh retarget module to utilize RGBD animation to create 3D facial animation by manipulating facial mesh with controller transformations, which are estimated from dense optical flow fields and blended together with geodesic-distance-based weights. Comprehensive experiments demonstrate the effectiveness of our proposed framework in generating impressive 3D facial animation results, highlighting its potential as a promising solution for the cost-effective and efficient production of facial animation in the metaverse.
Haoyu Wang 0009, Haozhe Wu, Junliang Xing, Jia Jia 0001
ACM Multimedia3
2023 Speech-Driven 3D Face Animation with Composite and Regional Facial Movements
abstract
Speech-driven 3D face animation poses significant challenges due to the intricacy and variability inherent in human facial movements. This paper emphasizes the importance of considering both the composite and regional natures of facial movements in speech-driven 3D face animation. The composite nature pertains to how speech-independent factors globally modulate speech-driven facial movements along the temporal dimension. Meanwhile, the regional nature alludes to the notion that facial movements are not globally correlated but are actuated by local musculature along the spatial dimension. It is thus indispensable to incorporate both natures for engendering vivid animation. To address the composite nature, we introduce an adaptive modulation module that employs arbitrary facial movements to dynamically adjust speech-driven facial movements across frames on a global scale. To accommodate the regional nature, our approach ensures that each constituent of the facial features for every frame focuses on the local spatial movements of 3D faces. Moreover, we present a non-autoregressive backbone for translating audio to 3D facial movements, which maintains high-frequency nuances of facial movements and facilitates efficient inference. Comprehensive experiments and user studies demonstrate that our method surpasses contemporary state-of-the-art approaches both qualitatively and quantitatively.
Haozhe Wu, Songtao Zhou, Jia Jia 0001, Junliang Xing
ACM Multimedia4
2023 Semantics2Hands: Transferring Hand Motion Semantics between Avatars
abstract
Human hands, the primary means of non-verbal communication, convey intricate semantics in various scenarios. Due to the high sensitivity of individuals to hand motions, even minor errors in hand motions can significantly impact the user experience. Real applications often involve multiple avatars with varying hand shapes, highlighting the importance of maintaining the intricate semantics of hand motions across the avatars. Therefore, this paper aims to transfer the hand motion semantics between diverse avatars based on their respective hand models. To address this problem, we introduce a novel anatomy-based semantic matrix (ASM) that encodes the semantics of hand motions. The ASM quantifies the positions of the palm and other joints relative to the local frame of the corresponding joint, enabling precise retargeting of hand motions. Subsequently, we obtain a mapping function from the source ASM to the target hand joint rotations by employing an anatomy-based semantics reconstruction network (ASRN). We train the ASRN using a semi-supervised learning strategy on the Mixamo and InterHand2.6M datasets. We evaluate our method in intra-domain and cross-domain hand motion retargeting tasks. The qualitative and quantitative results demonstrate the significant superiority of our ASRN over the state-of-the-arts. Code available at https://github.com/abcyzj/Semantics2Hand
Zijie Ye, Jia Jia 0001, Junliang Xing
ACM Multimedia3
2023 Automatic Grouping for Efficient Cooperative Multi-Agent Reinforcement Learning
abstract
Grouping is ubiquitous in natural systems and is essential for promoting efficiency in team coordination. This paper proposes a novel formulation of Group-oriented Multi-Agent Reinforcement Learning (GoMARL), which learns automatic grouping without domain knowledge for efficient cooperation. In contrast to existing approaches that attempt to directly learn the complex relationship between the joint action-values and individual utilities, we empower subgroups as a bridge to model the connection between small sets of agents and encourage cooperation among them, thereby improving the learning efficiency of the whole team. In particular, we factorize the joint action-values as a combination of group-wise values, which guide agents to improve their policies in a fine-grained fashion. We present an automatic grouping mechanism to generate dynamic groups and group action-values. We further introduce a hierarchical control for policy learning that drives the agents in the same group to specialize in similar policies and possess diverse strategies for various groups. Experiments on the StarCraft II micromanagement tasks and Google Research Football scenarios verify our method's effectiveness. Extensive component studies show how grouping works and enhances performance.
Yifan Zang 0001, Jinmin He, Kai Li 0022, Haobo Fu, Qiang Fu 0016, Junliang Xing, Jian Cheng 0001
NeurIPS6
2023 Attribute-guided transformer for robust person re-identification
abstract
Abstract Recent studies reveal the crucial role of local features in learning robust and discriminative representations for person re‐identification (Re‐ID). Existing approaches typically rely on external tasks, for example, semantic segmentation, or pose estimation, to locate identifiable parts of given images. However, they heuristically utilise the predictions from off‐the‐shelf models, which may be sub‐optimal in terms of both local partition and computational efficiency. They also ignore the mutual information with other inputs, which weakens the representation capabilities of local features. In this study, the authors put forward a novel Attribute‐guided Transformer (AiT), which explicitly exploits pedestrian attributes as semantic priors for discriminative representation learning. Specifically, the authors first introduce an attribute learning process, which generates a set of attention maps highlighting the informative parts of pedestrian images. Then, the authors design a Feature Diffusion Module (FDM) to iteratively inject attribute information into global feature maps, aiming at suppressing unnecessary noise and inferring attribute‐aware representations. Last, the authors propose a Feature Aggregation Module (FAM) to exploit mutual information for aggregating attribute characteristics from different images, enhancing the representation capabilities of feature embedding. Extensive experiments demonstrate the superiority of our AiT in learning robust and discriminative representations. As a result, the authors achieve competitive performance with state‐of‐the‐art methods on several challenging benchmarks without any bells and whistles.
Zhe Wang 0013, Jun Wang 0041, Junliang Xing
IET Comput. Vis.3
2023 Anti-UAV: A Large-Scale Benchmark for Vision-Based UAV Tracking
abstract
Unmanned Aerial Vehicles (UAV) have many applications in both commerce and recreation. However, irresponsibly operated UAVs will pose a threat to public safety. Therefore, developing our understanding of UAVs and their uses is of particular interest. This paper considers tracking UAVs, which provide multifaceted information around location, paths and trajectories. To facilitate research on this topic, we introduce a new benchmark, herein referred to as Anti-UAV, which provides a novel direction for UAV tracking with more than 300 video pairs containing over 580 k manually annotated bounding boxes. Addressing anti-UAV research challenges could help to design anti-UAV systems, which in turn may improve surveillance. Accordingly, we have proposed a simple yet effective approach, called dual-flow semantic consistency (DFSC) is proposed for UAV tracking. Modulated by the semantic flow across video sequences, tracker learns more robust class-level semantic information and obtains more discriminative instance-level features. Experiments highlight significant performance gain with the proposed approach over state-of-the-art trackers and the challenging aspects of Anti-UAV. The Anti-UAV benchmark and the code for the proposed approach have been made publicly available athttps://github.com/ucas-vg/Anti-UAVandhttps://github.com/ZhaoJ9014/Anti-UAV.
Kuiran Wang, Xiaoke Peng, Xuehui Yu, Qiang Wang 0051, Junliang Xing, Guorong Li, Guodong Guo, Qixiang Ye, Jianbin Jiao, Jian Zhao 0006, Zhenjun Han
IEEE Trans. Multim.6
2022 AutoCFR: Learning to Design Counterfactual Regret Minimization Algorithms
abstract
Counterfactual regret minimization (CFR) is the most commonly used algorithm to approximately solving two-player zero-sum imperfect-information games (IIGs). In recent years, a series of novel CFR variants such as CFR+, Linear CFR, DCFR have been proposed and have significantly improved the convergence rate of the vanilla CFR. However, most of these new variants are hand-designed by researchers through trial and error based on different motivations, which generally requires a tremendous amount of efforts and insights. This work proposes to meta-learn novel CFR algorithms through evolution to ease the burden of manual algorithm design. We first design a search language that is rich enough to represent many existing hand-designed CFR variants. We then exploit a scalable regularized evolution algorithm with a bag of acceleration techniques to efficiently search over the combinatorial space of algorithms defined by this language. The learned novel CFR algorithm can generalize to new IIGs not seen during training and performs on par with or better than existing state-of-the-art CFR variants. The code is available at https://github.com/rpSebastian/AutoCFR.
Hang Xu 0006, Kai Li 0022, Haobo Fu, Qiang Fu 0016, Junliang Xing
AAAI5
2022 AlphaHoldem: High-Performance Artificial Intelligence for Heads-Up No-Limit Poker via End-to-End Reinforcement Learning
abstract
Heads-up no-limit Texas hold’em (HUNL) is the quintessential game with imperfect information. Representative priorworks like DeepStack and Libratus heavily rely on counter-factual regret minimization (CFR) and its variants to tackleHUNL. However, the prohibitive computation cost of CFRiteration makes it difficult for subsequent researchers to learnthe CFR model in HUNL and apply it in other practical applications. In this work, we present AlphaHoldem, a high-performance and lightweight HUNL AI obtained with an end-to-end self-play reinforcement learning framework. The proposed framework adopts a pseudo-siamese architecture to directly learn from the input state information to the output actions by competing the learned model with its different historical versions. The main technical contributions include anovel state representation of card and betting information, amultitask self-play training loss function, and a new modelevaluation and selection metric to generate the final model.In a study involving 100,000 hands of poker, AlphaHoldemdefeats Slumbot and DeepStack using only one PC with threedays training. At the same time, AlphaHoldem only takes 2.9milliseconds for each decision-making using only a singleGPU, more than 1,000 times faster than DeepStack. We release the history data among among AlphaHoldem, Slumbot,and top human professionals in the author’s GitHub repository to facilitate further studies in this direction.
Enmin Zhao, Renye Yan, Jinqiu Li, Kai Li 0022, Junliang Xing
AAAI5
2022 Speedup Training Artificial Intelligence for Mahjong via Reward Variance Reduction
abstract
Despite significant breakthroughs in developing gaming artificial intelligence (AI), Mahjong remains quite challenging as a popular multi-player imperfect information game. Compared with games such as Go and Texas Hold’em, Mahjong has much more invisible information, unfixed game order, and a complicated scoring system, resulting in high randomness and variance of the rewarding signals during the reinforcement learning process. This paper presents a Mahjong AI by introducing Reward Variance Reduction (RVR) into a new self-play deep reinforcement learning algorithm. RVR handles the invisibility via a relative value network which leverages the global information to guide the model to converge to the optimal strategy under an oracle with perfect information. Moreover, RVR improves the training stability using an expected reward network to adapt to the complex, dynamic, and highly stochastic reward environment. Extensive experimental results show that RVR significantly reduces the variance in Mahjong AI training and improves the model performance. After only three days of self-play training on a single server with 8 GPUs, RVR defeats 62.5% opponents on the Botzone platform.
Jinqiu Li, Haobo Fu, Qiang Fu 0016, Enmin Zhao, Junliang Xing
CoG6
2022 Actor-Critic Policy Optimization in a Large-Scale Imperfect-Information Game
Haobo Fu, Weiming Liu 0004, Kai Li 0022, Junliang Xing, Bin Li 0025, Qiang Fu 0016, Wei Yang 0032
ICLR7
2022 Greedy when Sure and Conservative when Uncertain about the Opponents
abstract
We develop a new approach, named Greedy when Sure and Conservative when Uncertain (GSCU), to competing online against unknown and nonstationary opponents. GSCU improves in four aspects: 1) introduces a novel way of learning opponent policy embeddings offline; 2) trains offline a single best response (conditional additionally on our opponent policy embedding) instead of a finite set of separate best responses against any opponent; 3) computes online a posterior of the current opponent policy embedding, without making the discrete and ineffective decision which type the current opponent belongs to; and 4) selects online between a real-time greedy policy and a fixed conservative policy via an adversarial bandit algorithm, gaining a theoretically better regret than adhering to either. Experimental studies on popular benchmarks demonstrate GSCU’s superiority over the state-of-the-art methods. The code is available online at \url{https://github.com/YeTianJHU/GSCU}.
Haobo Fu, Hongxiang Yu, Weiming Liu 0004, Jiechao Xiong, Ying Wen 0001, Kai Li 0022, Junliang Xing, Qiang Fu 0016, Wei Yang 0032
ICML9
2022 Towards a Unified Benchmark for Reinforcement Learning in Sparse Reward Environments
Yongxin Kang, Enmin Zhao, Yifan Zang 0001, Kai Li 0022, Junliang Xing
ICONIP (4)5
2022 L2E: Learning to Exploit Your Opponent
abstract
Opponent modeling is essential to exploit sub-optimal opponents in strategic interactions. Most previous works focus on building explicit models to predict the opponents' styles or strategies, which require a large amount of data to train the model and lack adaptability to unknown opponents. In this work, we propose a novel Learning to Exploit (L2E) framework for implicit opponent modeling. L2E acquires the ability to exploit opponents through a few interactions with different opponents during training of a neural network and can quickly adapt to new opponents with unknown styles during testing. To automatically produce challenging and diverse opponents for training, we further present a novel opponent strategy generation algorithm. We evaluate L2E on two poker games and one grid soccer game, which are the commonly used benchmarks for opponent modeling. Comprehensive experimental results indicate that L2E rapidly adapts to diverse styles of unknown opponents.
Kai Li 0022, Hang Xu 0006, Yifan Zang 0001, Bo An 0001, Junliang Xing
IJCNN6
2022 GroupDancer: Music to Multi-People Dance Synthesis with Style Collaboration
abstract
Different people dance in different styles. So when multiple people dance together, the phenomenon of style collaboration occurs: people need to seek common points while reserving differences in various dancing periods. Thus, we introduce a novel Music-driven Group Dance Synthesis task. Compared with single-people dance synthesis explored by most previous works, modeling the style collaboration phenomenon and choreographing for multiple people are more complicated and challenging. Moreover, the lack of sufficient records for conducting multi-people choreography in prior datasets further aggravates this problem. To address these issues, we construct a rich-annotated 3D Multi-Dancer Choreography dataset (MDC) and newly devise a metric SCEU for style collaboration evaluation. To our best knowledge, MDC is the first 3D dance dataset that collects both individual and collaborated music-dance pairs. Based on MDC, we present a novel framework, GroupDancer, consisting of three stages: Dancer Collaboration, Motion Choreography and Motion Transition. The Dancer Collaboration stage determines when and which dancers should collaborate their dancing styles from music. Afterward, the Motion Choreography stage produces a motion sequence for each dancer. Finally, the Motion Transition stage fills the gaps between the motions to achieve fluent and natural group dance. To make GroupDancer trainable from end to end and able to synthesize group dance with style collaboration, we propose mixed training and selective updating strategies. Comprehensive evaluations on the MDC dataset demonstrate that the proposed GroupDancer model can synthesize quite satisfactory group dance synthesis results with style collaboration.
Zixuan Wang 0026, Jia Jia 0001, Haozhe Wu, Junliang Xing, Jinghe Cai, Guowen Chen
ACM Multimedia4
2022 Robust Face Alignment via Deep Progressive Reinitialization and Adaptive Error-Driven Learning
abstract
Regression-based face alignment involves learning a series of mapping functions to predict the true landmarks from an initial estimation of the alignment. Most existing approaches focus on learning efficacious mapping functions from some feature representations to improve performance. The issues related to the initial alignment estimation and the final learning objective, however, receive less attention. This work proposes a deep regression architecture with progressive reinitialization and a new error-driven learning loss function to explicitly address the above two issues. Given an image with a rough face detection result, the full face region is first mapped by a supervised spatial transformer network to a normalized form and trained to regress coarse positions of landmarks. Then, different face parts are further respectively reinitialized to their own normalized states, followed by another regression sub-network to refine the landmark positions. To deal with the inconsistent annotations in existing training datasets, we further propose an adaptive landmark-weighted loss function. It dynamically adjusts the importance of different landmarks according to their learning errors during training without depending on any hyper-parameters manually set by trial and error. A high level of robustness to annotation inconsistencies is thus achieved. The whole deep architecture permits training from end to end, and extensive experimental analyses and comparisons demonstrate its effectiveness and efficiency. The source code, trained models, and experimental results are made available at https://github.com/shaoxiaohu/Face_Alignment_DPR.git.
Xiaohu Shao, Junliang Xing, Jiangjing Lyu, Yu Shi 0003, Stephen J. Maybank
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 A Coulomb Force Inspired Loss Function for High-Performance Pedestrian Detection
abstract
Pedestrian detection has received considerable research interest due to its wide application and has made significant progress along with the development of deep neural networks. However, crowd occlusion still remains a significant challenge to current state-of-the-art pedestrian detectors due to the complication in formulating interactions between occluded instances. Inspired by the Coulomb force, we in this work set each proposal as a single electric charge and define the attractive and repulsive forces to model the interaction between ground truths and assigned proposals. This design is driven by two motivations: the attractive force pulls bounding boxes toward their assigned targets, aggregating them compactly around the ground truths. The repulsive force pushes bounding boxes away from other instances, preventing them from shifting to surrounding pedestrians. With this insight, we propose a novel bounding box regression loss and achieve more robust localization performance in crowded scenes without introducing any computational overhead. Extensive experimental evaluations on the CityPersons and CrowdHuman benchmarks demonstrate consistent state-of-the-art performance.
Zhe Wang 0013, Jun Wang 0041, Yezhou Yang, Junliang Xing
IEEE Signal Process. Lett.4
2022 Contrastive Context-Aware Learning for 3D High-Fidelity Mask Face Presentation Attack Detection
abstract
Face presentation attack detection (PAD) is essential to secure face recognition systems primarily from high-fidelity mask attacks. Most existing 3D mask PAD benchmarks suffer from several drawbacks: 1) a limited number of mask identities, types of sensors, and a total number of videos; 2) low-fidelity quality of facial masks. Basic deep models and remote photoplethysmography (rPPG) methods achieved acceptable performance on these benchmarks but still far from the needs of practical scenarios. To bridge the gap to real-world applications, we introduce a large-scale High-Fidelity Mask dataset, namely HiFiMask. Specifically, a total amount of 54,600 videos are recorded from 75 subjects with 225 realistic masks by 7 new kinds of sensors. Along with the dataset, we propose a novel Contrastive Context-aware Learning (CCL) framework. CCL is a new training methodology for supervised PAD tasks, which is able to learn by leveraging rich contexts accurately (e.g., subjects, mask material and lighting) among pairs of live faces and high-fidelity mask attacks. Extensive experimental evaluations on HiFiMask and three additional 3D mask datasets demonstrate the effectiveness of our method. The codes and dataset will be released soon.
Ajian Liu 0001, Zitong Yu, Jun Wan 0001, Anyang Su, Zichang Tan, Sergio Escalera, Junliang Xing, Yanyan Liang 0001, Guodong Guo, Zhen Lei 0001, Stan Z. Li
IEEE Trans. Inf. Forensics Secur.9
2021 Exploration via State influence Modeling
abstract
This paper studies the challenging problem of reinforcement learning (RL) in hard exploration tasks with sparse rewards. It focuses on the exploration stage before the agent gets the first positive reward, in which case, traditional RL algorithms with simple exploration strategies often work poorly. Unlike previous methods using some attribute of a single state as the intrinsic reward to encourage exploration, this work leverages the social influence between different states to permit more efficient exploration. It introduces a general intrinsic reward construction method to evaluate the social influence of states dynamically. Three kinds of social influence are introduced for a state: conformity, power, and authority. By measuring the state’s social influence, agents quickly find the focus state during the exploration process. The proposed RL framework with state social influence evaluation works well in hard exploration task. Extensive experimental analyses and comparisons in Grid Maze and many hard exploration Atari 2600 games demonstrate its high exploration efficiency.
Yongxin Kang, Enmin Zhao, Kai Li 0022, Junliang Xing
AAAI4
2021 Multi-View Face Recognition Using Deep Attention-Based Face Frontalization
abstract
Face frontalization has been widely used in face recognition to alleviate distribution discrepancy between multi-view faces. Given a profile face, existing models learn to synthesize a frontal face from the whole region indistinguishably, often resulting in unsatisfactory frontalization caused by a lack of synthetic focus and disturbances of trivial backgrounds. This paper proposes a novel Deep Attention-based Face Frontalization (DAFF) method to address the above issues explicitly. We first inject the 3D spatial prior of the input face into an encoder-decoder model. This process locates the discriminative foreground for decomposing meaningful convolutional embeddings. After that, we propose a novel objective that served as the generator’s geometric guidance to pay more attention to the target’s essential regions. Therefore, we can leverage the attentional constraints to perform recovery refinement at both embedding and texture levels. Extensive experiments show that DAFF achieves satisfactory frontalization and competitive recognition performance under constrained and in-the-wild benchmarks.
Xiaohu Shao, Junliang Xing, Ruihan Pan, Yu Shi 0003
ICME2
2021 Learning to Play Hard Exploration Games Using Graph-Guided Self-Navigation
abstract
This work considers the problem of deep reinforcement learning (RL) with long time dependencies and sparse rewards, as are found in many hard exploration games. A graph-based representation is proposed to allow an agent to perform self-navigation for environmental exploration. The graph representation not only effectively models the environment structure, but also efficiently traces the agent state changes and the corresponding actions. By encouraging the agent to earn a new influence-based curiosity reward for new game observations, the whole exploration task is divided into sub-tasks, which are effectively solved using a unified deep RL model. Experimental evaluations on hard exploration Atari Games demonstrate the effectiveness of the proposed method. The source code and learned models will be released to facilitate further studies on this problem.
Enmin Zhao, Renye Yan, Kai Li 0022, Lijuan Li 0002, Junliang Xing
IJCNN5
2021 Multiple object tracking: A literature review
Wenhan Luo, Junliang Xing, Anton Milan, Xiaoqin Zhang 0002, Wei Liu 0005, Tae-Kyun Kim 0001
Artif. Intell.2
2020 PedHunter: Occlusion Robust Pedestrian Detector in Crowded Scenes
abstract
Pedestrian detection in crowded scenes is a challenging problem, because occlusion happens frequently among different pedestrians. In this paper, we propose an effective and efficient detection network to hunt pedestrians in crowd scenes. The proposed method, namely PedHunter, introduces strong occlusion handling ability to existing region-based detection networks without bringing extra computations in the inference stage. Specifically, we design a mask-guided module to leverage the head information to enhance the feature representation learning of the backbone network. Moreover, we develop a strict classification criterion by improving the quality of positive samples during training to eliminate common false positives of pedestrian detection in crowded scenes. Besides, we present an occlusion-simulated data augmentation to enrich the pattern and quantity of occlusion samples to improve the occlusion robustness. As a consequent, we achieve state-of-the-art results on three pedestrian detection datasets including CityPersons, Caltech-USA and CrowdHuman. To facilitate further studies on the occluded pedestrian detection in surveillance scenes, we release a new pedestrian dataset, called SUR-PED, with a total of over 162k high-quality manually labeled instances in 10k images. The proposed dataset, source codes and trained models are available at https://github.com/ChiCheng123/PedHunter.
Cheng Chi 0003, Junliang Xing, Zhen Lei 0001, Stan Z. Li
AAAI3
2020 Relational Learning for Joint Head and Human Detection
abstract
Head and human detection have been rapidly improved with the development of deep convolutional neural networks. However, these two tasks are often studied separately without considering their inherent correlation, leading to that 1) head detection is often trapped in more false positives, and 2) the performance of human detector frequently drops dramatically in crowd scenes. To handle these two issues, we present a novel joint head and human detection network, namely JointDet, which effectively detects head and human body simultaneously. Moreover, we design a head-body relationship discriminating module to perform relational learning between heads and human bodies, and leverage this learned relationship to regain the suppressed human detections and reduce head false positives. To verify the effectiveness of the proposed method, we annotate head bounding boxes of the CityPersons and Caltech-USA datasets, and conduct extensive experiments on the CrowdHuman, CityPersons and Caltech-USA datasets. As a consequence, the proposed JointDet detector achieves state-of-the-art performance on these three benchmarks. To facilitate further studies on the head and human detection problem, all new annotations, source codes and trained models are available at https://github.com/ChiCheng123/JointDet.
Cheng Chi 0003, Junliang Xing, Zhen Lei 0001, Stan Z. Li
AAAI3
2020 Domain Adaptive Attention Learning for Unsupervised Person Re-Identification
abstract
Person re-identification (Re-ID) across multiple datasets is a challenging task due to two main reasons: the presence of large cross-dataset distinctions and the absence of annotated target instances. To address these two issues, this paper proposes a domain adaptive attention learning approach to reliably transfer discriminative representation from the labeled source domain to the unlabeled target domain. In this approach, a domain adaptive attention model is learned to separate the feature map into domain-shared part and domain-specific part. In this manner, the domain-shared part is used to capture transferable cues that can compensate cross-dataset distinctions and give positive contributions to the target task, while the domain-specific part aims to model the noisy information to avoid the negative transfer caused by domain diversity. A soft label loss is further employed to take full use of unlabeled target data by estimating pseudo labels. Extensive experiments on the Market-1501, DukeMTMC-reID and MSMT17 benchmarks demonstrate the proposed approach outperforms the state-of-the-arts.
Yangru Huang, Peixi Peng, Yi Jin 0001, Yidong Li, Junliang Xing
AAAI5
2020 RDSNet: A New Deep Architecture forReciprocal Object Detection and Instance Segmentation
abstract
Object detection and instance segmentation are two fundamental computer vision tasks. They are closely correlated but their relations have not yet been fully explored in most previous work. This paper presents RDSNet, a novel deep architecture for reciprocal object detection and instance segmentation. To reciprocate these two tasks, we design a two-stream structure to learn features on both the object level (i.e., bounding boxes) and the pixel level (i.e., instance masks) jointly. Within this structure, information from the two streams is fused alternately, namely information on the object level introduces the awareness of instance and translation variance to the pixel level, and information on the pixel level refines the localization accuracy of objects on the object level in return. Specifically, a correlation module and a cropping module are proposed to yield instance masks, as well as a mask based boundary refinement module for more accurate bounding boxes. Extensive experimental analyses and comparisons on the COCO dataset demonstrate the effectiveness and efficiency of RDSNet. The source code is available at https://github.com/wangsr126/RDSNet.
Shaoru Wang, Yongchao Gong, Junliang Xing, Lichao Huang, Chang Huang, Weiming Hu 0004
AAAI3
2020 Semantics-Guided Neural Networks for Efficient Skeleton-Based Human Action Recognition
abstract
Skeleton-based human action recognition has attracted great interest thanks to the easy accessibility of the human skeleton data. Recently, there is a trend of using very deep feedforward neural networks to model the 3D coordinates of joints without considering the computational efficiency. In this paper, we propose a simple yet effective semantics-guided neural network (SGN) for skeleton-based action recognition. We explicitly introduce the high level semantics of joints (joint type and frame index) into the network to enhance the feature representation capability. In addition, we exploit the relationship of joints hierarchically through two modules, i.e., a joint-level module for modeling the correlations of joints in the same frame and a framelevel module for modeling the dependencies of frames by taking the joints in the same frame as a whole. A strong baseline is proposed to facilitate the study of this field. With an order of magnitude smaller model size than most previous works, SGN achieves the state-of-the-art performance on the NTU60, NTU120, and SYSU datasets.
Pengfei Zhang 0005, Cuiling Lan, Wenjun Zeng 0001, Junliang Xing, Jianru Xue, Nanning Zheng 0001
CVPR4
2020 Hybrid Learning for Multi-agent Cooperation with Sub-optimal Demonstrations
abstract
This paper aims to learn multi-agent cooperation where each agent performs its actions in a decentralized way. In this case, it is very challenging to learn decentralized policies when the rewards are global and sparse. Recently, learning from demonstrations (LfD) provides a promising way to handle this challenge. However, in many practical tasks, the available demonstrations are often sub-optimal. To learn better policies from these sub-optimal demonstrations, this paper follows a centralized learning and decentralized execution framework and proposes a novel hybrid learning method based on multi-agent actor-critic. At first, the expert trajectory returns generated from demonstration actions are used to pre-train the centralized critic network. Then, multi-agent decisions are made by best response dynamics based on the critic and used to train the decentralized actor networks. Finally, the demonstrations are updated by the actor networks, and the critic and actor networks are learned jointly by running the above two steps alliteratively. We evaluate the proposed approach on a real-time strategy combat game. Experimental results show that the approach outperforms many competing demonstration-based methods.
Peixi Peng, Junliang Xing, Lili Cao
IJCAI2
2020 Potential Driven Reinforcement Learning for Hard Exploration Tasks
abstract
Experience replay plays a crucial role in Reinforcement Learning (RL), enabling the agent to remember and reuse experience from the past. Most previous methods sample experience transitions using simple heuristics like uniformly sampling or prioritizing those good ones. Since humans can learn from both good and bad experiences, more sophisticated experience replay algorithms need to be developed. Inspired by the potential energy in physics, this work introduces the artificial potential field into experience replay and develops Potentialized Experience Replay (PotER) as a new and effective sampling algorithm for RL in hard exploration tasks with sparse rewards. PotER defines a potential energy function for each state in experience replay and helps the agent to learn from both good and bad experiences using intrinsic state supervision. PotER can be combined with different RL algorithms as well as the self-imitation learning algorithm. Experimental analyses and comparisons on multiple challenging hard exploration environments have verified its effectiveness and efficiency.
Enmin Zhao, Shihong Deng, Yifan Zang 0001, Yongxin Kang, Kai Li 0022, Junliang Xing
IJCAI6
2020 Anchor-Free One-Stage Online Multi-object Tracking
Zongwei Zhou, Yangxi Li, Junliang Xing, Liang Li 0003, Weiming Hu 0004
PRCV (2)4
2020 Dual L1-Normalized Context Aware Tensor Power Iteration and Its Applications to Multi-object Tracking and Multi-graph Matching
abstract
Abstract The multi-dimensional assignment problem is universal for data association analysis such as data association-based visual multi-object tracking and multi-graph matching. In this paper, multi-dimensional assignment is formulated as a rank-1 tensor approximation problem. A dualL1-normalized context/hyper-context aware tensor power iteration optimization method is proposed. The method is applied to multi-object tracking and multi-graph matching. In the optimization method, tensor power iteration with the dual unit norm enables the capture of information across multiple sample sets. Interactions between sample associations are modeled as contexts or hyper-contexts which are combined with the global affinity into a unified optimization. The optimization is flexible for accommodating various types of contextual models. In multi-object tracking, the global affinity is defined according to the appearance similarity between objects detected in different frames. Interactions between objects are modeled as motion contexts which are encoded into the global association optimization. The tracking method integrates high order motion information and high order appearance variation. The multi-graph matching method carries out matching over graph vertices and structure matching over graph edges simultaneously. The matching consistency across multi-graphs is based on the high-order tensor optimization. Various types of vertex affinities and edge/hyper-edge affinities are flexibly integrated. Experiments on several public datasets, such as the MOT16 challenge benchmark, validate the effectiveness of the proposed methods.
Weiming Hu 0004, Xinchu Shi, Zongwei Zhou, Junliang Xing, Haibin Ling, Stephen J. Maybank
Int. J. Comput. Vis.4
2020 Recognizing Profile Faces by Imagining Frontal View
Jian Zhao 0006, Junliang Xing, Shuicheng Yan, Jiashi Feng
Int. J. Comput. Vis.2
2020 Tracking-by-Fusion via Gaussian Process Regression Extended to Transfer Learning
abstract
This paper presents a new Gaussian Processes (GPs)-based particle filter tracking framework. The framework non-trivially extends Gaussian process regression (GPR) to transfer learning, and, following the tracking-by-fusion strategy, integrates closely two tracking components, namely a GPs component and a CFs one. First, the GPs component analyzes and models the probability distribution of the object appearance by exploiting GPs. It categorizes the labeled samples into auxiliary and target ones, and explores unlabeled samples in transfer learning. The GPs component thus captures rich appearance information over object samples across time. On the other hand, to sample an initial particle set in regions of high likelihood through the direct simulation method in particle filtering, the powerful yet efficient correlation filters (CFs) are integrated, leading to the CFs component. In fact, the CFs component not only boosts the sampling quality, but also benefits from the GPs component, which provides re-weighted knowledge as latent variables for determining the impact of each correlation filter template from the auxiliary samples. In this way, the transfer learning based fusion enables effective interactions between the two components. Superior performance on four object tracking benchmarks (OTB-2015, Temple-Color, and VOT2015/2016), and in comparison with baselines and recent state-of-the-art trackers, has demonstrated clearly the effectiveness of the proposed framework.
Qiang Wang 0051, Junliang Xing, Haibin Ling, Weiming Hu 0004, Stephen J. Maybank
IEEE Trans. Pattern Anal. Mach. Intell.3
2020 Distractor-aware discrimination learning for online multiple object tracking
Zongwei Zhou, Wenhan Luo, Qiang Wang 0051, Junliang Xing, Weiming Hu 0004
Pattern Recognit.4
2020 Temporal-Spatial Mapping for Action Recognition
abstract
Deep learning models have enjoyed great success for image related computer vision tasks such as image classification and object detection. For video related tasks such as human action recognition, however, the advancements are not as significant yet. The main challenge is the lack of effective and efficient models in modeling the rich temporal-spatial information in a video. We introduce a simple yet effective operation, termed temporal-spatial mapping, for capturing the temporal evolution of the frames by jointly analyzing all the frames of a video. We propose a video level 2D feature representation by transforming the convolutional features of all frames to a 2D feature map, referred to as VideoMap. With each row being the vectorized feature representation of a frame, the temporal-spatial features are compactly represented, while the temporal dynamic evolution is also well embedded. Based on the VideoMap representation, we further propose a temporal attention model within a shallow convolutional neural network to efficiently exploit the temporal-spatial dynamics. The experiment results show that the proposed scheme achieves state-of-the-art performance, with 4.2% accuracy gain over the temporal segment network, a competing baseline method, on the challenging human action benchmark dataset HMDB51.
Cuiling Lan, Wenjun Zeng 0001, Junliang Xing, Xiaoyan Sun 0001, Jing-Yu Yang 0002
IEEE Trans. Circuits Syst. Video Technol.4
2020 Deep Spatial and Temporal Network for Robust Visual Object Tracking
abstract
There are two key components that can be leveraged for visual tracking: (a) object appearances; and (b) object motions. Many existing techniques have recently employed deep learning to enhance visual tracking due to its superior representation power and strong learning ability, where most of them employed object appearances but few of them exploited object motions. In this work, a deep spatial and temporal network (DSTN) is developed for visual tracking by explicitly exploiting both the object representations from each frame and their dynamics along multiple frames in a video, such that it can seamlessly integrate the object appearances with their motions to produce compact object appearances and capture their temporal variations effectively. Our DSTN method, which is deployed into a tracking pipeline in a coarse-to-fine form, can perceive the subtle differences on spatial and temporal variations of the target (object being tracked), and thus it benefits from both off-line training and online fine-tuning. We have also conducted our experiments over four largest tracking benchmarks, including OTB-2013, OTB-2015, VOT2015, and VOT2017, and our experimental results have demonstrated that our DSTN method can achieve competitive performance as compared with the state-of-the-art techniques. The source code, trained models, and all the experimental results of this work will be made public available to facilitate further studies on this problem.
Zhu Teng, Junliang Xing, Qiang Wang 0051, Baopeng Zhang, Jianping Fan 0001
IEEE Trans. Image Process.2
2020 Spatial Preserved Graph Convolution Networks for Person Re-identification
abstract
Person Re-identification is a very challenging task due to inter-class ambiguity caused by similar appearances, and large intra-class diversity caused by viewpoints, illuminations, and poses. To address these challenges, in this article, a graph convolution network based model for person re-identification is proposed to learn more discriminative feature embeddings, where a graph-structured relationship between person images and person parts are together integrated. Graph convolution networks extract common characteristics of the same person, while pyramid feature embedding exploits parts relations and learns stable representation with each person image. We achieve a very competitive performance respectively on three widely used datasets, indicating that the proposed approach significantly outperforms the baseline methods and achieves the state-of-the-art performance.
Zhaoju Li, Zongwei Zhou, Zhenjun Han, Junliang Xing, Jianbin Jiao
ACM Trans. Multim. Comput. Commun. Appl.5
2019 Selective Refinement Network for High Performance Face Detection
abstract
High performance face detection remains a very challenging problem, especially when there exists many tiny faces. This paper presents a novel single-shot face detector, named Selective Refinement Network (SRN), which introduces novel twostep classification and regression operations selectively into an anchor-based face detector to reduce false positives and improve location accuracy simultaneously. In particular, the SRN consists of two modules: the Selective Two-step Classification (STC) module and the Selective Two-step Regression (STR) module. The STC aims to filter out most simple negative anchors from low level detection layers to reduce the search space for the subsequent classifier, while the STR is designed to coarsely adjust the locations and sizes of anchors from high level detection layers to provide better initialization for the subsequent regressor. Moreover, we design a Receptive Field Enhancement (RFE) block to provide more diverse receptive field, which helps to better capture faces in some extreme poses. As a consequence, the proposed SRN detector achieves state-of-the-art performance on all the widely used face detection benchmarks, including AFW, PASCAL face, FDDB, and WIDER FACE datasets. Codes will be released to facilitate further studies on the face detection problem.
Cheng Chi 0003, Junliang Xing, Zhen Lei 0001, Stan Z. Li
AAAI3
2019 SiamRPN++: Evolution of Siamese Visual Tracking With Very Deep Networks
abstract
Siamese network based trackers formulate tracking as convolutional feature cross-correlation between target template and searching region. However, Siamese trackers still have accuracy gap compared with state-of-the-art algorithms and they cannot take advantage of feature from deep networks, such as ResNet-50 or deeper. In this work we prove the core reason comes from the lack of strict translation invariance. By comprehensive theoretical analysis and experimental validations, we break this restriction through a simple yet effective spatial aware sampling strategy and successfully train a ResNet-driven Siamese tracker with significant performance gain. Moreover, we propose a new model architecture to perform depth-wise and layer-wise aggregations, which not only further improves the accuracy but also reduces the model size. We conduct extensive ablation studies to demonstrate the effectiveness of the proposed tracker, which obtains currently the best results on four large tracking benchmarks, including OTB2015, VOT2018, UAV123, and LaSOT. Our model will be released to facilitate further studies based on this problem.
Bo Li 0114, Wei Wu 0021, Qiang Wang 0051, Fangyi Zhang, Junliang Xing
CVPR5
2019 Spatial Temporal Attentional Glimpse for Human Activity Classification in Video
abstract
Recently, the Convolutional Networks (ConvNet) has become the dominated approach to the human activity classification problem. We investigate current standard ConvNet architectures and pinpoint one of their main limitations: the spatial-temporal dependency is simply captured by global pooling operation, which may not well capture the complex long term spatial-temporal relationships in videos. For this work, we propose a Spatial Temporal Attentional Glimpse (STAG) module to overcome this shortcoming. Specifically, the input to this STAG module is a 3D tensor which is first processed by a spatial-temporal attention block. Spatial Temporal Glimpse block decomposes the resulting tensor into two low dimensional tensors and then fuses their operation results. The proposed STAG module is pluggable, easy to learn, and effective in computation. We conduct extended ablation studies to show that our model incorporated with the STAG block substantially improves the performance over the state-of-the-art. All the experimental results, the trained models, and the complete source codes will be released to facilitate further studies on this problem.
Jiangtao Kong, Rongchao Xu, Junliang Xing, Kai Li 0022
ICIP3
2019 Vehicle Re-Identification by Multi-Grain Learni
abstract
Vehicle re-identification (re-ID) is to identify the same vehicle captured by different cameras with non-overlapping views, which plays an important role in intelligent transportation system and traffic safety. Compared with face recognition and person re-ID tasks, it is difficult to train an effective vehicle re-ID model since different vehicles of the same vehicle model, such as Mercedes-Benz C300, may exhibit strong inter-class similarity. To handle this difficulty, we propose a multi-grain ranking loss with the auxiliary of vehicle model, which models the vehicle re-ID task as two sub-tasks with different granularities including matching vehicles in the same vehicle model and different vehicle models. In particular, we infer the vehicle model labels online for the unlabeled training samples by clustering. The experimental results on two benchmarks demonstrate the proposed method can achieve state-of-the-art performance.
Xiaoliang Yang, Congyan Lang, Peixi Peng, Junliang Xing
ICIP4
2019 Multi-View Learning for Vehicle Re-Identification
abstract
Vehicle re-identification (ReID) aims to identify a target vehicle in different cameras with non-overlapping views, and it plays an important role when the car licence plate recognition is unavailable or unreliable. Compared with face recognition and person ReID tasks, it is difficult to train an effective vehicle ReID model due to two reasons: the different views greatly affect the visual appearance of a vehicle, and different vehicles may exhibit fairly similar visual appearance when their images are captured from one unified single view. To handle these training difficulties, we introduce several latent groups to represent multiple views. Then, the vehicle ReID problem is modeled as two sub tasks, including matching vehicles in a same view and across different views. A fine-grain ranking loss and a relative coarse-grain ranking loss are proposed to each task respectively. Extensive experimental analyses and evaluations on two benchmarks demonstrate the proposed method can achieve state-of-the-art performance.
Weipeng Lin, Yidong Li, Xiaoliang Yang, Peixi Peng, Junliang Xing
ICME5
2019 Multi-Prototype Networks for Unconstrained Set-based Face Recognition
abstract
In this paper, we address the challenging unconstrained set-based face recognition problem where each subject face is instantiated by a set of media (images and videos) instead of a single image. Naively aggregating information from all the media within a set would suffer from the large intra-set variance caused by heterogeneous factors (e.g., varying media modalities, poses and illumination) and fail to learn discriminative face representations. A novel Multi-Prototype Network (MP- Net) model is thus proposed to learn multiple prototype face representations adaptively from the media sets. Each learned prototype is representative for the subject face under certain condition in terms of pose, illumination and media modality. Instead of handcrafting the set partition for prototype learn- ing, MPNet introduces a Dense SubGraph (DSG) learning sub-net that implicitly untangles inconsistent media and learns a number of representative prototypes. Qualitative and quantitative experiments clearly demonstrate the superiority of the proposed model over state-of-the-arts.
Jian Zhao 0006, Jianshu Li, Xiaoguang Tu, Fang Zhao 0006, Yuan Xin, Junliang Xing, Hengzhu Liu, Shuicheng Yan, Jiashi Feng
IJCAI6
2019 Learning Deep Decentralized Policy Network by Collective Rewards for Real-Time Combat Game
abstract
The task of real-time combat game is to coordinate multiple units to defeat their enemies controlled by the given opponent in a real-time combat scenario. It is difficult to design a high-level Artificial Intelligence (AI) program for such a task due to its extremely large state-action space and real-time requirements. This paper formulates this task as a collective decentralized partially observable Markov decision process, and designs a Deep Decentralized Policy Network (DDPN) to model the polices. To train DDPN effectively, a novel two-stage learning algorithm is proposed which combines imitation learning from opponent and reinforcement learning by no-regret dynamics. Extensive experimental results on various combat scenarios indicate that proposed method can defeat different opponent models and significantly outperforms many state-of-the-art approaches.
Peixi Peng, Junliang Xing, Lili Cao, Lisen Mu, Chang Huang
IJCAI2
2019 Realtime Human Segmentation in Video
Tairan Zhang, Congyan Lang, Junliang Xing
MMM (2)3
2019 MFAD: A Multi-modality Face Anti-spoofing Dataset
Bingqian Geng, Congyan Lang, Junliang Xing, Songhe Feng, Jun Wu 0007
PRICAI (2)3
2019 Global and Local Deep Feature Representation Fusion for Vehicle Re-Identification
abstract
This paper introduces our submission to the Grand Challenges on Vehicle Re-Identification (ReID) held in the VCIP 2019. Vehicle Re-Identification, which aims to retrieve images of a query vehicle from a large-scale vehicle database, is of great significance to the urban security and city management. Although significant progress has been made in the last decade, vehicle ReID in the wild remains a very challenging problem due to the large intra-class variations of one vehicle instance from viewpoint, illumination, occlusion patterns, and the possibly small inter-class differentiation of different vehicle instances (e.g., two different vehicles of the same model and color). To deal with these challenges, we in this work present a vehicle ReID framework that integrates both the global visual cues along with the local part-based cues to learn discriminative feature representations. In addition, the proposed framework also makes use of the extra information such as brands, models and colors to further improve the performance. Experimental results are performed on the VCIP 2019 VehicleReID dataset and the proposed framework achieved the second place in the competition.
Xing Zhao Lee, Jiangtao Kong, Chi Su, Junliang Xing, Shengmei Shen
VCIP5
2019 Joint Learning of Dictionary and Convolutional Network for Pedestrian Attribute Recognition
abstract
Pedestrian attribute recognition is to predict the presence of a set of attributes from a given image, and it plays an important role in video surveillance applications. Most existing works model the task as a multi-label classification problem. Although effective, they ignore the existence of correlations among attributes. In this work, to learn multiple attributes jointly, the attributes are modeled as a subspace and a dictionary is introduced to represent the subspace. Furthermore, to extract the convolutional features which are more suitable for attribute prediction, the dictionary is modeled as a network layer which is learned jointly with the convolutional network. Finally, a novel learning algorithm is proposed to optimize the dictionary and the convolutional network corporately. Extensive experimental analyses and evaluations on two largest pedestrian attribute benchmarks PETA and PA-100K demonstrate that the proposed method achieves state-of-the-art performance.
Yan Sha, Congyan Lang, Peixi Peng, Junliang Xing, Danxia Li
VCIP4
2019 Rank-1 Tensor Approximation for High-Order Association in Multi-target Tracking
Xinchu Shi, Haibin Ling, Weiming Hu 0004, Peng Chu, Junliang Xing
Int. J. Comput. Vis.6
2019 View Adaptive Neural Networks for High Performance Skeleton-Based Human Action Recognition
abstract
Skeleton-based human action recognition has recently attracted increasing attention thanks to the accessibility and the popularity of 3D skeleton data. One of the key challenges in action recognition lies in the large variations of action representations when they are captured from different viewpoints. In order to alleviate the effects of view variations, this paper introduces a novel view adaptation scheme, which automatically determines the virtual observation viewpoints over the course of an action in a learning based data driven manner. Instead of re-positioning the skeletons using a fixed human-defined prior criterion, we design two view adaptive neural networks, i.e., VA-RNN and VA-CNN, which are respectively built based on the recurrent neural network (RNN) with the Long Short-term Memory (LSTM) and the convolutional neural network (CNN). For each network, a novel view adaptation module learns and determines the most suitable observation viewpoints, and transforms the skeletons to those viewpoints for the end-to-end recognition with a main classification network. Ablation studies find that the proposed view adaptive models are capable of transforming the skeletons of various views to much more consistent virtual viewpoints. Therefore, the models largely eliminate the influence of the viewpoints, enabling the networks to focus on the learning of action-specific features and thus resulting in superior performance. In addition, we design a two-stream scheme (referred to as VA-fusion) that fuses the scores of the two networks to provide the final prediction, obtaining enhanced performance. Moreover, random rotation of skeleton sequences is employed to improve the robustness of view adaptation models and alleviate overfitting during training. Extensive experimental evaluations on five challenging benchmarks demonstrate the effectiveness of the proposed view-adaptive networks and superior performance over state-of-the-art approaches.
Pengfei Zhang 0005, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jianru Xue, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 3D-Aided Dual-Agent GANs for Unconstrained Face Recognition
abstract
Synthesizing realistic profile faces is beneficial for more efficiently training deep pose-invariant models for large-scale unconstrained face recognition, by augmenting the number of samples with extreme poses and avoiding costly annotation work. However, learning from synthetic faces may not achieve the desired performance due to the discrepancy betwedistributions of the synthetic and real face images. To narrow this gap, we propose a Dual-Agent Generative Adversarial Network (DA-GAN) model, which can improve the realism of a face simulator's output using unlabeled real faces while preserving the identity information during the realism refinement. The dual agents are specially designed for distinguishing real versus fake and identities simultaneously. In particular, we employ an off-the-shelf 3D face model as a simulator to generate profile face images with varying poses. DA-GAN leverages a fully convolutional network as the generator to generate high-resolution images and an auto-encoder as the discriminator with the dual agents. Besides the novel architecture, we make several key modifications to the standard GAN to preserve pose, texture as well as identity, and stabilize the training process: (i) a pose perception loss; (ii) an identity perception loss; (iii) an adversarial loss with a boundary equilibrium regularization term. Experimental results show that DA-GAN not only achieves outstanding perceptual results but also significantly outperforms state-of-the-arts on the large-scale and challenging NIST IJB-A and CFP unconstrained face recognition benchmarks. In addition, the proposed DA-GAN is also a promising new approach for solving generic transfer learning problems more effectively. DA-GAN is the foundation of our winning entry to the NIST IJB-A face recognition competition in which we secured the $1^{st}$ places on the tracks of verification and identification.
Jian Zhao 0006, Jianshu Li, Junliang Xing, Shuicheng Yan, Jiashi Feng
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 High performance person re-identification via a boosting ranking ensemble
Zhaoju Li, Zhenjun Han, Junliang Xing, Qixiang Ye, Xuehui Yu, Jianbin Jiao
Pattern Recognit.3
2019 Asymmetric 3D Convolutional Neural Networks for action recognition
Hao Yang 0010, Chunfeng Yuan, Bing Li 0001, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank
Pattern Recognit.5
2019 Multi-Modality Multi-Task Recurrent Neural Network for Online Action Detection
abstract
Online action detection is a brand new challenge and plays a critical role in visual surveillance analytics. It goes one step further than a conventional action recognition task, which recognizes human actions from well-segmented clips. Online action detection is desired to identify the action type and localize action positions on the fly from the untrimmed stream data. In this paper, we propose a multi-modality multi-task recurrent neural network, which incorporates both RGB and Skeleton networks. We design different temporal modeling networks to capture specific characteristics from various modalities. Then, a deep long short-term memory subnetwork is utilized effectively to capture the complex long-range temporal dynamics, naturally avoiding the conventional sliding window design and thus ensuring high computational efficiency. Constrained by a multi-task objective function in the training phase, this network achieves superior detection performance and is capable of automatically localizing the start and end points of actions more accurately. Furthermore, embedding subtask of regression provides the ability to forecast the action prior to its occurrence. We evaluate the proposed method and several other methods in action detection and forecasting on the online action detection data set and gaming action data set datasets. Experimental results demonstrate that our model achieves the state-of-the-art performance on both tasks.
Jiaying Liu 0001, Yanghao Li, Sijie Song, Junliang Xing, Cuiling Lan, Wenjun Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2019 Hierarchical Contextual Refinement Networks for Human Pose Estimation
abstract
Predicting human pose in the wild is a challenging problem due to high flexibility of joints and possible occlusion. Existing approaches generally tackle the difficulties either by holistic prediction or multi-stage processing, which suffer from poor performance for locating challenging joints or high computational cost. In this paper, we propose a new Hierarchical Contextual Refinement Network (HCRN) to robustly predict human poses in an efficient manner, where human body joints of different complexities are processed at different layers in a context hierarchy. Different from existing approaches, our proposed model predicts positions of joints from easy to difficult in a single stage through effectively exploiting informative contexts provided in the previous layer. Such approach offers two appealing advantages over state-of-the-arts: (1) more accurate than predicting all the joints together and (2) more efficient than multi-stage processing methods. We design a Contextual Refinement Unit (CRU) to implement the proposed model, which enables auto-diffusion of joint detection results to effectively transfer informative context from easy joints to difficult ones. In this way, difficult joints can be reliably detected even in presence of occlusion or severe distracting factors. Multiple CRUs are organized into a tree-structured hierarchy which is end-to-end trainable and does not require processing joints for multiple iterations. Comprehensive experiments evaluate the efficacy and efficiency of the proposed HCRN model to improve well-established baselines and achieve new state-of-the-art on multiple human pose estimation benchmarks.
Xuecheng Nie, Jiashi Feng, Junliang Xing, Shengtao Xiao, Shuicheng Yan
IEEE Trans. Image Process.3
2019 Spatial alignment network for facial landmark localization
Yidong Li, Junliang Xing, Hairong Dong 0001
World Wide Web3
2018 ZoomNet: Deep Aggregation Learning for High-Performance Small Pedestrian Detection
abstract
It remains very challenging for a single deep model to detect pedestrians of different sizes appears in an image. One typical remedy for the small pedestrian detection is to up-sample the input and pass it to the network multiple times. Unfortunately this strategy not only exponentially increases the computational cost but also probably impairs the model effectiveness. In this work, we present a deep architecture, refereed to as ZoomNet, which performs small pedestrian detection by deep aggregation learning without up-sampling the input. ZoomNet learns and aggregates deep feature representations at multiple levels and retains the spatial information of the pedestrian from different scales. ZoomNet also learns to cultivate the feature representations from the classification task to the detection task and obtains further performance improvements. Extensive experimental results demonstrate the state-of-the-art performance of ZoomNet. The source code of this work will be made public available to facilitate further studies on this problem.
Chong Shang, Haizhou Ai, Zijie Zhuang, Long Chen 0024, Junliang Xing
ACML5
2018 Deep Cost-Sensitive and Order-Preserving Feature Learning for Cross-Population Age Estimation
abstract
Facial age estimation from a face image is an important yet very challenging task in computer vision, since humans with different races and/or genders, exhibit quite different patterns in their facial aging processes. To deal with the influence of race and gender, previous methods perform age estimation within each population separately. In practice, however, it is often very difficult to collect and label sufficient data for each population. Therefore, it would be helpful to exploit an existing large labeled dataset of one (source) population to improve the age estimation performance on another (target) population with only a small labeled dataset available. In this work, we propose a Deep Cross-Population (DCP) age estimation model to achieve this goal. In particular, our DCP model develops a two-stage training strategy. First, a novel cost-sensitive multitask loss function is designed to learn transferable aging features by training on the source population. Second, a novel order-preserving pair-wise loss function is designed to align the aging features of the two populations. By doing so, our DCP model can transfer the knowledge encoded in the source population to the target population. Extensive experiments on the two of the largest benchmark datasets show that our DCP model outperforms several strong baseline methods and many state-of-the-art methods.
Kai Li 0022, Junliang Xing, Chi Su, Weiming Hu 0004, Stephen J. Maybank
CVPR2
2018 Learning Attentions: Residual Attentional Siamese Network for High Performance Online Visual Tracking
abstract
Offline training for object tracking has recently shown great potentials in balancing tracking accuracy and speed. However, it is still difficult to adapt an offline trained model to a target tracked online. This work presents a Residual Attentional Siamese Network (RASNet) for high performance object tracking. The RASNet model reformulates the correlation filter within a Siamese tracking framework, and introduces different kinds of the attention mechanisms to adapt the model without updating the model online. In particular, by exploiting the offline trained general attention, the target adapted residual attention, and the channel favored feature attention, the RASNet not only mitigates the over-fitting problem in deep network training, but also enhances its discriminative capacity and adaptability due to the separation of representation learning and discriminator learning. The proposed deep architecture is trained from end to end and takes full advantage of the rich spatial temporal information to achieve robust visual tracking. Experimental results on two latest benchmarks, OTB-2015 and VOT2017, show that the RASNet tracker has the state-of-the-art tracking accuracy while runs at more than 80 frames per second.
Qiang Wang 0051, Zhu Teng, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank
CVPR3
2018 Towards Pose Invariant Face Recognition in the Wild
abstract
Pose variation is one key challenge in face recognition. As opposed to current techniques for pose invariant face recognition, which either directly extract pose invariant features for recognition, or first normalize profile face images to frontal pose before feature extraction, we argue that it is more desirable to perform both tasks jointly to allow them to benefit from each other. To this end, we propose a Pose Invariant Model (PIM) for face recognition in the wild, with three distinct novelties. First, PIM is a novel and unified deep architecture, containing a Face Frontalization sub-Net (FFN) and a Discriminative Learning sub-Net (DLN), which are jointly learned from end to end. Second, FFN is a well-designed dual-path Generative Adversarial Network (GAN) which simultaneously perceives global structures and local details, incorporated with an unsupervised cross-domain adversarial training and a "learning to learn" strategy for high-fidelity and identity-preserving frontal view synthesis. Third, DLN is a generic Convolutional Neural Network (CNN) for face recognition with our enforced cross-entropy optimization strategy for learning discriminative yet generalized feature representation. Qualitative and quantitative experiments on both controlled and in-the-wild benchmarks demonstrate the superiority of the proposed model over the state-of-the-arts.
Jian Zhao 0006, Yu Cheng 0009, Yan Xu 0009, Jianshu Li, Fang Zhao 0006, Jayashree Karlekar, Sugiri Pranata, Shengmei Shen, Junliang Xing, Shuicheng Yan, Jiashi Feng
CVPR10
2018 Pose Partition Networks for Multi-person Pose Estimation
Xuecheng Nie, Jiashi Feng, Junliang Xing, Shuicheng Yan
ECCV (5)3
2018 Visual Tracking via Spatially Aligned Correlation Filters Network
Mengdan Zhang, Qiang Wang 0051, Junliang Xing, Peixi Peng, Weiming Hu 0004, Stephen J. Maybank
ECCV (3)3
2018 Skeleton-Indexed Deep Multi-Modal Feature Learning for High Performance Human Action Recognition
abstract
This paper presents a new framework for action recognition with multi-modal data. A skeleton-indexed feature learning procedure is developed to further exploit the detailed local features from RGB and optical flow videos. In particular, the proposed framework is built based on a deep Convolutional Network (ConvNet) and a Recurrent Neural Network (RNN) with Long Short Term Memory (LSTM). A skeleton-indexed transform layer is designed to automatically extract visual features around key joints, and a part-aggregated pooling is developed to uniformly regulate the visual features from different body parts and actors. Besides, several fusion schemes are explored to take advantage of multi-modal data. The proposed deep architecture is end-to-end trainable and can better incorporate different modalities to learn effective feature representations. Quantitative experiment results on two datasets, the NTU RGB+D dataset and the MSR dataset, demonstrate the excellent performance of our scheme over other state-of-the-arts. To our knowledge, the performance obtained by the proposed framework is currently the best on the challenging NTU RGB+D dataset.
Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001
ICME3
2018 Deep Age Estimation Model Stabilization from Images to Videos
abstract
Deep learning models for age estimation from a single image have significantly improved the state-of-the-art. However, when deploying a deep age estimation model from images directly to videos, it often suffers from the fluctuation issue, i.e., the estimated age varies a lot for face frames from the same person. To deal with this problem, this work presents a new deep age estimation model specifically designed for video facial age estimation, which produces very stable and accurate age estimation results. The proposed deep architecture for video facial age estimation incorporates a convolutional neural network with an attention mechanism, where the convolutional neural network extracts the facial features, and an attention block aggregates the facial feature vectors into a single feature representation for final age estimation. The whole model is trained by a novel loss function to guarantee both the accuracy of each frame and the stabilization of age estimation results of all the frames. To evaluate the proposed model for video facial age estimation, a new dataset is collected and annotated. Extensive experimental analyses and comparisons demonstrate the effectiveness of the proposed model and the state-of-the-art performances compared to many competing methods.
Zhipeng Ji, Congyan Lang, Kai Li 0023, Junliang Xing
ICPR4
2018 SPCNet: Scale Position Correlation Network for End-to-End Visual Tracking
abstract
We present a novel Scale Position Correlation Network (SPCNet) for learning to track objects robustly and efficiently. Different from most previous Correlation Filter (CF) based tracking models, SPCNet unifies the feature representation learning and CF based appearance modeling within one end-to-end learnable framework. In particular, SPCNet learns to track objects within a joint scale-position space, and is very effective in learning features for the accurate prediction of object scale and position. To learn our model from end to end, the SPCNet introduces a differentiable correlation filter layer into a Siamese architecture. Therefore, the localization error can be effectively back-propagated through the whole network, enabling fast adaptation of feature learning and appearance modeling for the objects to be tracked. Such task driven feature learning admits a very lightweight design that can be efficiently pre-trained. In addition, the dense appearance modeling in the joint scale-position space is also efficient. It benefits from the computation of gradients within the Fourier frequency domain. Such careful architecture design ensures that SPCNet is effective and efficient with a small model size. Extensive experimental analyses and evaluations on three largest benchmarks, OTB-2013, OTB-2015, and VOT2015, demonstrate its superiority over many state-of-the-art algorithms.
Qiang Wang 0051, Mengdan Zhang, Junliang Xing, Weiming Hu 0004
ICPR4
2018 Online Multi-Target Tracking with Tensor-Based High-Order Graph Matching
abstract
In this paper we formulate multi-target tracking (MTT) as a high-order graph matching problem and propose a l1-norm tensor power iteration solution. Concretely, the search for trajectory-observation correspondences in MTT task is cast as a hypergraph matching problem to maximize a multi-linear objective function over all permutations of the associations. This function is defined by a tensor representing the affinity between association tuples where pair-wise similarities, motion consistency and spatial structural information can be embedded expediently. To solve the matching problem, a dual-direction unit l1-norm constrained tensor power iteration algorithm is proposed. Additionally, as measuring the appearance affinity with features extracted from the rectangle patch, which is adopted in most methods, has a weak discrimination when bounding boxes overlap each other heavily, we present a deep pair-wise appearance similarity metric based on object mask in this paper where just the features from true target region are utilized. Experimental evaluation shows that our approach achieves an accuracy comparable to state-of-the-art online trackers. The source code of the proposed approach will be released to facilitate further studies on the MTT problem.
Zongwei Zhou, Junliang Xing, Mengdan Zhang, Weiming Hu 0004
ICPR2
2018 Do not Lose the Details: Reinforced Representation Learning for High Performance Visual Tracking
abstract
This work presents a novel end-to-end trainable CNN model for high performance visual object tracking. It learns both low-level fine-grained representations and a high-level semantic embedding space in a mutual reinforced way, and a multi-task learning strategy is proposed to perform the correlation analysis on representations from both levels. In particular, a fully convolutional encoder-decoder network is designed to reconstruct the original visual features from the semantic projections to preserve all the geometric information. Moreover, the correlation filter layer working on the fine-grained representations leverages a global context constraint for accurate object appearance modeling. The correlation filter in this layer is updated online efficiently without network fine-tuning. Therefore, the proposed tracker benefits from two complementary effects: the adaptability of the fine-grained correlation analysis and the generalization capability of the semantic embedding. Extensive experimental evaluations on four popular benchmarks demonstrate its state-of-the-art performance.
Qiang Wang 0051, Mengdan Zhang, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank
IJCAI3
2018 3D-Aided Deep Pose-Invariant Face Recognition
abstract
Learning from synthetic faces, though perhaps appealing for high data efficiency, may not bring satisfactory performance due to the distribution discrepancy of the synthetic and real face images. To mitigate this gap, we propose a 3D-Aided Deep Pose-Invariant Face Recognition Model (3D-PIM), which automatically recovers realistic frontal faces from arbitrary poses through a 3D face model in a novel way. Specifically, 3D-PIM incorporates a simulator with the aid of a 3D Morphable Model (3D MM) to obtain shape and appearance prior for accelerating face normalization learning, requiring less training data. It further leverages a global-local Generative Adversarial Network (GAN) with multiple critical improvements as a refiner to enhance the realism of both global structures and local details of the face simulator’s output using unlabelled real data only, while preserving the identity information. Qualitative and quantitative experiments on both controlled and in-the-wild benchmarks clearly demonstrate superiority of the proposed model over state-of-the-arts.
Jian Zhao 0006, Yu Cheng 0009, Jianshu Li, Yan Xu 0009, Jayashree Karlekar, Sugiri Pranata, Shengmei Shen, Junliang Xing, Shuicheng Yan, Jiashi Feng
IJCAI11
2018 Towards Robust and Accurate Multi-View and Partially-Occluded Face Alignment
abstract
Face alignment acts as an important task in computer vision. Regression-based methods currently dominate the approach to solving this problem, which generally employ a series of mapping functions from the face appearance to iteratively update the face shape hypothesis. One keypoint here is thus how to perform the regression procedure. In this work, we formulate this regression procedure as a sparse coding problem. We learn two relational dictionaries, one for the face appearance and the other one for the face shape, with coupled reconstruction coefficient to capture their underlying relationships. To deploy this model for face alignment, we derive the relational dictionaries in a stage-wised manner to perform close-loop refinement of themselves, i.e., the face appearance dictionary is first learned from the face shape dictionary and then used to update the face shape hypothesis, and the updated face shape dictionary from the shape hypothesis is in return used to refine the face appearance dictionary. To improve the model accuracy, we extend this model hierarchically from the whole face shape to face part shapes, thus both the global and local view variations of a face are captured. To locate facial landmarks under occlusions, we further introduce an occlusion dictionary into the face appearance dictionary to recover face shape from partially occluded face appearance. The occlusion dictionary is learned in a data driven manner from background images to represent a set of elemental occlusion patterns, a sparse combination of which models various practical partial face occlusions. By integrating all these technical innovations, we obtain a robust and accurate approach to locate facial landmarks under different face views and possibly severe occlusions for face images in the wild. Extensive experimental analyses and evaluations on different benchmark datasets, as well as two new datasets built by ourselves, have demonstrated the robustness and accuracy of our proposed model, especially for face images with large view variations and/or severe occlusions.
Junliang Xing, Zhiheng Niu, Junshi Huang, Weiming Hu 0004, Shuicheng Yan
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Multi-type attributes driven multi-camera person re-identification
Chi Su, Shiliang Zhang, Junliang Xing, Wen Gao 0001, Qi Tian 0001
Pattern Recognit.3
2018 FatRegion: A Fast Adaptive Tree-Structured Region Extraction Approach
abstract
Coherent image regions can be used as good features for many computer vision tasks, such as object tracking, segmentation, and recognition. Most of previous region extraction methods, however, are not suitable for online applications because of their either heavy computations or unsatisfactory results. We propose a seed-based region growing and merging approach to generate simultaneously coherent and discriminative image regions. We present a quadtree-based seed initialization algorithm to adaptively place seeds into different image areas and then grow them into regions by a color- and edge-guided growing procedure. To merge these regions in different levels, we propose to use the generalized boundary strength to measure the quality of region merging result. In addition, we present a region merging algorithm of linear time complexity to perform efficient and effective region merging. Overall, our new approach simultaneously holds these advantages: 1) it is extremely fast with linear complexity in both time and space, which takes less than 50 ms to process an HVGA image; 2) it can give a direct control of the region number and well adapt to image regions with various sizes and shapes; and 3) it provides a tree-structured representation of the regions and thus can model the image from multiple scales. We evaluate the proposed approach on the standard benchmarks with extensive comparisons with the state-of-the-art methods. The experimental results demonstrate its good comprehensive performances. Example applications using the extracted regions as features for online object tracking and multiclass object segmentation also exhibit its potential for many computer vision tasks.
Junliang Xing, Weiming Hu 0004, Haizhou Ai, Shuicheng Yan
IEEE Trans. Circuits Syst. Video Technol.1
2018 Deep Constrained Siamese Hash Coding Network and Load-Balanced Locality-Sensitive Hashing for Near Duplicate Image Detection
abstract
We construct a new efficient near duplicate image detection method using a hierarchical hash code learning neural network and load-balanced locality-sensitive hashing (LSH) indexing. We propose a deep constrained siamese hash coding neural network combined with deep feature learning. Our neural network is able to extract effective features for near duplicate image detection. The extracted features are used to construct a LSH-based index. We propose a load-balanced LSH method to produce load-balanced buckets in the hashing process. The load-balanced LSH significantly reduces the query time. Based on the proposed load-balanced LSH, we design an effective and feasible algorithm for near duplicate image detection. Extensive experiments on three benchmark data sets demonstrate the effectiveness of our deep siamese hash encoding network and load-balanced LSH.
Weiming Hu 0004, Yabo Fan, Junliang Xing, Zhaoquan Cai 0001, Stephen J. Maybank
IEEE Trans. Image Process.3
2018 Spatio-Temporal Attention-Based LSTM Networks for 3D Action Recognition and Detection
abstract
Human action analytics has attracted a lot of attention for decades in computer vision. It is important to extract discriminative spatio-temporal features to model the spatial and temporal evolutions of different actions. In this paper, we propose a spatial and temporal attention model to explore the spatial and temporal discriminative features for human action recognition and detection from skeleton data. We build our networks based on the recurrent neural networks with long short-term memory units. The learned model is capable of selectively focusing on discriminative joints of skeletons within each input frame and paying different levels of attention to the outputs of different frames. To ensure effective training of the network for action recognition, we propose a regularized cross-entropy loss to drive the learning process and develop a joint training strategy accordingly. Moreover, based on temporal attention, we develop a method to generate the action temporal proposals for action detection. We evaluate the proposed method on the SBU Kinect Interaction data set, the NTU RGB + D data set, and the PKU-MMD data set, respectively. Experiment results demonstrate the effectiveness of our proposed model on both action recognition and action detection.
Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001
IEEE Trans. Image Process.3
2017 An End-to-End Spatio-Temporal Attention Model for Human Action Recognition from Skeleton Data
abstract
Human action recognition is an important task in computer vision. Extracting discriminative spatial and temporal features to model the spatial and temporal evolutions of different actions plays a key role in accomplishing this task. In this work, we propose an end-to-end spatial and temporal attention model for human action recognition from skeleton data. We build our model on top of the Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM), which learns to selectively focus on discriminative joints of skeleton within each frame of the inputs and pays different levels of attention to the outputs of different frames. Furthermore, to ensure effective training of the network, we propose a regularized cross-entropy loss to drive the model learning process and develop a joint training strategy accordingly. Experimental results demonstrate the effectiveness of the proposed model, both on the small human action recognition dataset of SBU and the currently largest NTU dataset.
Sijie Song, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jiaying Liu 0001
AAAI3
2017 A Deep Regression Architecture with Two-Stage Re-initialization for High Performance Facial Landmark Detection
abstract
Regression based facial landmark detection methods usually learns a series of regression functions to update the landmark positions from an initial estimation. Most of existing approaches focus on learning effective mapping functions with robust image features to improve performance. The approach to dealing with the initialization issue, however, receives relatively fewer attentions. In this paper, we present a deep regression architecture with two-stage re-initialization to explicitly deal with the initialization problem. At the global stage, given an image with a rough face detection result, the full face region is firstly re-initialized by a supervised spatial transformer network to a canonical shape state and then trained to regress a coarse landmark estimation. At the local stage, different face parts are further separately re-initialized to their own canonical shape states, followed by another regression subnetwork to get the final estimation. Our proposed deep architecture is trained from end to end and obtains promising results using different kinds of unstable initialization. It also achieves superior performances over many competing algorithms.
Jiang-Jing Lv, Xiaohu Shao, Junliang Xing
CVPR3
2017 Pose-Driven Deep Convolutional Model for Person Re-identification
abstract
Feature extraction and matching are two crucial components in person Re-Identification (ReID). The large pose deformations and the complex view variations exhibited by the captured person images significantly increase the difficulty of learning and matching of the features from person images. To overcome these difficulties, in this work we propose a Pose-driven Deep Convolutional (PDC) model to learn improved feature extraction and matching models from end to end. Our deep architecture explicitly leverages the human part cues to alleviate the pose variations and learn robust feature representations from both the global image and different local parts. To match the features from global human body and local body parts, a pose driven feature weighting sub-network is further designed to learn adaptive feature fusions. Extensive experimental analyses and results on three popular datasets demonstrate significant performance improvements of our model over all published state-of-the-art methods.
Chi Su, Jianing Li 0001, Shiliang Zhang, Junliang Xing, Wen Gao 0001, Qi Tian 0001
ICCV4
2017 Human Pose Estimation Using Global and Local Normalization
abstract
In this paper, we address the problem of estimating the positions of human joints, i.e., articulated pose estimation. Recent state-of-the-art solutions model two key issues, joint detection and spatial configuration refinement, together using convolutional neural networks. Our work mainly focuses on spatial configuration refinement by reducing variations of human poses statistically, which is motivated by the observation that the scattered distribution of the relative locations of joints (e.g., the left wrist is distributed nearly uniformly in a circular area around the left shoulder) makes the learning of convolutional spatial models hard. We present a two-stage normalization scheme, human body normalization and limb normalization, to make the distribution of the relative joint locations compact, resulting in easier learning of convolutional spatial models and more accurate pose estimation. In addition, our empirical results show that incorporating multi-scale supervision and multi-scale fusion into the joint detection network is beneficial. Experiment results demonstrate that our method consistently outperforms state-of-the-art methods on the benchmarks.
Ke Sun 0009, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Dong Liu 0002, Jingdong Wang 0001
ICCV3
2017 Robust Object Tracking Based on Temporal and Spatial Deep Networks
abstract
Recently deep neural networks have been widely employed to deal with the visual tracking problem. In this work, we present a new deep architecture which incorporates the temporal and spatial information to boost the tracking performance. Our deep architecture contains three networks, a Feature Net, a Temporal Net, and a Spatial Net. The Feature Net extracts general feature representations of the target. With these feature representations, the Temporal Net encodes the trajectory of the target and directly learns temporal correspondences to estimate the object state from a global perspective. Based on the learning results of the Temporal Net, the Spatial Net further refines the object tracking state using local spatial object information. Extensive experiments on four of the largest tracking benchmarks, including VOT2014, VOT2016, OTB50, and OTB100, demonstrate competing performance of the proposed tracker over a number of state-of-the-art algorithms.
Zhu Teng, Junliang Xing, Qiang Wang 0051, Congyan Lang, Songhe Feng, Yi Jin 0001
ICCV2
2017 View Adaptive Recurrent Neural Networks for High Performance Human Action Recognition from Skeleton Data
abstract
Skeleton-based human action recognition has recently attracted increasing attention due to the popularity of 3D skeleton data. One main challenge lies in the large view variations in captured human actions. We propose a novel view adaptation scheme to automatically regulate observation viewpoints during the occurrence of an action. Rather than re-positioning the skeletons based on a human defined prior criterion, we design a view adaptive recurrent neural network (RNN) with LSTM architecture, which enables the network itself to adapt to the most suitable observation viewpoints from end to end. Extensive experiment analyses show that the proposed view adaptive RNN model strives to (1) transform the skeletons of various views to much more consistent viewpoints and (2) maintain the continuity of the action rather than transforming every frame to the same position with the same body orientation. Our model achieves significant improvement over the state-of-the-art approaches on three benchmark datasets.
Pengfei Zhang 0005, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Jianru Xue, Nanning Zheng 0001
ICCV3
2017 Weakly supervised multiscale-inception learning for web-scale face recognition
abstract
Supervised deep learning models like convolutional neural network (CNN) have shown very promising results for the face recognition problem, which often require a huge number of labeled face images. Since manually labeling a large training set is a very difficult and time-consuming task, it is very beneficial if the deep model can be trained from face samples with only weak annotations. In this paper, we propose a general framework to train a deep CNN model with weakly labeled facial images that are available on the Internet. Specifically, we first design a deep Multiscale-Inception CNN (MICNN) architecture to exploit the multi-scale information for face recognition. Then, we train an initial MICNN model with only a limited number of labeled samples. After that, we propose a dual-level sample selection strategy to further fine-tune the MICNN model with the weakly labeled samples from both the sample level and class level, which aims to skip outliers and select more samples from confusing class pairs during training. Extensive experimental results on the LFW and YTF benchmarks demonstrate the effectiveness of the proposed method.
Junliang Xing, Youji Feng, Xiaohu Shao, Kai Li 0022
ICIP2
2017 Hierarchical bilinear network for high performance face detection
abstract
Deep Convolutional Networks (DCNs) have achieved great success in face detection. Most architectures of the DCN-based methods, however, suffer from multiple separated steps and large-size models, which increase the training complexity and also slow down the testing speed. In this paper, we propose an efficient end-to-end architecture, called Hierarchical Bilinear Network (HBN), for fast and accurate face detection. It mainly consists of two parts: the Backbone Network and the Bilinear Network. The Backbone Network generates hierarchical feature maps for efficiently characterizing faces of different scales, while the Bilinear Network classifies the regions and regresses the face bounding-boxes on each feature map by introducing the Inception module and weights sharing. Benefited from the characters of the proposed architecture, it obtains a better comprehensive performance regarding the model effectiveness, running efficiency, and parameter size, compared with other DCN-based methods. Extensive experimental results demonstrate that our detector achieves competitive accuracy on both the FDDB database and the WIDER FACE database, while still runs in real time (about 69 FPS on a Titan Black GPU) with a tiny size (2.2 MB) model.
Jiang-Jing Lv, Xiaohu Shao, Junliang Xing
ICIP3
2017 SCNN: Sequential convolutional neural network for human action recognition in videos
abstract
Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) are two typical kinds of neural networks. While CNN models have achieved great success on image recognition due to their strong abilities in abstracting spatial information from multiple levels, RNN models have not achieved significant progress in video analyzing tasks (e.g. action recognition), although RNN can inherently model temporal dependencies from videos. In this work, we propose a Sequential Convolutional Neural Network, denoted as SCNN, to extract effective spatial-temporal features from videos, thus incorporating the strengths of both convolutional operation and recurrent operation. Our SCNN model extends RNN to directly process feature maps, rather than vectors flattened from feature maps, to keep spatial structures of the inputs. It replaces the full connections of RNN with convolutional connections to decrease parameter numbers, computational cost, and over-fitting risk. Moreover, we introduce asymmetric convolutional layers into SCNN to reduce parameter numbers and computational cost further. Our final SCNN deep architecture used for action recognition achieves very good performances on two challenging benchmarks, UCF-101 and HMDB-51, outperforming many state-of-the-art methods.
Hao Yang 0010, Chunfeng Yuan, Junliang Xing, Weiming Hu 0004
ICIP3
2017 Diversity encouraging ensemble of convolutional networks for high performance action recognition
abstract
We present a simple and effective ensemble method, Diversity Encouraging Ensemble (DEE), for deep convolutional networks to boost their performances. By training the convolutional network in two stages, we generate multiple component networks without adding any training cost. On the one hand, we modify the structure parameters of component networks in the training process to enlarge the diversities of the networks, which is found to be beneficial to improving the ensemble performance. On the other hand, we exploit monotonous decreasing learning rate schedule to accelerate the speed of deep network converging to different local minima, and we decrease the training time of integrating multiple networks to that of training a single network from traditional multi-step learning policy. We evaluate our ensemble method on two challenging action datasets, UCF-101 and HMDB-51, and obtain performance improvements from single deep network and other ensemble methods. Our results also outperform many state-of-the-art action recognition methods.
Hao Yang 0010, Chunfeng Yuan, Junliang Xing, Weiming Hu 0004
ICIP3
2017 Real-Time Deep Video SpaTial Resolution UpConversion SysTem (STRUCT++ Demo)
abstract
Image and video super-resolution (SR) has been explored for several decades. However, few works are integrated into practical systems for real-time image and video SR. In this work, we present a real-time deep video SpaTial Resolution UpConversion SysTem (STRUCT++). Our demo system achieves real-time performance (50 fps on CPU for CIF sequences and 45 fps on GPU for HDTV videos) and provides several functions: 1) batch processing; 2) full resolution comparison; 3) local region zooming in. These functions are convenient for super-resolution of a batch of videos (at most 10 videos in parallel), comparisons with other approaches and observations of local details of the SR results. The system is built on a Global context aggregation and Local queue jumping Network (GLNet). It has a thinner and deeper network structure to aggregate global context with an additional local queue jumping path to better model local structures of the signal. GLNet achieves state-of-the-art performance for real-time video SR.
Wenhan Yang, Shihong Deng, Yueyu Hu, Junliang Xing, Jiaying Liu 0001
ACM Multimedia4
2017 Towards human-like and transhuman perception in AI 2.0: a review
abstract
Perception is the interaction interface between an intelligent system and the real world. Without sophisticated and flexible perceptual capabilities, it is impossible to create advanced artificial intelligence (AI) systems. For the next-generation AI, called ‘AI 2.0’, one of the most significant features will be that AI is empowered with intelligent perceptual capabilities, which can simulate human brain’s mechanisms and are likely to surpass human brain in terms of performance. In this paper, we briefly review the state-of-the-art advances across different areas of perception, including visual perception, auditory perception, speech perception, and perceptual information processing and learning engines. On this basis, we envision several R&D trends in intelligent perception for the forthcoming era of AI 2.0, including: (1) human-like and transhuman active vision; (2) auditory perception and computation in an actual auditory setting; (3) speech perception and computation in a natural interaction setting; (4) autonomous learning of perceptual information; (5) large-scale perceptual information processing and learning platforms; and (6) urban omnidirectional intelligent perception and reasoning engines. We believe these research directions should be highlighted in the future plans for AI 2.0.
Yonghong Tian 0001, Xilin Chen 0001, Hongkai Xiong, Li-Rong Dai 0001, Jing Chen 0002, Junliang Xing, Jing Chen 0003, Xihong Wu, Weiming Hu 0004, Yu Hu 0003, Tiejun Huang 0001, Wen Gao 0001
Frontiers Inf. Technol. Electron. Eng.7
2017 Semi-Supervised Tensor-Based Graph Embedding Learning and Its Application to Visual Discriminant Tracking
abstract
An appearance model adaptable to changes in object appearance is critical in visual object tracking. In this paper, we treat an image patch as a two-order tensor which preserves the original image structure. We design two graphs for characterizing the intrinsic local geometrical structure of the tensor samples of the object and the background. Graph embedding is used to reduce the dimensions of the tensors while preserving the structure of the graphs. Then, a discriminant embedding space is constructed. We prove two propositions for finding the transformation matrices which are used to map the original tensor samples to the tensor-based graph embedding space. In order to encode more discriminant information in the embedding space, we propose a transfer-learning- based semi-supervised strategy to iteratively adjust the embedding space into which discriminative information obtained from earlier times is transferred. We apply the proposed semi-supervised tensor-based graph embedding learning algorithm to visual tracking. The new tracking algorithm captures an object's appearance characteristics during tracking and uses a particle filter to estimate the optimal object state. Experimental results on the CVPR 2013 benchmark dataset demonstrate the effectiveness of the proposed tracking algorithm.
Weiming Hu 0004, Junliang Xing, Chao Zhang 0089, Stephen J. Maybank
IEEE Trans. Pattern Anal. Mach. Intell.3
2017 D2C: Deep cumulatively and comparatively learning for human age estimation
Kai Li 0022, Junliang Xing, Weiming Hu 0004, Stephen J. Maybank
Pattern Recognit.2
2017 Diagnosing deep learning models for high accuracy age estimation from a single image
Junliang Xing, Kai Li 0022, Weiming Hu 0004, Chunfeng Yuan, Haibin Ling
Pattern Recognit.1
2017 Beyond Group: Multiple Person Tracking via Minimal Topology-Energy-Variation
abstract
Tracking multiple persons is a challenging task when persons move in groups and occlude each other. Existing group-based methods have extensively investigated how to make group division more accurately in a tracking-by-detection framework; however, few of them quantify the group dynamics from the perspective of targets' spatial topology or consider the group in a dynamic view. Inspired by the sociological properties of pedestrians, we propose a novel socio-topology model with a topology-energy function to factor the group dynamics of moving persons and groups. In this model, minimizing the topology-energy-variance in a two-level energy form is expected to produce smooth topology transitions, stable group tracking, and accurate target association. To search for the strong minimum in energy variation, we design the discrete group-tracklet jump moves embedded in the gradient descent method, which ensures that the moves reduce the energy variation of group and trajectory alternately in the varying topology dimension. Experimental results on both RGB and RGB-D data sets show the superiority of our proposed model for multiple person tracking in crowd scenes.
Shan Gao 0003, Qixiang Ye, Junliang Xing, Arjan Kuijper, Zhenjun Han, Jianbin Jiao, Xiangyang Ji
IEEE Trans. Image Process.3
2016 Co-Occurrence Feature Learning for Skeleton Based Action Recognition Using Regularized Deep LSTM Networks
abstract
Skeleton based action recognition distinguishes human actions using the trajectories of skeleton joints, which provide a very good representation for describing actions. Considering that recurrent neural networks (RNNs) with Long Short-Term Memory (LSTM) can learn feature representations and model long-term temporal dependencies automatically, we propose an end-to-end fully connected deep LSTM network for skeleton based action recognition. Inspired by the observation that the co-occurrences of the joints intrinsically characterize human actions, we take the skeleton as the input at each time slot and introduce a novel regularization scheme to learn the co-occurrence features of skeleton joints. To train the deep LSTM network effectively, we propose a new dropout algorithm which simultaneously operates on the gates, cells, and output responses of the LSTM neurons. Experimental results on three human action recognition datasets consistently demonstrate the effectiveness of the proposed model.
Wentao Zhu 0001, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Yanghao Li, Li Shen 0005, Xiaohui Xie
AAAI3
2016 Tensor Power Iteration for Multi-graph Matching
abstract
Due to its wide range of applications, matching between two graphs has been extensively studied and remains an active topic. By contrast, it is still under-exploited on how to jointly match multiple graphs, partly due to its intrinsic combinatorial intractability. In this work, we address this challenging problem in a principled way under the rank-1 tensor approximation framework. In particular, we formulate multi-graph matching as a combinational optimization problem with two main ingredients: unary matching over graph vertices and structure matching over graph edges, both of which across multiple graphs. Then we propose an efficient power iteration solution for the resulting NP-hard optimization problem. The proposed algorithm has several advantages: 1) the intrinsic matching consistency across multiple graphs based on the high-order tensor optimization, 2) the free employment of powerful high-order node affinity, 3) the flexible integration between various types of node affinities and edge/hyper-edge affinities. Experiments on diverse and challenging datasets validate the effectiveness of the proposed approach in comparison with state-of the-arts.
Xinchu Shi, Haibin Ling, Weiming Hu 0004, Junliang Xing
CVPR4
2016 Online Human Action Detection Using Joint Classification-Regression Recurrent Neural Networks
Yanghao Li, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Chunfeng Yuan, Jiaying Liu 0001
ECCV (7)3
2016 Deep Attributes Driven Multi-camera Person Re-identification
Chi Su, Shiliang Zhang, Junliang Xing, Wen Gao 0001, Qi Tian 0001
ECCV (2)3
2016 Robust Facial Landmark Detection via Recurrent Attentive-Refinement Networks
Shengtao Xiao, Jiashi Feng, Junliang Xing, Hanjiang Lai, Shuicheng Yan, Ashraf A. Kassim
ECCV (1)3
2016 Bootstrapping deep feature hierarchy for pornographic image recognition
abstract
Automatically recognizing pornographic images from the Web is a vital step to purify Internet environment. Inspired by the rapid developments of deep learning models, we present a deep architecture of convolutional neural network (CNN) for high accuracy pornographic image recognition. The proposed architecture is built upon existing CNNs which accepts input images of different sizes and incorporates features from different hierarchy to perform prediction. To effectively train the model, we propose a two-stage training strategy to learn the model parameters from scratch and end-to-end. During the training procedure, we also employ a hard negative sampling strategy to further reduce the false positive rate of the model. Experimental results on a large dataset demonstrate good performance of the proposed model and the effectiveness of our training strategies, with a considerable improvement over some traditional methods using hand-crafted features and deep learning method using mainstream CNN architecture.
Kai Li 0022, Junliang Xing, Bing Li 0001, Weiming Hu 0004
ICIP2
2016 Multi-Cue Illumination Estimation via a Tree-Structured Group Joint Sparse Representation
Bing Li 0001, Weihua Xiong, Weiming Hu 0004, Brian V. Funt, Junliang Xing
Int. J. Comput. Vis.5
2016 SERVE: Soft and Equalized Residual VEctors for image retrieval
Jun Li 0033, Chang Xu 0002, Mingming Gong, Junliang Xing, Wankou Yang, Changyin Sun 0001
Neurocomputing4
2015 Shape driven kernel adaptation in Convolutional Neural Network for robust facial trait recognition
abstract
One key challenge of facial trait recognition is the large non-rigid appearance variations due to some irrelevant real world factors, such as viewpoint and expression changes. In this paper, we explore how the shape information, i.e. facial landmark positions, can be explicitly deployed into the popular Convolutional Neural Network (CNN) architecture to disentangle such irrelevant non-rigid appearance variations. First, instead of using fixed kernels, we propose a kernel adaptation method to dynamically determine the convolutional kernels according to the spatial distribution of facial landmarks, which helps learning more robust features. Second, motivated by the intuition that different local facial regions may demand different adaptation functions, we further propose a tree-structured convolutional architecture to hierarchically fuse multiple local adaptive CNN subnetworks. Comprehensive experiments on WebFace, Morph II and MultiPIE databases well validate the effectiveness of the proposed kernel adaptation method and tree-structured convolutional architecture for facial trait recognition tasks, including identity, age and gender recognition. For all the tasks, the proposed architecture consistently achieves the state-of-the-art performances.
Shaoxin Li 0001, Junliang Xing, Zhiheng Niu, Shiguang Shan, Shuicheng Yan
CVPR2
2015 Local Subspace Collaborative Tracking
abstract
Subspace models have been widely used for appearance based object tracking. Most existing subspace based trackers employ a linear subspace to represent object appearances, which are not accurate enough to model large variations of objects. To address this, this paper presents a local subspace collaborative tracking method for robust visual tracking, where multiple linear and nonlinear subspaces are learned to better model the nonlinear relationship of object appearances. First, we retain a set of key samples and compute a set of local subspaces for each key sample. Then, we construct a hyper sphere to represent the local nonlinear subspace for each key sample. The hyper sphere of one key sample passes the local key samples and also is tangent to the local linear subspace of the specific key sample. In this way, we are able to represent the nonlinear distribution of the key samples and also approximate the local linear subspace near the specific key sample, so that local distributions of the samples can be represented more accurately. Experimental results on challenging video sequences demonstrate the effectiveness of our method.
Xiaoqin Zhang 0002, Weiming Hu 0004, Junliang Xing, Jiwen Lu, Jie Zhou 0001
ICCV4
2015 Load-balanced locality-sensitive hashing: A new method for efficient near duplicate image detection
abstract
Locality-Sensitive Hashing (LSH) is a mainstream method for the Near Duplicate Image Detection (NDID) problem. Previous LSH based methods, however, do not have a principled way to make the indexing structure generate the buckets of similar sizes, which will inevitably degrade the detection effectiveness and efficiency. In this work, we propose a Load-Balanced Locality-Sensitive Hashing (LBLSH) method with a new indexing structure to produce load-balanced buckets for the hashing process. As proved in the paper, the proposed LBLSH can guarantee load-balanced buckets in the hashing process and significantly reduce the query time and the storage space. Based on the proposed LBLSH method, we design an effective and feasible algorithm for the NDID problem. Extensive experiments on two benchmark datasets demonstrate the effectiveness and efficiency of our method.
Yabo Fan, Junliang Xing, Weiming Hu 0004
ICIP2
2015 Robust visual tracking using joint scale-spatial correlation filters
abstract
Scale adaptation is crucial to object tracking as the visual size of the target changes continuously. Many existing tracking algorithms, however, simply ignore scale changes either for the consideration of tracking efficiency or the lack of principle ways to scale estimation. In this work, we present an efficient and effective scale adaptive tracking algorithm by proposing a correlation filter based tracker in the joint spatial and scale space. We find that the exhaustive template searching in this joint space can be well modeled by a block-circulant matrix. With the properties of the block-circulant matrices, we prove that the expensive template matching can be transformed to efficient dot product in frequency domain by fast Fourier Transform. Based on these findings, our new tracker significantly improves the robustness and adaptability of previous competitive spatial correlation trackers. On the latest single object tracking benchmark, our tracker advances the state-of-the-art tracking results with a very large margin.
Mengdan Zhang, Junliang Xing, Weiming Hu 0004
ICIP2
2014 Multi-target Tracking with Motion Context in Tensor Power Iteration
abstract
Interactions between moving targets often provide discriminative clues for multiple target tracking (MTT), though many existing approaches ignore such interactions due to difficulty in effectively handling them. In this paper, we model interactions between neighbor targets by pair-wise motion context, and further encode such context into the global association optimization. To solve the resulting global non-convex maximization, we propose an effective and efficient power iteration framework. This solution enjoys two advantages for MTT: First, it allows us to combine the global energy accumulated from individual trajectories and the between-trajectory interaction energy into a united optimization, which can be solved by the proposed power iteration algorithm. Second, the framework is flexible to accommodate various types of pairwise context models and we in fact studied two different context models in this paper. For evaluation, we apply the proposed methods to four public datasets involving different challenging scenarios such as dense aerial borne traffic tracking, dense point set tracking, and semi-crowded pedestrian tracking. In all the experiments, our approaches demonstrate very promising results in comparison with state-of-the-art trackers.
Xinchu Shi, Haibin Ling, Weiming Hu 0004, Chunfeng Yuan, Junliang Xing
CVPR5
2014 Towards Multi-view and Partially-Occluded Face Alignment
abstract
We present a robust model to locate facial landmarks under different views and possibly severe occlusions. To build reliable relationships between face appearance and shape with large view variations, we propose to formulate face alignment as an l1-induced Stagewise Relational Dictionary (SRD) learning problem. During each training stage, the SRD model learns a relational dictionary to capture consistent relationships between face appearance and shape, which are respectively modeled by the pose-indexed image features and the shape displacements for current estimated landmarks. During testing, the SRD model automatically selects a sparse set of the most related shape displacements for the testing face and uses them to refine its shape iteratively. To locate facial landmarks under occlusions, we further propose to learn an occlusion dictionary to model different kinds of partial face occlusions. By deploying the occlusion dictionary into the SRD model, the alignment performance for occluded faces can be further improved. Our algorithm is simple, effective, and easy to implement. Extensive experiments on two benchmark datasets and two newly built datasets have demonstrated its superior performances over the state-of-the-art methods, especially for faces with large view variations and/or occlusions.
Junliang Xing, Zhiheng Niu, Junshi Huang, Weiming Hu 0004, Shuicheng Yan
CVPR1
2014 Transfer Learning Based Visual Tracking with Gaussian Processes Regression
Haibin Ling, Weiming Hu 0004, Junliang Xing
ECCV (3)4
2014 Scene transformation for detector adaptation
Liwei Liu 0006, Junliang Xing, Genquan Duan, Haizhou Ai
Pattern Recognit. Lett.2
2014 Circle & Search: Attribute-Aware Shoe Retrieval
abstract
Taking the shoe as a concrete example, we present an innovative product retrieval system that leverages object detection and retrieval techniques to support a brand-new online shopping experience in this article. The system, called Circle & Search, enables users to naturally indicate any preferred product by simply circling the product in images as the visual query, and then returns visually and semantically similar products to the users. The system is characterized by introducing attributes in both the detection and retrieval of the shoe. Specifically, we first develop an attribute-aware part-based shoe detection model. By maintaining the consistency between shoe parts and attributes, this shoe detector has the ability to model high-order relations between parts and thus the detection performance can be enhanced. Meanwhile, the attributes of this detected shoe can also be predicted as the semantic relations between parts. Based on the result of shoe detection, the system ranks all the shoes in the repository using an attribute refinement retrieval model that takes advantage of query-specific information and attribute correlation to provide an accurate and robust shoe retrieval. To evaluate this retrieval system, we build a large dataset with 17,151 shoe images, in which each shoe is annotated with 10 shoe attributes e.g., heel height, heel shape, sole shape, etc.). According to the experimental result and the user study, our Circle & Search system achieves promising shoe retrieval performance and thus significantly improves the users' online shopping experience.
Junshi Huang, Si Liu 0001, Junliang Xing, Tao Mei 0001, Shuicheng Yan
ACM Trans. Multim. Comput. Commun. Appl.3
2014 "Wow! You Are So Beautiful Today!"
abstract
Beauty e-Experts, a fully automatic system for makeover recommendation and synthesis, is developed in this work. The makeover recommendation and synthesis system simultaneously considers many kinds of makeover items on hairstyle and makeup. Given a user-provided frontal face image with short/bound hair and no/light makeup, the Beauty e-Experts system not only recommends the most suitable hairdo and makeup, but also synthesizes the virtual hairdo and makeup effects. To acquire enough knowledge for beauty modeling, we built the Beauty e-Experts Database, which contains 1,505 female photos with a variety of attributes annotated with different discrete values. We organize these attributes into two different categories, beauty attributes and beauty-related attributes. Beauty attributes refer to those values that are changeable during the makeover process and thus need to be recommended by the system. Beauty-related attributes are those values that cannot be changed during the makeup process but can help the system to perform recommendation. Based on this Beauty e-Experts Dataset, two problems are addressed for the Beauty e-Experts system: what to recommend and how to wear it, which describes a similar process of selecting hairstyle and cosmetics in daily life. For the what-to-recommend problem, we propose a multiple tree-structured supergraph model to explore the complex relationships among high-level beauty attributes, mid-level beauty-related attributes, and low-level image features. Based on this model, the most compatible beauty attributes for a given facial image can be efficiently inferred. For the how-to-wear-it problem, an effective and efficient facial image synthesis module is designed to seamlessly synthesize the recommended makeovers into the user facial image. We have conducted extensive experiments on testing images of various conditions to evaluate and analyze the proposed system. The experimental results well demonstrate the effectiveness and efficiency of the proposed system.
Luoqi Liu, Junliang Xing, Si Liu 0001, Shuicheng Yan
ACM Trans. Multim. Comput. Commun. Appl.2
2013 Multi-target Tracking by Rank-1 Tensor Approximation
abstract
In this paper we formulate multi-target tracking (MTT) as a rank-1 tensor approximation problem and propose an ℓ1norm tensor power iteration solution. In particular, a high order tensor is constructed based on trajectories in the time window, with each tensor element as the affinity of the corresponding trajectory candidate. The local assignment variables are the ℓ1normalized vectors, which are used to approximate the rank-1 tensor. Our approach provides a flexible and effective formulation where both pairwise and high-order association energies can be used expediently. We also show the close relation between our formulation and the multi-dimensional assignment (MDA) model. To solve the optimization in the rank-1 tensor approximation, we propose an algorithm that iteratively powers the intermediate solution followed by an ℓ1normalization. Aside from effectively capturing high-order motion information, the proposed solver runs efficiently with proved convergence. The experimental validations are conducted on two challenging datasets and our method demonstrates promising performances on both.
Xinchu Shi, Haibin Ling, Junliang Xing, Weiming Hu 0004
CVPR3
2013 Distance Map of Various Weights: A new feature for adaptive object tracking
abstract
In this paper, we propose a new feature, Distance Map of Various Weights (DMVW) based on distances between rows' textures, to perform tracking. The proposed new feature provides an effective object appearance model which is both illumination-invariant and robust to occlusion. We also develop a 2D PCA based method to effectively evaluate the new feature. We demonstrate the validity of the rows' or column's weights in computing 2D PCA subspaces. To balance the importance of local and global information, we define a coefficient to revise the locality extent of the proposed feature. A new method based on entropy of candidate state evaluation is proposed to select the most discriminative coefficient. Experimental results on challenging video sequences demonstrated the effectiveness of our method.
Junliang Xing, Xiaoqin Zhang 0002, Weiming Hu 0004
ICASSP2
2013 Adaptive cooperative tracking based on multi-graph embedding and Markov Random Field
abstract
Appearance model is of fundamental importance in a tracking algorithm. In this paper, we propose a new tracking method based on a cooperative object appearance model which incorporates both the discriminative and generative information. We represent the discriminative information with graph embedding (GE). To represent the local object appearance effectively, we divide the object and nearby background into patches. As the discriminative conditions around the 4 object boundaries are different, we divide the patches into 4 groups and perform GE for each group. Markov Random Filed (MRF) is designed to represent the generative information. We propose a novel MRF based method which not only considers the single patch's appearance but also the appearance relations between neighbor patches (not the relations between neighbor patches' states). The proposed cooperative appearance model can represent the object appearance's variation effectively and meanwhile discriminate the object from background robustly. Experimental results on challenging test sequences demonstrated the effectiveness of our method.
Junliang Xing, Xiaoqin Zhang 0002, Weiming Hu 0004
ICASSP2
2013 Discriminant Tracking Using Tensor Representation with Semi-supervised Improvement
abstract
Visual tracking has witnessed growing methods in object representation, which is crucial to robust tracking. The dominant mechanism in object representation is using image features encoded in a vector as observations to perform tracking, without considering that an image is intrinsically a matrix, or a 2^nd-order tensor. Thus approaches following this mechanism inevitably lose a lot of useful information, and therefore cannot fully exploit the spatial correlations within the 2D image ensembles. In this paper, we address an image as a 2^nd-order tensor in its original form, and find a discriminative linear embedding space approximation to the original nonlinear sub manifold embedded in the tensor space based on the graph embedding framework. We specially design two graphs for characterizing the intrinsic local geometrical structure of the tensor space, so as to retain more discriminant information when reducing the dimension along certain tensor dimensions. However, spatial correlations within a tensor are not limited to the elements along these dimensions. This means that some part of the discriminant information may not be encoded in the embedding space. We introduce a novel technique called semi-supervised improvement to iteratively adjust the embedding space to compensate for the loss of discriminant information, hence improving the performance of our tracker. Experimental results on challenging videos demonstrate the effectiveness and robustness of the proposed tracker.
Junliang Xing, Weiming Hu 0004, Stephen J. Maybank
ICCV2
2013 Robust Object Tracking with Online Multi-lifespan Dictionary Learning
abstract
Recently, sparse representation has been introduced for robust object tracking. By representing the object sparsely, i.e., using only a few templates via L1-norm minimization, these so-called L1-trackers exhibit promising tracking results. In this work, we address the object template building and updating problem in these L1-tracking approaches, which has not been fully studied. We propose to perform template updating, in a new perspective, as an online incremental dictionary learning problem, which is efficiently solved through an online optimization procedure. To guarantee the robustness and adaptability of the tracking algorithm, we also propose to build a multi-lifespan dictionary model. By building target dictionaries of different life spans, effective object observations can be obtained to deal with the well-known drifting problem in tracking and thus improve the tracking accuracy. We derive effective observation models both generatively and discriminatively based on the online multi-lifespan dictionary learning model and deploy them to the Bayesian sequential estimation framework to perform tracking. The proposed approach has been extensively evaluated on ten challenging video sequences. Experimental results demonstrate the effectiveness of the online learned templates, as well as the state-of-the-art tracking performance of the proposed approach.
Junliang Xing, Bing Li 0001, Weiming Hu 0004, Shuicheng Yan
ICCV1
2013 Online structure learning for robust object tracking
abstract
In this paper, we aim to track objects that undergo abrupt appearance changes and heavy occlusions. To address these problems, we propose an online structure learning algorithm which contains two layers, block-based online random forest classifiers (BORFs) and online structure models (OSMs). BORFs are able to handle occlusion problems since they model local appearances of the target. To further improve the accuracy and reliability, the algorithm utilizes relational models as context information to combine BORFs into the online structure models. Capturing the discriminative parts of targets with online learnt structures, OSMs help locate targets accurately even when they are heavily occluded. In addition, OSMs guide the block occlusion reasoning and the update scheme of BORFs and relational models, which can handle appearance changes and drift problems effectively. Experiments on challenging videos show that the proposed tracker performs better than several state-of-the-art algorithms which demonstrate the effectiveness of our approach.
Liwei Liu 0006, Junliang Xing, Haizhou Ai
ICIP2
2013 "Wow! you are so beautiful today!"
abstract
In this demo, we present Beauty e-Experts, a fully automatic system for hairstyle and facial makeup recommendation and synthesis. Given a user-provided frontal facial image with short/bound hair and no/light makeup, the Beauty e-Experts system can not only recommend the most suitable hairstyle and makeup, but also show the synthesis effects. Two problems are considered for the Beauty e-Experts system: what to recommend and how to wear, which describe a similar process of selecting and applying hairstyle and cosmetics in our daily life. For the what-to-recommend problem, we propose a multiple tree-structured super-graphs model to explore the complex relationships among the beauty attributes, beauty-related attributes and image features, and then based on this model, the most suitable beauty attributes for a given facial image can be efficiently inferred. For the how-to-wear problem, a facial image synthesis module is designed to seamlessly blend the recommended hairstyle and makeup into the user facial image. Extensive experimental evaluations and analysis on testing images well demonstrate the effectiveness of the proposed system.
Luoqi Liu, Si Liu 0001, Junliang Xing, Shuicheng Yan
ACM Multimedia4
2013 "Wow! you are so beautiful today!"
abstract
Beauty e-Experts, a fully automatic system for hairstyle and facial makeup recommendation and synthesis, is developed in this work. Given a user-provided frontal face image with short/bound hair and no/light makeup, the Beauty e-Experts system can not only recommend the most suitable hairdo and makeup, but also show the synthetic effects. To obtain enough knowledge for beauty modeling, we build the Beauty e-Experts Database, which contains 1,505 attractive female photos with a variety of beauty attributes and beauty-related attributes annotated. Based on this Beauty e-Experts Dataset, two problems are considered for the Beauty e-Experts system: what to recommend and how to wear, which describe a similar process of selecting hairstyle and cosmetics in our daily life. For the what-to-recommend problem, we propose a multiple tree-structured super-graphs model to explore the complex relationships among the high-level beauty attributes, mid-level beauty-related attributes and low-level image features, and then based on this model, the most compatible beauty attributes for a given facial image can be efficiently inferred. For the how-to-wear problem, an effective and efficient facial image synthesis module is designed to seamlessly synthesize the recommended hairstyle and makeup into the user facial image. Extensive experimental evaluations and analysis on testing images of various conditions well demonstrate the effectiveness of the proposed system.
Luoqi Liu, Junliang Xing, Si Liu 0001, Shuicheng Yan
ACM Multimedia3
2012 Semantic superpixel based vehicle tracking
Liwei Liu 0006, Junliang Xing, Haizhou Ai, Shihong Lao
ICPR2
2012 Hand posture recognition using finger geometric feature
Liwei Liu 0006, Junliang Xing, Haizhou Ai, Xiang Ruan
ICPR2
2012 A tracking based fast online complete video synopsis approach
Junliang Xing, Haizhou Ai, Shihong Lao
ICPR2
2012 Scene Aware Detection and Block Assignment Tracking in crowded scenes
Genquan Duan, Haizhou Ai, Junliang Xing, Shihong Lao
Image Vis. Comput.3
2011 Robust crowd counting using detection flow
abstract
Crowd counting which aims at obtaining the number of people within a scene is an important computer vision task. While most previous methods try to count people within one frame, this paper addresses this problem using the detection flow which is defined as a set of object detection responses along the temporal video sequence. We argue that counting based on detection flow provides a better way to estimate the crowd size with following merits: 1) it can greatly alleviate the common weakness of an object detector including miss detection and false alarms; 2) it is robust to temporal object occlusions and noises; 3) it is more competent to give specific descriptions of the crowd, e.g. crowd moving directions and target locations. Experiment results on PETS 2009 dataset demonstrate the potential of this method.
Junliang Xing, Haizhou Ai, Liwei Liu 0006, Shihong Lao
ICIP1
2011 Background subtraction through multiple life span modeling
abstract
Background subtraction plays a key role in many surveillance systems. A good background subtractor should not only be able to robustly detect targets under different situations (e.g. moving and static), but also to adaptively maintain the background model against various influences (e.g. dynamic scenes and noises). This paper proposes a novel background modeling approach with these good characteristics. By introducing the “life span” concept into a background model, different properties of the scene are obtained through different life span models. Specifically, three different models, i.e., the Long Life Span Model, the Middle Life Span Model, and the Short Life Span Model, are online adaptively built and updated in a collaborative manner. Output of the system gives an adaptive, robust, and efficient estimation of the foreground region which can facility many practical applications. Experiment results on lots of surveillance videos demonstrate the superiority of the proposed method over competing approaches.
Junliang Xing, Liwei Liu 0006, Haizhou Ai
ICIP1
2011 Multiple Player Tracking in Sports Video: A Dual-Mode Two-Way Bayesian Inference Approach With Progressive Observation Modeling
abstract
Multiple object tracking (MOT) is a very challenging task yet of fundamental importance for many practical applications. In this paper, we focus on the problem of tracking multiple players in sports video which is even more difficult due to the abrupt movements of players and their complex interactions. To handle the difficulties in this problem, we present a new MOT algorithm which contributes both in the observation modeling level and in the tracking strategy level. For the observation modeling, we develop a progressive observation modeling process that is able to provide strong tracking observations and greatly facilitate the tracking task. For the tracking strategy, we propose a dual-mode two-way Bayesian inference approach which dynamically switches between an offline general model and an online dedicated model to deal with single isolated object tracking and multiple occluded object tracking integrally by forward filtering and backward smoothing. Extensive experiments on different kinds of sports videos, including football, basketball, as well as hockey, demonstrate the effectiveness and efficiency of the proposed method.
Junliang Xing, Haizhou Ai, Liwei Liu 0006, Shihong Lao
IEEE Trans. Image Process.1
2010 Multiple Human Tracking Based on Multi-view Upper-Body Detection and Discriminative Learning
abstract
This paper focuses on the problem of tracking multiple humans in dense environments which is very challenging due to recurring occlusions between different humans. To cope with the difficulties it presents, an offline boosted multi-view upper-body detector is used to automatically initialize a new human trajectory and is capable of dealing with partial human occlusions. What is more, an online learning process is proposed to learn discriminative human observations, including discriminative interest points and color patches, to effectively track each human when even more occlusions occur. The offline and online observation models are neatly integrated into the particle filter framework to robustly track multiple highly interactive humans. Experiments results on CAVIAR dataset as well as many other challenging real-world cases demonstrate the effectiveness of the proposed method.
Junliang Xing, Haizhou Ai, Shihong Lao
ICPR1
2009 Multi-object tracking through occlusions by local tracklets filtering and global tracklets association with detection responses
abstract
This paper presents an online detection-based two-stage multi-object tracking method in dense visual surveillances scenarios with a single camera. In the local stage, a particle filter with observer selection that could deal with partial object occlusion is used to generate a set of reliable tracklets. In the global stage, the detection responses are collected from a temporal sliding window to deal with ambiguity caused by full object occlusion to generate a set of potential tracklets. The reliable tracklets generated in the local stage and the potential tracklets generated within the temporal sliding window are associated by Hungarian algorithm on a modified pairwise tracklets association cost matrix to get the global optimal association. This method is applied to the pedestrian class and evaluated on two challenging datasets. The experimental results prove the effectiveness of our method.
Junliang Xing, Haizhou Ai, Shihong Lao
CVPR1