EDBT 2026 Demo / reviewers in the wild / expert
Jingbo Wang 0003
dblp:10/1491-3
· DBLP profile ↗
42ranked-venue papers
7as first author
31since 2021 · last 2025
0000-0001-9700-6262ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 7 first-author · 25 since 2021Artificial intelligence and machine learning · 34 · 6 first-author · 26 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ScaMo: Exploring the Scaling Law in Autoregressive Motion Generation ModelabstractThe scaling law has been validated in various domains, such as natural language processing (NLP) and massive computer vision tasks; however, its application to motion generation remains largely unexplored. In this paper, we introduce a scalable motion generation framework that includes the motion tokenizer Motion FSQ-VAE and a text-prefix autoregressive transformer. Through comprehensive experiments, we observe the scaling behavior of this system. For the first time, we confirm the existence of scaling laws within the context of motion generation. Specifically, our results demonstrate that the normalized test loss of our prefix autoregressive models adheres to a logarithmic law in relation to compute budgets. Furthermore, we also confirm the power law between Non-Vocabulary Parameters, Vocabulary Parameters, and Data Tokens with respect to compute budgets respectively. Leveraging the scaling law, we predict the optimal transformer size, vocabulary size, and data requirements for a compute budget of 1e18. The test loss of the system, when trained with the optimal model size, vocabulary size, and required data, aligns precisely with the predicted test loss, thereby validating the scaling law. Project page: https://shunlinlu.github.io/ScaMo/ Shunlin Lu, Jingbo Wang 0003, Wenxun Dai, Junting Dong, Zhiyang Dou, Bo Dai 0002, Ruimao Zhang |
CVPR | 2 |
| 2025 | TokenHSI: Unified Synthesis of Physical Human-Scene Interactions through Task TokenizationabstractSynthesizing diverse and physically plausible Human-Scene Interactions (HSI) is pivotal for both computer animation and embodied AI. Despite encouraging progress, current methods mainly focus on developing separate controllers, each specialized for a specific interaction task. This significantly hinders the ability to tackle a wide variety of challenging HSI tasks that require the integration of multiple skills, e.g. sitting down while carrying an object (see Fig. 1). To address this issue, we present TokenHSI, a single, unified transformer-based policy capable of multi-skill unification and flexible adaptation. The key insight is to model the humanoid proprioception as a separate shared token and combine it with distinct task tokens via a masking mechanism. Such a unified policy enables effective knowledge sharing across skills, thereby facilitating the multi-task training. Moreover, our policy architecture supports variable length inputs, enabling flexible adaptation of learned skills to new scenarios. By training additional task tokenizers, we can not only modify the geometries of interaction targets but also coordinate multiple skills to address complex tasks. The experiments demonstrate that our approach can significantly improve versatility, adaptability, and extensibility in various HSI tasks. Liang Pan, Zeshi Yang, Zhiyang Dou, Wenjia Wang 0009, Buzhen Huang, Bo Dai 0002, Taku Komura, Jingbo Wang 0003 |
CVPR | 8 |
| 2025 | DRiVE: Diffusion-based Rigging Empowers Generation of Versatile and Expressive CharactersabstractRecent advances in generative models have enabled high-quality 3D character reconstruction from multi-modal. However, animating these generated characters remains a challenging task, especially for complex elements like garments and hair, due to the lack of large-scale datasets and effective rigging methods. To address this gap, we curate AnimeRig, a large-scale dataset with detailed skeleton and skinning annotations. Building upon this, we propose DRiVE, a novel framework for generating and rigging 3D human characters with intricate structures. Unlike existing methods, DRiVE utilizes a 3D Gaussian representation, facilitating efficient animation and high-quality rendering. We further introduce GSDiff, a 3D Gaussian-based diffusion module that predicts joint positions as spatial distributions, overcoming the limitations of regression-based approaches. Extensive experiments demonstrate that DRiVE achieves precise rigging results, enabling realistic dynamics for clothing and hair, and surpassing previous methods in both quality and versatility. Code, dataset and visualization results are available at https://DRiVEAvatar.github.io/. Junting Dong, Yurun Chen 0001, Shiwei Mao, Puhua Jiang, Jingbo Wang 0003, Bo Dai 0002, Ruqi Huang |
CVPR | 8 |
| 2025 | Go to Zero: Towards Zero-Shot Motion Generation with Million-Scale DataabstractGenerating diverse and natural human motion sequences based on textual descriptions constitutes a fundamental and challenging research area within the domains of computer vision, graphics, and robotics. Despite significant advancements in this field, current methodologies often face challenges regarding zero-shot generalization capabilities, largely attributable to the limited size of training datasets. Moreover, the lack of a comprehensive evaluation framework impedes the advancement of this task by failing to identify directions for improvement. In this work, we aim to push text-to-motion into a new era, that is, to achieve the generalization ability of zero-shot. To this end, firstly, we develop an efficient annotation pipeline and introduce MotionMillion-the largest human motion dataset to date, featuring over 2,000 hours and 2 million high-quality motion sequences. Additionally, we propose MotionMillion-Eval, the most comprehensive benchmark for evaluating zero-shot motion generation. Leveraging a scalable architecture, we scale our model to 7B parameters and validate its performance on MotionMillion-Eval. Our results demonstrate strong generalization to out-of-domain and complex compositional motions, marking a significant step toward zero-shot human motion generation. The code is available at https://github.com/VankouF/MotionMillion-Codes. Shunlin Lu, Minyue Dai, Runyi Yu 0003, Lixing Xiao, Zhiyang Dou, Junting Dong, Lizhuang Ma, Jingbo Wang 0003 |
ICCV | 9 |
| 2025 | ARMO: Autoregressive Rigging for Multi-Category ObjectsabstractRecent advancements in large-scale generative models have significantly improved the quality and diversity of 3D shape generation. However, most existing methods focus primarily on generating static 3D models, overlooking the potentially dynamic nature of certain shapes, such as humanoids, animals, and insects. To address this gap, we focus on rigging, a fundamental task in animation that establishes skeletal structures and skinning for 3D models. In this paper, we introduce OmniRig, the first large-scale rigging dataset, comprising 79,499 meshes with detailed skeleton and skinning information. Unlike traditional benchmarks that rely on predefined standard poses (e.g., A-pose, T-pose), our dataset embraces diverse shape categories, styles, and poses. Leveraging this rich dataset, we propose ARMO, a novel rigging framework that utilizes an autoregressive model to predict both joint positions and connectivity relationships in a unified manner. By treating the skeletal structure as a complete graph and discretizing it into tokens, we encode the joints using an auto-encoder to obtain a latent embedding and an autoregressive model to predict the tokens. A mesh-conditioned latent diffusion model is used to predict the latent embedding for conditional skeleton generation. Our method addresses the limitations of regression-based approaches, which often suffer from error accumulation and suboptimal connectivity estimation. Through extensive experiments on the OmniRig dataset, our approach achieves state-of-the-art performance in skeleton prediction, demonstrating improved generalization across diverse object categories. The code and dataset will be made public for academic use upon acceptance. Shiwei Mao, Keyi Chen 0013, Yurun Chen 0001, Shunlin Lu, Jingbo Wang 0003, Junting Dong, Ruqi Huang |
ICCV | 6 |
| 2025 | SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation
Wenjia Wang 0009, Liang Pan, Zhiyang Dou, Jidong Mei, Zhouyingcheng Liao, Yuke Lou, Yifan Wu 0039, Lei Yang 0045, Jingbo Wang 0003, Taku Komura |
ICCV | 9 |
| 2025 | MotionStreamer: Streaming Motion Generation via Diffusion-Based Autoregressive Model in Causal Latent Space
Lixing Xiao, Shunlin Lu, Huaijin Pi, Liang Pan, Yueer Zhou, Ziyong Feng, Xiaowei Zhou 0001, Sida Peng, Jingbo Wang 0003 |
ICCV | 10 |
| 2025 | 🎧MOSPA: Human Motion Generation Driven by Spatial AudioabstractEnabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite its significance, this task remains largely unexplored. Most previous works have primarily focused on mapping modalities like speech, audio, and music to generate human motion. As of yet, these models typically overlook the impact of spatial features encoded in spatial audio signals on human motion. To bridge this gap and enable high-quality modeling of human movements in response to spatial audio, we introduce the first comprehensive "Spatial Audio-Driven Human Motion" (SAM) dataset, which contains diverse and high-quality spatial audio and motion data. For benchmarking, we develop a simple yet effective diffusion-based generative framework for human "MOtion generation driven by SPatial Audio," termed MOSPA, which faithfully captures the relationship between body motion and spatial audio through an effective fusion mechanism. Once trained, MOSPA can generate diverse realistic human motions conditioned on varying spatial audio inputs. We perform a thorough investigation of the proposed dataset and conduct extensive experiments for benchmarking, where our method achieves state-of-the-art performance on this task. Our code and model are publicly available at https://github.com/xsy27/Mospa-Acoustic-driven-Motion-Generation.git Shuyang Xu, Zhiyang Dou, Mingyi Shi, Liang Pan, Leo Ho, Jingbo Wang 0003, Yuan Liu 0025, Cheng Lin 0001, Yuexin Ma, Wenping Wang 0001, Taku Komura |
NeurIPS | 6 |
| 2025 | Motion2Motion: Cross-topology Motion Transfer with Sparse CorrespondenceabstractThis work studies the challenge of transfer animations between characters whose skeletal topologies differ substantially. While many techniques have advanced retargeting techniques in decades, transfer motions across diverse topologies remains less-explored. The primary obstacle lies in the inherent topological inconsistency between source and target skeletons, which restricts the establishment of straightforward one-to-one bone correspondences. Besides, the current lack of large-scale paired motion datasets spanning different topological structures severely constrains the development of data-driven approaches. To address these limitations, we introduce Motion2Motion, a novel, training-free framework. Simply yet effectively, Motion2Motion works with only one or a few example motions on the target skeleton, by accessing a sparse set of bone correspondences between the source and target skeletons. Through comprehensive qualitative and quantitative evaluations, we demonstrate that Motion2Motion achieves efficient and reliable performance in both similar-skeleton and cross-species skeleton transfer scenarios. The practical utility of our approach is further evidenced by its successful integration in downstream applications and user interfaces, highlighting its potential for industrial applications. Code and data are available at https://lhchen.top/Motion2Motion. Zixin Yin, Zhiyang Dou, Xin Chen 0040, Jingbo Wang 0003, Taku Komura, Lei Zhang 0001 |
SIGGRAPH Asia | 6 |
| 2024 | PACER+: On-Demand Pedestrian Animation Controller in Driving ScenariosabstractWe address the challenge of content diversity and controllability in pedestrian simulation for driving scenarios. Recent pedestrian animation frameworks have a significant limitation wherein they primarily focus on either following trajectory [48] or the content of the reference video [60], consequently overlooking the potential diversity of human motion within such scenarios. This limitation restricts the ability to generate pedestrian behaviors that exhibit a wider range of variations and realistic motions and therefore re-stricts its usage to provide rich motion content for other components in the driving simulation system, e.g., suddenly changed motion to which the autonomous vehicle should respond. In our approach, we strive to surpass the limitation by showcasing diverse human motions obtained from various sources, such as generated human motions, in ad-dition to following the given trajectory. The fundamental contribution of our framework lies in combining the motion tracking task with trajectory following, which enables the tracking of specific motion parts (e.g., upper body) while simultaneously following the given trajectory by a single policy. This way, we significantly enhance both the diver-sity of simulated human motion within the given scenario and the controllability of the content, including language-based control. Our framework facilitates the generation of a wide range of human motions, contributing to greater re-alism and adaptability in pedestrian simulations for driving scenarios. Jingbo Wang 0003, Zhengyi Luo 0002, Ye Yuan 0007, Yixuan Li 0002, Bo Dai 0002 |
CVPR | 1 |
| 2024 | Cinematic Behavior Transfer via NeRF-based Differentiable FilmingabstractIn the evolving landscape of digital media and video pro-duction, the precise manipulation and reproduction of visual elements like camera movements and character actions are highly desired. Existing SLAM methods face limitations in dynamic scenes and human pose estimation often focuses on 2D projections, neglecting 3D statuses. To address these is-sues, we first introduce a reverse filming behavior estimation technique. It optimizes camera trajectories by leveraging NeRF as a differentiable renderer and refining SMPL tracks. We then introduce a cinematic transfer pipeline that is able to transfer various shot types to a new 2D video or a 3D vir-tual environment. The incorporation of 3D engine workflow enables superior rendering and control abilities, which also achieves a higher rating in the user study. Xuekun Jiang, Anyi Rao, Jingbo Wang 0003, Dahua Lin, Bo Dai 0002 |
CVPR | 3 |
| 2024 | MotionLCM: Real-Time Controllable Motion Generation via Latent Consistency Model
Wenxun Dai, Jingbo Wang 0003, Bo Dai 0002, Yansong Tang |
ECCV (16) | 3 |
| 2024 | TELA: Text to Layer-Wise 3D Clothed Human Generation
Junting Dong, Zehuan Huang, Xudong Xu, Jingbo Wang 0003, Sida Peng, Bo Dai 0002 |
ECCV (25) | 5 |
| 2024 | SemGrasp : Semantic Grasp Generation via Language Aligned Discretization
Kailin Li 0001, Jingbo Wang 0003, Lixin Yang 0001, Cewu Lu, Bo Dai 0002 |
ECCV (2) | 2 |
| 2024 | RoomTex: Texturing Compositional Indoor Scenes via Iterative Inpainting
Qi Wang 0105, Ruijie Lu, Xudong Xu, Jingbo Wang 0003, Michael Yu Wang, Bo Dai 0002, Dan Xu 0002 |
ECCV (68) | 4 |
| 2024 | EMDM: Efficient Motion Diffusion Model for Fast and High-Quality Motion Generation
Wenyang Zhou, Zhiyang Dou, Zeyu Cao, Zhouyingcheng Liao, Jingbo Wang 0003, Wenjia Wang 0009, Yuan Liu 0025, Taku Komura, Wenping Wang 0001, Lingjie Liu |
ECCV (2) | 5 |
| 2024 | Unified Human-Scene Interaction via Prompted Chain-of-ContactsabstractHuman-Scene Interaction (HSI) is a vital component of fields like embodied AI and virtual reality. Despite advancements in motion quality and physical plausibility, two pivotal factors, versatile interaction control and the development of a user-friendly interface, require further exploration before the practical application of HSI. This paper presents a unified HSI framework, UniHSI, which supports unified control of diverse interactions through language commands. The framework defines interaction as ``Chain of Contacts (CoC)", representing steps involving human joint-object part pairs. This concept is inspired by the strong correlation between interaction types and corresponding contact regions. Based on the definition, UniHSI constitutes a Large Language Model (LLM) Planner to translate language prompts into task plans in the form of CoC, and a Unified Controller that turns CoC into uniform task execution. To facilitate training and evaluation, we collect a new dataset named ScenePlan that encompasses thousands of task plans generated by LLMs based on diverse scenarios. Comprehensive experiments demonstrate the effectiveness of our framework in versatile task execution and generalizability to real scanned scenes. Zeqi Xiao, Jingbo Wang 0003, Jinkun Cao, Bo Dai 0002, Dahua Lin, Jiangmiao Pang |
ICLR | 3 |
| 2024 | Open-set Hierarchical Semantic Segmentation for 3D SceneabstractThe Segment-Anything Model (SAM) shows exceptional zero-shot capabilities for 2D images. Developing a similar model for 3D, however, is challenging due to limited datasets. In this paper, we introduce a zero-shot algorithm to segment a 3D scene into elements at various levels of detail, and further organize the results in a hierarchical tree structure. We propose a tree quality metric to evaluate the algorithm’s performance. Notably, our algorithm eliminates the need for 3D annotations. It uses robust 2D models to generate a 2D segmentation tree for each rendered image. Then, using graph neural networks, it aggregates these 2D trees to form a unified 3D segmentation tree. Extensive experiments on the PartNet dataset and complex 3D scenes validate the algorithm’s effectiveness. We release the source code at https://github.com/dnvtmf/OTS. Diwen Wan, Jiaxiang Tang, Jingbo Wang 0003, Xiaokang Chen, Lingyun Gan |
ICME | 3 |
| 2024 | InterControl: Zero-shot Human Interaction Generation by Controlling Every JointabstractText-conditioned motion synthesis has made remarkable progress with the emergence of diffusion models. However, the majority of these motion diffusion models are primarily designed for a single character and overlook multi-human interactions. In our approach, we strive to explore this problem by synthesizing human motion with interactions for a group of characters of any size in a zero-shot manner. The key aspect of our approach is the adaptation of human-wise interactions as pairs of human joints that can be either in contact or separated by a desired distance. In contrast to existing methods that necessitate training motion generation models on multi-human motion datasets with a fixed number of characters, our approach inherently possesses the flexibility to model human interactions involving an arbitrary number of individuals, thereby transcending the limitations imposed by the training data. We introduce a novel controllable motion generation method, InterControl, to encourage the synthesized motions maintaining the desired distance between joint pairs. It consists of a motion controller and an inverse kinematics guidance module that realistically and accurately aligns the joints of synthesized characters to the desired location. Furthermore, we demonstrate that the distance between joint pairs for human-wise interactions can be generated using an off-the-shelf Large Language Model (LLM). Experimental results highlight the capability of our framework to generate interactions with multiple human characters and its potential to work with off-the-shelf physics-based character simulators. Code is available at https://github.com/zhenzhiwang/intercontrol. Zhenzhi Wang 0001, Jingbo Wang 0003, Yixuan Li 0002, Dahua Lin, Bo Dai 0002 |
NeurIPS | 2 |
| 2024 | CooHOI: Learning Cooperative Human-Object Interaction with Manipulated Object DynamicsabstractEnabling humanoid robots to clean rooms has long been a pursued dream within humanoid research communities. However, many tasks require multi-humanoid collaboration, such as carrying large and heavy furniture together. Given the scarcity of motion capture data on multi-humanoid collaboration and the efficiency challenges associated with multi-agent learning, these tasks cannot be straightforwardly addressed using training paradigms designed for single-agent scenarios. In this paper, we introduce **Coo**perative **H**uman-**O**bject **I**nteraction (**CooHOI**), a framework designed to tackle the challenge of multi-humanoid object transportation problem through a two-phase learning paradigm: individual skill learning and subsequent policy transfer. First, a single humanoid character learns to interact with objects through imitation learning from human motion priors. Then, the humanoid learns to collaborate with others by considering the shared dynamics of the manipulated object using centralized training and decentralized execution (CTDE) multi-agent RL algorithms. When one agent interacts with the object, resulting in specific object dynamics changes, the other agents learn to respond appropriately, thereby achieving implicit communication and coordination between teammates. Unlike previous approaches that relied on tracking-based methods for multi-humanoid HOI, CooHOI is inherently efficient, does not depend on motion capture data of multi-humanoid interactions, and can be seamlessly extended to include more participants and a wide range of object types. Jiawei Gao 0004, Ziqin Wang, Zeqi Xiao, Jingbo Wang 0003, Jinkun Cao, Xiaolin Hu 0001, Si Liu 0001, Jifeng Dai, Jiangmiao Pang |
NeurIPS | 4 |
| 2024 | Task-Oriented Human-Object Interactions Generation with Implicit Neural RepresentationsabstractDigital human motion synthesis is a vibrant research field with applications in movies, AR/VR, and video games. Whereas methods were proposed to generate natural and realistic human motions, most only focus on modeling humans and largely ignore object movements. Generating task-oriented human-object interaction motions in simulation is challenging. For different intents of using the objects, humans conduct various motions, which requires the human first to approach the objects and then make them move consistently with the human instead of staying still. Also, to deploy in downstream applications, the synthesized motions are desired to be flexible in length, providing options to personalize the predicted motions for various purposes. To this end, we propose TOHO: Task-Oriented Human-Object Interactions Generation with Implicit Neural Representations, which generates full human-object interaction motions to conduct specific tasks, given only the task type, the object, and a starting human status. TOHO generates human-object motions in four steps: 1) it first estimates the object’s final position given the task intent; 2) it then generates keyframe poses grasping the objects; 3) after that, it infills the keyframes and generates continuous motions; 4) finally, it applies a compact closed-form object motion estimation to generate the object motion. Our method generates continuous motions that are parameterized only by the temporal coordinate, which allows for upsampling of the sequence to arbitrary frames and adjusting the motion speeds by designing the temporal coordinate vector. This work takes a step further toward general human-scene interaction simulation. Quanzhou Li, Jingbo Wang 0003, Chen Change Loy, Bo Dai 0002 |
WACV | 2 |
| 2024 | Interactive Character Control with Auto-Regressive Motion Diffusion ModelsabstractReal-time character control is an essential component for interactive experiences, with a broad range of applications, including physics simulations, video games, and virtual reality. The success of diffusion models for image synthesis has led to the use of these models for motion synthesis. However, the majority of these motion diffusion models are primarily designed for offline applications, where space-time models are used to synthesize an entire sequence of frames simultaneously with a pre-specified length. To enable real-time motion synthesis with diffusion model that allows time-varying controls, we propose A-MDM (Auto-regressive Motion Diffusion Model). Our conditional diffusion model takes an initial pose as input, and auto-regressively generates successive motion frames conditioned on the previous frame. Despite its streamlined network architecture, which uses simple MLPs, our framework is capable of generating diverse, long-horizon, and high-fidelity motion sequences. Furthermore, we introduce a suite of techniques for incorporating interactive controls into A-MDM, such as task-oriented sampling, in-painting, and hierarchical reinforcement learning (See Figure 1). These techniques enable a pre-trained A-MDM to be efficiently adapted for a variety of new downstream tasks. We conduct a comprehensive suite of experiments to demonstrate the effectiveness of A-MDM, and compare its performance against state-of-the-art auto-regressive methods. Yi Shi 0008, Jingbo Wang 0003, Xuekun Jiang, Bingkun Lin, Bo Dai 0002, Xue Bin Peng |
ACM Trans. Graph. | 2 |
| 2023 | DNA-Rendering: A Diverse Neural Actor Repository for High-Fidelity Human-centric RenderingabstractRealistic human-centric rendering plays a key role in both computer vision and computer graphics. Rapid progress has been made in the algorithm aspect over the years, yet existing human-centric rendering datasets and benchmarks are rather impoverished in terms of diversity (e.g., outfit's fabric/material, body's interaction with objects, and motion sequences), which are crucial for rendering effect. Researchers are usually constrained to explore and evaluate a small set of rendering problems on current datasets, while real-world applications require methods to be robust across different scenarios. In this work, we present DNA-Rendering, a large-scale, high-fidelity repository of human performance data for neural actor rendering. DNA-Rendering presents several appealing attributes. First, our dataset contains over 1500 human subjects, 5000 motion sequences, and 67.5M frames' data volume. Upon the massive collections, we provide human subjects with grand categories of pose actions, body shapes, clothing, accessories, hairdos, and object intersection, which ranges the geometry and appearance variances from everyday life to professional occasions. Second, we provide rich assets for each subject – 2D/3D human body keypoints, foreground masks, SMPLX models, cloth/accessory materials, multi-view images, and videos. These assets boost the current method's accuracy on downstream rendering tasks. Third, we construct a professional multi-view system to capture data, which contains 60 synchronous cameras with max 4096 × 3000 resolution, 15 fps speed, and stern camera calibration steps, ensuring high-quality resources for task training and evaluation.Along with the dataset, we provide a large-scale and quantitative benchmark in full-scale, with multiple tasks to evaluate the existing progress of novel view synthesis, novel pose animation synthesis, and novel identity rendering methods. In this manuscript, we describe our DNA-Rendering effort as a revealing of new observations, challenges, and future directions to human-centric rendering. The dataset, code, and benchmarks will be publicly available at https://dna-rendering.github.io/. Ruixiang Chen, Siming Fan, Wanqi Yin, Zhongang Cai, Jingbo Wang 0003, Yang Gao 0042, Zhengming Yu, Zhengyu Lin, Daxuan Ren, Lei Yang 0045, Ziwei Liu 0002, Chen Change Loy, Chen Qian 0006, Wayne Wu, Dahua Lin, Bo Dai 0002, Kwan-Yee Lin |
ICCV | 7 |
| 2023 | Learning Human Dynamics in Autonomous Driving ScenariosabstractSimulation has emerged as an indispensable tool for scaling and accelerating the development of self-driving systems. A critical aspect of this is simulating realistic and diverse human behavior and intent. In this work, we propose a holistic framework for learning physically plausible human dynamics from real driving scenarios, narrowing the gap between real and simulated human behavior in safety-critical applications. We show that state-of-the-art methods underperform in driving scenarios where video data is recorded from moving vehicles, and humans are frequently partially or fully occluded. Furthermore, existing methods often disregard the global scene where humans are situated, resulting in various motion artifacts like foot sliding, floating, or ground penetration. To address this challenge, we propose an approach that incorporates physics with a reinforcement learning-based motion controller to learn human dynamics for driving scenarios. Our framework can simulate physically plausible human dynamics that accurately match observed human motions and infill motions for occluded body parts, while improving the physical plausibility of the entire motion sequence. Experiments on the challenging Waymo Open Dataset show that our method outperforms state-of-the-art motion capture approaches significantly in recovering high-quality, physically plausible, and scene-aware human dynamics. Jingbo Wang 0003, Ye Yuan 0007, Zhengyi Luo 0002, Kevin Xie, Dahua Lin, Umar Iqbal 0001, Sanja Fidler, Sameh Khamis |
ICCV | 1 |
| 2022 | Not All Voxels Are Equal: Semantic Scene Completion from the Point-Voxel PerspectiveabstractWe revisit Semantic Scene Completion (SSC), a useful task to predict the semantic and occupancy representation of 3D scenes, in this paper. A number of methods for this task are always based on voxelized scene representations. Although voxel representations keep local structures of the scene, these methods suffer from heavy computation redundancy due to the existence of visible empty voxels when the network goes deeper. To address this dilemma, we propose our novel point-voxel aggregation network for this task. We first transfer the voxelized scenes to point clouds by removing these visible empty voxels and adopt a deep point stream to capture semantic information from the scene efficiently. Meanwhile, a light-weight voxel stream containing only two 3D convolution layers preserves local structures of the voxelized scenes. Furthermore, we design an anisotropic voxel aggregation operator to fuse the structure details from the voxel stream into the point stream, and a semantic-aware propagation module to enhance the up-sampling process in the point stream by semantic labels. We demonstrate that our model surpasses state-of-the-arts on two benchmarks by a large margin, with only the depth images as input. Jiaxiang Tang, Xiaokang Chen, Jingbo Wang 0003 |
AAAI | 3 |
| 2022 | Towards Diverse and Natural Scene-aware 3D Human Motion SynthesisabstractThe ability to synthesize long-term human motion sequences in real-world scenes can facilitate numerous applications. Previous approaches for scene-aware motion synthesis are constrained by pre-defined target objects or positions and thus limit the diversity of human-scene interactions for synthesized motions. In this paper, we focus on the problem of synthesizing diverse scene-aware human motions under the guidance of target action sequences. To achieve this, we first decompose the diversity of scene-aware human motions into three aspects, namely interaction diversity (e.g. sitting on different objects with different poses in the given scenes), path diversity (e.g. moving to the target locations following different paths), and the motion diversity (e.g. having various body movements during moving). Based on this factorized scheme, a hierarchical framework is proposed, with each sub-module responsible for modeling one aspect. We assess the effectiveness of our framework on two challenging datasets for scene-aware human motion synthesis. The experiment results show that the proposed framework remarkably outperforms previous methods in terms of diversity and naturalness. Jingbo Wang 0003, Yu Rong 0003, Sijie Yan, Dahua Lin, Bo Dai 0002 |
CVPR | 1 |
| 2022 | Point Scene Understanding via Disentangled Instance Mesh Reconstruction
Jiaxiang Tang, Xiaokang Chen, Jingbo Wang 0003 |
ECCV (32) | 3 |
| 2022 | Compressible-composable NeRF via Rank-residual DecompositionabstractNeural Radiance Field (NeRF) has emerged as a compelling method to represent 3D objects and scenes for photo-realistic rendering. However, its implicit representation causes difficulty in manipulating the models like the explicit mesh representation.Several recent advances in NeRF manipulation are usually restricted by a shared renderer network, or suffer from large model size. To circumvent the hurdle, in this paper, we present a neural field representation that enables efficient and convenient manipulation of models.To achieve this goal, we learn a hybrid tensor rank decomposition of the scene without neural networks. Motivated by the low-rank approximation property of the SVD algorithm, we propose a rank-residual learning strategy to encourage the preservation of primary information in lower ranks. The model size can then be dynamically adjusted by rank truncation to control the levels of detail, achieving near-optimal compression without extra optimization.Furthermore, different models can be arbitrarily transformed and composed into one scene by concatenating along the rank dimension.The growth of storage cost can also be mitigated by compressing the unimportant objects in the composed scene. We demonstrate that our method is able to achieve comparable rendering quality to state-of-the-art methods, while enabling extra capability of compression and composition.Code is available at https://github.com/ashawkey/CCNeRF. Jiaxiang Tang, Xiaokang Chen, Jingbo Wang 0003 |
NeurIPS | 3 |
| 2021 | Monocular 3D Reconstruction of Interacting Hands via Collision-Aware Factorized Refinementsabstract3D interacting hand reconstruction is essential to facilitate human-machine interaction and human behaviors understanding. Previous works in this field either rely on auxiliary inputs such as depth images or they can only handle a single hand if monocular single RGB images are used. Single-hand methods tend to generate collided hand meshes, when applied to closely interacting hands, since they cannot model the interactions between two hands explicitly. In this paper, we make the first attempt to reconstruct 3D interacting hands from monocular single RGB images. Our method can generate 3D hand meshes with both precise 3D poses and minimal collisions. This is made possible via a two-stage framework. Specifically, the first stage adopts a convolutional neural network to generate coarse predictions that tolerate collisions but encourage pose-accurate hand meshes. The second stage progressively ameliorates the collisions through a series of factorized refinements while retaining the preciseness of 3D poses. We carefully investigate potential implementations for the factorized refinement, considering the trade-off between efficiency and accuracy. Extensive quantitative and qualitative results on large-scale datasets such as InterHand2.6M demonstrate the effectiveness of the proposed approach. Yu Rong 0003, Jingbo Wang 0003, Ziwei Liu 0002, Chen Change Loy |
3DV | 2 |
| 2021 | Scene-Aware Generative Network for Human Motion SynthesisabstractWe revisit human motion synthesis, a task useful in various real-world applications, in this paper. Whereas a number of methods have been developed previously for this task, they are often limited in two aspects: 1) focus on the poses while leaving the location movement behind, and 2) ignore the impact of the environment on the human motion. In this paper, we propose a new framework, with the interaction between the scene and the human motion taken into account. Considering the uncertainty of human motion, we formulate this task as a generative task, whose objective is to generate plausible human motion conditioned on both the scene and the human’s initial position. This framework factorizes the distribution of human motions into a distribution of movement trajectories conditioned on scenes and that of body pose dynamics conditioned on both scenes and trajectories. We further derive a GAN-based learning approach, with discriminators to enforce the compatibility between the human motion and the contextual scene as well as the 3D-to-2D projection constraints. We assess the effectiveness of the proposed method on two challenging datasets, which cover both synthetic and real-world environments. Jingbo Wang 0003, Sijie Yan, Bo Dai 0002, Dahua Lin |
CVPR | 1 |
| 2021 | BiSeNet V2: Bilateral Network with Guided Aggregation for Real-Time Semantic Segmentation
Changqian Yu, Changxin Gao, Jingbo Wang 0003, Gang Yu 0002, Chunhua Shen, Nong Sang |
Int. J. Comput. Vis. | 3 |
| 2020 | Context Prior for Scene SegmentationabstractRecent works have widely explored the contextual dependencies to achieve more accurate segmentation results. However, most approaches rarely distinguish different types of contextual dependencies, which may pollute the scene understanding. In this work, we directly supervise the feature aggregation to distinguish the intra-class and interclass context clearly. Specifically, we develop a Context Prior with the supervision of the Affinity Loss. Given an input image and corresponding ground truth, Affinity Loss constructs an ideal affinity map to supervise the learning of Context Prior. The learned Context Prior extracts the pixels belonging to the same category, while the reversed prior focuses on the pixels of different classes. Embedded into a conventional deep CNN, the proposed Context Prior Layer can selectively capture the intra-class and inter-class contextual dependencies, leading to robust feature representation. To validate the effectiveness, we design an effective Context Prior Network (CPNet). Extensive quantitative and qualitative evaluations demonstrate that the proposed model performs favorably against state-of-the-art semantic segmentation approaches. More specifically, our algorithm achieves 46.3% mIoU on ADE20K, 53.9% mIoU on PASCAL-Context, and 81.3% mIoU on Cityscapes. Code is available at https://git.io/ContextPrior. Changqian Yu, Jingbo Wang 0003, Changxin Gao, Gang Yu 0002, Chunhua Shen, Nong Sang |
CVPR | 2 |
| 2020 | Bi-directional Cross-Modality Feature Propagation with Separation-and-Aggregation Gate for RGB-D Semantic Segmentation
Xiaokang Chen, Kwan-Yee Lin, Jingbo Wang 0003, Wayne Wu, Chen Qian 0006, Hongsheng Li 0001 |
ECCV (11) | 3 |
| 2020 | Motion Guided 3D Pose Estimation from Videos
Jingbo Wang 0003, Sijie Yan, Yuanjun Xiong, Dahua Lin |
ECCV (13) | 1 |
| 2020 | Malleable 2.5D Convolution: Learning Receptive Fields Along the Depth-Axis for RGB-D Scene Parsing
Yajie Xing, Jingbo Wang 0003 |
ECCV (19) | 2 |
| 2019 | An End-To-End Network for Panoptic SegmentationabstractPanoptic segmentation, which needs to assign a category label to each pixel and segment each object instance simultaneously, is a challenging topic. Traditionally, the existing approaches utilize two independent models without sharing features, which makes the pipeline inefficient to implement. In addition, a heuristic method is usually employed to merge the results. However, the overlapping relationship between object instances is difficult to determine without sufficient context information during the merging process. To address the problems, we propose a novel end-to-end Occlusion Aware Network (OANet) for panoptic segmentation, which can efficiently and effectively predict both the instance and stuff segmentation in a single network. Moreover, we introduce a novel spatial ranking module to deal with the occlusion problem between the predicted instances. Extensive experiments have been done to validate the performance of our proposed method and promising results have been achieved on the COCO Panoptic benchmark. Chao Peng 0001, Changqian Yu, Jingbo Wang 0003, Xu Liu 0017, Gang Yu 0002, Wei Jiang 0009 |
CVPR | 4 |
| 2019 | 2.5D Convolution for RGB-D Semantic SegmentationabstractConvolutional neural networks (CNN) have achieved great success in RGB semantic segmentation. RGB-D images provide additional depth information, which can improve segmentation performance. To take full advantages of the 3D geometry relations provided by RGB-D images, in this paper, we propose 2.5D convolution, which mimics one 3D convolution kernel by several masked 2D convolution kernels. Our 2.5D convolution can effectively process spatial relations between pixels in a manner similar to 3D convolution while still sampling pixels on 2D plane, and thus saves computational cost. And it can be seamlessly incorporated into pretrained CNNs. Experiments on two challenging RGB-D semantic segmentation benchmarks NYUDv2 and SUN-RGBD validate the effectiveness of our approach. Yajie Xing, Jingbo Wang 0003, Xiaokang Chen |
ICIP | 2 |
| 2019 | Coupling Two-Stream RGB-D Semantic Segmentation Network by Idempotent MappingsabstractIn RGB-D semantic segmentation tasks, it has been shown that HHA embeddings effectively encode rich depth features and using HHA together with RGB images can improve segmentation performance. In this paper, we propose a novel method to effectively integrate RGB and HHA features. By replacing identity mappings in ResNet-based two-stream network with idempotent mappings, we can couple the originally separated two branches to mix features from two modalities, while still keep the good information flow nature of ResNet. Moreover, our method does not bring any additional network blocks or parameters, and only needs very small modification on basic two-stream networks. We conduct experiments on two challenging RGB-D semantic segmentation datasets NYUDv2 and SUN-RGBD. The experiment results show that our method can significantly improve segmentation performance and our method achieves the state-of-the-art on these two datasets. Yajie Xing, Jingbo Wang 0003, Xiaokang Chen |
ICIP | 2 |
| 2018 | Learning a Discriminative Feature Network for Semantic SegmentationabstractMost existing methods of semantic segmentation still suffer from two aspects of challenges: intra-class inconsistency and inter-class indistinction. To tackle these two problems, we propose a Discriminative Feature Network (DFN), which contains two sub-networks: Smooth Network and Border Network. Specifically, to handle the intra-class inconsistency problem, we specially design a Smooth Network with Channel Attention Block and global average pooling to select the more discriminative features. Furthermore, we propose a Border Network to make the bilateral features of boundary distinguishable with deep semantic boundary supervision. Based on our proposed DFN, we achieve state-of-the-art performance 86.2% mean IOU on PASCAL VOC 2012 and 80.3% mean IOU on Cityscapes dataset. Changqian Yu, Jingbo Wang 0003, Chao Peng 0001, Changxin Gao, Gang Yu 0002, Nong Sang |
CVPR | 2 |
| 2018 | BiSeNet: Bilateral Segmentation Network for Real-Time Semantic Segmentation
Changqian Yu, Jingbo Wang 0003, Chao Peng 0001, Changxin Gao, Gang Yu 0002, Nong Sang |
ECCV (13) | 2 |
| 2018 | Global Context Encoding for Salient Objects DetectionabstractDeep convolutional neural networks (CNNs) have gained their reputation for the success in various tasks in computer vision, including salient objects detection. However, it remains a challenge that the CNNs have repeated downsample operators and always create low-resolution predictions, which tend to loss details and finer structure of images. To detect and segment the salient objects well, it is also necessary to merge high-level semantic information and low-level fine details simultaneously. Thus, we propose a novel network structure with stage-wise refinement sub-structures. In addition, we exploit the essence of salient objects detection by encoding the global image context in a specifically designed module, which is applied to every stage of the refinement structure. So the coarse saliency map generated from the base CNN can be refined with low-level feature and global context information step-by-step. Experimental results have demonstrated that the proposed method outperforms the state-of-the-art approaches on four benchmark datasets. Jingbo Wang 0003, Yajie Xing |
ICPR | 1 |
| 2018 | Attention Forest for Semantic Segmentation
Jingbo Wang 0003, Yajie Xing |
PRCV (1) | 1 |