EDBT 2026 Demo / reviewers in the wild / expert
Jiaman Li
dblp:211/7931
· DBLP profile ↗
22ranked-venue papers
9as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 8 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video GenerationabstractHuman-scene interaction (HSI) generation is crucial for applications in embodied AI, virtual reality, and robotics. Yet, existing methods cannot synthesize interactions in unseen environments such as in-the-wild scenes or reconstructed scenes, as they rely on paired 3D scenes and captured human motion data for training, which are unavailable for unseen environments. We present ZeroHSI, a novel approach that enables zero-shot 4D human-scene interaction synthesis, eliminating the need for training on any MoCap data. Our key insight is to distill human-scene interactions from state-of-the-art video generation models, which have been trained on vast amounts of natural human movements and interactions,and use differentiable rendering to reconstruct human-scene interactions. ZeroHSI can synthesize realistic human motions in both static scenes and environments with dynamic objects, without requiring any ground-truth motion data. We evaluate ZeroHSI on a curated dataset of different types of various indoor and outdoor scenes with different interaction prompts, demonstrating its ability to generate diverse and contextually appropriate human-scene interactions. Project page: https://awfuact.github.io/zerohsi. Hong-Xing Yu, Jiaman Li, Jiajun Wu 0001 |
3DV | 3 |
| 2026 | Trajectory and Latency-Aware DNN Task Offloading in Vehicular Edge Computing: A Deep Reinforcement Learning Approach
Yiyun Yang, Yukai Hao, Jiaman Li, Mian Ahmad Jan |
ICC | 3 |
| 2025 | Lifting Motion to the 3D World via 2D DiffusionabstractEstimating 3D motion from 2D observations is a longstanding research challenge. Prior work typically requires training on datasets containing ground truth 3D motions, limiting their applicability to activities well-represented in existing motion capture data. This dependency particularly hinders generalization to out-of-distribution scenarios or subjects where collecting 3D ground truth is challenging, such as complex athletic movements or animal motion. We introduce MVLift, a novel approach to predict global 3D motion—including both joint rotations and root trajectories in the world coordinate system—using only 2D pose sequences for training. Our multi-stage framework leverages 2D motion diffusion models to progressively generate consistent 2D pose sequences across multiple views, a key step in recovering accurate global 3D motion. MVLift generalizes across various domains, including human poses, human-object interactions, and animal poses. Despite not requiring 3D supervision, it outperforms prior work on five datasets, including those methods that require 3D supervision. Jiaman Li, C. Karen Liu, Jiajun Wu 0001 |
CVPR | 1 |
| 2025 | Hierarchical Decentralized Ring-Structured Federated Learning Approach for Collaborative Medical Image Analysis
Jiaman Li, Yujie Ye, Jing Lei 0007, Jianbo Du, Jiakai Wei, Celimuge Wu, Kok-Lim Alvin Yau |
GLOBECOM | 1 |
| 2025 | SFL-DCSA: Split Federated Learning for Breast Cancer Prediction with Dynamic Client Selection AggregationabstractAccompanied by the booming development of artificial intelligence technology, deep learning has been widely used in many fields of cancer, specially in cancer prediction. The dependence on training data for deep learning naturally raises privacy leakage concerns. Federated learning is a promising solution to these issues, but it is limited by the computational capacity of clients. Therefore, this paper has proposed a Split Federated Learning (SFL)-based breast cancer prediction scheme with dynamic client selection aggregation, by aggregating data from multiple healthcare organizations under the premise of privacy protection. First, the split learning is integrated into federated learning to predict breast canner, which contributes to protecting data privacy and reducing the computational burden on client devices. Then, the dynamic client selection aggregation is devised to lower the aggregation communication costs and improve communication efficiency by utilizing Long Short-Term Memory (LSTM) networks to evaluate the availability of each client devices in participating training. Finally, we have conducted extensive experiments on the CAMELYON16 dataset to evaluate the performance of our proposed scheme, and the experimental results have shown that our proposed scheme can converge faster and achieve the lower communication costs. Jiaman Li, Yiyun Yang, Rao Asad Mumtaz, Jianbo Du, Jiakai Wei, Kok-Lim Alvin Yau, Mian Ahmad Jan, Lei Liu 0031 |
ICC | 1 |
| 2025 | Human-Object Interaction from Human-level InstructionsabstractIntelligent agents must autonomously interact with the environments to perform daily tasks based on human-level instructions. They need a foundational understanding of the world to accurately interpret these instructions, along with precise low-level movement and interaction skills to execute the derived actions. In this work, we propose the first complete system for synthesizing physically plausible, long-horizon human-object interactions for object manipulation in contextual environments, driven by human-level instructions. We leverage large language models (LLMs) to interpret the input instructions into detailed execution plans. Unlike prior work, our system is capable of generating detailed finger-object interactions, in seamless coordination with full-body movements. We also train a policy to track generated motions in physics simulation via reinforcement learning (RL) to ensure physical plausibility of the motion. Our experiments demonstrate the effectiveness of our system in synthesizing realistic interactions with diverse objects in complex environments, highlighting its potential for real-world applications. Jiaman Li, Pei Xu 0005, C. Karen Liu |
ICCV | 2 |
| 2025 | Optimizing drone-captured maritime rescue image object detection through dataset rebalancing under sample constraints
Beigeng Zhao, Jiaman Li, Jiawen Zhao, Lizhi Yu, Jiren Liu |
Vis. Comput. | 2 |
| 2024 | Controllable Human-Object Interaction Synthesis
Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu 0001, Xavier Puig, C. Karen Liu |
ECCV (41) | 1 |
| 2023 | CIRCLE: Capture In Rich Contextual EnvironmentsabstractSynthesizing 3D human motion in a contextual, ecological environment is important for simulating realistic activities people perform in the real world. However, conventional optics-based motion capture systems are not suited for simultaneously capturing human movements and complex scenes. The lack of rich contextual 3D human motion datasets presents a roadblock to creating high-quality generative human motion models. We propose a novel motion acquisition system in which the actor perceives and operates in a highly contextual virtual world while being motion captured in the real world. Our system enables rapid collection of high-quality human motion in highly diverse scenes, without the concern of occlusion or the need for physical scene construction in the real world. We present CIRCLE, a dataset containing 10 hours of full-body reaching motion from 5 subjects across nine scenes, paired with ego-centric information of the environment represented in various forms, such as RGBD videos. We use this dataset to train a model that generates human motion conditioned on scene information. Leveraging our dataset, the model learns to use ego-centric scene information to achieve non-trivial reaching tasks in the context of complex 3D scenes. To download the data please visit our website. João Pedro Araújo 0001, Jiaman Li, Karthik Vetrivel, Rishi Agarwal, Jiajun Wu 0001, Deepak Gopinath, Alexander Clegg, C. Karen Liu |
CVPR | 2 |
| 2023 | Ego-Body Pose Estimation via Ego-Head Pose EstimationabstractEstimating 3D human motion from an egocentric video sequence plays a critical role in human behavior understanding and has various applications in VR/AR. However, naively learning a mapping between egocentric videos and human motions is challenging, because the user's body is often unobserved by the front-facing camera placed on the head of the user. In addition, collecting large-scale, high-quality datasets with paired egocentric videos and 3D human motions requires accurate motion capture devices, which often limit the variety of scenes in the videos to lab-like environments. To eliminate the need for paired egocentric video and human motions, we propose a new method, Ego-Body Pose Estimation via Ego-Head Pose Estimation (EgoEgo), which decomposes the problem into two stages, connected by the head motion as an intermediate representation. EgoEgo first integrates SLAM and a learning approach to estimate accurate head motion. Subsequently, leveraging the estimated head pose as input, EgoEgo utilizes conditional diffusion to generate multiple plausible full-body motions. This disentanglement of head and body pose eliminates the need for training datasets with paired egocentric videos and 3D human motion, enabling us to leverage large-scale egocentric video datasets and motion capture datasets separately. Moreover, for systematic benchmarking, we develop a synthetic dataset, AMASS-Replica-Ego-Syn (ARES), with paired egocentric videos and human motion. On both ARES and real data, our EgoEgo model performs significantly better than the current state-of-the-art methods. Jiaman Li, C. Karen Liu, Jiajun Wu 0001 |
CVPR | 1 |
| 2023 | Motion Question Answering via Modular Motion ProgramsabstractIn order to build artificial intelligence systems that can perceive and reason with human behavior in the real world, we must first design models that conduct complex spatio-temporal reasoning over motion sequences. Moving towards this goal, we propose the HumanMotionQA task to evaluate complex, multi-step reasoning abilities of models on long-form human motion sequences. We generate a dataset of question-answer pairs that require detecting motor cues in small portions of motion sequences, reasoning temporally about when events occur, and querying specific motion attributes. In addition, we propose NSPose, a neuro-symbolic method for this task that uses symbolic reasoning and a modular design to ground motion through learning motion concepts, attribute neural operators, and temporal relations. We demonstrate the suitability of NSPose for the HumanMotionQA task, outperforming all baseline methods. Mark Endo, Joy Hsu, Jiaman Li, Jiajun Wu 0001 |
ICML | 3 |
| 2023 | Object Motion Guided Human Motion SynthesisabstractModeling human behaviors in contextual environments has a wide range of applications in character animation, embodied AI, VR/AR, and robotics. In real-world scenarios, humans frequently interact with the environment and manipulate various objects to complete daily tasks. In this work, we study the problem of full-body human motion synthesis for the manipulation of large-sized objects. We propose Object MOtion guided human MOtion synthesis (OMOMO), a conditional diffusion framework that can generate full-body manipulation behaviors from only the object motion. Since naively applying diffusion models fails to precisely enforce contact constraints between the hands and the object, OMOMO learns two separate denoising processes to first predict hand positions from object motion and subsequently synthesize full-body poses based on the predicted hand positions. By employing the hand positions as an intermediate representation between the two denoising processes, we can explicitly enforce contact constraints, resulting in more physically plausible manipulation motions. With the learned model, we develop a novel system that captures full-body human manipulation motions by simply attaching a smartphone to the object being manipulated. Through extensive experiments, we demonstrate the effectiveness of our proposed pipeline and its ability to generalize to unseen objects. Additionally, as high-quality human-object interaction datasets are scarce, we collect a large-scale dataset consisting of 3D object geometry, object motion, and human motion. Our dataset contains human-object interaction motion for 15 objects, with a total duration of approximately 10 hours. Jiaman Li, Jiajun Wu 0001, C. Karen Liu |
ACM Trans. Graph. | 1 |
| 2022 | GIMO: Gaze-Informed Human Motion Prediction in Context
Yanchao Yang 0001, Kaichun Mo, Jiaman Li, Tao Yu 0007, Yebin Liu, C. Karen Liu, Leonidas J. Guibas |
ECCV (13) | 4 |
| 2022 | DenseGAP: Graph-Structured Dense Correspondence Learning with Anchor PointsabstractEstablishing dense correspondence between two images is a fundamental computer vision problem, which is typically tackled by matching local feature descriptors. However, without global awareness, such local features are often insufficient for disambiguating similar regions. And computing the pairwise feature correlation across images is both computation-expensive and memory-intensive. To make the local features aware of the global context and improve their matching accuracy, we introduce DenseGAP, a new solution for efficient Dense correspondence learning with a Graph-structured neural network conditioned on Anchor Points. Specifically, we first propose a graph structure that utilizes anchor points to provide sparse but reliable prior on inter- and intra-image context and propagates them to all image points via directed edges. We also design a graph-structured network to broadcast multi-level contexts via light-weighted message-passing layers and generate high-resolution feature maps at low memory cost. Finally, based on the predicted feature maps, we introduce a coarse-to-fine framework for accurate correspondence prediction using cycle consistency. Our feature descriptors capture both local and global information, thus enabling a continuous feature field for querying arbitrary points at high resolution. Through comprehensive ablative experiments and evaluations on large-scale indoor and outdoor datasets, we demonstrate that our method advances the state-of-the-art of correspondence learning on most benchmarks. Zhengfei Kuang, Jiaman Li, Mingming He |
ICPR | 2 |
| 2022 | Scene Synthesis from Human MotionabstractLarge-scale capture of human motion with diverse, complex scenes, while immensely useful, is often considered prohibitively costly. Meanwhile, human motion alone contains rich information about the scene they reside in and interact with. For example, a sitting human suggests the existence of a chair, and their leg position further implies the chair’s pose. In this paper, we propose to synthesize diverse, semantically reasonable, and physically plausible scenes based on human motion. Our framework, Scene Synthesis from HUMan MotiON (SUMMON), includes two steps. It first uses ContactFormer, our newly introduced contact predictor, to obtain temporally consistent contact labels from human motion. Based on these predictions, SUMMON then chooses interacting objects and optimizes physical plausibility losses; it further populates the scene with objects that do not interact with humans. Experimental results demonstrate that SUMMON synthesizes feasible, plausible, and diverse scenes and has the potential to generate extensive human-scene interaction data for the community. Sifan Ye, Yixing Wang, Jiaman Li, Dennis Park, C. Karen Liu, Huazhe Xu, Jiajun Wu 0001 |
SIGGRAPH Asia | 3 |
| 2021 | Task-Generic Hierarchical Human Motion Prior using VAEsabstractA deep generative model that describes human motions can benefit a wide range of fundamental computer vision and graphics tasks, such as providing robustness to video-based human pose estimation, predicting complete body movements for motion capture systems during occlusions, and assisting key frame animation with plausible movements. In this paper, we present a method for learning complex human motions independent of specific tasks using a combined global and local latent space to facilitate coarse and fine-grained modeling. Specifically, we propose a hierarchical motion variational autoencoder (HM-VAE) that consists of a 2-level hierarchical latent space. While the global latent space captures the overall global body motion, the local latent space enables to capture the refined poses of the different body parts. We demonstrate the effectiveness of our hierarchical motion variational autoencoder in a variety of tasks including video-based human pose estimation, motion completion from partial observations, and motion synthesis from sparse key-frames. Even though, our model has not been trained for any of these tasks specifically, it provides superior performance than task-specific alternatives. Our general-purpose human motion prior model can fix corrupted human body animations and generate complete movements from incomplete observations. Jiaman Li, Ruben Villegas, Duygu Ceylan, Jimei Yang, Zhengfei Kuang, Hao Li 0015 |
3DV | 1 |
| 2020 | Dynamic facial asset and rig generation from a single scanabstractThe creation of high-fidelity computer-generated (CG) characters for films and games is tied with intensive manual labor, which involves the creation of comprehensive facial assets that are often captured using complex hardware. To simplify and accelerate this digitization process, we propose a framework for the automatic generation of high-quality dynamic facial models, including rigs which can be readily deployed for artists to polish. Our framework takes a single scan as input to generate a set of personalized blendshapes, dynamic textures, as well as secondary facial components ( e.g. , teeth and eyeballs). Based on a facial database with over 4, 000 scans with pore-level details, varying expressions and identities, we adopt a self-supervised neural network to learn personalized blendshapes from a set of template expressions. We also model the joint distribution between identities and expressions, enabling the inference of a full set of personalized blendshapes with dynamic appearances from a single neutral input scan. Our generated personalized face rig assets are seamlessly compatible with professional production pipelines for facial animation and rendering. We demonstrate a highly robust and effective framework on a wide range of subjects, and showcase high-fidelity facial animations with automatically generated personalized dynamic textures. Jiaman Li, Zhengfei Kuang, Mingming He, Karl Bladin, Hao Li 0015 |
ACM Trans. Graph. | 1 |
| 2019 | Creative Flow+ DatasetabstractWe present the Creative Flow+ Dataset, the first diverse multi-style artistic video dataset richly labeled with per-pixel optical flow, occlusions, correspondences, segmentation labels, normals, and depth. Our dataset includes 3000 animated sequences rendered using styles randomly selected from 40 textured line styles and 38 shading styles, spanning the range between flat cartoon fill and wildly sketchy shading. Our dataset includes 124K+ train set frames and 10K test set frames rendered at 1500x1500 resolution, far surpassing the largest available optical flow datasets in size. While modern techniques for tasks such as optical flow estimation achieve impressive performance on realistic images and video, today there is no way to gauge their performance on non-photorealistic images. Creative Flow+ poses a new challenge to generalize real-world Computer Vision to messy stylized content. We show that learning-based optical flow methods fail to generalize to this data and struggle to compete with classical approaches, and invite new research in this area. Our dataset and a new optical flow benchmark will be publicly available at: www.cs.toronto.edu/creativeflow/. We further release the complete dataset creation pipeline, allowing the community to generate and stylize their own data on demand. Maria Shugrina, Ziheng Liang, Amlan Kar, Jiaman Li, Angad Singh, Karan Singh 0004, Sanja Fidler |
CVPR | 4 |
| 2019 | Full-Duplex OFDM Relaying Systems with Energy Harvesting in Multipath Fading ChannelsabstractIn this paper, the performance of in-band full- duplex OFDM relaying systems with energy- harvesting and self-interference cancellation in the polarization domain is analyzed. Specifically, we use the time switching-based relaying protocol to implement energy harvesting. The harvested energy is used by the relay to forward the transmitted information from the source. To cancel the self-interference, the polarization-enabled digital self-interference cancellation scheme is deployed at the relay. Our simulation results show that the full-duplex OFDM energy harvesting relaying system almost doubles the throughput, while maintaining the same bit error performance by a modest increase in the signal-to-noise ratio compared to the half-duplex OFDM energy harvesting relaying system. It is also revealed that the optimal time splitting factor should be less than 0.3 to maximize the full-duplex system throughput. Jiaman Li, Le Chung Tran, Farzad Safaei |
VTC Fall | 1 |
| 2018 | Learning to Act Properly: Predicting and Explaining Affordances From ImagesabstractWe address the problem of affordance reasoning in diverse scenes that appear in the real world. Affordances relate the agent's actions to their effects when taken on the surrounding objects. In our work, we take the egocentric view of the scene, and aim to reason about action-object affordances that respect both the physical world as well as the social norms imposed by the society. We also aim to teach artificial agents why some actions should not be taken in certain situations, and what would likely happen if these actions would be taken. We collect a new dataset that builds upon ADE20k [32], referred to as ADE-Affordance, which contains annotations enabling such rich visual reasoning. We propose a model that exploits Graph Neural Networks to propagate contextual information from the scene in order to perform detailed affordance reasoning about each object. Our model is showcased through various ablation studies, pointing to successes and challenges in this complex task. Ching-Yao Chuang, Jiaman Li, Antonio Torralba 0001, Sanja Fidler |
CVPR | 2 |
| 2018 | VirtualHome: Simulating Household Activities via ProgramsabstractIn this paper, we are interested in modeling complex activities that occur in a typical household. We propose to use programs, i.e., sequences of atomic actions and interactions, as a high level representation of complex tasks. Programs are interesting because they provide a non-ambiguous representation of a task, and allow agents to execute them. However, nowadays, there is no database providing this type of information. Towards this goal, we first crowd-source programs for a variety of activities that happen in people's homes, via a game-like interface used for teaching kids how to code. Using the collected dataset, we show how we can learn to extract programs directly from natural language descriptions or from videos. We then implement the most common atomic (inter)actions in the Unity3D game engine, and use our programs to "drive" an artificial agent to execute tasks in a simulated household environment. Our VirtualHome simulator allows us to create a large activity video dataset with rich ground-truth, enabling training and testing of video understanding models. We further showcase examples of our agent performing tasks in our VirtualHome based on language descriptions. Xavier Puig, Kevin Ra, Marko Boben, Jiaman Li, Tingwu Wang, Sanja Fidler, Antonio Torralba 0001 |
CVPR | 4 |
| 2018 | Differentiable Compositional Kernel Learning for Gaussian ProcessesabstractThe generalization properties of Gaussian processes depend heavily on the choice of kernel, and this choice remains a dark art. We present the Neural Kernel Network (NKN), a flexible family of kernels represented by a neural network. The NKN’s architecture is based on the composition rules for kernels, so that each unit of the network corresponds to a valid kernel. It can compactly approximate compositional kernel structures such as those used by the Automatic Statistician (Lloyd et al., 2014), but because the architecture is differentiable, it is end-to-end trainable with gradient- based optimization. We show that the NKN is universal for the class of stationary kernels. Empirically we demonstrate NKN’s pattern discovery and extrapolation abilities on several tasks that depend crucially on identifying the underlying structure, including time series and texture extrapolation, as well as Bayesian optimization. Shengyang Sun, Guodong Zhang 0006, Chaoqi Wang, Wenyuan Zeng, Jiaman Li, Roger B. Grosse |
ICML | 5 |