VLDB 2026 Research / reviewers in the wild / expert
Yixin Zhu 0001
dblp:91/1103-1
· DBLP profile ↗
109ranked-venue papers
2as first author
69since 2021 · last 2026
0000-0001-7024-1545ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 98 · 2 first-author · 63 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 first-author · 18 since 2021Systems, architecture and hardware · 26 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 11 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NarrativeLoom: Enhancing Creative Storytelling through Multi-Persona Collaborative Improvisation
Yongqian Peng, Siyu Zha, Chi Zhang 0017, Zixia Jia, Zilong Zheng, Yixin Zhu 0001 |
CHI | 8 |
| 2026 | AlphaChimp: Tracking and Behavior Recognition of Chimpanzees
Xiaoxuan Ma 0001, Yutang Lin 0001, Yuan Xu 0022, Stephan P. Kaufhold, Jack Terwilliger, Andres Meza 0001, Yixin Zhu 0001, Federico Rossano, Yizhou Wang 0001 |
Int. J. Comput. Vis. | 7 |
| 2026 | TacMan-Turbo: Proactive Tactile Control for Robust and Efficient Articulated Object ManipulationabstractAdept manipulation of articulated objects is essential for robots to operate successfully in human environments. Such manipulation requires both effectiveness—reliable operation despite uncertain object structures—and efficiency—swift execution with minimal redundant steps and smooth trajectories. Existing approaches struggle to achieve both objectives simultaneously: methods relying on predefined kinematic models lack robustness when encountering structural variations, while the tactile-informed approach achieves robust manipulation but sacrifices efficiency through reactive, step-by-step execute-and-recover cycles. To address this challenge, this paper introduces TacMan-Turbo, a proactive tactile control framework that unifies the cycles into a continuous control loop. Our key insight is to interpret tactile signals temporally in addition to spatially: instead of treating contact deviations merely as instantaneous error signals requiring immediate compensation, we analyze sequential tactile observations to reveal local kinematic information. This temporal perspective enables our controller to predict future object states and proactively modulate manipulation velocities, eliminating the previous execute-and-recover cycle. In evaluations across 200 diverse simulated articulated objects and real-world experiments, our approach maintains a near-perfect performance while significantly enhancing time efficiency, action efficiency, and trajectory smoothness (all p-values < 0.0001). These results demonstrate that tactile feedback serves a dual purpose: providing not only spatial error signals for reactive control but also critical temporal structural information that enables proactive prediction and control. By enabling robots to manipulate articulated objects both reliably and efficiently, this work advances robot capabilities for seamless operation in dynamic, human-centric environments. Zihang Zhao, Zhenghao Qi, Leiyao Cui, Zhi Han, Lecheng Ruan, Yixin Zhu 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2025 | A simulation-heuristics dual-process model for intuitive physics
Shiqian Li, Jiajun Yan, Bo Dai 0025, Yujia Peng, Chi Zhang 0017, Yixin Zhu 0001 |
CogSci | 7 |
| 2025 | Word Embeddings Track Social Group Changes Across 70 Years in China
Yongqian Peng, Yixin Zhu 0001 |
CogSci | 3 |
| 2025 | IMPROVISER: Multi-persona Co-creation System Enhances Story Creativity
Yongqian Peng, Chi Zhang 0017, Zixia Jia, Zilong Zheng, Yixin Zhu 0001 |
CogSci | 7 |
| 2025 | Probing and Inducing Combinational Creativity in Vision-Language Models
Yongqian Peng, Yuxuan Wang 0004, Yizhou Wang 0001, Chi Zhang 0017, Yixin Zhu 0001, Zilong Zheng |
CogSci | 7 |
| 2025 | GROVE: A Generalized Reward for Learning Open-Vocabulary Physical SkillabstractLearning open-vocabulary physical skills for simulated agents presents a significant challenge in Artificial Intelligence (AI). Current Reinforcement Learning (RL) approaches face critical limitations: manually designed rewards lack scalability across diverse tasks, while demonstration-based methods struggle to generalize beyond their training distribution. We introduce GROVE, a generalized reward framework that enables open-vocabulary physical skill learning without manual engineering or task-specific demonstrations. Our key insight is that Large Language Models (LLMs) and Vision Language Models (VLMs) provide complementary guidance—LLMs generate precise physical constraints capturing task requirements, while VLMs evaluate motion semantics and naturalness. Through an iterative design process, VLM-based feedback continuously refines LLM-generated constraints, creating a self-improving reward system. To bridge the domain gap between simulation and natural images, we develop Pose2CLIP, a lightweight mapper that efficiently projects agent poses directly into semantic feature space without computationally expensive rendering. Extensive experiments across diverse embodiments and learning paradigms demonstrate GROVE’s effectiveness, achieving 22.2% higher motion naturalness and 25.7% better task completion scores while training 8.4× faster than previous methods. These results establish a new foundation for scalable physical skill acquisition in simulated environments. Jieming Cui, Tengyu Liu, Jiale Yu, Ran Song 0001, Wei Zhang 0021, Yixin Zhu 0001, Siyuan Huang 0001 |
CVPR | 7 |
| 2025 | Dynamic Motion Blending for Versatile Motion EditingabstractText-guided motion editing enables high-level semantic control and iterative modifications beyond traditional keyframe animation. Existing methods rely on limited pre-collected training triplets (original motion, edited motion, and instruction), which severely hinders their versatility in diverse editing scenarios. We introduce MotionCutMix, an online data augmentation technique that dynamically generates training triplets by blending body part motions based on input text. While MotionCutMix effectively expands the training distribution, the compositional nature introduces increased randomness and potential body part incoordination. To model such a rich distribution, we present Mo-tionReFit, an auto-regressive diffusion model with a motion coordinator. The auto-regressive architecture facilitates learning by decomposing long sequences, while the motion coordinator mitigates the artifacts of motion composition. Our method handles both spatial and temporal motion edits directly from high-level human instructions, without relying on additional specifications or Large Language Models (LLMs). Through extensive experiments, we show that MotionReFit achieves state-of-the-art performance in text-guided motion editing. Ablation studies further verify that MotionCutMix significantly improves the model’s generalizability while maintaining training convergence. Ziye Yuan, Zimo He, Yixin Chen 0003, Tengyu Liu, Yixin Zhu 0001, Siyuan Huang 0001 |
CVPR | 7 |
| 2025 | PrimHOI: Compositional Human-Object Interaction via Reusable Primitives
Tengyu Liu, Yixin Zhu 0001, Mingtao Pei, Siyuan Huang 0001 |
ICCV | 3 |
| 2025 | Ag2x2: Robust Agent-Agnostic Visual Representations for Zero-Shot Bimanual ManipulationabstractBimanual manipulation, fundamental to human daily activities, remains a challenging task due to its inherent complexity of coordinated control. Recent advances have enabled zero-shot learning of single-arm manipulation skills through agent-agnostic visual representations derived from human videos; however, these methods overlook crucial agentspecific information necessary for bimanual coordination, such as end-effector positions. We propose Ag2x2, a computational framework for bimanual manipulation through coordination-aware visual representations that jointly encode object states and hand motion patterns while maintaining agent-agnosticism. Extensive experiments demonstrate that Ag2x2 achieves a 73.5% success rate across 13 diverse bimanual tasks from Bi-DexHands and PerAct2, including challenging scenarios with deformable objects like ropes. This performance outperforms baseline methods and even surpasses the success rate of policies trained with expert-engineered rewards. Furthermore, we show that representations learned through Ag2x2 can be effectively leveraged for imitation learning, establishing a scalable pipeline for skill acquisition without expert supervision. By maintaining robust performance across diverse tasks without human demonstrations or engineered rewards, Ag2x2 represents a step toward scalable learning of complex bimanual robotic skills. Ziyin Xiong, Yinghan Chen, Puhao Li, Yixin Zhu 0001, Tengyu Liu, Siyuan Huang 0001 |
IROS | 4 |
| 2025 | Taccel: Scaling Up Vision-based Tactile Robotics via High-performance GPU SimulationabstractTactile sensing is crucial for achieving human-level robotic capabilities in manipulation tasks. As a promising solution, Vision-based Tactile Sensors (VBTSs) offer high spatial resolution and cost-effectiveness, but present unique challenges in robotics for their complex physical characteristics and visual signal processing requirements. The lack of efficient and accurate simulation tools for VBTSs has significantly limited the scale and scope of tactile robotics research. We present Taccel, a high-performance simulation platform that integrates Incremental Potential Contact (IPC) and Affine Body Dynamics (ABD) to model robots, tactile sensors, and objects with both accuracy and unprecedented speed, achieving a total of 915 FPS with 4096 parallel environments. Unlike previous simulators that operate at sub-real-time speeds with limited parallelization, Taccel provides precise physics simulation and realistic tactile signals while supporting flexible robot-sensor configurations through user-friendly APIs. Through extensive validation in object recognition, robotic grasping, and articulated object manipulation, we demonstrate precise simulation and successful sim-to-real transfer. These capabilities position Taccel as a powerful tool for scaling up tactile robotics research and development, potentially transforming how robots interact with and understand their physical environment. Wenxin Du, Chang Yu 0005, Puhao Li, Zihang Zhao, Tengyu Liu, Chenfanfu Jiang, Yixin Zhu 0001, Siyuan Huang 0001 |
NeurIPS | 8 |
| 2025 | GlobalTomo: A global dataset for physics-ML seismic wavefield modeling and FWIabstractGlobal seismic tomography, taking advantage of seismic waves from natural earthquakes, provides essential insights into the earth's internal dynamics. Advanced Full-Waveform Inversion (FWI) techniques, whose aim is to meticulously interpret every detail in seismograms, confront formidable computational demands in forward modeling and adjoint simulations on a global scale. Recent advancements in Machine Learning (ML) offer a transformative potential for accelerating the computational efficiency of FWI and extending its applicability to larger scales. This work presents the first 3D global synthetic dataset tailored for seismic wavefield modeling and full-waveform tomography, referred to as the Global Tomography (GlobalTomo) dataset. This dataset is comprehensive, incorporating explicit wave physics and robust geophysical parameterization at realistic global scales, generated through state-of-the-art forward simulations optimized for 3D global wavefield calculations. Through extensive analysis and the establishment of ML baselines, we illustrate that ML approaches are particularly suitable for global FWI, overcoming its limitations with rapid forward modeling and flexible inversion strategies. This work represents a cross-disciplinary effort to enhance our understanding of the earth's interior through physics-ML modeling. Shiqian Li, Zhancun Mu, Shiji Xin, Zhixiang Dai, Kuangdai Leng, Rita Zhang, Yixin Zhu 0001 |
NeurIPS | 9 |
| 2025 | DrivAerStar: An Industrial-Grade CFD Dataset for Vehicle Aerodynamic OptimizationabstractVehicle aerodynamics optimization has become critical for automotive electrification, where drag reduction directly determines electric vehicle range and energy efficiency. Traditional approaches face an intractable trade-off: computationally expensive Computational Fluid Dynamics (CFD) simulations requiring weeks per design iteration, or simplified models that sacrifice production-grade accuracy. While machine learning offers transformative potential, existing datasets exhibit fundamental limitations -- inadequate mesh resolution, missing vehicle components, and validation errors exceeding 5% -- preventing deployment in industrial workflows. We present DrivAerStar, comprising 12,000 industrial-grade automotive CFD simulations generated using STAR-CCM+${}^{\textregistered}$ software. The dataset systematically explores three vehicle configurations through 20 Computer Aided Design (CAD) parameters via Free Form Deformation (FFD) algorithms, including complete engine compartments and cooling systems with realistic internal airflow. DrivAerStar achieves wind tunnel validation accuracy below 1.04% -- a five-fold improvement over existing datasets -- through refined mesh strategies with strict wall $y^+$ control. Benchmarks demonstrate that models trained on this data achieve production-ready accuracy while reducing computational costs from weeks to minutes. This represents the first dataset bridging academic machine learning research and industrial CFD practice, establishing a new standard for data-driven aerodynamic optimization in automotive development. Beyond automotive applications, DrivAerStar demonstrates a paradigm for integrating high-fidelity physics simulations with Artificial Intelligence (AI) across engineering disciplines where computational constraints currently limit innovation. Jiyan Qiu, Lyulin Kuang, Leiyao Cui, Shaotong Fu, Yixin Zhu 0001, Rita Zhang |
NeurIPS | 7 |
| 2025 | Heterogeneous Adversarial Play in Interactive EnvironmentsabstractSelf-play constitutes a fundamental paradigm for autonomous skill acquisition, whereby agents iteratively enhance their capabilities through self-directed environmental exploration. Conventional self-play frameworks exploit agent symmetry within zero-sum competitive settings, yet this approach proves inadequate for open-ended learning scenarios characterized by inherent asymmetry. Human pedagogical systems exemplify asymmetric instructional frameworks wherein educators systematically construct challenges calibrated to individual learners' developmental trajectories. The principal challenge resides in operationalizing these asymmetric, adaptive pedagogical mechanisms within artificial systems capable of autonomously synthesizing appropriate curricula without predetermined task hierarchies. Here we present Heterogeneous Adversarial Play (HAP), an adversarial Automatic Curriculum Learning framework that formalizes teacher-student interactions as a minimax optimization wherein task-generating instructor and problem-solving learner co-evolve through adversarial dynamics. In contrast to prevailing automatic curriculum learning methodologies that employ static curricula or unidirectional task selection mechanisms, HAP establishes a bidirectional feedback system wherein instructors continuously recalibrate task complexity in response to real-time learner performance metrics. Experimental validation across multi-task learning domains demonstrates that our framework achieves performance parity with SOTA baselines while generating curricula that enhance learning efficacy in both artificial agents and human subjects. Manjie Xu, Jiayu Zhan, Wei Liang 0008, Chi Zhang 0017, Yixin Zhu 0001 |
NeurIPS | 6 |
| 2025 | Integration of Robot and Scene Kinematics for Sequential Mobile Manipulation PlanningabstractWe present a Sequential Mobile Manipulation Planning (SMMP) framework that can solve long-horizon multi-step mobile manipulation tasks with coordinated whole-body motion, even when interacting with articulated objects. By abstracting environmental structures as kinematic models and integrating them with the robot's kinematics, we construct an Augmented Configuration Apace (A-Space) that unifies the previously separate task constraints for navigation and manipulation, while accounting for the joint reachability of the robot base, arm, and manipulated objects. This integration facilitates efficient planning within a tri-level framework: a task planner generates symbolic action sequences to model the evolution of A-Space, an optimization-based motion planner computes continuous trajectories within A-Space to achieve desired configurations for both the robot and scene elements, and an intermediate plan refinement stage selects action goals that ensure long-horizon feasibility. Our simulation studies first confirm that planning in A-Space achieves an 84.6% higher task success rate compared to baseline methods. Validation on real robotic systems demonstrates fluid mobile manipulation involving (i) seven types of rigid and articulated objects across 17 distinct contexts, and (ii) long-horizon tasks of up to 14 sequential steps. Our results highlight the significance of modeling scene kinematics into planning entities, rather than encoding task-specific constraints, offering a scalable and generalizable approach to complex robotic manipulation. Ziyuan Jiao, Yida Niu, Zeyu Zhang 0001, Yao Su 0001, Yixin Zhu 0001, Hangxin Liu, Song-Chun Zhu |
IEEE Trans. Robotics | 6 |
| 2025 | Tac-Man: Tactile-Informed Prior-Free Manipulation of Articulated ObjectsabstractIntegrating robots into human-centric environments such as homes, necessitates advanced manipulation skills as robotic devices will need to engage with articulated objects such as doors and drawers. Key challenges in robotic manipulation of articulated objects are the unpredictability and diversity of these objects' internal structures, which render models based on object kinematics priors, both explicit and implicit, and inadequate. Their reliability is significantly diminished by pre-interaction ambiguities, imperfect structural parameters, encounters with unknown objects, and unforeseen disturbances. Here, we present aprior-freestrategy, Tac-Man, focusing on maintaining stable robot-object contact during manipulation. Without relying on object priors, Tac-Man leverages tactile feedback to enable robots to proficiently handle a variety of articulated objects, including those with complex joints, even when influenced by unexpected disturbances. Demonstrated in both real-world experiments and extensive simulations, it consistently achieves near-perfect success in dynamic and varied settings, outperforming existing methods. Our results indicate that tactile sensing alone suffices for managing diverse articulated objects, offering greater robustness and generalization than prior-based approaches. This underscores the importance of detailed contact modeling in complex manipulation tasks, especially with articulated objects. Advancements in tactile-informed approaches significantly expand the scope of robotic applications in human-centric environments, particularly where accurate models are difficult to obtain. Zihang Zhao, Wanlin Li, Zhenghao Qi, Lecheng Ruan, Yixin Zhu 0001, Kaspar Althoefer |
IEEE Trans. Robotics | 6 |
| 2024 | Single-view 3D Scene Reconstruction with High-fidelity Shape and TextureabstractReconstructing detailed 3D scenes from single-view images remains a challenging task due to limitations in existing approaches, which primarily focus on geometric shape recovery, overlooking object appearances and fine shape details. To address these challenges, we propose a novel framework for simultaneous high-fidelity recovery of object shapes and textures from single-view images. Our approach utilizes the proposed Single-view neural implicit Shape and Radiance field (SSR) representations to leverage both explicit 3D shape supervision and volume rendering of color, depth, and surface normal images. To overcome shape-appearance ambiguity under partial observations, we introduce a two-stage learning curriculum incorporating both 3D and 2D supervisions. A distinctive feature of our framework is its ability to generate fine-grained textured meshes by seamlessly integrating rendering capabilities into the single-view 3D reconstruction model. This integration enables not only improved textured 3D object reconstruction by 27.7% and 11.6% on the 3D-FRONT and Pix3D datasets, respectively, but also supports the rendering of images from novel viewpoints. Beyond individual objects, our approach facilitates composing object-level representations into flexible scene representations, thereby enabling applications such as holistic scene understanding and 3D scene editing. We conduct extensive experiments to demonstrate the effectiveness of our method. Yixin Chen 0003, Junfeng Ni, Yaowei Zhang, Yixin Zhu 0001, Siyuan Huang 0001 |
3DV | 5 |
| 2024 | Evaluating and Modeling Social Intelligence: A Comparative Study of Human and AI Capabilities
Lixing Niu, Jiaheng Han, Yujia Peng, Yixin Zhu 0001, Lifeng Fan |
CogSci | 8 |
| 2024 | AnySkill: Learning Open-Vocabulary Physical Skill for Interactive AgentsabstractTraditional approaches in physics-based motion generation, centered around imitation learning and reward shaping, often struggle to adapt to new scenarios. To tackle this limitation, we propose AnySkill, a novel hierarchical method that learns physically plausible interactions following open-vocabulary instructions. Our approach begins by developing a set of atomic actions via a low-level controller trained via imitation learning. Upon receiving an open-vocabulary textual instruction, AnySkill employs a high-level policy that selects and integrates these atomic actions to maximize the CLIP similarity between the agent's rendered images and the text. An important feature of our method is the use of image-based rewards for the high-level policy, which allows the agent to learn interactions with objects without manual reward engineering. We demonstrate AnySkill's capability to generate realistic and natural motion sequences in response to unseen instructions of varying lengths, marking it the first method capable of open-vocabulary physical skill learning for interactive humanoid agents. Jieming Cui, Tengyu Liu, Nian Liu 0003, Yaodong Yang 0001, Yixin Zhu 0001, Siyuan Huang 0001 |
CVPR | 5 |
| 2024 | Scaling Up Dynamic Human-Scene Interaction ModelingabstractConfronting the challenges of data scarcity and advanced motion synthesis in HSI modeling, we introduce the TRUMANS (Tracking Human Actions in Scenes) dataset alongside a novel HSI motion synthesis method. TRUMANS stands as the most comprehensive motion-captured HSI dataset currently available, encompassing over 15 hours of human interactions across 100 indoor scenes. It intricately captures whole-body human motions and part-level object dynamics, focusing on the realism of contact. This dataset is further scaled up by transforming physical environments into exact virtual models and applying extensive augmentations to appearance and motion for both humans and objects while maintaining interaction fidelity. Utilizing TRUMANS, we devise a diffusion-based autoregressive model that efficiently generates Human-Scene Interaction (HSI) sequences of any length, taking into account both scene context and intended actions. In experiments, our approach shows remarkable zero-shot generalizability on a range of 3D scene datasets (e.g., PROX, Replica, ScanNet, ScanNet++), producing motions that closely mimic original motion-captured sequences, as confirmed by quantitative experiments and human studies. Xiaoxuan Ma 0001, Yixin Chen 0003, Tengyu Liu, Yixin Zhu 0001, Siyuan Huang 0001 |
CVPR | 8 |
| 2024 | Move as you Say, Interact as you can: Language-Guided Human Motion Generation with Scene AffordanceabstractDespite significant advancements in text-to-motion syn-thesis, generating language-guided human motion within 3D environments poses substantial challenges. These challenges stem primarily from (i) the absence of powerful generative models capable of jointly modeling natural language, 3D scenes, and human motion, and (ii) the generative models' in-tensive data requirements contrasted with the scarcity of comprehensive, high-quality, language-scene-motion datasets. To tackle these issues, we introduce a novel two-stage frame-work that employs scene affordance as an intermediate representation, effectively linking 3D scene grounding and conditional motion generation. Our framework comprises an Affordance Diffusion Model (ADM) for predicting ex-plicit affordance map and an Affordance-to-Motion Diffusion Model (AMDM) for generating plausible human motions. By leveraging scene affordance maps, our method overcomes the difficulty in generating human motion under multimodal condition signals, especially when training with limited data lacking extensive language-scene-motion pairs. Our exten-sive experiments demonstrate that our approach consistently outperforms all baselines on established benchmarks, in-cluding HumanML3D and HUMANISE. Additionally, we validate our model's exceptional generalization capabilities on a specially curated evaluation set featuring previously unseen descriptions and scenes. Yixin Chen 0003, Baoxiong Jia, Puhao Li, Jinlu Zhang 0001, Jingze Zhang, Tengyu Liu, Yixin Zhu 0001, Wei Liang 0008, Siyuan Huang 0001 |
CVPR | 8 |
| 2024 | Zero-Shot Image Feature Consensus with Deep Functional Maps
Xinle Cheng, Congyue Deng, Adam W. Harley, Yixin Zhu 0001, Leonidas J. Guibas |
ECCV (47) | 4 |
| 2024 | Neural-Symbolic Recursive Machine for Systematic GeneralizationabstractCurrent learning models often struggle with human-like systematic generalization, particularly in learning compositional rules from limited data and extrapolating them to novel combinations. We introduce the Neural-Symbolic Recursive Ma- chine ( NSR), whose core is a Grounded Symbol System ( GSS), allowing for the emergence of combinatorial syntax and semantics directly from training data. The NSR employs a modular design that integrates neural perception, syntactic parsing, and semantic reasoning. These components are synergistically trained through a novel deduction-abduction algorithm. Our findings demonstrate that NSR’s design, imbued with the inductive biases of equivariance and compositionality, grants it the expressiveness to adeptly handle diverse sequence-to-sequence tasks and achieve unparalleled systematic generalization. We evaluate NSR’s efficacy across four challenging benchmarks designed to probe systematic generalization capabilities: SCAN for semantic parsing, PCFG for string manipulation, HINT for arithmetic reasoning, and a compositional machine translation task. The results affirm NSR ’s superiority over contemporary neural and hybrid models in terms of generalization and transferability. Qing Li 0003, Yixin Zhu 0001, Yitao Liang, Ying Nian Wu, Song-Chun Zhu, Siyuan Huang 0001 |
ICLR | 2 |
| 2024 | I-PHYRE: Interactive Physical ReasoningabstractCurrent evaluation protocols predominantly assess physical reasoning in stationary scenes, creating a gap in evaluating agents' abilities to interact with dynamic events. While contemporary methods allow agents to modify initial scene configurations and observe consequences, they lack the capability to interact with events in real time. To address this, we introduce I-PHYRE, a framework that challenges agents to simultaneously exhibit intuitive physical reasoning, multi-step planning, and in-situ intervention. Here, intuitive physical reasoning refers to a quick, approximate understanding of physics to address complex problems; multi-step denotes the need for extensive sequence planning in I-PHYRE, considering each intervention can significantly alter subsequent choices; and in-situ implies the necessity for timely object manipulation within a scene, where minor timing deviations can result in task failure. We formulate four game splits to scrutinize agents' learning and generalization of essential principles of interactive physical reasoning, fostering learning through interaction with representative scenarios. Our exploration involves three planning strategies and examines several supervised and reinforcement agents' zero-shot generalization proficiency on I-PHYRE. The outcomes highlight a notable gap between existing learning algorithms and human performance, emphasizing the imperative for more research in enhancing agents with interactive physical reasoning capabilities. The environment and baselines will be made publicly available. Shiqian Li, Kewen Wu 0004, Chi Zhang 0017, Yixin Zhu 0001 |
ICLR | 4 |
| 2024 | SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous ManipulationabstractHumans demonstrate remarkable skill in transferring manipulation abilities across objects of varying shapes, poses, and appearances, a capability rooted in their understanding of semantic correspondences between different instances. To equip robots with a similar high-level comprehension, we present SparseDFF, a novel DFF for 3D scenes utilizing large 2D vision models to extract semantic features from sparse RGBD images, a domain where research is limited despite its relevance to many tasks with fixed-camera setups. SparseDFF generates view-consistent 3D DFFs, enabling efficient one-shot learning of dexterous manipulations by mapping image features to a 3D point cloud. Central to SparseDFF is a feature refinement network, optimized with a contrastive loss between views and a point-pruning mechanism for feature continuity. This facilitates the minimization of feature discrepancies w.r.t. end-effector parameters, bridging demonstrations and target manipulations. Validated in real-world scenarios with a dexterous hand, SparseDFF proves effective in manipulating both rigid and deformable objects, demonstrating significant generalization capabilities across object and scene variations. Qianxu Wang, Haotong Zhang 0005, Congyue Deng, Yang You 0004, Hao Dong 0003, Yixin Zhu 0001, Leonidas J. Guibas |
ICLR | 6 |
| 2024 | PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and EnvironmentsabstractRobotic manipulation with two-finger grippers is challenged by objects lacking distinct graspable features. Traditional pre-grasping methods, which typically involve repositioning objects or utilizing external aids like table edges, are limited in their adaptability across different object categories and environments. To overcome these limitations, we introduce PreAfford, a novel pre-grasping planning framework incorporating a point-level affordance representation and a relay training approach. Our method significantly improves adaptability, allowing effective manipulation across a wide range of environments and object types. When evaluated on the ShapeNet-v2 dataset, PreAfford not only enhances grasping success rates by 69% but also demonstrates its practicality through successful real-world experiments. These improvements highlight PreAfford’s potential to redefine standards for robotic handling of complex manipulation tasks in diverse settings. Kairui Ding, Boyuan Chen 0009, Ruihai Wu, Zongzheng Zhang, Huan-ang Gao, Siqi Li 0009, Guyue Zhou, Yixin Zhu 0001, Hao Dong 0003, Hao Zhao 0002 |
IROS | 9 |
| 2024 | Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action RepresentationsabstractAutonomous robotic systems capable of learning novel manipulation tasks are poised to transform industries from manufacturing to service automation. However, current methods (e.g., VIP and R3M) still face significant hurdles, notably the domain gap among robotic embodiments and the sparsity of successful task executions within specific action spaces, resulting in misaligned and ambiguous task representations. We introduce Ag2Manip (Agent-Agnostic representations for Manipulation), a framework aimed at addressing these challenges through two key innovations: (1) an agent-agnostic visual representation derived from human manipulation videos, with the specifics of embodiments obscured to enhance generalizability; and (2) an agent-agnostic action representation abstracting a robot’s kinematics to a universal agent proxy, emphasizing crucial interactions between end-effector and object. Ag2Manip has been empirically validated across simulated benchmarks, showing a 325% performance increase without relying on domain-specific demonstrations. Ablation studies further underline the essential contributions of the agent-agnostic visual and action representations to this success. Extending our evaluations to the real world, Ag2Manip significantly improves imitation learning success rates from 50% to 77.5%, demonstrating its effectiveness and generalizability across both simulated and real environments. Puhao Li, Tengyu Liu, Muzhi Han, Shu Wang 0002, Yixin Zhu 0001, Song-Chun Zhu, Siyuan Huang 0001 |
IROS | 7 |
| 2024 | PhyRecon: Physically Plausible Neural Scene ReconstructionabstractWe address the issue of physical implausibility in multi-view neural reconstruction. While implicit representations have gained popularity in multi-view 3D reconstruction, previous work struggles to yield physically plausible results, limiting their utility in domains requiring rigorous physical accuracy. This lack of plausibility stems from the absence of physics modeling in existing methods and their inability to recover intricate geometrical structures. In this paper, we introduce PHYRECON, the first approach to leverage both differentiable rendering and differentiable physics simulation to learn implicit surface representations. PHYRECON features a novel differentiable particle-based physical simulator built on neural implicit representations. Central to this design is an efficient transformation between SDF-based implicit representations and explicit surface points via our proposed Surface Points Marching Cubes (SP-MC), enabling differentiable learning with both rendering and physical losses. Additionally, PHYRECON models both rendering and physical uncertainty to identify and compensate for inconsistent and inaccurate monocular geometric priors. The physical uncertainty further facilitates physics-guided pixel sampling to enhance the learning of slender structures. By integrating these techniques, our model supports differentiable joint modeling of appearance, geometry, and physics. Extensive experiments demonstrate that PHYRECON significantly improves the reconstruction quality. Our results also exhibit superior physical stability in physical simulators, with at least a 40% improvement across all datasets, paving the way for future physics-based applications. Junfeng Ni, Yixin Chen 0003, Bohan Jing, Bo Dai 0025, Puhao Li, Yixin Zhu 0001, Song-Chun Zhu, Siyuan Huang 0001 |
NeurIPS | 8 |
| 2024 | Autonomous Character-Scene Interaction Synthesis from Text Instruction
Zimo He, Zi Wang 0014, Yixin Chen 0003, Siyuan Huang 0001, Yixin Zhu 0001 |
SIGGRAPH Asia | 7 |
| 2023 | Diffusion-based Generation, Optimization, and Planning in 3D ScenesabstractWe introduce the SceneDiffuser, a conditional generative model for 3D scene understanding. SceneDiffuser provides a unified model for solving scene-conditioned generation, optimization, and planning. In contrast to prior work, SceneDiffuser is intrinsically scene-aware, physics-based, and goal-oriented. With an iterative sampling strategy, SceneDiffuser jointly formulates the scene-aware generation, physics-based optimization, and goal-oriented planning via a diffusion-based denoising process in a fully differentiable fashion. Such a design alleviates the discrepancies among different modules and the posterior collapse of previous scene-conditioned generative models. We evaluate the SceneDiffuser on various 3D scene understanding tasks, including human pose and motion generation, dexterous grasp generation, path planning for 3D navigation, and motion planning for robot arms. The results show significant improvements compared with previous models, demonstrating the tremendous potential of the SceneDiffuser for the broad community of 3D scene understanding. Siyuan Huang 0001, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu 0001, Wei Liang 0008, Song-Chun Zhu |
CVPR | 6 |
| 2023 | X-VoE: Measuring eXplanatory Violation of Expectation in Physical EventsabstractIntuitive physics is pivotal for human understanding of the physical world, enabling prediction and interpretation of events even in infancy. Nonetheless, replicating this level of intuitive physics in artificial intelligence (AI) remains a formidable challenge. This study introduces X-VoE, a comprehensive benchmark dataset, to assess AI agents’ grasp of intuitive physics. Built on the developmental psychology-rooted Violation of Expectation (VoE) paradigm, X-VoE establishes a higher bar for the explanatory capacities of intuitive physics models. Each VoE scenario within X-VoE encompasses three distinct settings, probing models’ comprehension of events and their underlying explanations. Beyond model evaluation, we present an explanation-based learning system that captures physics dynamics and infers occluded object states solely from visual sequences, without explicit occlusion labels. Experimental outcomes highlight our model’s alignment with human commonsense when tested against X-VoE. A remarkable feature is our model’s ability to visually expound VoE events by reconstructing concealed scenes. Concluding, we discuss the findings’ implications and outline future research directions. Through X-VoE, we catalyze the advancement of AI endowed with human-like intuitive physics capabilities. Bo Dai 0025, Linge Wang, Baoxiong Jia, Zeyu Zhang 0001, Song-Chun Zhu, Chi Zhang 0017, Yixin Zhu 0001 |
ICCV | 7 |
| 2023 | Full-Body Articulated Human-Object InteractionabstractFine-grained capture of 3D Human-Object Interactions (HOIs) enhances human activity comprehension and supports various downstream visual tasks. However, previous models often assume that humans interact with rigid objects using only a few body parts, constraining their applicability. In this paper, we address the intricate challenge of Full-Body Articulated Human-Object Interaction (f-AHOI), where complete human bodies interact with articulated objects having interconnected movable joints. We introduce CHAIRS, an extensive motion-captured f-AHOI dataset comprising 17.3 hours of diverse interactions involving 46 participants and 81 articulated as well as rigid sittable objects. The CHAIRS provides 3D meshes of both humans and articulated objects throughout the interactive sequences, offering realistic and physically plausible full-body interactions. We demonstrate the utility of CHAIRS through object pose estimation. Leveraging the geometric relationships inherent in HOI, we propose a pioneering model that employs human pose estimation to address articulated object pose and shape estimation within whole-body interactions. Given an image and an estimated human pose, our model reconstructs the object’s pose and shape, refining the reconstruction based on a learned interaction prior. Across two evaluation scenarios, our model significantly outperforms baseline methods. Additionally, we showcase the significance of CHAIRS in a downstream task involving human pose generation conditioned on interacting with articulated objects. We anticipate that the availability of CHAIRS will advance the community’s understanding of finer-grained interactions. Tengyu Liu, Zhexuan Cao, Jieming Cui, Yixin Chen 0003, He Wang 0010, Yixin Zhu 0001, Siyuan Huang 0001 |
ICCV | 8 |
| 2023 | A Minimalist Dataset for Systematic Generalization of Perception, Syntax, and Semantics
Qing Li 0003, Siyuan Huang 0001, Yining Hong, Yixin Zhu 0001, Ying Nian Wu, Song-Chun Zhu |
ICLR | 4 |
| 2023 | Understanding Embodied Reference with Touch-Line Transformer
Yang Li 0178, Xiaoxue Chen, Hao Zhao 0002, Jiangtao Gong, Guyue Zhou, Federico Rossano, Yixin Zhu 0001 |
ICLR | 7 |
| 2023 | MEWL: Few-shot multimodal word learning with referential uncertaintyabstractWithout explicit feedback, humans can rapidly learn the meaning of words. Children can acquire a new word after just a few passive exposures, a process known as fast mapping. This word learning capability is believed to be the most fundamental building block of multimodal understanding and reasoning. Despite recent advancements in multimodal learning, a systematic and rigorous evaluation is still missing for human-like word learning in machines. To fill in this gap, we introduce the MachinE Word Learning (MEWL) benchmark to assess how machines learn word meaning in grounded visual scenes. MEWL covers human’s core cognitive toolkits in word learning: cross-situational reasoning, bootstrapping, and pragmatic learning. Specifically, MEWL is a few-shot benchmark suite consisting of nine tasks for probing various word learning capabilities. These tasks are carefully designed to be aligned with the children’s core abilities in word learning and echo the theories in the developmental literature. By evaluating multimodal and unimodal agents’ performance with a comparative analysis of human performance, we notice a sharp divergence in human and machine word learning. We further discuss these differences between humans and machines and call for human-like few-shot word learning in machines. Guangyuan Jiang, Manjie Xu, Shiji Xin, Wei Liang 0008, Yujia Peng, Chi Zhang 0017, Yixin Zhu 0001 |
ICML | 7 |
| 2023 | On the Complexity of Bayesian GeneralizationabstractWe examine concept generalization at a large scale in the natural visual spectrum. Established computational modes (*i.e.*, rule-based or similarity-based) are primarily studied isolated, focusing on confined and abstract problem spaces. In this work, we study these two modes when the *problem space* scales up and when the *complexity* of concepts becomes diverse. At the **representational level**, we investigate how the complexity varies when a visual concept is mapped to the representation space. Prior literature has shown that two types of complexities (Griffiths & Tenenbaum, 2003) build an inverted-U relation (Donderi, 2006; Sun & Firestone, 2021). Leveraging *Representativeness of Attribute* (RoA), we computationally confirm: Models use attributes with high RoA to describe visual concepts, and the description length falls in an inverted-U relation with the increment in visual complexity. At the **computational level**, we examine how the complexity of representation affects the shift between the rule- and similarity-based generalization. We hypothesize that category-conditioned visual modeling estimates the co-occurrence frequency between visual and categorical attributes, thus potentially serving as the prior for the natural visual world. Experimental results show that representations with relatively high subjective complexity outperform those with relatively low subjective complexity in rule-based generalization, while the trend is the opposite in similarity-based generalization. Yu-Zhe Shi, Manjie Xu, John E. Hopcroft, Kun He 0001, Josh Tenenbaum, Song-Chun Zhu, Ying Nian Wu, Wenjuan Han, Yixin Zhu 0001 |
ICML | 9 |
| 2023 | GenDexGrasp: Generalizable Dexterous GraspingabstractGenerating dexterous grasping has been a long-standing and challenging robotic task. Despite recent progress, existing methods primarily suffer from two issues. First, most prior art focuses on a specific type of robot hand, lacking generalizable capability of handling unseen ones. Second, prior arts oftentimes fail to rapidly generate diverse grasps with a high success rate. To jointly tackle these challenges with a unified solution, we propose the GenDexGrasp, a novel hand-agnostic grasping algorithm for generalizable grasping. GenDexGrasp is trained on our proposed large-scale multi-hand grasping dataset MultiDex synthesized with force closure optimization. By leveraging the contact map as a hand-agnostic intermediate representation, GenDexGrasp efficiently generates diverse and plausible grasping poses with a high success rate and can transfer among diverse multi-fingered robotic hands. Compared with previous methods, GenDexGrasp achieves a three-way trade-off among success rate, inference speed, and diversity. Puhao Li, Tengyu Liu, Yiran Geng, Yixin Zhu 0001, Yaodong Yang 0001, Siyuan Huang 0001 |
ICRA | 5 |
| 2023 | Rearrange Indoor Scenes for Human-Robot Co-ActivityabstractWe present an optimization-based framework for rearranging indoor furniture to accommodate human-robot co-activities better. The rearrangement aims to afford sufficient accessible space for robot activities without compromising everyday human activities. To retain human activities, our algorithm preserves the functional relations among furniture by integrating spatial and semantic co-occurrence extracted from SUNCG and ConceptNet, respectively. By defining the robot's accessible space by the amount of open space it can traverse and the number of objects it can reach, we formulate the rearrangement for human-robot co-activity as an optimization problem, solved by adaptive simulated annealing (ASA) and covariance matrix adaptation evolution strategy (CMA-ES). Our experiments on the SUNCG dataset quantitatively show that rearranged scenes provide a robot with 14% more accessible space and 30% more objects to interact with on average. The quality of the rearranged scenes is qualitatively validated by a human study, indicating the efficacy of the proposed strategy. Weiqi Wang 0004, Zihang Zhao, Ziyuan Jiao, Yixin Zhu 0001, Song-Chun Zhu, Hangxin Liu |
ICRA | 4 |
| 2023 | Learning a Causal Transition Model for Object CuttingabstractCutting objects into desired fragments is challenging for robots due to the spatially unstructured nature of fragments and the complex one-to-many object fragmentation caused by actions. We present a novel approach to model object fragmentation using an attributed stochastic grammar. This grammar abstracts fragment states as node variables and captures causal transitions in object fragmentation through production rules. We devise a probabilistic framework to learn this grammar from human demonstrations. The planning process for object cutting involves inferring an optimal parse tree of desired fragments using the learned grammar, with parse tree productions corresponding to cutting actions. We employ Monte Carlo Tree Search (MCTS) to efficiently approximate the optimal parse tree and generate a sequence of executable cutting actions. The experiments demonstrate the efficacy in planning for object-cutting tasks, both in simulation and on a physical robot. The proposed approach outperforms several baselines by demonstrating superior generalization to novel setups, thanks to the compositionality of the grammar model. Zeyu Zhang 0001, Muzhi Han, Baoxiong Jia, Ziyuan Jiao, Yixin Zhu 0001, Song-Chun Zhu, Hangxin Liu |
IROS | 5 |
| 2023 | Sequential Manipulation Planning for Over-Actuated Unmanned Aerial ManipulatorsabstractWe investigate the sequential manipulation planning problem for unmanned aerial manipulators (UAMs). Unlike prior work that primarily focuses on one-step manipulation tasks, sequential manipulations require coordinated motions of a UAM's floating base, the manipulator, and the object being manipulated, entailing a unified kinematics and dynamics model for motion planning under designated constraints. By leveraging a virtual kinematic chain (VKC)-based motion planning framework that consolidates components' kinematics into one chain, the sequential manipulation task of a UAM can be planned as a whole, yielding more coordinated motions. Integrating the kinematics and dynamics models with a hierarchical control framework, we demonstrate, for the first time, an over-actuated UAM achieves a series of new sequential manipulation capabilities in both simulation and experiment. Yao Su 0001, Ziyuan Jiao, Meng Wang 0051, Chi Chu, Yixin Zhu 0001, Hangxin Liu |
IROS | 7 |
| 2023 | Part-level Scene Reconstruction Affords Robot InteractionabstractExisting methods for reconstructing interactive scenes primarily focus on replacing reconstructed objects with CAD models retrieved from a limited database, resulting in significant discrepancies between the reconstructed and observed scenes. To address this issue, our work introduces a part-level reconstruction approach that reassembles objects using primitive shapes. This enables us to precisely replicate the observed physical scenes and simulate robot interactions with both rigid and articulated objects. By segmenting reconstructed objects into semantic parts and aligning primitive shapes to these parts, we assemble them as CAD models while estimating kinematic relations, including parent-child contact relations, joint types, and parameters. Specifically, we derive the optimal primitive alignment by solving a series of optimization problems, and estimate kinematic relations based on part semantics and geometry. Our experiments demonstrate that part-level scene reconstruction outperforms object-level reconstruction by accurately capturing finer details and improving precision. These reconstructed part-level interactive scenes provide valuable kinematic information for various robotic applications; we showcase the feasibility of certifying mobile manipulation planning in these interactive scenes before executing tasks in the physical world. Zeyu Zhang 0001, Lexing Zhang, Zaijin Wang, Ziyuan Jiao, Muzhi Han, Yixin Zhu 0001, Song-Chun Zhu, Hangxin Liu |
IROS | 6 |
| 2023 | ProBio: A Protocol-guided Multimodal Dataset for Molecular Biology LababstractThe challenge of replicating research results has posed a significant impediment to the field of molecular biology. The advent of modern intelligent systems has led to notable progress in various domains. Consequently, we embarked on an investigation of intelligent monitoring systems as a means of tackling the issue of the reproducibility crisis. Specifically, we first curate a comprehensive multimodal dataset, named ProBio, as an initial step towards this objective. This dataset comprises fine-grained hierarchical annotations intended for the purpose of studying activity understanding in BioLab. Next, we devise two challenging benchmarks, transparent solution tracking and multimodal action recognition, to emphasize the unique characteristics and difficulties associated with activity understanding in BioLab settings. Finally, we provide a thorough experimental evaluation of contemporary video understanding models and highlight their limitations in this specialized domain to identify potential avenues for future research. We hope \dataset with associated benchmarks may garner increased focus on modern AI techniques in the realm of molecular biology. Jieming Cui, Ziren Gong, Baoxiong Jia, Siyuan Huang 0001, Zilong Zheng, Jianzhu Ma, Yixin Zhu 0001 |
NeurIPS | 7 |
| 2023 | Evaluating and Inducing Personality in Pre-trained Language ModelsabstractStandardized and quantified evaluation of machine behaviors is a crux of understanding LLMs. In this study, we draw inspiration from psychometric studies by leveraging human personality theory as a tool for studying machine behaviors. Originating as a philosophical quest for human behaviors, the study of personality delves into how individuals differ in thinking, feeling, and behaving. Toward building and understanding human-like social machines, we are motivated to ask: Can we assess machine behaviors by leveraging human psychometric tests in a **principled** and **quantitative** manner? If so, can we induce a specific personality in LLMs? To answer these questions, we introduce the Machine Personality Inventory (MPI) tool for studying machine behaviors; MPI follows standardized
personality tests, built upon the Big Five Personality Factors (Big Five) theory and personality assessment inventories. By systematically evaluating LLMs with MPI, we provide the first piece of evidence demonstrating the efficacy of MPI in studying LLMs behaviors. We further devise a Personality Prompting (P$^2$) method to induce LLMs with specific personalities in a **controllable** way, capable of producing diverse and verifiable behaviors. We hope this work sheds light on future studies by adopting personality as the essential indicator for various downstream tasks, and could further motivate research into equally intriguing human-like machine behaviors. Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang 0017, Yixin Zhu 0001 |
NeurIPS | 6 |
| 2023 | ChimpACT: A Longitudinal Dataset for Understanding Chimpanzee BehaviorsabstractUnderstanding the behavior of non-human primates is crucial for improving animal welfare, modeling social behavior, and gaining insights into distinctively human and phylogenetically shared behaviors. However, the lack of datasets on non-human primate behavior hinders in-depth exploration of primate social interactions, posing challenges to research on our closest living relatives. To address these limitations, we present ChimpACT, a comprehensive dataset for quantifying the longitudinal behavior and social relations of chimpanzees within a social group. Spanning from 2015 to 2018, ChimpACT features videos of a group of over 20 chimpanzees residing at the Leipzig Zoo, Germany, with a particular focus on documenting the developmental trajectory of one young male, Azibo. ChimpACT is both comprehensive and challenging, consisting of 163 videos with a cumulative 160,500 frames, each richly annotated with detection, identification, pose estimation, and fine-grained spatiotemporal behavior labels. We benchmark representative methods of three tracks on ChimpACT: (i) tracking and identification, (ii) pose estimation, and (iii) spatiotemporal action detection of the chimpanzees. Our experiments reveal that ChimpACT offers ample opportunities for both devising new methods and adapting existing ones to solve fundamental computer vision tasks applied to chimpanzee groups, such as detection, pose estimation, and behavior analysis, ultimately deepening our comprehension of communication and sociality in non-human primates. Xiaoxuan Ma 0001, Stephan P. Kaufhold, Jiajun Su, Wentao Zhu 0004, Jack Terwilliger, Andres Meza 0001, Yixin Zhu 0001, Federico Rossano, Yizhou Wang 0001 |
NeurIPS | 7 |
| 2023 | Active Reasoning in an Open-World EnvironmentabstractRecent advances in vision-language learning have achieved notable success on *complete-information* question-answering datasets through the integration of extensive world knowledge. Yet, most models operate *passively*, responding to questions based on pre-stored knowledge. In stark contrast, humans possess the ability to *actively* explore, accumulate, and reason using both newfound and existing information to tackle *incomplete-information* questions. In response to this gap, we introduce **Conan**, an interactive open-world environment devised for the assessment of *active reasoning*. **Conan** facilitates active exploration and promotes multi-round abductive inference, reminiscent of rich, open-world settings like Minecraft. Diverging from previous works that lean primarily on single-round deduction via instruction following, **Conan** compels agents to actively interact with their surroundings, amalgamating new evidence with prior knowledge to elucidate events from incomplete observations. Our analysis on \bench underscores the shortcomings of contemporary state-of-the-art models in active exploration and understanding complex scenarios. Additionally, we explore *Abduction from Deduction*, where agents harness Bayesian rules to recast the challenge of abduction as a deductive process. Through **Conan**, we aim to galvanize advancements in active reasoning and set the stage for the next generation of artificial intelligence agents adept at dynamically engaging in environments. Manjie Xu, Guangyuan Jiang, Wei Liang 0008, Chi Zhang 0017, Yixin Zhu 0001 |
NeurIPS | 5 |
| 2023 | Interactive Visual Reasoning under UncertaintyabstractOne of the fundamental cognitive abilities of humans is to quickly resolve uncertainty by generating hypotheses and testing them via active trials. Encountering a novel phenomenon accompanied by ambiguous cause-effect relationships, humans make hypotheses against data, conduct inferences from observation, test their theory via experimentation, and correct the proposition if inconsistency arises. These iterative processes persist until the underlying mechanism becomes clear. In this work, we devise the IVRE (pronounced as "ivory") environment for evaluating artificial agents' reasoning ability under uncertainty. IVRE is an interactive environment featuring rich scenarios centered around Blicket detection. Agents in IVRE are placed into environments with various ambiguous action-effect pairs and asked to determine each object's role. They are encouraged to propose effective and efficient experiments to validate their hypotheses based on observations and actively gather new information. The game ends when all uncertainties are resolved or the maximum number of trials is consumed. By evaluating modern artificial agents in IVRE, we notice a clear failure of today's learning methods compared to humans. Such inefficacy in interactive reasoning ability under uncertainty calls for future research in building human-like intelligence. Manjie Xu, Guangyuan Jiang, Wei Liang 0008, Chi Zhang 0017, Yixin Zhu 0001 |
NeurIPS | 5 |
| 2022 | What Is the point? a Theory of Mind Model of Relevance
Stephanie Stacy, Annya L. Dahmani, Boxuan Jiang, Federico Rossano, Yixin Zhu 0001, Tao Gao 0004 |
CogSci | 6 |
| 2022 | Learning Algebraic Representation for Systematic Generalization in Abstract Reasoning
Chi Zhang 0017, Sirui Xie, Baoxiong Jia, Ying Nian Wu, Song-Chun Zhu, Yixin Zhu 0001 |
ECCV (39) | 6 |
| 2022 | Latent Diffusion Energy-Based Model for Interpretable Text ModellingabstractLatent space Energy-Based Models (EBMs), also known as energy-based priors, have drawn growing interests in generative modeling. Fueled by its flexibility in the formulation and strong modeling power of the latent space, recent works built upon it have made interesting attempts aiming at the interpretability of text modeling. However, latent space EBMs also inherit some flaws from EBMs in data space; the degenerate MCMC sampling quality in practice can lead to poor generation quality and instability in training, especially on data with complex latent structures. Inspired by the recent efforts that leverage diffusion recovery likelihood learning as a cure for the sampling issue, we introduce a novel symbiosis between the diffusion models and latent space EBMs in a variational learning framework, coined as the latent diffusion energy-based model. We develop a geometric clustering-based regularization jointly with the information bottleneck to further improve the quality of the learned latent space. Experiments on several challenging tasks demonstrate the superior performance of our model on interpretable text modeling over strong counterparts. Peiyu Yu, Sirui Xie, Xiaojian Ma 0001, Baoxiong Jia, Bo Pang 0004, Ruiqi Gao, Yixin Zhu 0001, Song-Chun Zhu, Ying Nian Wu |
ICML | 7 |
| 2022 | Sequential Manipulation Planning on Scene GraphabstractWe devise a 3D scene graph representation, contact graph+(cg+), for efficient sequential manipulation planning. Augmented with predicate-like attributes, this contact graph-based representation abstracts scene layouts with succinct geometric information and valid robot-scene interactions. Goal configurations, naturally specified on contact graphs, can be produced by a genetic algorithm with a stochastic optimization method. A task plan is then initialized by computing the Graph Editing Distance (GED) between the initial contact graph and the goal configuration, which generates graph edit operations corresponding to possible robot actions. We finalize the task plan by imposing constraints to regulate the temporal feasibility of graph edit operations, ensuring valid task and motion correspondences. In a series of simulated and real experiments, robots successfully complete complex sequential object rearrangement tasks that are difficult to specify using conventional planning language like Planning Domain Definition Language (PDDL), demonstrating high potential of planning sequential manipulation tasks on cg+. Ziyuan Jiao, Yida Niu, Zeyu Zhang 0001, Song-Chun Zhu, Yixin Zhu 0001, Hangxin Liu |
IROS | 5 |
| 2022 | Downwash-aware Control Allocation for Over-actuated UAV PlatformsabstractTracking position and orientation independently affords more agile maneuver for over-actuated multirotor Unmanned Aerial Vehicles (UAVs) while introducing undesired downwash effects; downwash flows generated by thrust generators may counteract others due to close proximity, which significantly threatens the stability of the platform. The complexity of modeling aerodynamic airflow challenges control algorithms from properly compensating for such a side effect. Leveraging the input redundancies in over-actuated UAVs, we tackle this issue with a novel control allocation framework that considers downwash effects and explores the entire allocation space for an optimal solution. This optimal solution avoids downwash effects while providing high thrust efficiency within the hardware constraints. To the best of our knowledge, ours is the first formal derivation to investigate the downwash effects on over-actuated UAVs. We verify our framework on different hardware configurations in both simulation and experiment. Yao Su 0001, Chi Chu, Meng Wang 0051, Yixin Zhu 0001, Hangxin Liu |
IROS | 6 |
| 2022 | On the Learning Mechanisms in Physical ReasoningabstractIs dynamics prediction indispensable for physical reasoning? If so, what kind of roles do the dynamics prediction modules play during the physical reasoning process? Most studies focus on designing dynamics prediction networks and treating physical reasoning as a downstream task without investigating the questions above, taking for granted that the designed dynamics prediction would undoubtedly help the reasoning process. In this work, we take a closer look at this assumption, exploring this fundamental hypothesis by comparing two learning mechanisms: Learning from Dynamics (LfD) and Learning from Intuition (LfI). In the first experiment, we directly examine and compare these two mechanisms. Results show a surprising finding: Simple LfI is better than or on par with state-of-the-art LfD. This observation leads to the second experiment with Ground-truth Dynamics (GD), the ideal case of LfD wherein dynamics are obtained directly from a simulator. Results show that dynamics, if directly given instead of approximated, would achieve much higher performance than LfI alone on physical reasoning; this essentially serves as the performance upper bound. Yet practically, LfD mechanism can only predict Approximate Dynamics (AD) using dynamics learning modules that mimic the physical laws, making the following downstream physical reasoning modules degenerate into the LfI paradigm; see the third experiment. We note that this issue is hard to mitigate, as dynamics prediction errors inevitably accumulate in the long horizon. Finally, in the fourth experiment, we note that LfI, the extremely simpler strategy when done right, is more effective in learning to solve physical reasoning problems. Taken together, the results on the challenging benchmark of PHYRE show that LfI is, if not better, as good as LfD with bells and whistles for dynamics prediction. However, the potential improvement from LfD, though challenging, remains lucrative. Shiqian Li, Kewen Wu 0004, Chi Zhang 0017, Yixin Zhu 0001 |
NeurIPS | 4 |
| 2022 | Emergent Graphical Conventions in a Visual Communication GameabstractHumans communicate with graphical sketches apart from symbolic languages. Primarily focusing on the latter, recent studies of emergent communication overlook the sketches; they do not account for the evolution process through which symbolic sign systems emerge in the trade-off between iconicity and symbolicity. In this work, we take the very first step to model and simulate this process via two neural agents playing a visual communication game; the sender communicates with the receiver by sketching on a canvas. We devise a novel reinforcement learning method such that agents are evolved jointly towards successful communication and abstract graphical conventions. To inspect the emerged conventions, we define three key properties -- iconicity, symbolicity, and semanticity -- and design evaluation methods accordingly. Our experimental results under different controls are consistent with the observation in studies of human graphical conventions. Of note, we find that evolved sketches can preserve the continuum of semantics under proper environmental pressures. More interestingly, co-evolved agents can switch between conventionalized and iconic communication based on their familiarity with referents. We hope the present research can pave the path for studying emergent communication with the modality of sketches. Shuwen Qiu, Sirui Xie, Lifeng Fan, Tao Gao 0004, Jungseock Joo, Song-Chun Zhu, Yixin Zhu 0001 |
NeurIPS | 7 |
| 2022 | HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesabstractLearning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characters of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality and lack semantics. To fill in the gap, we propose a large-scale and semantic-rich synthetic HSI dataset, denoted as HUMANISE, by aligning the captured human motion sequences with various 3D indoor scenes. We automatically annotate the aligned motions with language descriptions that depict the action and the individual interacting objects; e.g., sit on the armchair near the desk. HUMANIZE thus enables a new generation task, language-conditioned human motion generation in 3D scenes. The proposed task is challenging as it requires joint modeling of the 3D scene, human motion, and natural language. To tackle this task, we present a novel scene-and-language conditioned generative model that can produce 3D human motions of the desirable action interacting with the specified objects. Our experiments demonstrate that our model generates diverse and semantically consistent human motions in 3D scenes. Yixin Chen 0003, Tengyu Liu, Yixin Zhu 0001, Wei Liang 0008, Siyuan Huang 0001 |
NeurIPS | 4 |
| 2022 | Scene Reconstruction with Functional Objects for Robot Autonomy
Muzhi Han, Zeyu Zhang 0001, Ziyuan Jiao, Xu Xie 0001, Yixin Zhu 0001, Song-Chun Zhu, Hangxin Liu |
Int. J. Comput. Vis. | 5 |
| 2021 | Individual vs. Joint Perception: a Pragmatic Model of Pointing as Smithian Helping
Stephanie Stacy, Adelpha Chan, Chuyu Wei, Federico Rossano, Yixin Zhu 0001, Tao Gao 0004 |
CogSci | 6 |
| 2021 | Sharing is not Needed: Modeling Animal Coordinated Hunting with Reinforcement Learning
Minglu Zhao, Annya L. Dahmani, Ross Richard Perry, Yixin Zhu 0001, Federico Rossano, Tao Gao 0004 |
CogSci | 5 |
| 2021 | ACRE: Abstract Causal REasoning Beyond CovariationabstractCausal induction, i.e., identifying unobservable mechanisms that lead to the observable relations among variables, has played a pivotal role in modern scientific discovery, especially in scenarios with only sparse and limited data. Humans, even young toddlers, can induce causal relationships surprisingly well in various settings despite its notorious difficulty. However, in contrast to the commonplace trait of human cognition is the lack of a diagnostic benchmark to measure causal induction for modern Artificial Intelligence (AI) systems. Therefore, in this work, we introduce the Abstract Causal REasoning (ACRE) dataset for systematic evaluation of current vision systems in causal induction. Motivated by the stream of research on causal discovery in Blicket experiments, we query a visual reasoning system with the following four types of questions in either an independent scenario or an interventional scenario: direct, indirect, screening-off, and backward-blocking, intentionally going beyond the simple strategy of inducing causal relationships by covariation. By analyzing visual reasoning architectures on this testbed, we notice that pure neural models tend towards an associative strategy under their chance-level performance, whereas neuro-symbolic combinations struggle in backward-blocking reasoning. These deficiencies call for future research in models with a more comprehensive capability of causal induction. Chi Zhang 0017, Baoxiong Jia, Mark Edmonds, Song-Chun Zhu, Yixin Zhu 0001 |
CVPR | 5 |
| 2021 | Abstract Spatial-Temporal Reasoning via Probabilistic Abduction and Execution
Chi Zhang 0017, Baoxiong Jia, Song-Chun Zhu, Yixin Zhu 0001 |
CVPR | 4 |
| 2021 | Learning Triadic Belief Dynamics in Nonverbal Communication From VideosabstractHumans possess a unique social cognition capability [43], [20]; nonverbal communication can convey rich social information among agents. In contrast, such crucial social characteristics are mostly missing in the existing scene understanding literature. In this paper, we incorporate different nonverbal communication cues (e.g., gaze, human poses, and gestures) to represent, model, learn, and infer agents’ mental states from pure visual inputs. Crucially, such a mental representation takes the agent’s belief into account so that it represents what the true world state is and infers the beliefs in each agent’s mental state, which may differ from the true world states. By aggregating different beliefs and true world states, our model essentially forms "five minds" during the interactions between two agents. This "five minds" model differs from prior works that infer beliefs in an infinite recursion; instead, agents’ beliefs are converged into a "common mind" [31], [47]. Based on this representation, we further devise a hierarchical energy-based model that jointly tracks and predicts all five minds. From this new perspective, a social event is interpreted by a series of nonverbal communication and belief dynamics, which transcends the classic keyframe video summary. In the experiments, we demonstrate that using such a social account provides a better video summary on videos with rich social interactions compared with state-of-the-art keyframe video summary methods. Lifeng Fan, Shuwen Qiu, Zilong Zheng, Tao Gao 0004, Song-Chun Zhu, Yixin Zhu 0001 |
CVPR | 6 |
| 2021 | YouRefIt: Embodied Reference Understanding with Language and GestureabstractWe study the machine’s understanding of embodied reference: One agent uses both language and gesture to refer to an object to another agent in a shared physical environment. Of note, this new visual task requires understanding multimodal cues with perspective-taking to identify which object is being referred to. To tackle this problem, we introduce YouRefIt, a new crowd-sourced dataset of embodied reference collected in various physical scenes; the dataset contains 4,195 unique reference clips in 432 indoor scenes. To the best of our knowledge, this is the first embodied reference dataset that allows us to study referring expressions in daily physical scenes to understand referential behavior, human communication, and human-robot interaction. We further devise two benchmarks for image-based and video-based embodied reference understanding. Comprehensive baselines and extensive experiments provide the very first result of machine perception on how the referring expressions and gestures affect the embodied reference understanding. Our results provide essential evidence that gestural cues are as critical as language cues in understanding the embodied reference. Yixin Chen 0003, Qing Li 0003, Deqian Kong, Yik Lun Kei, Song-Chun Zhu, Tao Gao 0004, Yixin Zhu 0001, Siyuan Huang 0001 |
ICCV | 7 |
| 2021 | Spatio-temporal Self-Supervised Representation Learning for 3D Point CloudsabstractTo date, various 3D scene understanding tasks still lack practical and generalizable pre-trained models, primarily due to the intricate nature of 3D scene understanding tasks and their immense variations introduced by camera views, lighting, occlusions, etc. In this paper, we tackle this challenge by introducing a spatio-temporal representation learning (STRL) framework, capable of learning from unlabeled 3D point clouds in a self-supervised fashion. Inspired by how infants learn from visual data in the wild, we explore the rich spatio-temporal cues derived from the 3D data. Specifically, STRL takes two temporally-correlated frames from a 3D point cloud sequence as the input, transforms it with the spatial data augmentation, and learns the invariant representation self-supervisedly. To corroborate the efficacy of STRL, we conduct extensive experiments on three types (synthetic, indoor, and outdoor) of datasets. Experimental results demonstrate that, compared with supervised learning methods, the learned self-supervised representation facilitates various models to attain comparable or even better performances while capable of generalizing pre-trained models to downstream tasks, including 3D shape classification, 3D object detection, and 3D semantic segmentation. Moreover, the spatio-temporal contextual cues embedded in 3D point clouds significantly improve the learned representations. Siyuan Huang 0001, Yichen Xie 0002, Song-Chun Zhu, Yixin Zhu 0001 |
ICCV | 4 |
| 2021 | Congestion-aware Multi-agent Trajectory Prediction for Collision AvoidanceabstractPredicting agents’ future trajectories plays a crucial role in modern AI systems, yet it is challenging due to intricate interactions exhibited in multi-agent systems, especially when it comes to collision avoidance. To address this challenge, we propose to learn congestion patterns as contextual cues explicitly and devise a novel "Sense–Learn–Reason–Predict" framework by exploiting advantages of three different doctrines of thought, which yields the following desirable benefits: (i) Representing congestion as contextual cues via latent factors subsumes the concept of social force commonly used in physics- based approaches and implicitly encodes the distance as a cost, similar to the way a planning-based method models the environment. (ii) By decomposing the learning phases into two stages, a "student" can learn contextual cues from a "teacher" while generating collision-free trajectories. To make the framework computationally tractable, we formulate it as an optimization problem and derive an upper bound by leveraging the variational parametrization. In experiments, we demonstrate that the proposed model is able to generate collision- free trajectory predictions in a synthetic dataset designed for collision avoidance evaluation and remains competitive on the commonly used NGSIM US-101 highway dataset. Source code and dataset tools can be accessed via Github. Xu Xie 0001, Chi Zhang 0017, Yixin Zhu 0001, Ying Nian Wu, Song-Chun Zhu |
ICRA | 3 |
| 2021 | Reconstructing Interactive 3D Scenes by Panoptic Mapping and CAD Model AlignmentsabstractIn this paper, we rethink the problem of scene reconstruction from an embodied agent’s perspective: While the classic view focuses on the reconstruction accuracy, our new perspective emphasizes the underlying functions and constraints such that the reconstructed scenes provide actionable information for simulating interactions with agents. Here, we address this challenging problem by reconstructing an interactive scene using RGB-D data stream, which captures (i) the semantics and geometry of objects and layouts by a 3D volumetric panoptic mapping module, and (ii) object affordance and contextual relations by reasoning over physical common sense among objects, organized by a graph-based scene representation. Crucially, this reconstructed scene replaces the object meshes in the dense panoptic map with part-based articulated CAD models for finer-grained robot interactions. In the experiments, we demonstrate that (i) our panoptic mapping module outperforms previous state-of-the-art methods, (ii) a high-performant physical reasoning procedure that matches, aligns, and replaces objects’ meshes with best-fitted CAD models, and (iii) reconstructed scenes are physically plausible and naturally afford actionable interactions; without any manual labeling, they are seamlessly imported to ROS-based simulators and virtual environments for complex robot task executions.1 Muzhi Han, Zeyu Zhang 0001, Ziyuan Jiao, Xu Xie 0001, Yixin Zhu 0001, Song-Chun Zhu, Hangxin Liu |
ICRA | 5 |
| 2021 | Consolidating Kinematic Models to Promote Coordinated Mobile ManipulationsabstractWe construct a Virtual Kinematic Chain (VKC) that readily consolidates the kinematics of the mobile base, the arm, and the object to be manipulated in mobile manipulations. Accordingly, a mobile manipulation task is represented by altering the state of the constructed VKC, which can be converted to a motion planning problem, formulated and solved by trajectory optimization. This new VKC perspective of mobile manipulation allows a service robot to (i) produce well-coordinated motions, suitable for complex household environments, and (ii) perform intricate multi-step tasks while interacting with multiple objects without an explicit definition of intermediate goals. In simulated experiments, we validate these advantages by comparing the VKC-based approach with baselines that solely optimize individual components. The results manifest that VKC-based joint modeling and planning promote task success rates and produce more efficient trajectories. Ziyuan Jiao, Zeyu Zhang 0001, David K. Han, Song-Chun Zhu, Yixin Zhu 0001, Hangxin Liu |
IROS | 6 |
| 2021 | Efficient Task Planning for Mobile Manipulation: a Virtual Kinematic Chain PerspectiveabstractWe present a Virtual Kinematic Chain (VKC) perspective, a simple yet effective method, to improve task planning efficacy for mobile manipulation. By consolidating the kinematics of the mobile base, the arm, and the object being manipulated collectively as a whole, this novel VKC perspective naturally defines abstract actions and eliminates unnecessary predicates in describing intermediate poses. As a result, these advantages simplify the design of the planning domain and significantly reduce the search space and branching factors in solving planning problems. In experiments, we implement a task planner using Planning Domain Definition Language (PDDL) with VKC. Compared with conventional domain definition, our VKC-based domain definition is more efficient in both planning time and memory. In addition, abstract actions perform better in producing feasible motion plans and trajectories. We further scale up the VKC-based task planner in complex mobile manipulation tasks. Taken together, these results demonstrate that task planning using VKC for mobile manipulation is not only natural and effective but also introduces new capabilities. Ziyuan Jiao, Zeyu Zhang 0001, Weiqi Wang 0004, David K. Han, Song-Chun Zhu, Yixin Zhu 0001, Hangxin Liu |
IROS | 6 |
| 2021 | Communicative Learning with Natural Gestures for Embodied Navigation Agents with Human-in-the-SceneabstractHuman-robot collaboration is an essential re-search topic in artificial intelligence (AI), enabling researchers to devise cognitive AI systems and affords an intuitive means for users to interact with the robot. Of note, communication plays a central role. To date, prior studies in embodied agent navigation have only demonstrated that human languages facilitate communication by instructions in natural languages. Nevertheless, a plethora of other forms of communication is left unexplored. In fact, human communication originated in gestures and oftentimes is delivered through multimodal cues, e.g., “go there” with a pointing gesture. To bridge the gap and fill in the missing dimension of communication in embodied agent navigation, we propose investigating the effects of using gestures as the communicative interface instead of verbal cues. Specifically, we develop a VR-based 3D simulation environment, named Gesture-based THOR (Ges-THOR), based on AI2-THOR platform. In this virtual environment, a human player is placed in the same virtual scene and shepherds the artificial agent using only gestures. The agent is tasked to solve the navigation problem guided by natural gestures with unknown semantics; we do not use any predefined gestures due to the diversity and versatile nature of human gestures. We argue that learning the semantics of natural gestures is mutually beneficial to learning the navigation task—learn to communicate and communicate to learn. In a series of experiments, we demonstrate that human gesture cues, even without predefined semantics, improve the object-goal navigation for an embodied agent, outperforming various state-of-the-art methods. Cheng-Ju Wu, Yixin Zhu 0001, Jungseock Joo |
IROS | 3 |
| 2021 | Unsupervised Foreground Extraction via Deep Region CompetitionabstractWe present Deep Region Competition (DRC), an algorithm designed to extract foreground objects from images in a fully unsupervised manner. Foreground extraction can be viewed as a special case of generic image segmentation that focuses on identifying and disentangling objects from the background. In this work, we rethink the foreground extraction by reconciling energy-based prior with generative image modeling in the form of Mixture of Experts (MoE), where we further introduce the learned pixel re-assignment as the essential inductive bias to capture the regularities of background regions. With this modeling, the foreground-background partition can be naturally found through Expectation-Maximization (EM). We show that the proposed method effectively exploits the interaction between the mixture components during the partitioning process, which closely connects to region competition, a seminal approach for generic image segmentation. Experiments demonstrate that DRC exhibits more competitive performances on complex real-world data and challenging multi-object scenes compared with prior methods. Moreover, we show empirically that DRC can potentially generalize to novel foreground objects even from categories unseen during training. Peiyu Yu, Sirui Xie, Xiaojian Ma 0001, Yixin Zhu 0001, Ying Nian Wu, Song-Chun Zhu |
NeurIPS | 4 |
| 2020 | Theory-Based Causal Transfer: Integrating Instance-Level Induction and Abstract-Level Structure LearningabstractLearning transferable knowledge across similar but different settings is a fundamental component of generalized intelligence. In this paper, we approach the transfer learning challenge from a causal theory perspective. Our agent is endowed with two basic yet general theories for transfer learning: (i) a task shares a common abstract structure that is invariant across domains, and (ii) the behavior of specific features of the environment remain constant across domains. We adopt a Bayesian perspective of causal theory induction and use these theories to transfer knowledge between environments. Given these general theories, the goal is to train an agent by interactively exploring the problem space to (i) discover, form, and transfer useful abstract and structural knowledge, and (ii) induce useful knowledge from the instance-level attributes observed in the environment. A hierarchy of Bayesian structures is used to model abstract-level structural causal knowledge, and an instance-level associative learning scheme learns which specific objects can be used to induce state changes through interaction. This model-learning scheme is then integrated with a model-based planner to achieve a task in the OpenLock environment, a virtual “escape room” with a complex hierarchy that requires agents to reason about an abstract, generalized causal structure. We compare performances against a set of predominate model-free reinforcement learning (RL) algorithms. RL agents showed poor ability transferring learned knowledge across different trials. Whereas the proposed model revealed similar performance trends as human learners, and more importantly, demonstrated transfer behavior across trials and learning situations.1 Mark Edmonds, Xiaojian Ma 0001, Siyuan Qi, Yixin Zhu 0001, Hongjing Lu, Song-Chun Zhu |
AAAI | 4 |
| 2020 | Machine Number Sense: A Dataset of Visual Arithmetic Problems for Abstract and Relational ReasoningabstractAs a comprehensive indicator of mathematical thinking and intelligence, the number sense (Dehaene 2011) bridges the induction of symbolic concepts and the competence of problem-solving. To endow such a crucial cognitive ability to machine intelligence, we propose a dataset, Machine Number Sense (MNS), consisting of visual arithmetic problems automatically generated using a grammar model—And-Or Graph (AOG). These visual arithmetic problems are in the form of geometric figures: each problem has a set of geometric shapes as its context and embedded number symbols. Solving such problems is not trivial; the machine not only has to recognize the number, but also to interpret the number with its contexts, shapes, and relations (e.g., symmetry) together with proper operations. We benchmark the MNS dataset using four predominant neural network models as baselines in this visual reasoning task. Comprehensive experiments show that current neural-network-based models still struggle to understand number concepts and relational operations. We show that a simple brute-force search algorithm could work out some of the problems without context information. Crucially, taking geometric context into account by an additional perception module would provide a sharp performance gain with fewer search steps. Altogether, we call for attention in fusing the classic search-based algorithms with modern neural networks to discover the essential number concepts in future research. Wenhe Zhang, Chi Zhang 0017, Yixin Zhu 0001, Song-Chun Zhu |
AAAI | 3 |
| 2020 | LEMMA: A Multi-view Dataset for LEarning Multi-agent Multi-task Activities
Baoxiong Jia, Yixin Chen 0003, Siyuan Huang 0001, Yixin Zhu 0001, Song-Chun Zhu |
ECCV (26) | 4 |
| 2020 | Joint Inference of States, Robot Knowledge, and Human (False-)BeliefsabstractAiming to understand how human (false-)belief- a core socio-cognitive ability-would affect human interactions with robots, this paper proposes to adopt a graphical model to unify the representation of object states, robot knowledge, and human (false-)beliefs. Specifically, a parse graph (pg) is learned from a single-view spatiotemporal parsing by aggregating various object states along the time; such a learned representation is accumulated as the robot's knowledge. An inference algorithm is derived to fuse individual pg from all robots across multi-views into a joint pg, which affords more effective reasoning and inference capability to overcome the errors originated from a single view. In the experiments, through the joint inference over pgs, the system correctly recognizes human (false-)belief in various settings and achieves better cross-view accuracy on a challenging small object tracking dataset. Hangxin Liu, Lifeng Fan, Zilong Zheng, Tao Gao 0004, Yixin Zhu 0001, Song-Chun Zhu |
ICRA | 6 |
| 2020 | Congestion-aware Evacuation Routing using Augmented Reality DevicesabstractWe present a congestion-aware routing solution for indoor evacuation, which produces real-time individual-customized evacuation routes among multiple destinations while keeping tracks of all evacuees’ locations. A population density map, obtained on-the-fly by aggregating locations of evacuees from user-end Augmented Reality (AR) devices, is used to model the congestion distribution inside a building. To efficiently search the evacuation route among all destinations, a variant of A⋆algorithm is devised to obtain the optimal solution in a single pass. In a series of simulated studies, we show that the proposed algorithm is more computationally optimized compared to classic path planning algorithms; it generates a more time-efficient evacuation route for each individual that minimizes the overall congestion. A complete system using AR devices is implemented for a pilot study in real-world environments, demonstrating the efficacy of the proposed approach. Zeyu Zhang 0001, Hangxin Liu, Ziyuan Jiao, Yixin Zhu 0001, Song-Chun Zhu |
ICRA | 4 |
| 2020 | Graph-based Hierarchical Knowledge Representation for Robot Task Transfer from Virtual to Physical WorldabstractWe study the hierarchical knowledge transfer problem using a cloth-folding task, wherein the agent is first given a set of human demonstrations in the virtual world using an Oculus Headset, and later transferred and validated on a physical Baxter robot. We argue that such an intricate robot task transfer across different embodiments is only realizable if an abstract and hierarchical knowledge representation is formed to facilitate the process, in contrast to prior literature of sim2real in a reinforcement learning setting. Specifically, the knowledge in both the virtual and physical worlds are measured by information entropy built on top of a graph-based representation, so that the problem of task transfer becomes the minimization of the relative entropy between the two worlds. An And-Or-Graph (AOG) is introduced to represent the knowledge, induced from the human demonstrations performed across six virtual scenarios inside the Virtual Reality (VR). During the transfer, the success of a physical Baxter robot platform across all six tasks demonstrates the efficacy of the graph-based hierarchical knowledge representation. Zhenliang Zhang 0002, Yixin Zhu 0001, Song-Chun Zhu |
IROS | 2 |
| 2020 | Human-Robot Interaction in a Shared Augmented Reality WorkspaceabstractWe design and develop a new shared Augmented Reality (AR) workspace for Human-Robot Interaction (HRI), which establishes a bi-directional communication between human agents and robots. In a prototype system, the shared AR workspace enables a shared perception, so that a physical robot not only perceives the virtual elements in its own view but also infers the utility of the human agent-the cost needed to perceive and interact in AR-by sensing the human agent's gaze and pose. Such a new HRI design also affords a shared manipulation, wherein the physical robot can control and alter virtual objects in AR as an active agent; crucially, a robot can proactively interact with human agents, instead of purely passively executing received commands. In experiments, we design a resource collection game that qualitatively demonstrates how a robot perceives, processes, and manipulates in AR and quantitatively evaluates the efficacy of HRI using the shared AR workspace. We further discuss how the system can potentially benefit future HRI studies that are otherwise challenging. Shuwen Qiu, Hangxin Liu, Zeyu Zhang 0001, Yixin Zhu 0001, Song-Chun Zhu |
IROS | 4 |
| 2020 | IQ-MPM: an interface quadrature material point method for non-sticky strongly two-way coupled nonlinear solids and fluidsabstractWe propose a novel scheme for simulating two-way coupled interactions between nonlinear elastic solids and incompressible fluids. The key ingredient of this approach is a ghost matrix operator-splitting scheme for strongly coupled nonlinear elastica and incompressible fluids through the weak form of their governing equations. This leads to a stable and efficient method handling large time steps under the CFL limit while using a single monolithic solve for the coupled pressure fields, even in the case with highly nonlinear elastic solids. The use of the Material Point Method (MPM) is essential in the designing of the scheme, it not only preserves discretization consistency with the hybrid Lagrangian-Eulerian fluid solver, but also works naturally with our novel interface quadrature (IQ) discretization for free-slip boundary conditions. While traditional MPM suffers from sticky numerical artifacts, our framework naturally supports discontinuous tangential velocities at the solid-fluid interface. Our IQ discretization results in an easy-to-implement, fully particle-based treatment of the interfacial boundary, avoiding the additional complexities associated with intermediate level set or explicit mesh representations. The efficacy of the proposed scheme is verified by various challenging simulations with fluid-elastica interactions. Yu Fang 0010, Ziyin Qu, Minchen Li, Yixin Zhu 0001, Mridul Aanjaneya, Chenfanfu Jiang |
ACM Trans. Graph. | 5 |
| 2020 | A massively parallel and scalable multi-CPU material point methodabstractHarnessing the power of modern multi-GPU architectures, we present a massively parallel simulation system based on the Material Point Method (MPM) for simulating physical behaviors of materials undergoing complex topological changes, self-collision, and large deformations. Our system makes three critical contributions. First, we introduce a new particle data structure that promotes coalesced memory access patterns on the GPU and eliminates the need for complex atomic operations on the memory hierarchy when writing particle data to the grid. Second, we propose a kernel fusion approach using a new Grid-to-Particles-to-Grid ( G2P2G ) scheme, which efficiently reduces GPU kernel launches, improves latency, and significantly reduces the amount of global memory needed to store particle data. Finally, we introduce optimized algorithmic designs that allow for efficient sparse grids in a shared memory context, enabling us to best utilize modern multi-GPU computational platforms for hybrid Lagrangian-Eulerian computational patterns. We demonstrate the effectiveness of our method with extensive benchmarks, evaluations, and dynamic simulations with elastoplasticity, granular media, and fluid dynamics. In comparisons against an open-source and heavily optimized CPU-based MPM codebase [Fang et al. 2019] on an elastic sphere colliding scene with particle counts ranging from 5 to 40 million, our GPU MPM achieves over 100x per-time-step speedup on a workstation with an Intel 8086K CPU and a single Quadro P6000 GPU, exposing exciting possibilities for future MPM simulations in computer graphics and computational science. Moreover, compared to the state-of-the-art GPU MPM method [Hu et al. 2019a], we not only achieve 2x acceleration on a single GPU but our kernel fusion strategy and Array-of-Structs-of-Array ( AoSoA ) data structure design also generalizes to multi-GPU systems. Our multi-GPU MPM exhibits near-perfect weak and strong scaling with 4 GPUs, enabling performant and large-scale simulations on a 1024 3 grid with close to 100 million particles with less than 4 minutes per frame on a single 4-GPU workstation and 134 million particles with less than 1 minute per frame on an 8-GPU workstation. Yuxing Qiu, Stuart R. Slattery, Yu Fang 0010, Minchen Li, Song-Chun Zhu, Yixin Zhu 0001, Min Tang 0001, Dinesh Manocha, Chenfanfu Jiang |
ACM Trans. Graph. | 7 |
| 2019 | Mirroring without Overimitation: Learning Functionally Equivalent Manipulation ActionsabstractThis paper presents a mirroring approach, inspired by the neuroscience discovery of the mirror neurons, to transfer demonstrated manipulation actions to robots. Designed to address the different embodiments between a human (demonstrator) and a robot, this approach extends the classic robot Learning from Demonstration (LfD) in the following aspects:i) It incorporates fine-grained hand forces collected by a tactile glove in demonstration to learn robot’s fine manipulative actions; ii) Through model-free reinforcement learning and grammar induction, the demonstration is represented by a goal-oriented grammar consisting of goal states and the corresponding forces to reach the states, independent of robot embodiments; iii) A physics-based simulation engine is applied to emulate various robot actions and mirrors the actions that are functionally equivalent to the human’s in the sense of causing the same state changes by exerting similar forces. Through this approach, a robot reasons about which forces to exert and what goals to achieve to generate actions (i.e., mirroring), rather than strictly mimicking demonstration (i.e., overimitation). Thus the embodiment difference between a human and a robot is naturally overcome. In the experiment, we demonstrate the proposed approach by teaching a real Baxter robot with a complex manipulation task involving haptic feedback—opening medicine bottles. Hangxin Liu, Chi Zhang 0017, Yixin Zhu 0001, Chenfanfu Jiang, Song-Chun Zhu |
AAAI | 3 |
| 2019 | MetaStyle: Three-Way Trade-off among Speed, Flexibility, and Quality in Neural Style TransferabstractAn unprecedented booming has been witnessed in the research area of artistic style transfer ever since Gatys et al. introduced the neural method. One of the remaining challenges is to balance a trade-off among three critical aspects—speed, flexibility, and quality: (i) the vanilla optimization-based algorithm produces impressive results for arbitrary styles, but is unsatisfyingly slow due to its iterative nature, (ii) the fast approximation methods based on feed-forward neural networks generate satisfactory artistic effects but bound to only a limited number of styles, and (iii) feature-matching methods like AdaIN achieve arbitrary style transfer in a real-time manner but at a cost of the compromised quality. We find it considerably difficult to balance the trade-off well merely using a single feed-forward step and ask, instead, whether there exists an algorithm that could adapt quickly to any style, while the adapted model maintains high efficiency and good image quality. Motivated by this idea, we propose a novel method, coined MetaStyle, which formulates the neural style transfer as a bilevel optimization problem and combines learning with only a few post-processing update steps to adapt to a fast approximation model with satisfying artistic effects, comparable to the optimization-based methods for an arbitrary style. The qualitative and quantitative analysis in the experiments demonstrates that the proposed approach achieves high-quality arbitrary artistic style transfer effectively, with a good trade-off among speed, flexibility, and quality. Chi Zhang 0017, Yixin Zhu 0001, Song-Chun Zhu |
AAAI | 2 |
| 2019 | Decomposing Human Causal Learning: Bottom-up Associative Learning and Top-down Schema Reasoning
Mark Edmonds, Siyuan Qi, Yixin Zhu 0001, James Kubricht, Song-Chun Zhu, Hongjing Lu |
CogSci | 3 |
| 2019 | RAVEN: A Dataset for Relational and Analogical Visual REasoNingabstractDramatic progress has been witnessed in basic vision tasks involving low-level perception, such as object recognition, detection, and tracking. Unfortunately, there is still enormous performance gap between artificial vision systems and human intelligence in terms of higher-level vision problems, especially ones involving reasoning. Earlier attempts in equipping machines with high-level reasoning have hovered around Visual Question Answering (VQA), one typical task associating vision and language understanding. In this work, we propose a new dataset, built in the context of Raven's Progressive Matrices (RPM) and aimed at lifting machine intelligence by associating vision with structural, relational, and analogical reasoning in a hierarchical representation. Unlike previous works in measuring abstract reasoning using RPM, we establish a semantic link between vision and reasoning by providing structure representation. This addition enables a new type of abstract reasoning by jointly operating on the structure representation. Machine reasoning ability using modern computer vision is evaluated in this newly proposed dataset. Additionally, we also provide human performance as a reference. Finally, we show consistent improvement across all models by incorporating a simple neural module that combines visual understanding and structure reasoning. Chi Zhang 0017, Feng Gao 0013, Baoxiong Jia, Yixin Zhu 0001, Song-Chun Zhu |
CVPR | 4 |
| 2019 | Holistic++ Scene Understanding: Single-View 3D Holistic Scene Parsing and Human Pose Estimation With Human-Object Interaction and Physical CommonsenseabstractWe propose a new 3D holistic++scene understanding problem, which jointly tackles two tasks from a single-view image: (i) holistic scene parsing and reconstruction-3D estimations of object bounding boxes, camera pose, and room layout, and (ii) 3D human pose estimation. The intuition behind is to leverage the coupled nature of these two tasks to improve the granularity and performance of scene understanding. We propose to exploit two critical and essential connections between these two tasks: (i) human-object interaction (HOI) to model the fine-grained relations between agents and objects in the scene, and (ii) physical commonsense to model the physical plausibility of the reconstructed scene. The optimal configuration of the 3D scene, represented by a parse graph, is inferred using Markov chain Monte Carlo (MCMC), which efficiently traverses through the non-differentiable joint solution space. Experimental results demonstrate that the proposed algorithm significantly improves the performance of the two tasks on three datasets, showing an improved generalization ability. Yixin Chen 0003, Siyuan Huang 0001, Yixin Zhu 0001, Siyuan Qi, Song-Chun Zhu |
ICCV | 4 |
| 2019 | High-Fidelity Grasping in Virtual Reality using a Glove-based SystemabstractThis paper presents a design that jointly provides hand pose sensing, hand localization, and haptic feedback to facilitate real-time stable grasps in Virtual Reality (VR). The design is based on an easy-to-replicate glove-based system that can reliably perform (i) a high-fidelity hand pose sensing in real time through a network of 15 IMUs, and (ii) the hand localization using a Vive Tracker. The supported physics-based simulation in VR is capable of detecting collisions and contact points for virtual object manipulation, which drives the collision event to trigger the physical vibration motors on the glove to signal the user, providing a better realism inside virtual environments. A caging-based approach using collision geometry is integrated to determine whether a grasp is stable. In the experiment, we showcase successful grasps of virtual objects with large geometry variations. Comparing to the popular LeapMotion sensor, we demonstrate the proposed glove-based design yields a higher success rate in various tasks in VR. We hope such a glove-based system can simplify the data collection of human manipulations with VR. Hangxin Liu, Zhenliang Zhang 0002, Xu Xie 0001, Yixin Zhu 0001, Yue Liu 0005, Yongtian Wang, Song-Chun Zhu |
ICRA | 4 |
| 2019 | Self-Supervised Incremental Learning for Sound Source Localization in Complex Indoor EnvironmentabstractThis paper presents an incremental learning framework for mobile robots localizing the human sound source using a microphone array in a complex indoor environment consisting of multiple rooms. In contrast to conventional approaches that leverage direction-of-arrival (DOA) estimation, the framework allows a robot to accumulate training data and improve the performance of the prediction model over time using an incremental learning scheme. Specifically, we use implicit acoustic features obtained from an auto-encoder together with the geometry features from the map for training. A self-supervision process is developed such that the model ranks the priority of rooms to explore and assigns the ground truth label to the collected data, updating the learned model on-the-fly. The framework does not require pre-collected data and can be directly applied to real-world scenarios without any human supervisions or interventions. In experiments, we demonstrate that the prediction accuracy reaches 67% using about 20 training samples and eventually achieves 90% accuracy within 120 samples, surpassing prior classification-based methods with explicit GCC-PHAT features. Hangxin Liu, Zeyu Zhang 0001, Yixin Zhu 0001, Song-Chun Zhu |
ICRA | 3 |
| 2019 | Learning Virtual Grasp with Failed Demonstrations via Bayesian Inverse Reinforcement LearningabstractWe propose Bayesian Inverse Reinforcement Learning with Failure (BIRLF), which makes use of failed demonstrations that were often ignored or filtered in previous methods due to the difficulties to incorporate them in addition to the successful ones. Specifically, we leverage halfspaces derived from policy optimality conditions to incorporate failed demonstrations under Bayesian Inverse Reinforcement Learning (BIRL) framework. Under the continuous control setting, the reward function and policy are learned in an alternative manner, both of which are estimated by function approximators to guarantee the learning ability. Our approach is formulated as a model-free Inverse Reinforcement Learning (IRL) method that naturally accommodates more complex environments with continuous state and action spaces. In experiments, we demonstrate the proposed method in a virtual grasping task, achieving a significant performance boost compared to existing methods. Xu Xie 0001, ChangYang Li, Chi Zhang 0017, Yixin Zhu 0001, Song-Chun Zhu |
IROS | 4 |
| 2019 | PerspectiveNet: 3D Object Detection from a Single RGB Image via Perspective PointsabstractDetecting 3D objects from a single RGB image is intrinsically ambiguous, thus requiring appropriate prior knowledge and intermediate representations as constraints to reduce the uncertainties and improve the consistencies between the 2D image plane and the 3D world coordinate. To address this challenge, we propose to adopt perspective points as a new intermediate representation for 3D object detection, defined as the 2D projections of local Manhattan 3D keypoints to locate an object; these perspective points satisfy geometric constraints imposed by the perspective projection. We further devise PerspectiveNet, an end-to-end trainable model that simultaneously detects the 2D bounding box, 2D perspective points, and 3D object bounding box for each object from a single RGB image. PerspectiveNet yields three unique advantages: (i) 3D object bounding boxes are estimated based on perspective points, bridging the gap between 2D and 3D bounding boxes without the need of category-specific 3D shape priors. (ii) It predicts the perspective points by a template-based method, and a perspective loss is formulated to maintain the perspective constraints. (iii) It maintains the consistency between the 2D perspective points and 3D bounding boxes via a differentiable projective function. Experiments on SUN RGB-D dataset show that the proposed method significantly outperforms existing RGB-based approaches for 3D object detection. Siyuan Huang 0001, Yixin Chen 0003, Siyuan Qi, Yixin Zhu 0001, Song-Chun Zhu |
NeurIPS | 5 |
| 2019 | Learning Perceptual Inference by Contrastingabstract“Thinking in pictures,” [1] i.e., spatial-temporal reasoning, effortless and instantaneous for humans, is believed to be a significant ability to perform logical induction and a crucial factor in the intellectual history of technology development. Modern Artificial Intelligence (AI), fueled by massive datasets, deeper models, and mighty computation, has come to a stage where (super-)human-level performances are observed in certain specific tasks. However, current AI's ability in “thinking in pictures” is still far lacking behind. In this work, we study how to improve machines' reasoning ability on one challenging task of this kind: Raven's Progressive Matrices (RPM). Specifically, we borrow the very idea of “contrast effects” from the field of psychology, cognition, and education to design and train a permutation-invariant model. Inspired by cognitive studies, we equip our model with a simple inference module that is jointly trained with the perception backbone. Combining all the elements, we propose the Contrastive Perceptual Inference network (CoPINet) and empirically demonstrate that CoPINet sets the new state-of-the-art for permutation-invariant models on two major datasets. We conclude that spatial-temporal reasoning depends on envisaging the possibilities consistent with the relations between objects and can be solved from pixel-level inputs. Chi Zhang 0017, Baoxiong Jia, Feng Gao 0013, Yixin Zhu 0001, Hongjing Lu, Song-Chun Zhu |
NeurIPS | 4 |
| 2018 | Tracking Occluded Objects and Recovering Incomplete Trajectories by Reasoning About Containment Relations and Human ActionsabstractThis paper studies a challenging problem of tracking severely occluded objects in long video sequences. The proposed method reasons about the containment relations and human actions, thus infers and recovers occluded objects identities while contained or blocked by others. There are two conditions that lead to incomplete trajectories: i) Contained. The occlusion is caused by a containment relation formed between two objects, e.g., an unobserved laptop inside a backpack forms containment relation between the laptop and the backpack. ii) Blocked. The occlusion is caused by other objects blocking the view from certain locations, during which the containment relation does not change. By explicitly distinguishing these two causes of occlusions, the proposed algorithm formulates tracking problem as a network flow representation encoding containment relations and their changes. By assuming all the occlusions are not spontaneously happened but only triggered by human actions, an MAP inference is applied to jointly interpret the trajectory of an object by detection in space and human actions in time. To quantitatively evaluate our algorithm, we collect a new occluded object dataset captured by Kinect sensor, including a set of RGB-D videos and human skeletons with multiple actors, various objects, and different changes of containment relations. In the experiments, we show that the proposed method demonstrates better performance on tracking occluded objects compared with baseline methods. Wei Liang 0008, Yixin Zhu 0001, Song-Chun Zhu |
AAAI | 2 |
| 2018 | Human Causal Transfer: Challenges for Deep Reinforcement Learning
Mark Edmonds, James Kubricht, Colin Summers, Yixin Zhu 0001, Brandon Rothrock, Song-Chun Zhu, Hongjing Lu |
CogSci | 4 |
| 2018 | Human-Centric Indoor Scene Synthesis Using Stochastic GrammarabstractWe present a human-centric method to sample and synthesize 3D room layouts and 2D images thereof, to obtain large-scale 2D/3D image data with the perfect per-pixel ground truth. An attributed spatial And-Or graph (S-AOG) is proposed to represent indoor scenes. The S-AOG is a probabilistic grammar model, in which the terminal nodes are object entities including room, furniture, and supported objects. Human contexts as contextual relations are encoded by Markov Random Fields (MRF) on the terminal nodes. We learn the distributions from an indoor scene dataset and sample new layouts using Monte Carlo Markov Chain. Experiments demonstrate that the proposed method can robustly sample a large variety of realistic room layouts based on three criteria: (i) visual realism comparing to a state-of-the-art room arrangement method, (ii) accuracy of the affordance maps with respect to ground-truth, and (ii) the functionality and naturalness of synthesized rooms evaluated by human subjects. Siyuan Qi, Yixin Zhu 0001, Siyuan Huang 0001, Chenfanfu Jiang, Song-Chun Zhu |
CVPR | 2 |
| 2018 | Holistic 3D Scene Parsing and Reconstruction from a Single RGB Image
Siyuan Huang 0001, Siyuan Qi, Yixin Zhu 0001, Yinxue Xiao, Yuanlu Xu, Song-Chun Zhu |
ECCV (7) | 3 |
| 2018 | Interactive Robot Knowledge Patching Using Augmented RealityabstractWe present a novel Augmented Reality (AR) approach, through Microsoft HoloLens, to address the challenging problems of diagnosing, teaching, and patching interpretable knowledge of a robot. A Temporal And-Or graph (T-AOG) of opening bottles is learned from human demonstration and programmed to the robot. This representation yields a hierarchical structure that captures the compositional nature of the given task, which is highly interpretable for the users. By visualizing the knowledge structure represented by a T-AOG and the decision making process by parsing the T-AOG, the user can intuitively understand what the robot knows, supervise the robot's action planner, and monitor visually latent robot states (e.g., the force exerted during interactions). Given a new task, through such comprehensive visualizations of robot's inner functioning, users can quickly identify the reasons of failures, interactively teach the robot with a new action, and patch it to the current knowledge structure. In this way, the robot is capable of solving similar but new tasks only through minor modifications provided by the users interactively. This process demonstrates the interpretability of our knowledge representation and the effectiveness of the AR interface. Hangxin Liu, Yaofang Zhang, Wenwen Si, Xu Xie 0001, Yixin Zhu 0001, Song-Chun Zhu |
ICRA | 5 |
| 2018 | Unsupervised Learning of Hierarchical Models for Hand-Object InteractionsabstractContact forces of the hand are visually unobservable, but play a crucial role in understanding hand-object interactions. In this paper, we propose an unsupervised learning approach for manipulation event segmentation and manipulation event parsing. The proposed framework incorporates hand pose kinematics and contact forces using a low-cost easy-to-replicate tactile glove. We use a temporal grammar model to capture the hierarchical structure of events, integrating extracted force vectors from the raw sensory input of poses and forces. The temporal grammar is represented as a temporal And-Or graph (T-AOG), which can be induced in an unsupervised manner. We obtain the event labeling sequences by measuring the similarity between segments using the Dynamic Time Alignment Kernel (DTAK). Experimental results show that our method achieves high accuracy in manipulation event segmentation, recognition and parsing by utilizing both pose and force data. Xu Xie 0001, Hangxin Liu, Mark Edmonds, Feng Gao 0013, Siyuan Qi, Yixin Zhu 0001, Brandon Rothrock, Song-Chun Zhu |
ICRA | 6 |
| 2018 | Cooperative Holistic Scene Understanding: Unifying 3D Object, Layout, and Camera Pose EstimationabstractHolistic 3D indoor scene understanding refers to jointly recovering the i) object bounding boxes, ii) room layout, and iii) camera pose, all in 3D. The existing methods either are ineffective or only tackle the problem partially. In this paper, we propose an end-to-end model that simultaneously solves all three tasks in real-time given only a single RGB image. The essence of the proposed method is to improve the prediction by i) parametrizing the targets (e.g., 3D boxes) instead of directly estimating the targets, and ii) cooperative training across different modules in contrast to training these modules individually. Specifically, we parametrize the 3D object bounding boxes by the predictions from several modules, i.e., 3D camera pose and object attributes. The proposed method provides two major advantages: i) The parametrization helps maintain the consistency between the 2D image and the 3D world, thus largely reducing the prediction variances in 3D coordinates. ii) Constraints can be imposed on the parametrization to train different modules simultaneously. We call these constraints "cooperative losses" as they enable the joint training and inference. We employ three cooperative losses for 3D bounding boxes, 2D projections, and physical constraints to estimate a geometrically consistent and physically plausible 3D scene. Experiments on the SUN RGB-D dataset shows that the proposed method significantly outperforms prior approaches on 3D layout estimation, 3D object detection, 3D camera pose estimation, and holistic scene understanding. Siyuan Huang 0001, Siyuan Qi, Yinxue Xiao, Yixin Zhu 0001, Ying Nian Wu, Song-Chun Zhu |
NeurIPS | 4 |
| 2018 | Spatially Perturbed Collision Sounds Attenuate Perceived Causality in 3D Launching EventsabstractWhen a moving object collides with an object at rest, people immediately perceive a causal event: i.e., the first object has launched the second object forwards. However, when the second object's motion is delayed, or is accompanied by a collision sound, causal impressions attenuate and strengthen. Despite a rich literature on causal perception, researchers have exclusively utilized 2D visual displays to examine the launching effect. It remains unclear whether people are equally sensitive to the spatiotemporal properties of observed collisions in the real world. The present study first examined whether previous findings in causal perception with audiovisual inputs can be extended to immersive 3D virtual environments. We then investigated whether perceived causality is influenced by variations in the spatial position of an auditory collision indicator. We found that people are able to localize sound positions based on auditory inputs in VR environments, and spatial discrepancy between the estimated position of the collision sound and the visually observed impact location attenuates perceived causality. Duotun Wang, James Kubricht, Yixin Zhu 0001, Wei Liang 0008, Song-Chun Zhu, Chenfanfu Jiang, Hongjing Lu |
VR | 3 |
| 2018 | Configurable 3D Scene Synthesis and 2D Image Rendering with Per-pixel Ground Truth Using Stochastic Grammars
Chenfanfu Jiang, Siyuan Qi, Yixin Zhu 0001, Siyuan Huang 0001, Jenny Lin, Lap-Fai Yu, Demetri Terzopoulos, Song-Chun Zhu |
Int. J. Comput. Vis. | 3 |
| 2018 | A moving least squares material point method with displacement discontinuity and two-way rigid body couplingabstractIn this paper, we introduce the Moving Least Squares Material Point Method (MLS-MPM). MLS-MPM naturally leads to the formulation of Affine Particle-In-Cell (APIC) [Jiang et al. 2015] and Polynomial Particle-In-Cell [Fu et al. 2017] in a way that is consistent with a Galerkin-style weak form discretization of the governing equations. Additionally, it enables a new stress divergence discretization that effortlessly allows all MPM simulations to run two times faster than before. We also develop a Compatible Particle-In-Cell (CPIC) algorithm on top of MLS-MPM. Utilizing a colored distance field representation and a novel compatibility condition for particles and grid nodes, our framework enables the simulation of various new phenomena that are not previously supported by MPM, including material cutting, dynamic open boundaries, and two-way coupling with rigid bodies. MLS-MPM with CPIC is easy to implement and friendly to performance optimization. Yuanming Hu, Yu Fang 0010, Ziheng Ge, Ziyin Qu, Yixin Zhu 0001, Andre Pradhana Tampubolon, Chenfanfu Jiang |
ACM Trans. Graph. | 5 |
| 2017 | Consistent Probabilistic Simulation Underlying Human Judgment in Substance Dynamics
James Kubricht, Yixin Zhu 0001, Chenfanfu Jiang, Demetri Terzopoulos, Song-Chun Zhu, Hongjing Lu |
CogSci | 2 |
| 2017 | Visuomotor Adaptation and Sensory Recalibration in Reversed Hand Movement Task
Jenny Lin, Yixin Zhu 0001, James Kubricht, Song-Chun Zhu, Hongjing Lu |
CogSci | 2 |
| 2017 | Feeling the force: Integrating force and pose for fluent discovery through imitation learning to open medicine bottlesabstractLearning complex robot manipulation policies for real-world objects is challenging, often requiring significant tuning within controlled environments. In this paper, we learn a manipulation model to execute tasks with multiple stages and variable structure, which typically are not suitable for most robot manipulation approaches. The model is learned from human demonstration using a tactile glove that measures both hand pose and contact forces. The tactile glove enables observation of visually latent changes in the scene, specifically the forces imposed to unlock the child-safety mechanisms of medicine bottles. From these observations, we learn an action planner through both a top-down stochastic grammar model (And-Or graph) to represent the compositional nature of the task sequence and a bottom-up discriminative model from the observed poses and forces. These two terms are combined during planning to select the next optimal action. We present a method for transferring this human-specific knowledge onto a robot platform and demonstrate that the robot can perform successful manipulations of unseen objects with similar task structure. Mark Edmonds, Feng Gao 0013, Xu Xie 0001, Hangxin Liu, Siyuan Qi, Yixin Zhu 0001, Brandon Rothrock, Song-Chun Zhu |
IROS | 6 |
| 2017 | A glove-based system for studying hand-object manipulation via joint pose and force sensingabstractWe present a design of an easy-to-replicate glove-based system that can reliably perform simultaneous hand pose and force sensing in real time, for the purpose of collecting human hand data during fine manipulative actions. The design consists of a sensory glove that is capable of jointly collecting data of finger poses, hand poses, as well as forces on palm and each phalanx. Specifically, the sensory glove employs a network of 15 IMUs to measure the rotations between individual phalanxes. Hand pose is then reconstructed using forward kinematics. Contact forces on the palm and each phalanx are measured by 6 customized force sensors made from Velostat, a piezoresistive material whose force-voltage relation is investigated. We further develop an open-source software pipeline consisting of drivers and processing code and a system for visualizing hand actions that is compatible with the popular Raspberry Pi architecture. In our experiment, we conduct a series of evaluations that quantitatively characterize both individual sensors and the overall system, proving the effectiveness of the proposed design. Hangxin Liu, Xu Xie 0001, Matt Millar, Mark Edmonds, Feng Gao 0013, Yixin Zhu 0001, Veronica J. Santos, Brandon Rothrock, Song-Chun Zhu |
IROS | 6 |
| 2017 | The Martian: Examining Human Physical Judgments across Virtual Gravity FieldsabstractThis paper examines how humans adapt to novel physical situations with unknown gravitational acceleration in immersive virtual environments. We designed four virtual reality experiments with different tasks for participants to complete: strike a ball to hit a target, trigger a ball to hit a target, predict the landing location of a projectile, and estimate the flight duration of a projectile. The first two experiments compared human behavior in the virtual environment with real-world performance reported in the literature. The last two experiments aimed to test the human ability to adapt to novel gravity fields by measuring their performance in trajectory prediction and time estimation tasks. The experiment results show that: 1) based on brief observation of a projectile's initial trajectory, humans are accurate at predicting the landing location even under novel gravity fields, and 2) humans' time estimation in a familiar earth environment fluctuates around the ground truth flight duration, although the time estimation in unknown gravity fields indicates a bias toward earth's gravity. Tian Ye 0005, Siyuan Qi, James Kubricht, Yixin Zhu 0001, Hongjing Lu, Song-Chun Zhu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2016 | Proposal of the Second Workshop on Physical and Social Scene Understanding
Tao Gao 0004, Chenfanfu Jiang, Yixin Zhu 0001, Yibiao Zhao, Lap-Fai Yu |
CogSci | 3 |
| 2016 | Probabilistic Simulation Predicts Human Performance on Viscous Fluid-Pouring Problem
James Kubricht, Chenfanfu Jiang, Yixin Zhu 0001, Song-Chun Zhu, Demetri Terzopoulos, Hongjing Lu |
CogSci | 3 |
| 2016 | Inferring Forces and Learning Human Utilities from VideosabstractWe propose a notion of affordance that takes into account physical quantities generated when the human body interacts with real-world objects, and introduce a learning framework that incorporates the concept of human utilities, which in our opinion provides a deeper and finer-grained account not only of object affordance but also of people's interaction with objects. Rather than defining affordance in terms of the geometric compatibility between body poses and 3D objects, we devise algorithms that employ physicsbased simulation to infer the relevant forces/pressures acting on body parts. By observing the choices people make in videos (particularly in selecting a chair in which to sit) our system learns the comfort intervals of the forces exerted on body parts (while sitting). We account for people's preferences in terms of human utilities, which transcend comfort intervals to account also for meaningful tasks within scenes and spatiotemporal constraints in motion planning, such as for the purposes of robot task planning. Yixin Zhu 0001, Chenfanfu Jiang, Yibiao Zhao, Demetri Terzopoulos, Song-Chun Zhu |
CVPR | 1 |
| 2016 | What Is Where: Inferring Containment Relations from Videos
Wei Liang 0008, Yibiao Zhao, Yixin Zhu 0001, Song-Chun Zhu |
IJCAI | 3 |
| 2015 | Evaluating Human Cognition of Containing Relations with Physical Simulation
Wei Liang 0008, Yibiao Zhao, Yixin Zhu 0001, Song-Chun Zhu |
CogSci | 3 |
| 2015 | Understanding tools: Task-oriented object modeling, learning and recognitionabstractIn this paper, we present a new framework - task-oriented modeling, learning and recognition which aims at understanding the underlying functions, physics and causality in using objects as “tools”. Given a task, such as, cracking a nut or painting a wall, we represent each object, e.g. a hammer or brush, in a generative spatio-temporal representation consisting of four components: i) an affordance basis to be grasped by hand; ii) a functional basis to act on a target object (the nut), iii) the imagined actions with typical motion trajectories; and iv) the underlying physical concepts, e.g. force, pressure, etc. In a learning phase, our algorithm observes only one RGB-D video, in which a rational human picks up one object (i.e. tool) among a number of candidates to accomplish the task. From this example, our algorithm learns the essential physical concepts in the task (e.g. forces in cracking nuts). In an inference phase, our algorithm is given a new set of objects (daily objects or stones), and picks the best choice available together with the inferred affordance basis, functional basis, imagined human actions (sequence of poses), and the expected physical quantity that it will produce. From this new perspective, any objects can be viewed as a hammer or a shovel, and object recognition is not merely memorizing typical appearance examples for each category but reasoning the physical mechanisms in various tasks to achieve generalization. Yixin Zhu 0001, Yibiao Zhao, Song-Chun Zhu |
CVPR | 1 |