John Z. Zhang

dblp:66/3756 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 8 since 2021Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021Theory of computation · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Real-Time Whole-Body Control of Legged Robots with Model-Predictive Path Integral Control
abstract
This paper presents a system for enabling real-time synthesis of whole-body locomotion and manipulation policies for real-world legged robots. Motivated by recent advancements in robot simulation, we leverage the efficient parallelization capabilities of the MuJoCo simulator on a multi-core CPU to achieve fast sampling over the robot state and action trajectories. Our results show surprisingly effective real-world locomotion and manipulation capabilities with a very simple control strategy. We demonstrate our approach on several hardware and simulation experiments: robust locomotion over flat and uneven terrains, climbing over a box whose height is comparable to the robot, and pushing a box to a goal position. To our knowledge, this is the first successful deployment of whole-body sampling-based MPC on real-world legged robot hardware. Experiment videos and code can be found at: whole-body-mppi.github.io.
Juan Alvarez-Padilla, John Z. Zhang, Sofia Kwok, John M. Dolan, Zachary Manchester
ICRA2
2025 Wallbounce: Push Wall to Navigate with Contact-Implicit MPC
abstract
In this work, we introduce a framework that enables highly maneuverable locomotion using non-periodic contacts. This task is challenging for traditional optimization and planning methods to handle due to difficulties in specifying contact mode sequences in real-time. To address this, we use a bi-level contact-implicit planner and hybrid model predictive controller to draft and execute a motion plan. We investigate how this method allows us to plan arm contact events on the shmoobot, a smaller ballbot, which uses an inverse mouseball drive to achieve dynamic balancing with a low number of actuators. Through multiple experiments we show how the arms allow for acceleration, deceleration and dynamic obstacle avoidance that are not achievable with the mouseball drive alone. This demonstrates how a holistic approach to locomotion can increase the control authority of unique robot morpohologies without additional hardware by leveraging robot arms that are typically used only for manipulation. Project website: https://cmushmoobot.github.io/Wallbounce
Cunxi Dai, John Z. Zhang, Arun L. Bishop, Zachary Manchester, Ralph Hollis
ICRA3
2025 Robots with Attitude: Singularity-Free Quaternion-Based Model-Predictive Control for Agile Legged Robots
abstract
We present a model-predictive control (MPC) framework for legged robots that avoids the singularities associated with common three-parameter attitude representations like Euler angles during large-angle rotations. Our method parameterizes the robot's attitude with singularity-free unit quaternions and makes modifications to the iterative linear-quadratic regulator (iLQR) algorithm to deal with the resulting geometry. The derivation of our algorithm requires only elementary calculus and linear algebra, deliberately avoiding the abstraction and notation of Lie groups. We demonstrate the performance and computational efficiency of quaternion MPC in several experiments on quadruped and humanoid robots.
Zixin Zhang 0003, John Z. Zhang, Zachary Manchester
ICRA2
2024 ReLU-QP: A GPU-Accelerated Quadratic Programming Solver for Model-Predictive Control
abstract
We present ReLU-QP, a GPU-accelerated solver for quadratic programs (QPs) that is capable of solving high-dimensional control problems at real-time rates. ReLU-QP is derived by exactly reformulating the Alternating Direction Method of Multipliers (ADMM) algorithm for solving QPs as a deep, weight-tied neural network with rectified linear unit (ReLU) activations. This reformulation enables the deployment of ReLU-QP on GPUs using standard machine-learning toolboxes. We evaluate the performance of ReLU-QP across three model-predictive control (MPC) benchmarks: stabilizing random linear dynamical systems with control limits, balancing an Atlas humanoid robot on a single foot, and performing a whole-body pick-up motion on a quadruped equipped with a six-degree-of-freedom arm. These benchmarks indicate that ReLU-QP is competitive with state-of-the-art CPU-based solvers for small-to-medium-scale problems and offers order-of-magnitude speed improvements for larger-scale problems.
Arun L. Bishop, John Z. Zhang, Swaminathan Gurumurthy, Kevin Tracy, Zachary Manchester
ICRA2
2024 Enhancing Music Genre Classification Using Augmented Features Ensemble Learning Technique
Raad Shariat, John Z. Zhang
PRICAI (1)2
2024 Fast Contact-Implicit Model Predictive Control
abstract
In this article, we present a general approach for controlling robotic systems that make and break contact with their environments. Contact-implicit model predictive control (CI-MPC) generalizes linear MPC to contact-rich settings by utilizing a bilevel planning formulation with lower level contact dynamics formulated as time-varying linear complementarity problems (LCPs) computed using strategic Taylor approximations about a reference trajectory. These dynamics enable the upper level planning problem to reason about contact timing and forces, and generate entirely new contact-mode sequences online. To achieve reliable and fast numerical convergence, we devise a structure-exploiting interior-point solver for these LCP contact dynamics and a custom trajectory optimizer for the tracking problem. We demonstrate real-time solution rates for CI-MPC and the ability to generate and track nonperiodic behaviors in hardware experiments on a quadrupedal robot. We also show that the controller is robust to model mismatch and can respond to disturbances by discovering and exploiting new contact modes across a variety of robotic systems in simulation, including a pushbot, planar hopper, planar quadruped, and planar biped.
Simon Le Cleac'h, Taylor A. Howell, Chi-Yen Lee, John Z. Zhang, Arun L. Bishop, Mac Schwager, Zachary Manchester
IEEE Trans. Robotics5
2024 Heterogeneous Policy Networks for Composite Robot Team Communication and Coordination
abstract
High-performing human–human teams learn intelligent and efficient communication and coordination strategies to maximize their joint utility. These teams implicitly understand the different roles of heterogeneous team members and adapt their communication protocols accordingly. Multiagent reinforcement learning (MARL) has attempted to develop computational methods for synthesizing such joint coordination–communication strategies, but emulating heterogeneous communication patterns across agents with different state, action, and observation spaces has remained a challenge. Without properly modeling agent heterogeneity, as in prior MARL work that leverages homogeneous graph networks, communication becomes less helpful and can even deteriorate the team's performance. In the past, we proposed heterogeneous policy networks (HetNet) to learn efficient and diverse communication models for coordinating cooperative heterogeneous teams. In this extended work, we extend HetNet to support scaling heterogeneous robot teams. Building on heterogeneous graph-attention networks, we show that HetNet not only facilitates learning heterogeneous collaborative policies, but also enables end-to-end training for learning highly efficient binarized messaging. Our empirical evaluation shows that HetNet sets a new state-of-the-art in learning coordination and communication strategies for heterogeneous multiagent teams by achieving an 5.84% to 707.65% performance improvement over the next-best baseline across multiple domains while simultaneously achieving a 200× reduction in the required communication bandwidth.
Esmaeil Seraj, Rohan R. Paleja, Luis Pimentel, Kin Man Lee, Zheyuan Wang, Matthew Sklar, John Z. Zhang, Zahi M. Kakish, Matthew C. Gombolay
IEEE Trans. Robotics8
2023 PPR: Physically Plausible Reconstruction from Monocular Videos
abstract
Given monocular videos, we build 3D models of articulated objects and environments whose 3D configurations satisfy dynamics and contact constraints. At its core, our method leverages differentiable physics simulation to aid visual reconstructions. We couple differentiable physics simulation with differentiable rendering via coordinate descent, which enables end-to-end optimization of, not only 3D reconstructions, but also physical system parameters from videos. We demonstrate the effectiveness of physics-informed reconstruction on monocular videos of quadruped animals and humans. It reduces reconstruction artifacts (e.g., scale ambiguity, unbalanced poses, and foot swapping) that are challenging to address by visual cues alone, and produces better foot contact estimation.
Gengshan Yang, John Z. Zhang, Zachary Manchester, Deva Ramanan
ICCV3
2022 Fast Aquatic Swimmer Optimization with Differentiable Projective Dynamics and Neural Network Hydrodynamic Models
abstract
Aquatic locomotion is a classic fluid-structure interaction (FSI) problem of interest to biologists and engineers. Solving the fully coupled FSI equations for incompressible Navier-Stokes and finite elasticity is computationally expensive. Optimizing robotic swimmer design within such a system generally involves cumbersome, gradient-free procedures on top of the already costly simulation. To address this challenge we present a novel, fully differentiable hybrid approach to FSI that combines a 2D direct numerical simulation for the deformable solid structure of the swimmer and a physics-constrained neural network surrogate to capture hydrodynamic effects of the fluid. For the deformable solid simulation of the swimmer’s body, we use state-of-the-art techniques from the field of computer graphics to speed up the finite-element method (FEM). For the fluid simulation, we use a U-Net architecture trained with a physics-based loss function to predict the flow field at each time step. The pressure and velocity field outputs from the neural network are sampled around the boundary of our swimmer using an immersed boundary method (IBM) to compute its swimming motion accurately and efficiently. We demonstrate the computational efficiency and differentiability of our hybrid simulator on a 2D carangiform swimmer. Due to differentiability, the simulator can be used for computational design of controls for soft bodies immersed in fluids via direct gradient-based optimization.
Elvis Nava, John Z. Zhang, Mike Yan Michelis, Tao Du 0001, Pingchuan Ma 0002, Benjamin F. Grewe, Wojciech Matusik, Robert K. Katzschmann
ICML2
2022 Sim2Real for Soft Robotic Fish via Differentiable Simulation
abstract
Accurate simulation of soft mechanisms under dynamic actuation is critical for the design of soft robots. We address this gap with our differentiable simulation tool by learning the material parameters of our soft robotic fish. On the example of a soft robotic fish, we demonstrate an experimentally-verified, fast optimization pipeline for learning the material parameters from quasi-static data via differentiable simulation and apply it to the prediction of dynamic performance. Our method identifies physically plausible Young's moduli for various soft silicone elastomers and stiff acetal copolymers used in creation of our three different robotic fish tail designs. We show that our method is compatible with varying internal geometry of the actuators, such as the number of hollow cavities. Our framework allows high fidelity prediction of dynamic behavior for composite bi-morph bending structures in real hardware to millimeter-accuracy and within 3% error normalized to actuator length. We provide a differentiable and robust estimate of the thrust force using a neural network thrust predictor; this estimate allows for accurate modeling of our experimental setup measuring bollard pull. This work presents a prototypical hardware and simulation problem solved using our differentiable framework; the framework can be applied to higher dimensional parameter inference, learning control policies, and computational design due to its differentiable character.
John Z. Zhang, Pingchuan Ma 0002, Elvis Nava, Tao Du 0001, Philip Arm, Wojciech Matusik, Robert K. Katzschmann
IROS1
2021 DiffAqua: a differentiable computational design pipeline for soft underwater swimmers with shape interpolation
abstract
The computational design of soft underwater swimmers is challenging because of the high degrees of freedom in soft-body modeling. In this paper, we present a differentiable pipeline for co-designing a soft swimmer's geometry and controller. Our pipeline unlocks gradient-based algorithms for discovering novel swimmer designs more efficiently than traditional gradient-free solutions. We propose Wasserstein barycenters as a basis for the geometric design of soft underwater swimmers since it is differentiable and can naturally interpolate between bio-inspired base shapes via optimal transport. By combining this design space with differentiable simulation and control, we can efficiently optimize a soft underwater swimmer's performance with fewer simulations than baseline methods. We demonstrate the efficacy of our method on various design problems such as fast, stable, and energy-efficient swimming and demonstrate applicability to multi-objective design.
Pingchuan Ma 0002, Tao Du 0001, John Z. Zhang, Kui Wu 0003, Andrew Spielberg, Robert K. Katzschmann, Wojciech Matusik
ACM Trans. Graph.3
2015 An Empirical Study on Structured Dichotomies in Music Genre Classification
abstract
Ensemble learning approaches have been gaining popularity in the non-trivial task of multi-class classification. Some of these, including 1-against-1, 1-against-all, and dichotomy-based methods, are based on decomposing the class space of a multi-class task into a set of binary-class ones. In this work, we investigate whether they could help improve genre classification in music. In particular, we explore various dichotomy structures of binary classifiers in music data. In addition to the existing ones, we propose several strategies to build new binary tree structures. We base our approach on the observation that people find it easy to distinguish between certain classes and difficult between others. In our investigation, we use several base classifiers that are common in the literature and conduct series of empirical experiments on two benchmarking music datasets. We report the initial results of our investigation in this paper.
Tom Arjannikov, John Z. Zhang
ICMLA2
2012 Improving Neural Networks Classification through Chaining
Khobaib Zaamout, John Z. Zhang
ICANN (2)2
2011 Enhancing multi-label music genre classification through ensemble techniques
abstract
In the field of Music Information Retrieval (MIR), multi-label genre classification is the problem of assigning one or more genre labels to a music piece. In this work, we propose a set of ensemble techniques, which are specific to the task of multi-label genre classification. Our goal is to enhance classification performance by combining multiple classifiers. In addition, we also investigate some existing ensemble techniques from machine learning. The effectiveness of these techniques is demonstrated through a set of empirical experiments and various related issues are discussed. To the best of our knowledge, there has been limited work on applying ensemble techniques to multi-label genre classification in the literature and we consider the results in this work as our initial efforts toward this end. The significance of our work has two folds: (1) proposing a set of ensemble techniques specific to music genre classification and (2) shedding light on further research along this direction.
Chris Sanden, John Z. Zhang
SIGIR2
2010 Finding the Minimum-Distance Schedule for a Boundary Searcher with a Flashlight
Tsunehiko Kameda, Ichiro Suzuki, John Z. Zhang
LATIN3
2009 Surveillance of a polygonal area by a mobile searcher from the boundary: Searchability testing
abstract
We study the surveillance of a polygonal area by a robot, which is equipped with a flashlight and moves along the polygon boundary. Its aim is to illuminate any intruder who can move faster than the moving flashlight beam, trying to avoid detection. We propose an O(n)-time algorithm for testing if it is possible for such a robot to always detect any intruder in a given polygon, where n is the number of vertices of the given polygon. This improves upon the best previous time complexity of O(n log n).
Binay K. Bhattacharya, Tsunehiko Kameda, John Z. Zhang
ICRA3
2009 The Two-Guard Polygon Walk Problem
John Z. Zhang
TAMC1
2008 A Linear-Time Algorithm for Finding All Door Locations That Make a Room Searchable
John Z. Zhang, Tsunehiko Kameda
TAMC1
2006 Where to Build a Door
abstract
A room is a simple polygon with a prespecified point, called the door, on its boundary. Search starts at the door, and must detect all intruders that may be in the room, while making sure that no intruder escapes through the door during the search. Depending on where the door is placed, the intruders may be able to avoid detection. We present an efficient algorithm that can determine all the intervals on the boundary where the door should be placed in order for the polygon to be searchable by two guards on the boundary who keep mutual visibility, or a single searcher with a flashlight. Our algorithm works in O(n log n) time, where n is the number of vertices of the given polygon
John Z. Zhang, Tsunehiko Kameda
IROS1