Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Matthew Chang

dblp:56/2174 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
7since 2021 · last 2025
0000-0002-9486-1228ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Reinforcement learning · 28% 3D vision · 26% Robot manipulation · 14%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › pose estimation
3d hand pose estimation
1.022024
3D Hand Pose Estimation in Everyday Egocentric Images · ECCV (78) 2024
3D Reconstruction of Objects in Hands Without Real World 3D Supervision · ECCV (78) 2024
Robotics › Robot manipulation › human-robot interaction
human-robot collaboration
0.912025
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks · ICLR 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
task planning
0.912025
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks · ICLR 2025
Human-robot interaction
collaborative task
0.912025
PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks · ICLR 2025
Computer vision › 3D vision
3d reconstruction
0.812024
3D Reconstruction of Objects in Hands Without Real World 3D Supervision · ECCV (78) 2024
Robotics › Robot navigation and mapping
embodied navigation
0.812024
GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation · CVPR 2024
Computer vision › 3D vision › 3d reconstruction › object reconstruction
hand-object reconstruction
0.812024
3D Reconstruction of Objects in Hands Without Real World 3D Supervision · ECCV (78) 2024
Machine learning › Reinforcement learning
imitation learning
0.712023
One-shot Visual Imitation via Attributed Waypoints and Demonstration Augmentation · ICRA 2023
Robotics › Motion planning and robot control
robot learning
0.712023
One-shot Visual Imitation via Attributed Waypoints and Demonstration Augmentation · ICRA 2023
Computer vision › Video understanding and tracking › video reconstruction
video inpainting
0.712023
Look Ma, No Hands! Agent-Environment Factorization of Egocentric Videos · NeurIPS 2023
Machine learning › Reinforcement learning
offline reinforcement learning
0.612022
Learning Value Functions from Undirected State-only Experience · ICLR 2022
Machine learning › Reinforcement learning
value function estimation
0.612022
Learning Value Functions from Undirected State-only Experience · ICLR 2022
Machine learning › Reinforcement learning › imitation learning
learning from observation
0.412020
Semantic Visual Navigation by Watching YouTube Videos · NeurIPS 2020
Robotics › Robot navigation and mapping › visual navigation
semantic visual navigation
0.412020
Semantic Visual Navigation by Watching YouTube Videos · NeurIPS 2020
Computer vision › 3D vision
egocentric vision
0.212024
3D Hand Pose Estimation in Everyday Egocentric Images · ECCV (78) 2024

Methods — techniques the papers use, named apart from their topics

simulation-in-the-loop · 1.7large language model · 1.7fine-tuning · 1.7self-supervision · 0.8reinforcement learning · 0.8pose estimation · 0.8modular navigation · 0.8egocentric images · 0.8differentiable rendering · 0.8data augmentation · 0.7
YearPublicationVenuePosition
2025 PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks
abstract
We present a benchmark for Planning And Reasoning Tasks in humaN-Robot collaboration (PARTNR) designed to study human-robot coordination in household activities. PARTNR tasks exhibit characteristics of everyday tasks, such as spatial, temporal, and heterogeneous agent capability constraints. We employ a semi-automated task generation pipeline using Large Language Models (LLMs), incorporating simulation-in-the-loop for the grounding and verification. PARTNR stands as the largest benchmark of its kind, comprising 100,000 natural language tasks, spanning 60 houses and 5,819 unique objects. We analyze state-of-the-art LLMs on PARTNR tasks, across the axes of planning, perception and skill execution. The analysis reveals significant limitations in SoTA models, such as poor coordination and failures in task tracking and recovery from errors. When LLMs are paired with 'real' humans, they require 1.5x as many steps as two humans collaborating and 1.1x more steps than a single human, underscoring the potential for improvement in these models. We further show that fine-tuning smaller LLMs with planning data can achieve performance on par with models 9 times larger, while being 8.6x faster at inference. Overall, PARTNR highlights significant challenges facing collaborative embodied agents and aims to drive research in this direction.
Matthew Chang, Gunjan Chhablani, Alexander Clegg, Mikael Dallaire Cote, Ruta Desai, Michal Hlavac, Vladimir Karashchuk, Jacob Krantz, Roozbeh Mottaghi, Priyam Parashar, Siddharth Patki, Ishita Prasad, Xavier Puig, Akshara Rai, Ram Ramrakhya, Daniel Tran, Joanne Truong, John M. Turner, Eric Undersander, Tsung-Yen Yang
ICLR1
2024 GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation
abstract
The Embodied AI community has made significant strides in visual navigation tasks, exploring targets from 3D coordinates, objects, language descriptions, and images. However, these navigation models often handle only a single input modality as the target. With the progress achieved so for, it is time to move towards universal navigation models capable of handling various goal types, enabling more effective user interaction with robots. To facilitate this goal, we propose GOAT-Bench, a benchmark for the universal navigation task referred to as GO to AnyThing (GOAT). In this task, the agent is directed to navigate to a sequence of targets specified by the category name, language description, or image in an open-vocabulary fashion. We benchmark monolithic RL and modular methods on the GOAT task, analyzing their performance across modalities, the role of explicit and implicit scene memories, their robustness to noise in goal specifications, and the impact of memory in lifelong scenarios.
Mukul Khanna, Ram Ramrakhya, Gunjan Chhablani, Sriram Yenamandra, Théophile Gervet, Matthew Chang, Zsolt Kira, Devendra Singh Chaplot, Dhruv Batra, Roozbeh Mottaghi
CVPR6
2024 3D Reconstruction of Objects in Hands Without Real World 3D Supervision
Matthew Chang, Matthew Jin, Ruisen Tu, Saurabh Gupta 0001
ECCV (78)2
2024 3D Hand Pose Estimation in Everyday Egocentric Images
Ruisen Tu, Matthew Chang, Saurabh Gupta 0001
ECCV (78)3
2023 One-shot Visual Imitation via Attributed Waypoints and Demonstration Augmentation
abstract
In this paper, we analyze the behavior of existing techniques and design new solutions for the problem of one-shot visual imitation. In this setting, an agent must solve a novel instance of a novel task given just a single visual demonstration. Our analysis reveals that current methods fall short because of three errors: the DAgger problem arising from purely offline training, last centimeter errors in interacting with objects, and mis-fitting to the task context rather than to the actual task. This motivates the design of our modular approach where we a) separate out task inference (what to do) from task execution (how to do it), and b) develop data augmentation and generation techniques to mitigate mis-fitting. The former allows us to leverage hand-crafted motor primitives for task execution which side-steps the DAgger problem and last centimeter errors, while the latter gets the model to focus on the task rather than the task context. Our model gets 100% and 48% success rates on two recent benchmarks, improving upon the current state-of-the-art by absolute 90% and 20% respectively.
Matthew Chang, Saurabh Gupta 0001
ICRA1
2023 Look Ma, No Hands! Agent-Environment Factorization of Egocentric Videos
abstract
The analysis and use of egocentric videos for robotics tasks is made challenging by occlusion and the visual mismatch between the human hand and a robot end-effector. Past work views the human hand as a nuisance and removes it from the scene. However, the hand also provides a valuable signal for learning. In this work, we propose to extract a factored representation of the scene that separates the agent (human hand) and the environment. This alleviates both occlusion and mismatch while preserving the signal, thereby easing the design of models for downstream robotics tasks. At the heart of this factorization is our proposed Video Inpainting via Diffusion Model (VIDM) that leverages both a prior on real-world images (through a large-scale pre-trained diffusion model) and the appearance of the object in earlier frames of the video (through attention). Our experiments demonstrate the effectiveness of VIDM at improving the in-painting quality in egocentric videos and the power of our factored representation for numerous tasks: object detection, 3D reconstruction of manipulated objects, and learning of reward functions, policies, and affordances from videos.
Matthew Chang, Saurabh Gupta 0001
NeurIPS1
2022 Learning Value Functions from Undirected State-only Experience
Matthew Chang, Saurabh Gupta 0001
ICLR1
2020 Semantic Visual Navigation by Watching YouTube Videos
abstract
Semantic cues and statistical regularities in real-world environment layouts can improve efficiency for navigation in novel environments. This paper learns and leverages such semantic cues for navigating to objects of interest in novel environments, by simply watching YouTube videos. This is challenging because YouTube videos don't come with labels for actions or goals, and may not even showcase optimal behavior. Our method tackles these challenges through the use of Q-learning on pseudo-labeled transition quadruples (image, action, next image, reward). We show that such off-policy Q-learning from passive data is able to learn meaningful semantic cues for navigation. These cues, when used in a hierarchical navigation policy, lead to improved efficiency at the ObjectGoal task in visually realistic simulations. We observe a relative improvement of 15-83% over end-to-end RL, behavior cloning, and classical methods, while using minimal direct interaction.
Matthew Chang, Saurabh Gupta 0001
NeurIPS1
2009 Using phrases as features in email classification
Matthew Chang, Chung Keung Poon
J. Syst. Softw.1
2008 Efficient phrase querying with common phrase index
Matthew Chang, Chung Keung Poon
Inf. Process. Manag.1
2006 Efficient Phrase Querying with Common Phrase Index
Matthew Chang, Chung Keung Poon
ECIR1
2005 Catching the Picospams
Matthew Chang, Chung Keung Poon
ISMIS1
2003 An Email Classifier Based on Resemblance
Chung Keung Poon, Matthew Chang
ISMIS2