VLDB 2026 Research / reviewers in the wild / expert
Matthew Chang
dblp:56/2174
· DBLP profile ↗
13ranked-venue papers
9as first author
7since 2021 · last 2025
0000-0002-9486-1228ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorSystems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Reinforcement learning · 28% 3D vision · 26% Robot manipulation · 14% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-robot interaction · 100% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › pose estimation
3d hand pose estimation |
1.0 | 2 | 2024 | 3D Hand Pose Estimation in Everyday Egocentric Images · ECCV (78) 2024 3D Reconstruction of Objects in Hands Without Real World 3D Supervision · ECCV (78) 2024 |
Robotics › Robot manipulation › human-robot interaction
human-robot collaboration |
0.9 | 1 | 2025 | PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks · ICLR 2025 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
task planning |
0.9 | 1 | 2025 | PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks · ICLR 2025 |
Human-robot interaction
collaborative task |
0.9 | 1 | 2025 | PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent Tasks · ICLR 2025 |
Computer vision › 3D vision
3d reconstruction |
0.8 | 1 | 2024 | 3D Reconstruction of Objects in Hands Without Real World 3D Supervision · ECCV (78) 2024 |
Robotics › Robot navigation and mapping
embodied navigation |
0.8 | 1 | 2024 | GOAT-Bench: A Benchmark for Multi-Modal Lifelong Navigation · CVPR 2024 |
Computer vision › 3D vision › 3d reconstruction › object reconstruction
hand-object reconstruction |
0.8 | 1 | 2024 | 3D Reconstruction of Objects in Hands Without Real World 3D Supervision · ECCV (78) 2024 |
Machine learning › Reinforcement learning
imitation learning |
0.7 | 1 | 2023 | One-shot Visual Imitation via Attributed Waypoints and Demonstration Augmentation · ICRA 2023 |
Robotics › Motion planning and robot control
robot learning |
0.7 | 1 | 2023 | One-shot Visual Imitation via Attributed Waypoints and Demonstration Augmentation · ICRA 2023 |
Computer vision › Video understanding and tracking › video reconstruction
video inpainting |
0.7 | 1 | 2023 | Look Ma, No Hands! Agent-Environment Factorization of Egocentric Videos · NeurIPS 2023 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.6 | 1 | 2022 | Learning Value Functions from Undirected State-only Experience · ICLR 2022 |
Machine learning › Reinforcement learning
value function estimation |
0.6 | 1 | 2022 | Learning Value Functions from Undirected State-only Experience · ICLR 2022 |
Machine learning › Reinforcement learning › imitation learning
learning from observation |
0.4 | 1 | 2020 | Semantic Visual Navigation by Watching YouTube Videos · NeurIPS 2020 |
Robotics › Robot navigation and mapping › visual navigation
semantic visual navigation |
0.4 | 1 | 2020 | Semantic Visual Navigation by Watching YouTube Videos · NeurIPS 2020 |
Computer vision › 3D vision
egocentric vision |
0.2 | 1 | 2024 | 3D Hand Pose Estimation in Everyday Egocentric Images · ECCV (78) 2024 |
Methods — techniques the papers use, named apart from their topics
simulation-in-the-loop · 1.7large language model · 1.7fine-tuning · 1.7self-supervision · 0.8reinforcement learning · 0.8pose estimation · 0.8modular navigation · 0.8egocentric images · 0.8differentiable rendering · 0.8data augmentation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PARTNR: A Benchmark for Planning and Reasoning in Embodied Multi-agent TasksabstractWe present a benchmark for Planning And Reasoning Tasks in humaN-Robot collaboration (PARTNR) designed to study human-robot coordination in household activities. PARTNR tasks exhibit characteristics of everyday tasks, such as spatial, temporal, and heterogeneous agent capability constraints. We employ a semi-automated task generation pipeline using Large Language Models (LLMs), incorporating simulation-in-the-loop for the grounding and verification. PARTNR stands as the largest benchmark of its kind, comprising 100,000 natural language tasks, spanning 60 houses and 5,819 unique objects. We analyze state-of-the-art LLMs on PARTNR tasks, across the axes of planning, perception and skill execution. The analysis reveals significant limitations in SoTA models, such as poor coordination and failures in task tracking and recovery from errors. When LLMs are paired with 'real' humans, they require 1.5x as many steps as two humans collaborating and 1.1x more steps than a single human, underscoring the potential for improvement in these models. We further show that fine-tuning smaller LLMs with planning data can achieve performance on par with models 9 times larger, while being 8.6x faster at inference. Overall, PARTNR highlights significant challenges facing collaborative embodied agents and aims to drive research in this direction. Matthew Chang, Gunjan Chhablani, Alexander Clegg, Mikael Dallaire Cote, Ruta Desai, Michal Hlavac, Vladimir Karashchuk, Jacob Krantz, Roozbeh Mottaghi, Priyam Parashar, Siddharth Patki, Ishita Prasad, Xavier Puig, Akshara Rai, Ram Ramrakhya, Daniel Tran, Joanne Truong, John M. Turner, Eric Undersander, Tsung-Yen Yang |
ICLR | 1 |
| 2024 | GOAT-Bench: A Benchmark for Multi-Modal Lifelong NavigationabstractThe Embodied AI community has made significant strides in visual navigation tasks, exploring targets from 3D coordinates, objects, language descriptions, and images. However, these navigation models often handle only a single input modality as the target. With the progress achieved so for, it is time to move towards universal navigation models capable of handling various goal types, enabling more effective user interaction with robots. To facilitate this goal, we propose GOAT-Bench, a benchmark for the universal navigation task referred to as GO to AnyThing (GOAT). In this task, the agent is directed to navigate to a sequence of targets specified by the category name, language description, or image in an open-vocabulary fashion. We benchmark monolithic RL and modular methods on the GOAT task, analyzing their performance across modalities, the role of explicit and implicit scene memories, their robustness to noise in goal specifications, and the impact of memory in lifelong scenarios. Mukul Khanna, Ram Ramrakhya, Gunjan Chhablani, Sriram Yenamandra, Théophile Gervet, Matthew Chang, Zsolt Kira, Devendra Singh Chaplot, Dhruv Batra, Roozbeh Mottaghi |
CVPR | 6 |
| 2024 | 3D Reconstruction of Objects in Hands Without Real World 3D Supervision
Matthew Chang, Matthew Jin, Ruisen Tu, Saurabh Gupta 0001 |
ECCV (78) | 2 |
| 2024 | 3D Hand Pose Estimation in Everyday Egocentric Images
Ruisen Tu, Matthew Chang, Saurabh Gupta 0001 |
ECCV (78) | 3 |
| 2023 | One-shot Visual Imitation via Attributed Waypoints and Demonstration AugmentationabstractIn this paper, we analyze the behavior of existing techniques and design new solutions for the problem of one-shot visual imitation. In this setting, an agent must solve a novel instance of a novel task given just a single visual demonstration. Our analysis reveals that current methods fall short because of three errors: the DAgger problem arising from purely offline training, last centimeter errors in interacting with objects, and mis-fitting to the task context rather than to the actual task. This motivates the design of our modular approach where we a) separate out task inference (what to do) from task execution (how to do it), and b) develop data augmentation and generation techniques to mitigate mis-fitting. The former allows us to leverage hand-crafted motor primitives for task execution which side-steps the DAgger problem and last centimeter errors, while the latter gets the model to focus on the task rather than the task context. Our model gets 100% and 48% success rates on two recent benchmarks, improving upon the current state-of-the-art by absolute 90% and 20% respectively. Matthew Chang, Saurabh Gupta 0001 |
ICRA | 1 |
| 2023 | Look Ma, No Hands! Agent-Environment Factorization of Egocentric VideosabstractThe analysis and use of egocentric videos for robotics tasks is made challenging by occlusion and the visual mismatch between the human hand and a robot end-effector. Past work views the human hand as a nuisance and removes it from the scene. However, the hand also provides a valuable signal for learning. In this work, we propose to extract a factored representation of the scene that separates the agent (human hand) and the environment. This alleviates both occlusion and mismatch while preserving the signal, thereby easing the design of models for downstream robotics tasks. At the heart of this factorization is our proposed Video Inpainting via Diffusion Model (VIDM) that leverages both a prior on real-world images (through a large-scale pre-trained diffusion model) and the appearance of the object in earlier frames of the video (through attention). Our experiments demonstrate the effectiveness of VIDM at improving the in-painting quality in egocentric videos and the power of our factored representation for numerous tasks: object detection, 3D reconstruction of manipulated objects, and learning of reward functions, policies, and affordances from videos. Matthew Chang, Saurabh Gupta 0001 |
NeurIPS | 1 |
| 2022 | Learning Value Functions from Undirected State-only Experience
Matthew Chang, Saurabh Gupta 0001 |
ICLR | 1 |
| 2020 | Semantic Visual Navigation by Watching YouTube VideosabstractSemantic cues and statistical regularities in real-world environment layouts can improve efficiency for navigation in novel environments. This paper learns and leverages such semantic cues for navigating to objects of interest in novel environments, by simply watching YouTube videos. This is challenging because YouTube videos don't come with labels for actions or goals, and may not even showcase optimal behavior. Our method tackles these challenges through the use of Q-learning on pseudo-labeled transition quadruples (image, action, next image, reward). We show that such off-policy Q-learning from passive data is able to learn meaningful semantic cues for navigation. These cues, when used in a hierarchical navigation policy, lead to improved efficiency at the ObjectGoal task in visually realistic simulations. We observe a relative improvement of 15-83% over end-to-end RL, behavior cloning, and classical methods, while using minimal direct interaction. Matthew Chang, Saurabh Gupta 0001 |
NeurIPS | 1 |
| 2009 | Using phrases as features in email classification
Matthew Chang, Chung Keung Poon |
J. Syst. Softw. | 1 |
| 2008 | Efficient phrase querying with common phrase index
Matthew Chang, Chung Keung Poon |
Inf. Process. Manag. | 1 |
| 2006 | Efficient Phrase Querying with Common Phrase Index
Matthew Chang, Chung Keung Poon |
ECIR | 1 |
| 2005 | Catching the Picospams
Matthew Chang, Chung Keung Poon |
ISMIS | 1 |
| 2003 | An Email Classifier Based on Resemblance
Chung Keung Poon, Matthew Chang |
ISMIS | 2 |