VLDB 2026 Research / reviewers in the wild / expert
Mathew Monfort
dblp:160/9991
· DBLP profile ↗
14ranked-venue papers
5as first author
4since 2021 · last 2024
0000-0001-6373-5520ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
11 papers |
Video understanding and tracking · 29% Reinforcement learning · 14% Learning theory · 9% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% | |
| Human-computer interaction and pervasive computing
2 papers |
Human-robot interaction · 100% |
Topics — the 23 heaviest of 28, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
action recognition |
1.4 | 3 | 2022 | Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2022 Moments in Time Dataset: One Million Videos for Event Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020 Reasoning About Human-Object Interactions Through Dual Attention Networks · ICCV 2019 |
Machine learning › Reinforcement learning › imitation learning › inverse reinforcement learning
inverse optimal control |
0.9 | 4 | 2017 | Goal-predictive robotic teleoperation from noisy sensors · ICRA 2017 Softstar: Heuristic-Guided Probabilistic Inference · NIPS 2015 Graph-Based Inverse Optimal Control for Robot Manipulation · IJCAI 2015 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.8 | 1 | 2024 | Precise Model Benchmarking with Only a Few Observations · EMNLP 2024 |
Multimedia analysis and retrieval › audio-visual learning
audio-visual representation learning |
0.5 | 1 | 2021 | Spoken Moments: Learning Joint Audio-Visual Representations From Video Descriptions · CVPR 2021 |
Natural language and speech › Information extraction and text analysis › event analysis
event understanding |
0.4 | 1 | 2020 | Moments in Time Dataset: One Million Videos for Event Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020 |
Computer vision › Video understanding and tracking › action recognition
human-object interaction recognition |
0.4 | 1 | 2019 | Reasoning About Human-Object Interactions Through Dual Attention Networks · ICCV 2019 |
Knowledge, reasoning and agents › Multi-agent systems › agent interaction › multi-agent interaction
multi-agent interaction modeling |
0.4 | 1 | 2019 | Multi-Agent Tensor Fusion for Contextual Trajectory Prediction · CVPR 2019 |
Robotics › Autonomous driving
trajectory prediction |
0.4 | 1 | 2019 | Multi-Agent Tensor Fusion for Contextual Trajectory Prediction · CVPR 2019 |
Human-robot interaction
teleoperation |
0.3 | 1 | 2017 | Goal-predictive robotic teleoperation from noisy sensors · ICRA 2017 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
empirical bayes |
0.2 | 1 | 2024 | Precise Model Benchmarking with Only a Few Observations · EMNLP 2024 |
Mathematical optimization
stratified sampling |
0.2 | 1 | 2024 | A Framework for Efficient Model Evaluation Through Stratification, Sampling, and Estimation · ECCV (88) 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
heuristic search |
0.2 | 1 | 2015 | Softstar: Heuristic-Guided Probabilistic Inference · NIPS 2015 |
Machine learning › Reinforcement learning › imitation learning
inverse reinforcement learning |
0.2 | 1 | 2015 | Softstar: Heuristic-Guided Probabilistic Inference · NIPS 2015 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
probabilistic search |
0.2 | 1 | 2015 | Softstar: Heuristic-Guided Probabilistic Inference · NIPS 2015 |
Robotics › Motion planning and robot control
robot learning |
0.2 | 1 | 2015 | Graph-Based Inverse Optimal Control for Robot Manipulation · IJCAI 2015 |
Robotics › Motion planning and robot control
trajectory optimization |
0.2 | 1 | 2015 | Intent Prediction and Trajectory Forecasting via Predictive Inverse Linear-Quadratic Regulation · AAAI 2015 |
Human-robot interaction › human behavior modeling
human motion prediction |
0.2 | 1 | 2015 | Intent Prediction and Trajectory Forecasting via Predictive Inverse Linear-Quadratic Regulation · AAAI 2015 |
Human-robot interaction
intent prediction |
0.2 | 1 | 2015 | Intent Prediction and Trajectory Forecasting via Predictive Inverse Linear-Quadratic Regulation · AAAI 2015 |
Machine learning › Learning paradigms
multi-label classification |
0.2 | 1 | 2022 | Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › Video understanding and tracking › action recognition
multimodal action recognition |
0.1 | 1 | 2020 | Moments in Time Dataset: One Million Videos for Event Understanding · IEEE Trans. Pattern Anal. Mach. Intell. 2020 |
Computer vision › Segmentation and scene understanding › object segmentation
affordance segmentation |
0.1 | 1 | 2019 | Reasoning About Human-Object Interactions Through Dual Attention Networks · ICCV 2019 |
Machine learning › Time series and sequential data › time series modeling
probabilistic forecasting |
0.1 | 1 | 2019 | Multi-Agent Tensor Fusion for Contextual Trajectory Prediction · CVPR 2019 |
Robotics › Motion planning and robot control › robot learning › manipulation task learning
goal-conditioned manipulation |
0.1 | 1 | 2017 | Goal-predictive robotic teleoperation from noisy sensors · ICRA 2017 |
Methods — techniques the papers use, named apart from their topics
stratification · 1.5sampling · 1.5estimation · 1.5inverse optimal control · 0.8regression modeling · 0.8empirical bayes · 0.8model interpretation · 0.6long-tail learning · 0.6contrastive learning · 0.5adaptive mean margin · 0.5convolutional fusion · 0.4adversarial loss · 0.4depth camera · 0.3probabilistic trajectory modeling · 0.2inverse linear-quadratic regulation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | A Framework for Efficient Model Evaluation Through Stratification, Sampling, and Estimation
Riccardo Fogliato, Pratik Patil, Mathew Monfort, Pietro Perona |
ECCV (88) | 3 |
| 2024 | Precise Model Benchmarking with Only a Few ObservationsabstractHow can we precisely estimate a large language model's (LLM) accuracy on questions belonging to a specific topic within a larger questionanswering dataset?The standard direct estimator, which averages the model's accuracy on the questions in each subgroup, may exhibit high variance for subgroups (topics) with small sample sizes.Synthetic regression modeling, which leverages the model's accuracy on questions about other topics, may yield biased estimates that are too unreliable for large subgroups.We prescribe a simple yet effective solution: an empirical Bayes (EB) estimator that balances direct and regression estimates for each subgroup separately, improving the precision of subgroup-level estimates of model performance.Our experiments on multiple datasets show that this approach consistently provides more precise estimates of the LLM performance compared to the direct and regression approaches, achieving substantial reductions in the mean squared error.Confidence intervals for EB estimates also have nearnominal coverage and are narrower compared to those for the direct estimator.Additional experiments on tabular and vision data validate the benefits of this EB approach. Riccardo Fogliato, Pratik Patil, Nil-Jana Akpinar, Mathew Monfort |
EMNLP | 4 |
| 2022 | Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video UnderstandingabstractVideos capture events that typically contain multiple sequential, and simultaneous, actions even in the span of only a few seconds. However, most large-scale datasets built to train models for action recognition in video only provide a single label per video. Consequently, models can be incorrectly penalized for classifying actions that exist in the videos but are not explicitly labeled and do not learn the full spectrum of information present in each video in training. Towards this goal, we present the Multi-Moments in Time dataset (M-MiT) which includes over two million action labels for over one million three second videos. This multi-label dataset introduces novel challenges on how to train and analyze models for multi-action detection. Here, we present baseline results for multi-action recognition using loss functions adapted for long tail multi-label learning, provide improved methods for visualizing and interpreting models trained for multi-label action detection and show the strength of transferring models trained on M-MiT to smaller datasets. Mathew Monfort, Bowen Pan, Kandan Ramakrishnan, Alex Andonian, Barry A. McNamara, Alex Lascelles, Quanfu Fan, Dan Gutfreund, Rogério Feris, Aude Oliva |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Spoken Moments: Learning Joint Audio-Visual Representations From Video DescriptionsabstractWhen people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information describing the important high-level details (what, where, who and how) of the observed event and exclude background information that is deemed unimportant to the observer. With this in mind, the descriptions people generate for videos of different dynamic events can greatly improve our understanding of the key information of interest in each video. These descriptions can be captured in captions that provide expanded attributes for video labeling (e.g. actions/objects/scenes/sentiment/etc.) while allowing us to gain new insight into what people find important or necessary to summarize specific events. Existing caption datasets for video understanding are either small in scale or restricted to a specific domain. To address this, we present the Spoken Moments (S-MiT) dataset of 500k spoken captions each attributed to a unique short video depicting a broad range of different events. We collect our descriptions using audio recordings to ensure that they remain as natural and concise as possible while allowing us to scale the size of a large classification dataset. In order to utilize our proposed dataset, we present a novel Adaptive Mean Margin (AMM) approach to contrastive learning and evaluate our models on video/caption retrieval on multiple datasets. We show that our AMM approach consistently improves our results and that models trained on our Spoken Moments dataset generalize better than those trained on other video-caption datasets.http://moments.csail.mit.edu/spoken.html Mathew Monfort, SouYoung Jin, Alexander H. Liu, David F. Harwath, Rogério Feris, James R. Glass, Aude Oliva |
CVPR | 1 |
| 2020 | We Have So Much in Common: Modeling Semantic Relational Set Abstractions in Videos
Alex Andonian, Camilo Fosco, Mathew Monfort, Allen Lee, Rogério Feris, Carl Vondrick, Aude Oliva |
ECCV (18) | 3 |
| 2020 | Moments in Time Dataset: One Million Videos for Event UnderstandingabstractWe present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions occurring in 3 second videos poses many challenges: meaningful events do not include only people, but also objects, animals, and natural phenomena; visual and auditory events can be symmetrical in time ("opening" is "closing" in reverse), and either transient or sustained. We describe the annotation process of our dataset (each video is tagged with one action or activity label among 339 different classes), analyze its scale and diversity in comparison to other large-scale video datasets for action recognition, and report results of several baseline models addressing separately, and jointly, three modalities: spatial, temporal and auditory. The Moments in Time dataset, designed to have a large coverage and diversity of events in both visual and auditory modalities, can serve as a new challenge to develop models that scale to the level of complexity and abstract reasoning that a human processes on a daily basis. Mathew Monfort, Carl Vondrick, Aude Oliva, Alex Andonian, Bolei Zhou, Kandan Ramakrishnan, Sarah Adel Bargal, Tom Yan, Lisa M. Brown, Quanfu Fan, Dan Gutfreund |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Multi-Agent Tensor Fusion for Contextual Trajectory PredictionabstractAccurate prediction of others' trajectories is essential for autonomous driving. Trajectory prediction is challenging because it requires reasoning about agents' past movements, social interactions among varying numbers and kinds of agents, constraints from the scene context, and the stochasticity of human behavior. Our approach models these interactions and constraints jointly within a novel Multi-Agent Tensor Fusion (MATF) network. Specifically, the model encodes multiple agents' past trajectories and the scene context into a Multi-Agent Tensor, then applies convolutional fusion to capture multiagent interactions while retaining the spatial structure of agents and the scene context. The model decodes recurrently to multiple agents' future trajectories, using adversarial loss to learn stochastic predictions. Experiments on both highway driving and pedestrian crowd datasets show that the model achieves state-of-the-art prediction accuracy. Tianyang Zhao 0004, Mathew Monfort, Wongun Choi, Chris L. Baker, Yibiao Zhao, Yizhou Wang 0001, Ying Nian Wu |
CVPR | 3 |
| 2019 | Reasoning About Human-Object Interactions Through Dual Attention NetworksabstractObjects are entities we act upon, where the functionality of an object is determined by how we interact with it. In this work we propose a Dual Attention Network model which reasons about human-object interactions. The dual-attentional framework weights the important features for objects and actions respectively. As a result, the recognition of objects and actions mutually benefit each other. The proposed model shows competitive classification performance on the human-object interaction dataset Something-Something. Besides, it can perform weak spatiotemporal localization and affordance segmentation, despite being trained only with video-level labels. The model not only finds when an action is happening and which object is being manipulated, but also identifies which part of the object is being interacted with. Tete Xiao, Quanfu Fan, Dan Gutfreund, Mathew Monfort, Aude Oliva, Bolei Zhou |
ICCV | 4 |
| 2017 | Goal-predictive robotic teleoperation from noisy sensorsabstractRobotic teleoperation from a human operator's pose demonstrations provides an intuitive and effective means of control that has been made feasible by improvements in sensor technologies in recent years. However, the imprecision of low-cost depth cameras and the difficulty of calibrating a frame of reference for the operator introduce inefficiencies in this process when performing tasks that require interactions with objects in the robot's workspace. We develop a goal-predictive teleoperation system that aids in “de-noising” the controls of the operator to be more goal-directed. Our approach uses inverse optimal control to predict the intended object of interaction from the current motion trajectory in real time and then adapts the degree of autonomy between the operator's demonstrations and autonomous completion of the predicted task. We evaluate our approach using the Microsoft Kinect depth camera as our input sensor to control a Rethink Robotics Baxter robot. Christopher Schultz, Sanket Gaurav, Mathew Monfort, Lingfei Zhang, Brian D. Ziebart |
ICRA | 3 |
| 2016 | Robust Covariate Shift RegressionabstractIn many learning settings, the source data available to train a regression model differs from the target data it encounters when making predictions due to input distribution shift. Appropriately dealing with this situation remains an important challenge. Existing methods attempt to “reweight” the source data samples to better represent the target domain, but this introduces strong inductive biases that are highly extrapolative and can often err greatly in practice. We propose a robust approach for regression under covariate shift that embraces the uncertainty resulting from sample selection bias by producing regression models that are explicitly robust to it. We demonstrate the benefits of our approach on a number of regression tasks. Xiangli Chen, Mathew Monfort, Anqi Liu 0001, Brian D. Ziebart |
AISTATS | 2 |
| 2016 | Adversarial Inverse Optimal Control for General Imitation Learning Losses and Embodiment Transfer
Xiangli Chen, Mathew Monfort, Brian D. Ziebart |
UAI | 2 |
| 2015 | Intent Prediction and Trajectory Forecasting via Predictive Inverse Linear-Quadratic RegulationabstractTo facilitate interaction with people, robots must not only recognize current actions, but also infer a person's intentions and future behavior. Recent advances in depth camera technology have significantly improved human motion tracking. However, the inherent high dimensionality of interacting with the physical world makes efficiently forecasting human intention and future behavior a challenging task. Predictive methods that estimate uncertainty are therefore critical for supporting appropriate robotic responses to the many ambiguities posed within the human-robot interaction setting. We address these two challenges, high dimensionality and uncertainty, by employing predictive inverse optimal control methods to estimate a probabilistic model of human motion trajectories. Our inverse optimal control formulation estimates quadratic cost functions that best rationalize observed trajectories framed as solutions to linear-quadratic regularization problems. The formulation calibrates its uncertainty from observed motion trajectories, and is efficient in high-dimensional state spaces with linear dynamics. We demonstrate its effectiveness on a task of anticipating the future trajectories, target locations and activity intentions of hand motions. Mathew Monfort, Anqi Liu 0001, Brian D. Ziebart |
AAAI | 1 |
| 2015 | Graph-Based Inverse Optimal Control for Robot Manipulation
Arunkumar Byravan, Mathew Monfort, Brian D. Ziebart, Byron Boots, Dieter Fox |
IJCAI | 2 |
| 2015 | Softstar: Heuristic-Guided Probabilistic InferenceabstractRecent machine learning methods for sequential behavior prediction estimate the motives of behavior rather than the behavior itself. This higher-level abstraction improves generalization in different prediction settings, but computing predictions often becomes intractable in large decision spaces. We propose the Softstar algorithm, a softened heuristic-guided search technique for the maximum entropy inverse optimal control model of sequential behavior. This approach supports probabilistic search with bounded approximation error at a significantly reduced computational cost when compared to sampling based methods. We present the algorithm, analyze approximation guarantees, and compare performance with simulation-based inference on two distinct complex decision tasks. Mathew Monfort, Brenden M. Lake, Brian D. Ziebart, Patrick Lucey, Josh Tenenbaum |
NIPS | 1 |