Ming Xu 0015

dblp:43/3362-15 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2025
0000-0002-6478-0582ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2025 Context-Dependent Anomaly Action Recognition
abstract
We explore the problem of unsupervised anomaly action recognition, focusing on identifying abnormal human behavior that takes into account the surrounding context. Unlike conventional approaches that rely solely on observing the human action, we recognize that anomalies can be context-dependent—texting while driving, for example, may be considered an anomaly, but not if the car is parked safely and the driver is waiting to pick up a passenger. To this end, we propose a simple framework that processes concurrent data streams of human action and contextual information. To learn the concept of normal behavior, we employ normalizing flows to model the joint probability distribution over pre-defined normal actioncontext sample pairs. Our method is flexible in admitting the use of different action and context streams, including the ability to leverage pretrained models for extracting features describing the action and context. We conduct a series of experiments on publicly available data, demonstrating the effectiveness of our approach in identifying context-dependent anomaly actions, including application to driving scenarios. Our code is available at https://github.com/knbandit/CDAAR.
Nutthadech Banditakkarakul, Ming Xu 0015, Akshay Asthana, Liang Zheng 0001, Stephen Gould
FG2
2025 Flexible Geometric Guidance for Probabilistic Human Pose Estimation with Diffusion Models
abstract
3D human pose estimation from 2D images is a challenging problem due to depth ambiguity and occlusion. Because of these challenges the task is underdetermined, where there exists multiple—possibly infinite—poses that are plausible given the image. Despite this, many prior works assume the existence of a deterministic mapping and estimate a single pose given an image. Furthermore, methods based on machine learning require a large amount of paired 2D-3D data to train and suffer from generalization issues to unseen scenarios. To address both of these issues, we propose a framework for pose estimation using diffusion models, which enables sampling from a probability distribution over plausible poses which are consistent with a 2D image. Our approach falls under the guidance framework for conditional generation, and guides samples from an unconditional diffusion model, trained only on 3D data, using the gradients of the heatmaps from a 2D keypoint detector. We evaluate our method on the Human 3.6M dataset under best-of- m multiple hypothesis evaluation, showing state-of-the-art performance among methods which do not require paired 2D-3D data for training. We additionally evaluate the generalization ability using the MPI-INF-3DHP and 3DPW datasets and demonstrate competitive performance. Finally, we demonstrate the flexibility of our framework by using it for novel tasks including pose generation and pose completion, without the need to train bespoke conditional models. We make code available at https://github.com/fsnelgar/diffusion_pose.
Francis Snelgar, Ming Xu 0015, Stephen Gould, Liang Zheng 0001, Akshay Asthana
FG2
2025 Can We Predict Performance of Large Models across Vision-Language Tasks?
abstract
Evaluating large vision-language models (LVLMs) is very expensive, due to high computational cost and the wide variety of tasks. The good news is that if we already have some observed performance scores, we may be able to infer unknown ones. In this study, we propose a new framework for predicting unknown performance scores based on observed ones from other LVLMs or tasks. We first formulate the performance prediction as a matrix completion task. Specifically, we construct a sparse performance matrix $\boldsymbol{R}$, where each entry $R_{mn}$ represents the performance score of the $m$-th model on the $n$-th dataset. By applying probabilistic matrix factorization (PMF) with Markov chain Monte Carlo (MCMC), we can complete the performance matrix, i.e., predict unknown scores. Additionally, we estimate the uncertainty of performance prediction based on MCMC. Practitioners can evaluate their models on untested tasks with higher uncertainty first, which quickly reduces the prediction errors. We further introduce several improvements to enhance PMF for scenarios with sparse observed performance scores. Our experiments demonstrate the accuracy of PMF in predicting unknown scores, the reliability of uncertainty estimates in ordering evaluations, and the effectiveness of our enhancements for handling sparse data. Our code is available at https://github.com/Qinyu-Allen-Zhao/CrossPred-LVLM.
Qinyu Zhao, Ming Xu 0015, Kartik Gupta, Akshay Asthana, Liang Zheng 0001, Stephen Gould
ICML2
2024 Temporally Consistent Unbalanced Optimal Transport for Unsupervised Action Segmentation
abstract
We propose a novel approach to the action segmentation task for long, untrimmed videos, based on solving an opti-mal transport problem. By encoding a temporal consistency prior into a Gromov- Wasserstein problem, we are able to decode a temporally consistent segmentation from a noisy affinity/matching cost matrix between video frames and action classes. Unlike previous approaches, our method does not require knowing the action order for a video to attain temporal consistency. Furthermore, our resulting (fused) Gromov- Wasserstein problem can be efficiently solved on GPUs using a few iterations of projected mirror descent. We demonstrate the effectiveness of our method in an unsu-pervised learning setting, where our method is used to gen-erate pseudo-labels for self-training. We evaluate our seg-mentation approach and unsupervised learning pipeline on the Breakfast, 50-Salads, YouTube Instructions and Desk-top Assembly datasets, yielding state-of-the-art results for the unsupervised video action segmentation task.
Ming Xu 0015, Stephen Gould
CVPR1
2024 VLAD-BuFF: Burst-Aware Fast Feature Aggregation for Visual Place Recognition
Ahmad Khaliq, Ming Xu 0015, Stephen Hausler, Michael Milford, Sourav Garg
ECCV (44)2
2024 The First to Know: How Token Distributions Reveal Hidden Knowledge in Large Vision-Language Models?
Qinyu Zhao, Ming Xu 0015, Kartik Gupta, Akshay Asthana, Liang Zheng 0001, Stephen Gould
ECCV (48)2
2024 Towards Optimal Feature-Shaping Methods for Out-of-Distribution Detection
abstract
Feature shaping refers to a family of methods that exhibit state-of-the-art performance for out-of-distribution (OOD) detection. These approaches manipulate the feature representation, typically from the penultimate layer of a pre-trained deep learning model, so as to better differentiate between in-distribution (ID) and OOD samples. However, existing feature-shaping methods usually employ rules manually designed for specific model architectures and OOD datasets, which consequently limit their generalization ability. To address this gap, we first formulate an abstract optimization framework for studying feature-shaping methods. We then propose a concrete reduction of the framework with a simple piecewise constant shaping function and show that existing feature-shaping methods approximate the optimal solution to the concrete optimization problem. Further, assuming that OOD data is inaccessible, we propose a formulation that yields a closed-form solution for the piecewise constant shaping function, utilizing solely the ID data. Through extensive experiments, we show that the feature-shaping function optimized by our method improves the generalization ability of OOD detection across a large variety of datasets and model architectures. Our code is available at https://github.com/Qinyu-Allen-Zhao/OptFSOOD.
Qinyu Zhao, Ming Xu 0015, Kartik Gupta, Akshay Asthana, Liang Zheng 0001, Stephen Gould
ICLR2
2023 Deep Declarative Dynamic Time Warping for End-to-End Learning of Alignment Paths
Ming Xu 0015, Sourav Garg, Michael Milford, Stephen Gould
ICLR1
2021 Patch-NetVLAD: Multi-Scale Fusion of Locally-Global Descriptors for Place Recognition
abstract
Visual Place Recognition is a challenging task for robotics and autonomous systems, which must deal with the twin problems of appearance and viewpoint change in an always changing world. This paper introduces Patch-NetVLAD, which provides a novel formulation for combining the advantages of both local and global descriptor methods by deriving patch-level features from NetVLAD residuals. Unlike the fixed spatial neighborhood regime of existing local keypoint features, our method enables aggregation and matching of deep-learned local features defined over the feature-space grid. We further introduce a multi-scale fusion of patch features that have complementary scales (i.e. patch sizes) via an integral feature space and show that the fused features are highly invariant to both condition (season, structure, and illumination) and viewpoint (translation and rotation) changes. Patch-NetVLAD achieves state-of-the-art visual place recognition results in computationally limited scenarios, validated on a range of challenging real-world datasets, including winning the Facebook Mapillary Visual Place Recognition Challenge at ECCV2020. It is also adaptable to user requirements, with a speed-optimised version operating over an order of magnitude faster than the state-of-the-art. By combining superior performance with improved computational efficiency in a configurable framework, Patch-NetVLAD is well suited to enhance both stand-alone place recognition capabilities and the overall performance of SLAM systems.
Stephen Hausler, Sourav Garg, Ming Xu 0015, Michael Milford, Tobias Fischer 0001
CVPR3
2019 Variance reduction properties of the reparameterization trick
abstract
The reparameterization trick is widely used in variational inference as it yields more accurate estimates of the gradient of the variational objective than alternative approaches such as the score function method. Although there is overwhelming empirical evidence in the literature showing its success, there is relatively little research exploring why the reparameterization trick is so effective. We explore this under the idealized assumptions that the variational approximation is a mean-field Gaussian density and that the log of the joint density of the model parameters and the data is a quadratic function that depends on the variational mean. From this, we show that the marginal variances of the reparameterization gradient estimator are smaller than those of the score function gradient estimator. We apply the result of our idealized analysis to real-world examples.
Ming Xu 0015, Matias Quiroz, Robert Kohn, Scott A. Sisson
AISTATS1
2019 Hierarchical Encoding of Sequential Data With Compact and Sub-Linear Storage Cost
Huu Le, Ming Xu 0015, Tuan Hoang, Michael Milford
ICCV2