VLDB 2026 Research / reviewers in the wild / expert
Mingrui Wu
dblp:70/213
· DBLP profile ↗
26ranked-venue papers
16as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 15 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating and Mitigating Relationship Hallucinations in Large Vision-Language Models
Mingrui Wu, Jiayi Ji, Hao Fei 0001, Xiaoshuai Sun, Rongrong Ji |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Parametric Chunk Quantization Algorithm for Fast Passive Emitter Localization
Mingrui Wu, Hao Huan, Ran Tao 0003 |
IEEE Signal Process. Lett. | 1 |
| 2025 | ParamΔ for Direct Mixing: Post-Train Large Language Model At Zero Cost
Mingrui Wu, Karthik Prasad, Yuandong Tian, Zechun Liu |
ICLR | 2 |
| 2025 | MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language ModelsabstractDespite growing interest in hallucination in Multimodal Large Language Models, existing studies primarily focus on single-image settings, leaving hallucination in multi-image scenarios largely unexplored. To address this gap, we conduct the first systematic study of hallucinations in multi-image MLLMs and propose MIHBench, a benchmark specifically tailored for evaluating object-related hallucinations across multiple images. MIHBench comprises three core tasks: Multi-Image Object Existence Hallucination, Multi-Image Object Count Hallucination, and Object Identity Consistency Hallucination, targeting semantic understanding across object existence, quantity reasoning, and cross-view identity consistency. Through extensive evaluation, we identify key factors associated with the occurrence of multi-image hallucinations, including: a progressive relationship between the number of image inputs and the likelihood of hallucination occurrences; a strong correlation between single-image hallucination tendencies and those observed in multi-image contexts; and the influence of same-object image ratios and the positional placement of negative samples within image sequences on the occurrence of object identity consistency hallucination. To address these challenges, we propose a Dynamic Attention Balancing mechanism that adjusts inter-image attention distributions while preserving the overall visual attention proportion. Experiments across multiple state-of-the-art MLLMs demonstrate that our method effectively reduces hallucination occurrences and enhances semantic integration and reasoning stability in multi-image scenarios. Mingrui Wu, Zixiang Jin, Jiayi Ji, Xiaoshuai Sun, Liujuan Cao, Rongrong Ji |
ACM Multimedia | 2 |
| 2025 | TraDiffusion: Trajectory-Based Training-Free Image Generation
Mingrui Wu, Oucheng Huang, Jiayi Ji, Jianzhuang Liu, Xiaoshuai Sun, Liujuan Cao, Rongrong Ji |
Int. J. Comput. Vis. | 1 |
| 2024 | Toward Open-Set Human Object Interaction DetectionabstractThis work is oriented toward the task of open-set Human Object Interaction (HOI) detection. The challenge lies in identifying completely new, out-of-domain relationships, as opposed to in-domain ones which have seen improvements in zero-shot HOI detection. To address this challenge, we introduce a simple Disentangled HOI Detection (DHD) model for detecting novel relationships by integrating an open-set object detector with a Visual Language Model (VLM). We utilize a disentangled image-text contrastive learning metric for training and connect the bottom-up visual features to text embeddings through lightweight unary and pair-wise adapters. Our model can benefit from the open-set object detector and the VLM to detect novel action categories and combine actions with novel object categories. We further present the VG-HOI dataset, a comprehensive benchmark with over 17k HOI relationships for open-set scenarios. Experimental results show that our model can detect unknown action classes and combine unknown object classes. Furthermore, it can generalize to over 17k HOI classes while being trained on just 600 HOI classes. Mingrui Wu, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji |
AAAI | 1 |
| 2024 | Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language ModelsabstractThe issue of hallucinations is a prevalent concern in existing Large Vision-Language Models (LVLMs). Previous efforts have primarily focused on investigating object hallucinations, which can be easily alleviated by introducing object detectors. However, these efforts neglect hallucinations in inter-object relationships, which is essential for visual comprehension. In this work, we introduce R-Bench, a novel benchmark for evaluating Vision Relationship Hallucination. R-Bench features image-level questions that focus on the existence of relationships and instance-level questions that assess local visual comprehension. We identify three types of relationship co-occurrences that lead to hallucinations: relationship-relationship, subject-relationship, and relationship-object. The visual instruction tuning dataset's long-tail distribution significantly impacts LVLMs' understanding of visual relationships. Additionally, our analysis reveals that current LVLMs tend to overlook visual content, overly rely on the common sense knowledge of Large Language Models (LLMs), and struggle with spatial relationship reasoning based on contextual information. Mingrui Wu, Jiayi Ji, Oucheng Huang, Yuhang Wu 0004, Xiaoshuai Sun, Rongrong Ji |
ICML | 1 |
| 2024 | ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language ModelsabstractIn this work, we propose a training-free method to inject visual prompts into Multimodal Large Language Models (MLLMs) through learnable latent variable optimization. We observe that attention, as the core module of MLLMs, connects text prompt tokens and visual tokens, ultimately determining the final results. Our approach involves adjusting visual tokens from the MLP output during inference, controlling the attention response to ensure text prompt tokens attend to visual tokens in referring regions. We optimize a learnable latent variable based on an energy function, enhancing the strength of referring regions in the attention map. This enables detailed region description and reasoning without the need for substantial training costs or model retraining. Our method offers a promising direction for integrating referring abilities into MLLMs, and supports referring with box, mask, scribble and point. The results demonstrate that our method exhibits out-of-domain generalization and interpretability. Mingrui Wu, Xinyue Cai, Jiayi Ji, Oucheng Huang, Gen Luo, Hao Fei 0001, Guannan Jiang, Xiaoshuai Sun, Rongrong Ji |
NeurIPS | 1 |
| 2023 | End-to-End Zero-Shot HOI Detection via Vision and Language Knowledge DistillationabstractMost existing Human-Object Interaction (HOI) Detection methods rely heavily on full annotations with predefined HOI categories, which is limited in diversity and costly to scale further. We aim at advancing zero-shot HOI detection to detect both seen and unseen HOIs simultaneously. The fundamental challenges are to discover potential human-object pairs and identify novel HOI categories. To overcome the above challenges, we propose a novel End-to-end zero-shot HOI Detection (EoID) framework via vision-language knowledge distillation. We first design an Interactive Score module combined with a Two-stage Bipartite Matching algorithm to achieve interaction distinguishment for human-object pairs in an action-agnostic manner. Then we transfer the distribution of action probability from the pretrained vision-language teacher as well as the seen ground truth to the HOI model to attain zero-shot HOI classification. Extensive experiments on HICO-Det dataset demonstrate that our model discovers potential interactive pairs and enables the recognition of unseen HOIs. Finally, our method outperforms the previous SOTA under various zero-shot settings. Moreover, our method is generalizable to large-scale object detection data to further scale up the action sets. The source code is available at: https://github.com/mrwu-mac/EoID. Mingrui Wu, Jiaxin Gu, Yunhang Shen, Mingbao Lin, Chao Chen 0026, Xiaoshuai Sun |
AAAI | 1 |
| 2022 | DIFNet: Boosting Visual Information Flow for Image CaptioningabstractCurrent Image Captioning (IC) methods predict textual words sequentially based on the input visual information from the visual feature extractor and the partially generated sentence information. However, for most cases, the partially generated sentence may dominate the target word prediction due to the insufficiency of visual information, making the generated descriptions irrelevant to the content of the given image. In this paper, we propose a Dual Information Flow Network (DIFNet11Source code is available at: https://github.com/mrwu-mac/DIFNet) to address this issue, which takes segmentation feature as another visual information source to enhance the contribution of visual information for prediction. To maximize the use of two information flows, we also propose an effective feature fusion module termed Iterative Independent Layer Normalization (IILN) which can condense the most relevant inputs while retraining modality-specific information in each flow. Experiments show that our method is able to enhance the dependence of prediction on visual information, making word prediction more focused on the visual content, and thus achieves new state-of-the-art performance on the MSCOCO dataset, e.g., 136.2 CIDEr on COCO Karpathy test split. Mingrui Wu, Xuying Zhang, Xiaoshuai Sun, Yiyi Zhou, Chao Chen 0026, Jiaxin Gu, Xing Sun 0001, Rongrong Ji |
CVPR | 1 |
| 2022 | PINE: Universal Deep Embedding for Graph Nodes via Partial Permutation Invariant Set FunctionsabstractGraph node embedding aims at learning a vector representation for all nodes given a graph. It is a central problem in many machine learning tasks (e.g., node classification, recommendation, community detection). The key problem in graph node embedding lies in how to define the dependence to neighbors. Existing approaches specify (either explicitly or implicitly) certain dependencies on neighbors, which may lead to loss of subtle but important structural information within the graph and other dependencies among neighbors. This intrigues us to ask the question: can we design a model to give the adaptive flexibility of dependencies to each node's neighborhood. In this paper, we propose a novel graph node embedding method (named PINE) via a novel notion of partial permutation invariant set function, to capture any possible dependence. Our method 1) can learn an arbitrary form of the representation function from the neighborhood, without losing any potential dependence structures, and 2) is applicable to both homogeneous and heterogeneous graph embedding, the latter of which is challenged by the diversity of node types. Furthermore, we provide theoretical guarantee for the representation capability of our method for general homogeneous and heterogeneous graphs. Empirical evaluation results on benchmark data sets show that our proposed PINE method outperforms the state-of-the-art approaches on producing node vectors for various learning tasks of both homogeneous and heterogeneous graphs. Shupeng Gui, Xiangliang Zhang 0001, Pan Zhong, Mingrui Wu, Jieping Ye, Zhengdao Wang, Ji Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | Classification over Clustering: Augmenting Text Representation with Clusters Helps!
Xiaoye Tan, Rui Yan 0001, Chongyang Tao, Mingrui Wu |
NLPCC (2) | 4 |
| 2012 | Differentially private search log sanitization with optimal output utilityabstractWeb search logs contain extremely sensitive data, as evidenced by the recent AOL incident. However, storing and analyzing search logs can be very useful for many purposes (i.e. investigating human behavior). Thus, an important research question is how to privately sanitize search logs. Several search log anonymization techniques have been proposed with concrete privacy models. However, in all of these solutions, the output utility of the techniques is only evaluated rather than being maximized in any fashion. Indeed, for effective search log anonymization, it is desirable to derive the outputs with optimal utility while meeting the privacy standard. In this paper, we propose utility-maximizing sanitization based on the rigorous privacy standard of differential privacy, in the context of search logs. Specifically, we utilize optimization models to maximize the output utility of the sanitization for different applications, while ensuring that the production process satisfies differential privacy. An added benefit is that our novel randomization strategy maintains the schema integrity in the output search logs. A comprehensive evaluation on real search logs validates the approach and demonstrates its robustness and scalability. Yuan Hong 0001, Jaideep Vaidya, Haibing Lu, Mingrui Wu |
EDBT | 4 |
| 2011 | Hierarchical Classification via Orthogonal Transfer
Dengyong Zhou, Mingrui Wu |
ICML | 3 |
| 2010 | Gradient descent optimization of smoothed information retrieval metrics
Olivier Chapelle, Mingrui Wu |
Inf. Retr. | 2 |
| 2009 | Smoothing DCG for learning to rank: a novel approach using smoothed hinge functionsabstractDiscounted cumulative gain (DCG) is widely used for evaluating ranking functions. It is therefore natural to learn a ranking function that directly optimizes DCG. However, DCG is non-smooth, rendering gradient-based optimization algorithms inapplicable. To remedy this, smoothed versions of DCG have been proposed but with only partial success. In this paper, we first present analysis that shows it is ineffective using the gradient of the smoothed DCG to drive the optimization algorithm. We then propose a novel approach, SHF-SDCG, for smoothing DCG by using smoothed hinge functions (SHF). It has the advantage of seamlessly transition from driving the optimization mimicking pairwise learning when the ranking function does not fit the data well, to driving the optimization using DCG when the ranking function becomes more accurate. SHF-SDCG is then extended to REG-SHF-SDCG, an algorithm which gradually transits from pointwise and pairwise to listwise learning. Finally experimental results are provided to validate the effectiveness of SHF-SDCG and REG-SHF-SDCG. Mingrui Wu, Yi Chang 0001, Zhaohui Zheng 0001, Hongyuan Zha |
CIKM | 1 |
| 2009 | A Small Sphere and Large Margin Approach for Novelty Detection Using Training Data with OutliersabstractWe present a small sphere and large margin approach for novelty detection problems, where the majority of training data are normal examples. In addition, the training data also contain a small number of abnormal examples or outliers. The basic idea is to construct a hypersphere that contains most of the normal examples, such that the volume of this sphere is as small as possible, while at the same time the margin between the surface of this sphere and the outlier training data is as large as possible. This can result in a closed and tight boundary around the normal data. To build such a sphere, we only need to solve a convex optimization problem that can be efficiently solved with the existing software packages for training \nu\hbox{-}Support Vector Machines. Experimental results are provided to validate the effectiveness of the proposed algorithm. Mingrui Wu, Jieping Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Learning subspace kernels for classificationabstractKernel methods have been applied successfully in many data mining tasks. Subspace kernel learning was recently proposed to discover an effective low-dimensional subspace of a kernel feature space for improved classification. In this paper, we propose to construct a subspace kernel using the Hilbert-Schmidt Independence Criterion (HSIC). We show that the optimal subspace kernel can be obtained efficiently by solving an eigenvalue problem. One limitation of the existing subspace kernel learning formulations is that the kernel learning and classification are independent and the subspace kernel may not be optimally adapted for classification. To overcome this limitation, we propose a joint optimization framework, in which we learn the subspace kernel and subsequent classifiers simultaneously. In addition, we propose a novel learning formulation that extracts an uncorrelated subspace kernel to reduce the redundant information in a subspace kernel. Following the idea from multiple kernel learning, we extend the proposed formulations to the case when multiple kernels are available and need to be combined. We show that the integration of subspace kernels can be formulated as a semidefinite program (SDP) which is computationally expensive. To improve the efficiency of the SDP formulation, we propose an equivalent semi-infinite linear program (SILP) formulation which can be solved efficiently by the column generation technique. Experimental results on a collection of benchmark data sets demonstrate the effectiveness of the proposed algorithms. Shuiwang Ji, Betul Ceran, Qi Li 0001, Mingrui Wu, Jieping Ye |
KDD | 5 |
| 2007 | Local learning projectionsabstractThis paper presents a Local Learning Projection (LLP) approach for linear dimensionality reduction. We first point out that the well known Principal Component Analysis (PCA) essentially seeks the projection that has the minimal global estimation error. Then we propose a dimensionality reduction algorithm that leads to the projection with the minimal local estimation error, and elucidate its advantages for classification tasks. We also indicate that LLP keeps the local information in the sense that the projection value of each point can be well estimated based on its neighbors and their projection values. Experimental results are provided to validate the effectiveness of the proposed algorithm. Mingrui Wu, Kai Yu 0001, Shipeng Yu, Bernhard Schölkopf |
ICML | 1 |
| 2007 | A Subspace Kernel for Nonlinear Feature Extraction
Mingrui Wu, Jason D. R. Farquhar |
IJCAI | 1 |
| 2007 | Discriminative K-means for ClusteringabstractWe present a theoretical study on the discriminative clustering framework, recently proposed for simultaneous subspace selection via linear discriminant analysis (LDA) and clustering. Empirical results have shown its favorable performance in comparison with several other popular clustering algorithms. However, the inherent relationship between subspace selection and clustering in this framework is not well understood, due to the iterative nature of the algorithm. We show in this paper that this iterative subspace selection and clustering is equivalent to kernel K-means with a specific kernel Gram matrix. This provides significant and new insights into the nature of this subspace selection procedure. Based on this equivalence relationship, we propose the Discriminative K-means (DisKmeans) algorithm for simultaneous LDA subspace selection and clustering, as well as an automatic parameter estimation procedure. We also present the nonlinear extension of DisKmeans using kernels. We show that the learning of the kernel matrix over a convex set of pre-specified kernel matrices can be incorporated into the clustering formulation. The connection between DisKmeans and several other clustering algorithms is also analyzed. The presented theories and algorithms are evaluated through experiments on a collection of benchmark data sets. Jieping Ye, Mingrui Wu |
NIPS | 3 |
| 2006 | Supervised probabilistic principal component analysisabstractPrincipal component analysis (PCA) has been extensively applied in data mining, pattern recognition and information retrieval for unsupervised dimensionality reduction. When labels of data are available, e.g., in a classification or regression task, PCA is however not able to use this information. The problem is more interesting if only part of the input data are labeled, i.e., in a semi-supervised setting. In this paper we propose a supervised PCA model called SPPCA and a semi-supervised PCA model called S2PPCA, both of which are extensions of a probabilistic PCA model. The proposed models are able to incorporate the label information into the projection phase, and can naturally handle multiple outputs (i.e., in multi-task learning problems). We derive an efficient EM learning algorithm for both models, and also provide theoretical justifications of the model behaviors. SPPCA and S2PPCA are compared with other supervised projection methods on various learning tasks, and show not only promising performance but also good scalability. Shipeng Yu, Kai Yu 0001, Volker Tresp, Hans-Peter Kriegel, Mingrui Wu |
KDD | 5 |
| 2006 | A Local Learning Approach for ClusteringabstractWe present a local learning approach for clustering. The basic idea is that a good clustering result should have the property that the cluster label of each data point can be well predicted based on its neighboring data and their cluster labels, using current supervised learning methods. An optimization problem is formulated such that its solution has the above property. Relaxation and eigen-decomposition are applied to solve this optimization problem. We also briefly investigate the parameter selection issue and provide a simple parameter selection method for the proposed algorithm. Experimental results are provided to validate the effectiveness of the proposed approach. Mingrui Wu, Bernhard Schölkopf |
NIPS | 1 |
| 2006 | A Direct Method for Building Sparse Kernel Learning AlgorithmsabstractMany kernel learning algorithms, including support vector machines, result in a kernel machine, such as a kernel classifier, whose key component is a weight vector in a feature space implicitly introduced by a positive definite kernel function. This weight vector is usually obtained by solving a convex optimization problem. Based on this fact we present a direct method to build sparse kernel learning algorithms by adding one more constraint to the original convex optimization problem, such that the sparseness of the resulting kernel machine is explicitly controlled while at the same time performance is kept as high as possible. A gradient based approach is provided to solve this modified optimization problem. Applying this method to the support vectom machine results in a concrete algorithm for building sparse large margin classifiers. These classifiers essentially find a discriminating subspace that can be spanned by a small number of vectors, and in this subspace, the different classes of data are linearly well separated. Experimental results over several classification benchmarks demonstrate the effectiveness of our approach. Mingrui Wu, Bernhard Schölkopf, Gökhan H. Bakir |
J. Mach. Learn. Res. | 1 |
| 2005 | Building Sparse Large Margin ClassifiersabstractThis paper presents an approach to build Sparse Large Margin Classifiers (SLMC) by adding one more constraint to the standard Support Vector Machine (SVM) training problem. The added constraint explicitly controls the sparseness of the classifier and an approach is provided to solve the formulated problem. When considering the dual of this problem. it can be seen that building an SLMC is equivalent to constructing an SVM with a modified kernel function. Further analysis of this kernel function indicates that the proposed approach essentially finds a discriminating subspace that can be spanned by a small number of vectors, and in this subspace different classes of data are linearly well separated. Experimental results over several classification benchmarks show that in most cases the proposed approach outperforms the state-of-art sparse learning algorithms. Mingrui Wu, Bernhard Schölkopf, Gökhan H. Bakir |
ICML | 1 |
| 2000 | A Neural Network Based Classifier for Handwritten Chinese Character RecognitionabstractIn this paper, a geometrical approach for building neural networks is proposed. With the proposed approach, it is very easy to construct an efficient neural classifier to solve the handwritten Chinese character recognition problem, as well as other pattern recognition problems of large scale. Experiments are conducted to evaluate the performance of the proposed approach and results obtained are promising. Mingrui Wu, Bo Zhang 0010, Ling Zhang 0001 |
ICPR | 1 |