VLDB 2026 Research / reviewers in the wild / expert
Haoyu Lei
dblp:296/4654
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Generative modeling · 25% Trustworthy machine learning · 20% Optimization for machine learning · 18% |
Topics — the 12 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.9 | 2 | 2026 | Boosting Cross-problem Generalization in Diffusion-Based Neural Combinatorial Solver via Inference Time Adaptation · AAAI 2026 SPARKE: Scalable Prompt-Aware Diversity and Novelty Guidance in Diffusion Models via RKE Score · NeurIPS 2025 |
Machine learning › Optimization for machine learning
combinatorial optimization |
1.0 | 1 | 2026 | Boosting Cross-problem Generalization in Diffusion-Based Neural Combinatorial Solver via Inference Time Adaptation · AAAI 2026 |
Machine learning › Optimization for machine learning › combinatorial optimization
neural combinatorial optimization |
1.0 | 1 | 2026 | Boosting Cross-problem Generalization in Diffusion-Based Neural Combinatorial Solver via Inference Time Adaptation · AAAI 2026 |
Computer vision › Vision and language › vision-language model
contrastive vision-language model |
0.9 | 1 | 2025 | Boosting the visual interpretability of CLIP via adversarial fine-tuning · ICLR 2025 |
Machine learning › Deep learning architectures and training
diversity-aware sampling |
0.9 | 1 | 2025 | SPARKE: Scalable Prompt-Aware Diversity and Novelty Guidance in Diffusion Models via RKE Score · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
hallucination |
0.9 | 1 | 2025 | MESH - Understanding Videos Like Human: Measuring Hallucinations in Large Video Models · ACM Multimedia 2025 |
Robotics › Robot manipulation › learning from demonstration
imitation learning for manipulation |
0.9 | 1 | 2025 | Two-Steps Diffusion Policy for Robotic Manipulation via Genetic Denoising · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Boosting the visual interpretability of CLIP via adversarial fine-tuning · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | SPARKE: Scalable Prompt-Aware Diversity and Novelty Guidance in Diffusion Models via RKE Score · NeurIPS 2025 |
Machine learning › Transfer learning and domain adaptation
cross-task generalization |
0.3 | 1 | 2026 | Boosting Cross-problem Generalization in Diffusion-Based Neural Combinatorial Solver via Inference Time Adaptation · AAAI 2026 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness › adversarial training
adversarial fine-tuning |
0.3 | 1 | 2025 | Boosting the visual interpretability of CLIP via adversarial fine-tuning · ICLR 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.3 | 1 | 2025 | Boosting the visual interpretability of CLIP via adversarial fine-tuning · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
diffusion model · 1.9training-free guidance · 1.0inference time adaptation · 1.0population-based sampling · 0.9norm regularization · 0.9network dissection · 0.9genetic denoising · 0.9feature attribution · 0.9conditional entropy · 0.9adversarial fine-tuning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Cross-problem Generalization in Diffusion-Based Neural Combinatorial Solver via Inference Time AdaptationabstractDiffusion-based Neural Combinatorial Optimization (NCO) has demonstrated effectiveness in solving NP-complete (NPC) problems by learning discrete diffusion models for solution generation, eliminating hand-crafted domain knowledge. Despite their success, existing NCO methods face significant challenges in both cross-scale and cross-problem generalization, and high training costs compared to traditional solvers. While recent studies on diffusion models have introduced training-free guidance approaches that leverage pre-defined guidance functions for conditional generation, such methodologies have not been extensively explored in combinatorial optimization. To bridge this gap, we propose a training-free inference time adaptation framework (DIFU-Ada) that enables both the zero-shot cross-problem transfer and cross-scale generalization capabilities of diffusion-based NCO solvers without requiring additional training. We provide theoretical analysis that helps understanding the cross-problem transfer capability. Our experimental results demonstrate that a diffusion solver, trained exclusively on the Traveling Salesman Problem (TSP), can achieve competitive zero-shot transfer performance across different problem scales on TSP variants, such as Prize Collecting TSP (PCTSP) and the Orienteering Problem (OP), through inference time adaptation. Haoyu Lei, Kaiwen Zhou 0001, Yinchuan Li, Zhitang Chen, Farzan Farnia |
AAAI | 1 |
| 2026 | On the Fragility of AI-Based Channel Decoders under Small Channel PerturbationsabstractRecent advances in deep learning have led to AI-based error correction decoders that report empirical performance improvements over traditional belief-propagation (BP) decoding on AWGN channels. While such gains are promising, a fundamental question remains: where do these improvements come from, and what cost is paid to achieve them? In this work, we study this question through the lens of robustness to distributional shifts at the channel output. We evaluate both input-dependent adversarial perturbations (FGM and projected gradient methods under $\ell_2$ constraints) and universal adversarial perturbations that apply a single norm-bounded shift to all received vectors. Our results show that recent AI decoders, including ECCT and CrossMPT, could suffer significant performance degradation under such perturbations, despite superior nominal performance under i.i.d. AWGN. Moreover, adversarial perturbations transfer relatively strongly between AI decoders but weakly to BP-based decoders, and universal perturbations are substantially more harmful than random perturbations of equal norm. These numerical findings suggest a potential robustness cost and higher sensitivity to channel distribution underlying recent AI decoding gains. Haoyu Lei, Mohammad Jalali, Chin Wa Lau, Farzan Farnia |
ISIT | 1 |
| 2026 | Genetic algorithm for bi-objective optimization under min-product fuzzy relation inequality constraints in supply chains
Haoyu Lei, Yushu Feng, Xiaopeng Yang 0001, Qianyu Shu |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | The lexicographic maximum solution subject to min-product fuzzy relation inequalities with absent coefficients
Haoyu Lei, Xiaopeng Yang 0001, Qianyu Shu |
Fuzzy Sets Syst. | 1 |
| 2025 | Boosting the visual interpretability of CLIP via adversarial fine-tuningabstractCLIP has achieved great success in visual representation learning and is becoming an important plug-in component for many large multi-modal models like LLaVA and DALL-E. However, the lack of interpretability caused by the intricate image encoder architecture and training process restricts its wider use in high-stake decision making applications. In this work, we propose an unsupervised adversarial fine-tuning (AFT) with norm-regularization to enhance the visual interpretability of CLIP. We provide theoretical analysis showing that AFT has implicit regularization that enforces the image encoder to encode the input features sparsely, directing the network's focus towards meaningful features. Evaluations by both feature attribution techniques and network dissection offer convincing evidence that the visual interpretability of CLIP has significant improvements. With AFT, the image encoder prioritizes pertinent input features, and the neuron within the encoder exhibits better alignment with human-understandable concepts. Moreover, these effects are generalizable to out-of-distribution datasets and can be transferred to downstream tasks. Additionally, AFT enhances the visual interpretability of derived large vision-language models that incorporate the pre-trained CLIP an integral component. The code of this paper is available at [the CLIP_AFT GitHub repository](https://github.com/peterant330/CLIP_AFT). Shizhan Gong, Haoyu Lei, Qi Dou 0001, Farzan Farnia |
ICLR | 2 |
| 2025 | MESH - Understanding Videos Like Human: Measuring Hallucinations in Large Video ModelsabstractLarge Video Models (LVMs) build on the semantic capabilities of Large Language Models (LLMs) and vision modules by integrating temporal information to better understand dynamic video content. Despite their progress, LVMs are prone to hallucinations-producing inaccurate or irrelevant descriptions. Current benchmarks for video hallucination depend heavily on manual categorization of video content, neglecting the perception-based processes through which humans naturally interpret videos. We introduce MESH, a benchmark designed to evaluate hallucinations in LVMs systematically. MESH uses a Question-Answering framework with binary and multi-choice formats incorporating target and trap instances. It follows a bottom-up approach, evaluating basic objects, coarse-to-fine subject features, and subject-action pairs, aligning with human video understanding. We demonstrate that MESH offers an effective and comprehensive approach for identifying hallucinations in video understanding. Our evaluations show that while LVMs excel at recognizing basic objects and features, their susceptibility to hallucinations increases markedly when handling fine details or aligning multiple actions involving various subjects in longer videos. The benchmark is available at MESH-Benchmark. Garry Yang, Zizhe Chen, Man Hon Wong 0001, Haoyu Lei, Yongqiang Chen 0002, Zhenguo Li, Kaiwen Zhou 0001, James Cheng |
ACM Multimedia | 4 |
| 2025 | Two-Steps Diffusion Policy for Robotic Manipulation via Genetic DenoisingabstractDiffusion models, such as diffusion policy, have achieved state-of-the-art results in robotic manipulation by imitating expert demonstrations. While diffusion models were originally developed for vision tasks like image and video generation, many of their inference strategies have been directly transferred to control domains without adaptation. In this work, we show that by tailoring the denoising process to the specific characteristics of embodied AI tasks—particularly the structured, low-dimensional nature of action distributions---diffusion policies can operate effectively with as few as 5 neural function evaluations (NFE).
Building on this insight, we propose a population-based sampling strategy, genetic denoising, which enhances both performance and stability by selecting denoising trajectories with low out-of-distribution risk. Our method solves challenging tasks with only 2 NFE while improving or matching performance. We evaluate our approach across 14 robotic manipulation tasks from D4RL and Robomimic, spanning multiple action horizons and inference budgets. In over 2 million evaluations, our method consistently outperforms standard diffusion-based policies, achieving up to 20\% performance gains with significantly fewer inference steps. Mateo Clémente, Leo Maxime Brunswic, Rui Heng Yang, Yasser H. Khalil, Haoyu Lei, Amir Rasouli, Yinchuan Li |
NeurIPS | 6 |
| 2025 | SPARKE: Scalable Prompt-Aware Diversity and Novelty Guidance in Diffusion Models via RKE ScoreabstractDiffusion models have demonstrated remarkable success in high-fidelity image synthesis and prompt-guided generative modeling. However, ensuring adequate diversity in generated samples of prompt-guided diffusion models remains a challenge, particularly when the prompts span a broad semantic spectrum and the diversity of generated data needs to be evaluated in a prompt-aware fashion across semantically similar prompts. Recent methods have introduced guidance via diversity measures to encourage more varied generations. In this work, we extend the diversity measure-based approaches by proposing the *S*calable *P*rompt-*A*ware *R*eny *K*ernel *E*ntropy Diversity Guidance (*SPARKE*) method for prompt-aware diversity guidance. SPARKE utilizes conditional entropy for diversity guidance, which dynamically conditions diversity measurement on similar prompts and enables prompt-aware diversity control. While the entropy-based guidance approach enhances prompt-aware diversity, its reliance on the matrix-based entropy scores poses computational challenges in large-scale generation settings. To address this, we focus on the special case of \textit{Conditional latent RKE Score Guidance}, reducing entropy computation and gradient-based optimization complexity from the $\mathcal{O}(n^3)$ of general entropy measures to $\mathcal{O}(n)$. The reduced computational complexity allows for diversity-guided sampling over potentially thousands of generation rounds on different prompts. We numerically test the SPARKE method on several text-to-image diffusion models, demonstrating that the proposed method improves the prompt-aware diversity of the generated data without incurring significant computational costs. We release our code on the project page: [https://mjalali.github.io/SPARKE/](https://mjalali.github.io/SPARKE). Mohammad Jalali, Haoyu Lei, Amin Gohari, Farzan Farnia |
NeurIPS | 2 |
| 2025 | Min-product fuzzy relation inequalities with absent variables and their weighted max-min optimization in supply chain system
Haoyu Lei, Qianyu Shu, Xiaopeng Yang 0001 |
Fuzzy Sets Syst. | 1 |
| 2024 | On the Inductive Biases of Demographic Parity-based Fair Learning AlgorithmsabstractFair supervised learning algorithms assigning labels with little dependence on a sensitive attribute have attracted great attention in the machine learning community. While the demographic parity (DP) notion has been frequently used to measure a model’s fairness in training fair classifiers, several studies in the literature suggest potential impacts of enforcing DP in fair learning algorithms. In this work, we analytically study the effect of standard DP-based regularization methods on the conditional distribution of the predicted label given the sensitive attribute. Our analysis shows that an imbalanced training dataset with a non-uniform distribution of the sensitive attribute could lead to a classification rule biased toward the sensitive attribute outcome holding the majority of training data. To control such inductive biases in DP-based fair learning, we propose a sensitive attribute-based distributionally robust optimization (SA-DRO) method improving robustness against the marginal distribution of the sensitive attribute. Finally, we present several numerical results on the application of DP-based learning methods to standard centralized and distributed learning problems. The empirical findings support our theoretical results on the inductive biases in DP-based fair learning algorithms and the debiasing effects of the proposed SA-DRO method. The project code is available at [github.com/lh218/Fairness-IB.git](https://github.com/lh218/Fairness-IB.git). Haoyu Lei, Amin Gohari, Farzan Farnia |
UAI | 1 |
| 2021 | An Efficient Alternating Direction Method for Graph Learning from Smooth SignalsabstractWe consider the problem of identifying the graph topology from a set of smooth graph signals. A well-known approach to this problem is minimizing the Dirichlet energy accompanied with some Frobenius norm regularization. Recent works have incorporated the logarithmic barrier on the node degrees to improve the overall graph connectivity without compromising graph sparsity, which is shown to be quite effective in enhancing the quality of the learned graphs. Although a primal-dual algorithm has been proposed in the literature to solve this type of graph learning formulations, it lacks a rigorous convergence analysis and appears to have a slow empirical performance. In this paper, we cast the graph learning formulation as a nonsmooth, strictly convex optimization problem and develop an efficient alternating direction method of multipliers to solve it. We show that our algorithm converges to the global minimum with arbitrary initialization. We conduct extensive experiments on various synthetic and real-world graphs, the results of which show that our method exhibits sharp linear convergence and is substantially faster than the commonly adopted primal-dual method. Chaorui Yao, Haoyu Lei, Anthony Man-Cho So |
ICASSP | 3 |