EDBT 2026 Demo / reviewers in the wild / expert
Mingye Zhu
dblp:272/5830
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Language models and text generation · 68% Trustworthy machine learning · 16% Optimization for machine learning · 7% |
Topics — the 12 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › alignment
preference alignment |
2.5 | 3 | 2025 | Leveraging robust optimization for llm alignment under distribution shifts · NeurIPS 2025 On-the-fly Preference Alignment via Principle-Guided Decoding · ICLR 2025 FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization · EMNLP 2024 |
Natural language and speech › Language models and text generation
alignment |
2.0 | 3 | 2025 | Leveraging robust optimization for llm alignment under distribution shifts · NeurIPS 2025 Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models · NeurIPS 2025 On-the-fly Preference Alignment via Principle-Guided Decoding · ICLR 2025 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
1.0 | 1 | 2026 | In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback · AAAI 2026 |
Natural language and speech › Language models and text generation
decoding |
0.9 | 1 | 2025 | On-the-fly Preference Alignment via Principle-Guided Decoding · ICLR 2025 |
Natural language and speech › Language models and text generation › alignment
inference-time alignment |
0.9 | 1 | 2025 | On-the-fly Preference Alignment via Principle-Guided Decoding · ICLR 2025 |
Machine learning › Trustworthy machine learning
robustness |
0.9 | 1 | 2025 | Leveraging robust optimization for llm alignment under distribution shifts · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › robustness › distribution shift
robustness to distribution shift |
0.9 | 1 | 2025 | Leveraging robust optimization for llm alignment under distribution shifts · NeurIPS 2025 |
Machine learning › Optimization for machine learning
constrained optimization |
0.8 | 1 | 2024 | FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization · EMNLP 2024 |
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards |
0.3 | 1 | 2026 | In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-Feedback · AAAI 2026 |
Computer vision › Segmentation and scene understanding › semantic segmentation › remote sensing image segmentation
aerial image segmentation |
0.2 | 1 | 2023 | Disentangling the Benefits of Self-Supervised Learning to Deployment-Driven Downstream Tasks of Satellite Images (Student Abstract) · AAAI 2023 |
Computer vision › Image recognition and object detection › image classification › remote sensing image classification
satellite imagery classification |
0.2 | 1 | 2023 | Disentangling the Benefits of Self-Supervised Learning to Deployment-Driven Downstream Tasks of Satellite Images (Student Abstract) · AAAI 2023 |
Computer vision › Image recognition and object detection
scene recognition |
0.2 | 1 | 2023 | Disentangling the Benefits of Self-Supervised Learning to Deployment-Driven Downstream Tasks of Satellite Images (Student Abstract) · AAAI 2023 |
Methods — techniques the papers use, named apart from their topics
constrained optimization · 1.6self-feedback · 1.0in-token optimization · 1.0importance weighting · 1.0reward shaping · 0.9resampling algorithm · 0.9importance sampling · 0.9distribution-aware optimization · 0.9classifier-based calibration · 0.9autoregressive alignment module · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | In-Token Rationality Optimization: Towards Accurate and Concise LLM Reasoning via Self-FeedbackabstractTraining Large Language Models (LLMs) for chain-of-thought reasoning presents a significant challenge: supervised fine-tuning on a single "golden" rationale hurts generalization as it penalizes equally valid alternatives, whereas reinforcement learning with verifiable rewards struggles with credit assignment and prohibitive computational cost. To tackle these limitations, we introduce InTRO (In-Token Rationality Optimization), a new framework that enables both token-level exploration and self-feedback for accurate and concise reasoning. Instead of directly optimizing an intractable objective over all valid reasoning paths, InTRO leverages correction factors—token-wise importance weights estimated by the information discrepancy between the generative policy and its answer-conditioned counterpart, for informative next-token selection. This approach allows the model to perform token-level exploration and receive self-generated feedback within a single forward pass, ultimately encouraging accurate and concise rationales. Across six math-reasoning benchmarks, InTRO consistently outperforms other baselines, raising solution accuracy by up to 20% relative to the base model. Its chains of thought are also notably more concise, exhibiting reduced verbosity. Beyond this, InTRO enables cross-domain transfer, successfully adapting to out-of-domain reasoning tasks that extend beyond the realm of mathematics, demonstrating robust generalization. Mingye Zhu, Yi Liu 0148, Zheren Fu, Quan Wang 0002, Yongdong Zhang 0001 |
AAAI | 1 |
| 2025 | On-the-fly Preference Alignment via Principle-Guided DecodingabstractWith the rapidly expanding landscape of large language models, aligning model generations with human values and preferences is becoming increasingly important. Popular alignment methods, such as Reinforcement Learning from Human Feedback, have shown significant success in guiding models with greater control. However, these methods require considerable computational resources, which is inefficient, and substantial collection of training data to accommodate the diverse and pluralistic nature of human preferences, which is impractical. These limitations significantly constrain the scope and efficacy of both task-specific and general preference alignment methods. In this work, we introduce On-the-fly Preference Alignment via Principle-Guided Decoding (OPAD) to directly align
model outputs with human preferences during inference, eliminating the need for fine-tuning. Our approach involves first curating a surrogate solution to an otherwise infeasible optimization problem and then designing a principle-guided reward function based on this surrogate. The final decoding policy is derived by maximizing this customized reward, which exploits the discrepancy between the
constrained policy and its unconstrained counterpart. OPAD directly modifies the model’s predictions during inference, ensuring principle adherence without incurring the computational overhead of retraining or fine-tuning. Experiments show that OPAD achieves competitive or superior performance in both general and personalized alignment tasks, demonstrating its efficiency and effectiveness compared to state-of-the-art baselines. Mingye Zhu, Yi Liu 0148, Lei Zhang 0119, Junbo Guo, Zhendong Mao 0001 |
ICLR | 1 |
| 2025 | Leveraging Importance Sampling to Detach Alignment Modules from Large Language ModelsabstractThe widespread adoption of large language models (LLMs) across industries has increased the demand for high-quality and customizable outputs. However, traditional alignment methods often require retraining large pretrained models, making it difficult to quickly adapt and optimize LLMs for diverse applications. To address this limitation, we propose a novel \textit{Residual Alignment Model} (\textit{RAM}) that formalizes the alignment process as a type of importance sampling. In this framework, the unaligned upstream model serves as the proposal distribution, while the alignment process is framed as secondary sampling based on an autoregressive alignment module that acts as an estimator of the importance weights. This design enables a natural detachment of the alignment module from the target aligned model, improving flexibility and scalability. Based on this model, we derive an efficient sequence-level training strategy for the alignment module, which operates independently of the proposal module. Additionally, we develop a resampling algorithm with iterative token-level decoding to address the common first-token latency issue in comparable methods. Experimental evaluations on two leading open-source LLMs across diverse tasks, including instruction following, domain adaptation, and preference optimization, demonstrate that our approach consistently outperforms baseline models. Yi Liu 0148, Dianqing Liu, Mingye Zhu, Junbo Guo, Yongdong Zhang 0001, Zhendong Mao 0001 |
NeurIPS | 3 |
| 2025 | Leveraging robust optimization for llm alignment under distribution shiftsabstractPreference alignment methods are increasingly critical for steering large language models (LLMs) to generate outputs consistent with human values. While recent approaches often rely on synthetic data generated by LLMs for scalability and cost-efficiency reasons, this reliance can introduce distributional shifts that undermine the nuanced representation of human preferences needed for desirable outputs. In this paper, we propose a novel distribution-aware optimization framework that improves preference alignment despite such shifts. Our approach first leverages well-learned classifiers to assign a calibration value to each training sample, quantifying its alignment with the target human-preferred distribution. These values are then incorporated into a robust optimization objective that minimizes the worst-case loss over regions of the data space most relevant to human preferences. By explicitly focusing optimization on the target distribution, our approach mitigates the impact of distributional mismatch and improves the generation of responses that better reflect intended values. Mingye Zhu, Yi Liu 0148, Zheren Fu, Yongdong Zhang 0001, Zhendong Mao 0001 |
NeurIPS | 1 |
| 2024 | FlipGuard: Defending Preference Alignment against Update Regression with Constrained OptimizationabstractRecent breakthroughs in preference alignment have significantly improved Large Language Models' ability to generate texts that align with human preferences and values.However, current alignment metrics typically emphasize the post-hoc overall improvement, while overlooking a critical aspect: regression, which refers to the backsliding on previously correctly-handled data after updates.This potential pitfall may arise from excessive fine-tuning on already well-aligned data, which subsequently leads to over-alignment and degeneration.To address this challenge, we propose FlipGuard, a constrained optimization approach to detect and mitigate update regression with focal attention.Specifically, FlipGuard identifies performance degradation using a customized reward characterization and strategically enforces a constraint to encourage conditional congruence with the pre-aligned model during training.Comprehensive experiments demonstrate that FlipGuard effectively alleviates update regression while demonstrating excellent overall performance, with the added benefit of knowledge preservation while aligning preferences. Mingye Zhu, Yi Liu 0148, Quan Wang 0002, Junbo Guo, Zhendong Mao 0001 |
EMNLP | 1 |
| 2023 | Disentangling the Benefits of Self-Supervised Learning to Deployment-Driven Downstream Tasks of Satellite Images (Student Abstract)abstractIn this paper, we investigate the benefits of self-supervised learning (SSL) to downstream tasks of satellite images. Unlike common student academic projects, this work focuses on the advantages of the SSL for deployment-driven tasks which have specific scenarios with low or high-spatial resolution images. Our preliminary experiments demonstrate the robust benefits of the SSL trained by medium-resolution (10m) images to both low-resolution (100m) scene classification case (4.25%↑) and very high-resolution (5cm) aerial image segmentation case (1.96%↑), respectively. Zhuo Deng 0001, Yibing Wei, Mingye Zhu, Junchi Zhou, Zhenjie Cao, Jui-Hsin Lai |
AAAI | 3 |
| 2021 | Leveraging probabilistic circuits for nonparametric multi-output regressionabstractInspired by recent advances in the field of expert-based approximations of Gaussian processes (GPs), we present an expert-based approach to large-scale multi-output regression using single-output GP experts. Employing a deeply structured mixture of single-output GPs encoded via a probabilistic circuit allows us to capture correlations between multiple output dimensions accurately. By recursively partitioning the covariate space and the output space, posterior inference in our model reduces to inference on single-output GP experts, which only need to be conditioned on a small subset of the observations. We show that inference can be performed exactly and efficiently in our model, that it can capture correlations between output dimensions and, hence, often outperforms approaches that do not incorporate inter-output correlations, as demonstrated on several data sets in terms of the negative log predictive density. Zhongjie Yu 0001, Mingye Zhu, Martin Trapp 0001, Arseny Skryagin, Kristian Kersting |
UAI | 2 |