EDBT 2026 Demo / reviewers in the wild / expert
Mingye Gao
dblp:294/9762
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Reinforcement learning · 41% Trustworthy machine learning · 29% Efficient and distributed learning · 22% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reward learning
reward modeling |
1.9 | 2 | 2026 | ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models · AAAI 2026 RuleAdapter: Dynamic Rules for training Safety Reward Models in RLHF · ICML 2025 |
Machine learning › Efficient and distributed learning
data selection |
1.0 | 1 | 2026 | Selection of LLM Fine-Tuning Data Based on Orthogonal Rules · AAAI 2026 |
Machine learning › Efficient and distributed learning › data selection
data selection for fine-tuning |
1.0 | 1 | 2026 | Selection of LLM Fine-Tuning Data Based on Orthogonal Rules · AAAI 2026 |
Machine learning › Reinforcement learning
reward design |
1.0 | 1 | 2026 | ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models · AAAI 2026 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | RuleAdapter: Dynamic Rules for training Safety Reward Models in RLHF · ICML 2025 |
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
0.9 | 1 | 2025 | RuleAdapter: Dynamic Rules for training Safety Reward Models in RLHF · ICML 2025 |
Machine learning › Trustworthy machine learning › fairness
demographic bias |
0.8 | 1 | 2024 | Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias · NeurIPS 2024 |
Machine learning › Trustworthy machine learning
fairness |
0.8 | 1 | 2024 | Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model |
0.8 | 1 | 2024 | Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › AI safety › safety alignment
LLM safety alignment |
0.3 | 1 | 2026 | ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
benchmark framework · 1.5alignment method · 1.5rule-based scoring · 1.0entropy-guided reward composition · 1.0determinantal point process · 1.0mutual information · 0.9maximum discrepancy · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
Xupeng Chen, Jingxuan Fan, Eric Hanchen Jiang, Mingye Gao |
AAAI | 5 |
| 2026 | Selection of LLM Fine-Tuning Data Based on Orthogonal RulesabstractHigh-quality training data is critical to the performance of large language models (LLMs). Recent work has explored using LLMs to rate and select data based on a small set of human-designed criteria (rules), but these approaches often rely heavily on heuristics, lack principled metrics for rule evaluation, and generalize poorly to new tasks. We propose a novel rule-based data selection framework that introduces a metric based on the orthogonality of rule score vectors to evaluate and select complementary rules. Our automated pipeline first uses LLMs to generate diverse rules covering multiple aspects of data quality, then rates samples according to these rules and applies the determinantal point process (DPP) to select the most independent rules. These rules are then used to score the full dataset, and high-scoring samples are selected for downstream tasks such as LLM fine-tuning. We evaluate our framework in two experiment setups: (1) alignment with ground-truth ratings and (2) performance of LLMs fine-tuned on the selected data. Experiments across IMDB, Medical, Math, and Code domains demonstrate that our DPP-based rule selection consistently improves both rating accuracy and downstream model performance over strong baselines. Mingye Gao, Chang Yue |
AAAI | 2 |
| 2026 | Heart Rate Monitoring Using Continuous-Wave Radar in Home EnvironmentabstractRadar-based, contactless in-home vital sign monitoring can significantly improve the healthcare of the rapidly growing aging population by enabling early detection of clinical events and continuous tracking of physiological states, while maintaining privacy, comfort, and user compliance. We introduce a unique experimental environment, collected for over a year – consisting of populated test homes, equipped with arrays of continuous-wave radar sensors placed on ceilings and walls, with occupants engaged in everyday activities within the homes. We applied and compared multiple analytical methods to extract heart rate from the radar data, using combinations of normalized cross-correlation (NXC), machine learning with individual sensors, and sensor fusion. Our results show that sensor fusion leads to the lowest mean absolute errors, reduced by over 30% from single-sensor machine learning models and over 80% from NXC. We additionally observed that a dense sensor array design enhances signal quality and reduces noise by more than 15%, providing more accurate heart rate monitoring even during dynamic activities, such as jumping and walking. These results highlight the effectiveness of sensor fusion and dense array designs in real-world settings. Our novel experimental platform and analytical results represent a meaningful step toward scalable application of radar technology in real residential settings for non-contact health monitoring. Inbar Chityat, Mingye Gao, Xiang Zhang 0038, Daniel Copeland, Buntoku Mori, Mina Okitsu, Brian Anthony 0001 |
IEEE Internet Things J. | 2 |
| 2025 | RuleAdapter: Dynamic Rules for training Safety Reward Models in RLHFabstractReinforcement Learning from Human Feedback (RLHF) is widely used to align models with human preferences, particularly to enhance the safety of responses generated by LLMs. This method traditionally relies on choosing preferred responses from response pairs. However, due to variations in human opinions and the difficulty of making an overall comparison of two responses, there is a growing shift towards a fine-grained annotation approach, assessing responses based on multiple specific metrics or rules. Selecting and applying these rules efficiently while accommodating the diversity of preference data remains a significant challenge. In this paper, we introduce a dynamic approach that adaptively selects the most critical rules for each pair of responses. We develop a mathematical framework that leverages the maximum discrepancy between each paired responses and theoretically show that this strategy optimizes the mutual information between the rule-based labeling and the hidden ground-truth preferences. We then train an 8B reward model using the adaptively labeled preference dataset and evaluate its performance on RewardBench. As of May 25, 2025, our model achieved the highest safety performance on the leaderboard, outperforming various larger models. Mingye Gao, Jingxuan Fan |
ICML | 2 |
| 2024 | Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model BiasabstractLarge language models (LLMs) are increasingly essential in processing natural languages, yet their application is frequently compromised by biases and inaccuracies originating in their training data.In this study, we introduce \textbf{Cross-Care}, the first benchmark framework dedicated to assessing biases and real world knowledge in LLMs, specifically focusing on the representation of disease prevalence across diverse demographic groups.We systematically evaluate how demographic biases embedded in pre-training corpora like $ThePile$ influence the outputs of LLMs.We expose and quantify discrepancies by juxtaposing these biases against actual disease prevalences in various U.S. demographic groups.Our results highlight substantial misalignment between LLM representation of disease prevalence and real disease prevalence rates across demographic subgroups, indicating a pronounced risk of bias propagation and a lack of real-world grounding for medical applications of LLMs.Furthermore, we observe that various alignment methods minimally resolve inconsistencies in the models' representation of disease prevalence across different languages.For further exploration and analysis, we make all data and a data visualization tool available at: \url{www.crosscare.net}. Shan Chen 0004, Jack Gallifant, Mingye Gao, Nikolaj Munch, Ajay Muthukkumar, Arvind Rajan, Jaya Kolluri, Amelia Fiske, Janna Hastings, Hugo J. W. L. Aerts, Brian Anthony 0001, Leo A. Celi, William G. La Cava, Danielle S. Bitterman |
NeurIPS | 3 |
| 2022 | Cooperative Self-training of Machine Reading ComprehensionabstractHongyin Luo, Shang-Wen Li, Mingye Gao, Seunghak Yu, James Glass. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Hongyin Luo, Shang-Wen Li 0001, Mingye Gao, Seunghak Yu, James R. Glass |
NAACL-HLT | 3 |