Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Mingye Gao

dblp:294/9762 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 41% Trustworthy machine learning · 29% Efficient and distributed learning · 22%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › reward learning
reward modeling
1.922026
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models · AAAI 2026
RuleAdapter: Dynamic Rules for training Safety Reward Models in RLHF · ICML 2025
Machine learning › Efficient and distributed learning
data selection
1.012026
Selection of LLM Fine-Tuning Data Based on Orthogonal Rules · AAAI 2026
Machine learning › Efficient and distributed learning › data selection
data selection for fine-tuning
1.012026
Selection of LLM Fine-Tuning Data Based on Orthogonal Rules · AAAI 2026
Machine learning › Reinforcement learning
reward design
1.012026
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models · AAAI 2026
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
RuleAdapter: Dynamic Rules for training Safety Reward Models in RLHF · ICML 2025
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
RuleAdapter: Dynamic Rules for training Safety Reward Models in RLHF · ICML 2025
Machine learning › Trustworthy machine learning › fairness
demographic bias
0.812024
Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias · NeurIPS 2024
Machine learning › Trustworthy machine learning
fairness
0.812024
Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias · NeurIPS 2024
Natural language and speech › Language models and text generation
large language model
0.812024
Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias · NeurIPS 2024
Machine learning › Trustworthy machine learning › AI safety › safety alignment
LLM safety alignment
0.312026
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models · AAAI 2026

Methods — techniques the papers use, named apart from their topics

benchmark framework · 1.5alignment method · 1.5rule-based scoring · 1.0entropy-guided reward composition · 1.0determinantal point process · 1.0mutual information · 0.9maximum discrepancy · 0.9
YearPublicationVenuePosition
2026 ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
Xupeng Chen, Jingxuan Fan, Eric Hanchen Jiang, Mingye Gao
AAAI5
2026 Selection of LLM Fine-Tuning Data Based on Orthogonal Rules
abstract
High-quality training data is critical to the performance of large language models (LLMs). Recent work has explored using LLMs to rate and select data based on a small set of human-designed criteria (rules), but these approaches often rely heavily on heuristics, lack principled metrics for rule evaluation, and generalize poorly to new tasks. We propose a novel rule-based data selection framework that introduces a metric based on the orthogonality of rule score vectors to evaluate and select complementary rules. Our automated pipeline first uses LLMs to generate diverse rules covering multiple aspects of data quality, then rates samples according to these rules and applies the determinantal point process (DPP) to select the most independent rules. These rules are then used to score the full dataset, and high-scoring samples are selected for downstream tasks such as LLM fine-tuning. We evaluate our framework in two experiment setups: (1) alignment with ground-truth ratings and (2) performance of LLMs fine-tuned on the selected data. Experiments across IMDB, Medical, Math, and Code domains demonstrate that our DPP-based rule selection consistently improves both rating accuracy and downstream model performance over strong baselines.
Mingye Gao, Chang Yue
AAAI2
2026 Heart Rate Monitoring Using Continuous-Wave Radar in Home Environment
abstract
Radar-based, contactless in-home vital sign monitoring can significantly improve the healthcare of the rapidly growing aging population by enabling early detection of clinical events and continuous tracking of physiological states, while maintaining privacy, comfort, and user compliance. We introduce a unique experimental environment, collected for over a year – consisting of populated test homes, equipped with arrays of continuous-wave radar sensors placed on ceilings and walls, with occupants engaged in everyday activities within the homes. We applied and compared multiple analytical methods to extract heart rate from the radar data, using combinations of normalized cross-correlation (NXC), machine learning with individual sensors, and sensor fusion. Our results show that sensor fusion leads to the lowest mean absolute errors, reduced by over 30% from single-sensor machine learning models and over 80% from NXC. We additionally observed that a dense sensor array design enhances signal quality and reduces noise by more than 15%, providing more accurate heart rate monitoring even during dynamic activities, such as jumping and walking. These results highlight the effectiveness of sensor fusion and dense array designs in real-world settings. Our novel experimental platform and analytical results represent a meaningful step toward scalable application of radar technology in real residential settings for non-contact health monitoring.
Inbar Chityat, Mingye Gao, Xiang Zhang 0038, Daniel Copeland, Buntoku Mori, Mina Okitsu, Brian Anthony 0001
IEEE Internet Things J.2
2025 RuleAdapter: Dynamic Rules for training Safety Reward Models in RLHF
abstract
Reinforcement Learning from Human Feedback (RLHF) is widely used to align models with human preferences, particularly to enhance the safety of responses generated by LLMs. This method traditionally relies on choosing preferred responses from response pairs. However, due to variations in human opinions and the difficulty of making an overall comparison of two responses, there is a growing shift towards a fine-grained annotation approach, assessing responses based on multiple specific metrics or rules. Selecting and applying these rules efficiently while accommodating the diversity of preference data remains a significant challenge. In this paper, we introduce a dynamic approach that adaptively selects the most critical rules for each pair of responses. We develop a mathematical framework that leverages the maximum discrepancy between each paired responses and theoretically show that this strategy optimizes the mutual information between the rule-based labeling and the hidden ground-truth preferences. We then train an 8B reward model using the adaptively labeled preference dataset and evaluate its performance on RewardBench. As of May 25, 2025, our model achieved the highest safety performance on the leaderboard, outperforming various larger models.
Mingye Gao, Jingxuan Fan
ICML2
2024 Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias
abstract
Large language models (LLMs) are increasingly essential in processing natural languages, yet their application is frequently compromised by biases and inaccuracies originating in their training data.In this study, we introduce \textbf{Cross-Care}, the first benchmark framework dedicated to assessing biases and real world knowledge in LLMs, specifically focusing on the representation of disease prevalence across diverse demographic groups.We systematically evaluate how demographic biases embedded in pre-training corpora like $ThePile$ influence the outputs of LLMs.We expose and quantify discrepancies by juxtaposing these biases against actual disease prevalences in various U.S. demographic groups.Our results highlight substantial misalignment between LLM representation of disease prevalence and real disease prevalence rates across demographic subgroups, indicating a pronounced risk of bias propagation and a lack of real-world grounding for medical applications of LLMs.Furthermore, we observe that various alignment methods minimally resolve inconsistencies in the models' representation of disease prevalence across different languages.For further exploration and analysis, we make all data and a data visualization tool available at: \url{www.crosscare.net}.
Shan Chen 0004, Jack Gallifant, Mingye Gao, Nikolaj Munch, Ajay Muthukkumar, Arvind Rajan, Jaya Kolluri, Amelia Fiske, Janna Hastings, Hugo J. W. L. Aerts, Brian Anthony 0001, Leo A. Celi, William G. La Cava, Danielle S. Bitterman
NeurIPS3
2022 Cooperative Self-training of Machine Reading Comprehension
abstract
Hongyin Luo, Shang-Wen Li, Mingye Gao, Seunghak Yu, James Glass. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Hongyin Luo, Shang-Wen Li 0001, Mingye Gao, Seunghak Yu, James R. Glass
NAACL-HLT3