Andre Wang He

dblp:318/3206 · also Andre He · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 40% Information extraction and text analysis · 30% Machine translation · 30%
Theoretical computer science
1 paper
Automated reasoning and model checking · 100%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
historical linguistics
0.712023
Neural Unsupervised Reconstruction of Protolanguage Word Forms · ACL (1) 2023
Natural language and speech › Machine translation
unsupervised machine translation
0.712023
Neural Unsupervised Reconstruction of Protolanguage Word Forms · ACL (1) 2023
Automated reasoning and model checking › theorem proving
formal theorem proving
0.312025
Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

unlikeliness reward · 1.7pass@n evaluation · 1.7GRPO · 1.7neural model · 0.7monotonic alignment · 0.7expectation-maximization · 0.7
YearPublicationVenuePosition
2025 Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
abstract
Reinforcement learning is emerging as a primary driver for improving language model reasoning capabilities.A fundamental question is whether current reinforcement learning algorithms-such as Group Relative Policy Optimization (GRPO), the de facto standard algorithm used to improve language model reasoning-merely sharpen the base model's distribution around problems it can already solve.We investigate this question in the context of formal theorem proving, which has access to a perfect verifier.We identify a degenerate rank bias in GRPO in which highly probable trajectories are reinforced and rare ones are neglected.This results in distribution sharpening: the model can solve some problems with fewer samples, but underperforms simply sampling more solutions from the original model.To overcome GRPO's rank bias we introduce unlikeliness reward, a simple method for explicitly up-weighting rare but correct solutions.We show that unlikeliness reward mitigates rank bias and improves pass@N across a large range of N in both synthetic and real theorem proving settings.We also uncover an unexpected link between rank bias and a seemingly mundane hyperparameterthe number of updates per batch-that leads to a second, complementary mitigation.We combine our insights into a revised GRPO training recipe for formal theorem proving, yielding an open pipeline that achieves competitive performance to DeepSeek-Prover-V1.5-RL on the miniF2F-test benchmark.We release our implementation at https://github.com/ AndreHe02/rewarding-unlikely-release.
Andre Wang He, Daniel Fried, Sean Welleck
EMNLP1
2023 Neural Unsupervised Reconstruction of Protolanguage Word Forms
abstract
We present a state-of-the-art neural approach to the unsupervised reconstruction of ancient word forms.Previous work in this domain used expectation-maximization to predict simple phonological changes between ancient word forms and their cognates in modern languages.We extend this work with neural models that can capture more complicated phonological and morphological changes.At the same time, we preserve the inductive biases from classical methods by building monotonic alignment constraints into the model and deliberately underfitting during the maximization step.We evaluate our performance on the task of reconstructing Latin from a dataset of cognates across five Romance languages, achieving a notable reduction in edit distance from the target word forms compared to previous methods.
Andre Wang He, Nicholas Tomlin, Daniel Klein 0001
ACL (1)1
2020 Improve Unseen Domain Generalization via Enhanced Local Color Transformation
Jianhao Xiong, Andre Wang He, Congxin Liu, ZongYuan Ge
MICCAI (2)2