VLDB 2026 Research / reviewers in the wild / expert
Andre Wang He
dblp:318/3206 · also Andre He
· DBLP profile ↗
3ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 40% Information extraction and text analysis · 30% Machine translation · 30% | |
| Theoretical computer science
1 paper |
Automated reasoning and model checking · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
historical linguistics |
0.7 | 1 | 2023 | Neural Unsupervised Reconstruction of Protolanguage Word Forms · ACL (1) 2023 |
Natural language and speech › Machine translation
unsupervised machine translation |
0.7 | 1 | 2023 | Neural Unsupervised Reconstruction of Protolanguage Word Forms · ACL (1) 2023 |
Automated reasoning and model checking › theorem proving
formal theorem proving |
0.3 | 1 | 2025 | Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
unlikeliness reward · 1.7pass@n evaluation · 1.7GRPO · 1.7neural model · 0.7monotonic alignment · 0.7expectation-maximization · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rewarding the Unlikely: Lifting GRPO Beyond Distribution SharpeningabstractReinforcement learning is emerging as a primary driver for improving language model reasoning capabilities.A fundamental question is whether current reinforcement learning algorithms-such as Group Relative Policy Optimization (GRPO), the de facto standard algorithm used to improve language model reasoning-merely sharpen the base model's distribution around problems it can already solve.We investigate this question in the context of formal theorem proving, which has access to a perfect verifier.We identify a degenerate rank bias in GRPO in which highly probable trajectories are reinforced and rare ones are neglected.This results in distribution sharpening: the model can solve some problems with fewer samples, but underperforms simply sampling more solutions from the original model.To overcome GRPO's rank bias we introduce unlikeliness reward, a simple method for explicitly up-weighting rare but correct solutions.We show that unlikeliness reward mitigates rank bias and improves pass@N across a large range of N in both synthetic and real theorem proving settings.We also uncover an unexpected link between rank bias and a seemingly mundane hyperparameterthe number of updates per batch-that leads to a second, complementary mitigation.We combine our insights into a revised GRPO training recipe for formal theorem proving, yielding an open pipeline that achieves competitive performance to DeepSeek-Prover-V1.5-RL on the miniF2F-test benchmark.We release our implementation at https://github.com/ AndreHe02/rewarding-unlikely-release. Andre Wang He, Daniel Fried, Sean Welleck |
EMNLP | 1 |
| 2023 | Neural Unsupervised Reconstruction of Protolanguage Word FormsabstractWe present a state-of-the-art neural approach to the unsupervised reconstruction of ancient word forms.Previous work in this domain used expectation-maximization to predict simple phonological changes between ancient word forms and their cognates in modern languages.We extend this work with neural models that can capture more complicated phonological and morphological changes.At the same time, we preserve the inductive biases from classical methods by building monotonic alignment constraints into the model and deliberately underfitting during the maximization step.We evaluate our performance on the task of reconstructing Latin from a dataset of cognates across five Romance languages, achieving a notable reduction in edit distance from the target word forms compared to previous methods. Andre Wang He, Nicholas Tomlin, Daniel Klein 0001 |
ACL (1) | 1 |
| 2020 | Improve Unseen Domain Generalization via Enhanced Local Color Transformation
Jianhao Xiong, Andre Wang He, Congxin Liu, ZongYuan Ge |
MICCAI (2) | 2 |