Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jiawei Ge 0003

dblp:250/3739-3 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 36% Language models and text generation · 30% Learning theory · 17%
Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 80% Approximation and online algorithms · 20%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › evaluation of language models
benchmark construction
0.912025
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations · ICML 2025
Natural language and speech › Language models and text generation
mathematical reasoning
0.912025
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations · ICML 2025
Natural language and speech › Language models and text generation › evaluation of language models › reasoning evaluation
mathematical reasoning benchmark
0.912025
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations · ICML 2025
Machine learning › Trustworthy machine learning › robustness
robustness to perturbation
0.912025
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations · ICML 2025
Algorithmic game theory and mechanism design
equilibrium computation
0.912025
Securing Equal Share: A Principled Approach for Learning Multiplayer Symmetric Games · ICML 2025
Algorithmic game theory and mechanism design
multi-player games
0.912025
Securing Equal Share: A Principled Approach for Learning Multiplayer Symmetric Games · ICML 2025
Algorithmic game theory and mechanism design › regret minimization
no-regret learning in games
0.912025
Securing Equal Share: A Principled Approach for Learning Multiplayer Symmetric Games · ICML 2025
Approximation and online algorithms
online learning
0.912025
Securing Equal Share: A Principled Approach for Learning Multiplayer Symmetric Games · ICML 2025
Algorithmic game theory and mechanism design
regret minimization
0.912025
Securing Equal Share: A Principled Approach for Learning Multiplayer Symmetric Games · ICML 2025
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift
0.812024
Maximum Likelihood Estimation is All You Need for Well-Specified Covariate Shift · ICLR 2024
Machine learning › Trustworthy machine learning › robustness
distribution shift
0.812024
Optimal Aggregation of Prediction Intervals under Unsupervised Domain Shift · NeurIPS 2024
Machine learning › Learning theory
minimax optimality
0.812024
Maximum Likelihood Estimation is All You Need for Well-Specified Covariate Shift · ICLR 2024
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.812024
Maximum Likelihood Estimation is All You Need for Well-Specified Covariate Shift · ICLR 2024
Machine learning › Trustworthy machine learning
uncertainty estimation
0.812024
Optimal Aggregation of Prediction Intervals under Unsupervised Domain Shift · NeurIPS 2024
Machine learning › Representation and self-supervised learning › pre-training
unsupervised pre-training
0.812024
On the Provable Advantage of Unsupervised Pretraining · ICLR 2024

Methods — techniques the papers use, named apart from their topics

maximum likelihood estimation · 1.5no-regret learning · 0.9lower bound analysis · 0.9in-context learning · 0.9hard perturbation · 0.9measure-preserving transformation · 0.8maximum weighted likelihood estimation · 0.8finite-sample bounds · 0.8empirical risk minimization · 0.8density ratio · 0.8
YearPublicationVenuePosition
2025 Securing Equal Share: A Principled Approach for Learning Multiplayer Symmetric Games
abstract
This paper examines multiplayer symmetric constant-sum games with more than two players in a competitive setting, such as Mahjong, Poker, and various board and video games. In contrast to two-player zero-sum games, equilibria in multiplayer games are neither unique nor non-exploitable, failing to provide meaningful guarantees when competing against opponents who play different equilibria or non-equilibrium strategies. This gives rise to a series of long-lasting fundamental questions in multiplayer games regarding suitable objectives, solution concepts, and principled algorithms. This paper takes an initial step towards addressing these challenges by focusing on the natural objective of equal share—securing an expected payoff of $C/n$ in an $n$-player symmetric game with a total payoff of $C$. We rigorously identify the theoretical conditions under which achieving an equal share is tractable and design a series of efficient algorithms, inspired by no-regret learning, that provably attain approximate equal share across various settings. Furthermore, we provide complementary lower bounds that justify the sharpness of our theoretical results. Our experimental results highlight worst-case scenarios where meta-algorithms from prior state-of-the-art systems for multiplayer games fail to secure an equal share, while our algorithm succeeds, demonstrating the effectiveness of our approach.
Jiawei Ge 0003, Yuanhao Wang 0001, Chi Jin 0001
ICML1
2025 MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
abstract
Large language models have demonstrated impressive performance on challenging mathematical reasoning tasks, which has triggered the discussion of whether the performance is achieved by true reasoning capability or memorization. To investigate this question, prior work has constructed mathematical benchmarks when questions undergo simple perturbations – modifications that still preserve the underlying reasoning patterns of the solutions. However, no work has explored hard perturbations, which fundamentally change the nature of the problem so that the original solution steps do not apply. To bridge the gap, we construct MATH-P-Simple and MATH-P-Hard via simple perturbation and hard perturbation, respectively. Each consists of 279 perturbed math problems derived from level-5 (hardest) problems in the MATH dataset (Hendrycks et al., 2021). We observe significant performance drops on MATH-P-Hard across various models, including o1-mini (-16.49%) and gemini-2.0-flash-thinking (-12.9%). We also raise concerns about a novel form of memorization where models blindly apply learned problem-solving skills without assessing their applicability to modified contexts. This issue is amplified when using original problems for in-context learning. We call for research efforts to address this challenge, which is critical for developing more robust and reliable reasoning models. The project is available at https://math-perturb.github.io/.
Kaixuan Huang, Jiacheng Guo, Jiawei Ge 0003, Tianle Cai, Hui Yuan 0002, Runzhe Wang, Ming Yin 0003, Shange Tang, Yangsibo Huang, Chi Jin 0001, Chiyuan Zhang, Mengdi Wang 0001
ICML5
2024 On the Provable Advantage of Unsupervised Pretraining
abstract
Unsupervised pretraining, which learns a useful representation using a large amount of unlabeled data to facilitate the learning of downstream tasks, is a critical component of modern large-scale machine learning systems. Despite its tremendous empirical success, the rigorous theoretical understanding of why unsupervised pretraining generally helps remains rather limited---most existing results are restricted to particular methods or approaches for unsupervised pretraining with specialized structural assumptions. This paper studies a generic framework, where the unsupervised representation learning task is specified by an abstract class of latent variable models $\Phi$ and the downstream task is specified by a class of prediction functions $\Psi$. We consider a natural approach of using Maximum Likelihood Estimation (MLE) for unsupervised pretraining and Empirical Risk Minimization (ERM) for learning downstream tasks. We prove that, under a mild ``informative'' condition, our algorithm achieves an excess risk of $\\tilde{\\mathcal{O}}(\sqrt{\mathcal{C}\_\Phi/m} + \sqrt{\mathcal{C}\_\Psi/n})$ for downstream tasks, where $\mathcal{C}\_\Phi, \mathcal{C}\_\Psi$ are complexity measures of function classes $\Phi, \Psi$, and $m, n$ are the number of unlabeled and labeled data respectively. Comparing to the baseline of $\tilde{\mathcal{O}}(\sqrt{\mathcal{C}\_{\Phi \circ \Psi}/n})$ achieved by performing supervised learning using only the labeled data, our result rigorously shows the benefit of unsupervised pretraining when $m \gg n$ and $\mathcal{C}\_{\Phi\circ \Psi} > \mathcal{C}\_\Psi$. This paper further shows that our generic framework covers a wide range of approaches for unsupervised pretraining, including factor models, Gaussian mixture models, and contrastive learning.
Jiawei Ge 0003, Shange Tang, Jianqing Fan, Chi Jin 0001
ICLR1
2024 Maximum Likelihood Estimation is All You Need for Well-Specified Covariate Shift
abstract
A key challenge of modern machine learning systems is to achieve Out-of-Distribution (OOD) generalization---generalizing to target data whose distribution differs from that of source data. Despite its significant importance, the fundamental question of ``what are the most effective algorithms for OOD generalization'' remains open even under the standard setting of covariate shift. This paper addresses this fundamental question by proving that, surprisingly, classical Maximum Likelihood Estimation (MLE) purely using source data (without any modification) achieves the *minimax* optimality for covariate shift under the *well-specified* setting. That is, *no* algorithm performs better than MLE in this setting (up to a constant factor), justifying MLE is all you need. Our result holds for a very rich class of parametric models, and does not require any boundedness condition on the density ratio. We illustrate the wide applicability of our framework by instantiating it to three concrete examples---linear regression, logistic regression, and phase retrieval. This paper further complement the study by proving that, under the *misspecified setting*, MLE is no longer the optimal choice, whereas Maximum Weighted Likelihood Estimator (MWLE) emerges as minimax optimal in certain scenarios.
Jiawei Ge 0003, Shange Tang, Jianqing Fan, Cong Ma 0001, Chi Jin 0001
ICLR1
2024 Optimal Aggregation of Prediction Intervals under Unsupervised Domain Shift
abstract
As machine learning models are increasingly deployed in dynamic environments, it becomes paramount to assess and quantify uncertainties associated with distribution shifts. A distribution shift occurs when the underlying data-generating process changes, leading to a deviation in the model's performance. The prediction interval, which captures the range of likely outcomes for a given prediction, serves as a crucial tool for characterizing uncertainties induced by their underlying distribution. In this paper, we propose methodologies for aggregating prediction intervals to obtain one with minimal width and adequate coverage on the target domain under unsupervised domain shift, under which we have labeled samples from a related source domain and unlabeled covariates from the target domain. Our analysis encompasses scenarios where the source and the target domain are related via i) a bounded density ratio, and ii) a measure-preserving transformation. Our proposed methodologies are computationally efficient and easy to implement. Beyond illustrating the performance of our method through real-world datasets, we also delve into the theoretical details. This includes establishing rigorous theoretical guarantees, coupled with finite sample bounds, regarding the coverage and width of our prediction intervals. Our approach excels in practical applications and is underpinned by a solid theoretical framework, ensuring its reliability and effectiveness across diverse contexts.
Jiawei Ge 0003, Debarghya Mukherjee, Jianqing Fan
NeurIPS1