EDBT 2026 Demo / reviewers in the wild / expert
Shange Tang
dblp:255/5774
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Learning theory · 28% Language models and text generation · 22% Trustworthy machine learning · 22% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
out-of-distribution generalization |
1.6 | 2 | 2025 | Benign Overfitting in Out-of-Distribution Generalization of Linear Models · ICLR 2025 Maximum Likelihood Estimation is All You Need for Well-Specified Covariate Shift · ICLR 2024 |
Natural language and speech › Language models and text generation › evaluation of language models
benchmark construction |
0.9 | 1 | 2025 | MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations · ICML 2025 |
Machine learning › Learning theory › overfitting
benign overfitting |
0.9 | 1 | 2025 | Benign Overfitting in Out-of-Distribution Generalization of Linear Models · ICLR 2025 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.9 | 1 | 2025 | MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations · ICML 2025 |
Natural language and speech › Language models and text generation › evaluation of language models › reasoning evaluation
mathematical reasoning benchmark |
0.9 | 1 | 2025 | MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations · ICML 2025 |
Machine learning › Kernel, tree and ensemble methods › linear model
principal component regression |
0.9 | 1 | 2025 | Benign Overfitting in Out-of-Distribution Generalization of Linear Models · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › least squares regression
ridge regression |
0.9 | 1 | 2025 | Benign Overfitting in Out-of-Distribution Generalization of Linear Models · ICLR 2025 |
Machine learning › Trustworthy machine learning › robustness
robustness to perturbation |
0.9 | 1 | 2025 | MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations · ICML 2025 |
Machine learning › Learning theory
statistical learning theory |
0.9 | 1 | 2025 | Benign Overfitting in Out-of-Distribution Generalization of Linear Models · ICLR 2025 |
Machine learning › Transfer learning and domain adaptation › domain shift
covariate shift |
0.8 | 1 | 2024 | Maximum Likelihood Estimation is All You Need for Well-Specified Covariate Shift · ICLR 2024 |
Machine learning › Learning theory
minimax optimality |
0.8 | 1 | 2024 | Maximum Likelihood Estimation is All You Need for Well-Specified Covariate Shift · ICLR 2024 |
Machine learning › Representation and self-supervised learning › pre-training
unsupervised pre-training |
0.8 | 1 | 2024 | On the Provable Advantage of Unsupervised Pretraining · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
maximum likelihood estimation · 1.5ridge regression · 0.9principal component regression · 0.9in-context learning · 0.9hard perturbation · 0.9maximum weighted likelihood estimation · 0.8empirical risk minimization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Benign Overfitting in Out-of-Distribution Generalization of Linear ModelsabstractBenign overfitting refers to the phenomenon where an over-parameterized model fits the training data perfectly, including noise in the data, but still generalizes well to the unseen test data. While prior work provides some theoretical understanding of this phenomenon under the in-distribution setup, modern machine learning often operates in a more challenging Out-of-Distribution (OOD) regime, where the target (test) distribution can be rather different from the source (training) distribution. In this work, we take an initial step towards understanding benign overfitting in the OOD regime by focusing on the basic setup of over-parameterized linear models under covariate shift. We provide non-asymptotic guarantees proving that benign overfitting occurs in standard ridge regression, even under the OOD regime when the target covariance satisfies certain structural conditions. We identify several vital quantities relating to source and target covariance, which govern the performance of OOD generalization. Our result is sharp, which provably recovers prior in-distribution benign overfitting guarantee (Tsigler & Bartlett, 2023), as well as under-parameterized OOD guarantee (Ge et al., 2024) when specializing to each setup. Moreover, we also present theoretical results for a more general family of target covariance matrix, where standard ridge regression only achieves a slow statistical rate of $\mathcal{O}(1/\sqrt{n})$ for the excess risk, while Principal Component Regression (PCR) is guaranteed to achieve the fast rate $\mathcal{O}(1/n)$, where $n$ is the number of samples. Shange Tang, Jiayun Wu, Jianqing Fan, Chi Jin 0001 |
ICLR | 1 |
| 2025 | MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard PerturbationsabstractLarge language models have demonstrated impressive performance on challenging mathematical reasoning tasks, which has triggered the discussion of whether the performance is achieved by true reasoning capability or memorization. To investigate this question, prior work has constructed mathematical benchmarks when questions undergo simple perturbations – modifications that still preserve the underlying reasoning patterns of the solutions. However, no work has explored hard perturbations, which fundamentally change the nature of the problem so that the original solution steps do not apply. To bridge the gap, we construct MATH-P-Simple and MATH-P-Hard via simple perturbation and hard perturbation, respectively. Each consists of 279 perturbed math problems derived from level-5 (hardest) problems in the MATH dataset (Hendrycks et al., 2021). We observe significant performance drops on MATH-P-Hard across various models, including o1-mini (-16.49%) and gemini-2.0-flash-thinking (-12.9%). We also raise concerns about a novel form of memorization where models blindly apply learned problem-solving skills without assessing their applicability to modified contexts. This issue is amplified when using original problems for in-context learning. We call for research efforts to address this challenge, which is critical for developing more robust and reliable reasoning models. The project is available at https://math-perturb.github.io/. Kaixuan Huang, Jiacheng Guo, Jiawei Ge 0003, Tianle Cai, Hui Yuan 0002, Runzhe Wang, Ming Yin 0003, Shange Tang, Yangsibo Huang, Chi Jin 0001, Chiyuan Zhang, Mengdi Wang 0001 |
ICML | 13 |
| 2025 | Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving of InequalitiesabstractLLM-based formal proof assistants (e.g., in Lean) hold great promise for automating mathematical discovery. But beyond syntactic correctness, do these systems truly understand mathematical structure as humans do? We investigate this question in context of mathematical inequalities---specifically the prover's ability to recognize that the given problem simplifies by applying a known inequality such as AM/GM. Specifically, we are interested in their ability to do this in a {\em compositional setting} where multiple inequalities must be applied as part of a solution. We introduce \ineqcomp, a benchmark built from elementary inequalities through systematic transformations, including variable duplication, algebraic rewriting, and multi-step composition. Although these problems remain easy for humans, we find that most provers---including Goedel, STP, and Kimina-7B---struggle significantly. DeepSeek-Prover-V2-7B shows relative robustness, but still suffers a 20\% performance drop (pass@32). Even for DeepSeek-Prover-V2-671B model, the gap between compositional variants and seed problems exists, implying that simply scaling up the model size alone does not fully solve the compositional weakness. Strikingly, performance remains poor for all models even when formal proofs of the constituent parts are provided in context, revealing that the source of weakness is indeed in compositional reasoning. Our results expose a persisting gap between the generalization behavior of current AI provers and human mathematical intuition. All data and evaluation code can be found at \url{https://github.com/haoyuzhao123/LeanIneqComp}. Yihan Geng, Shange Tang, Bohan Lyu 0001, Hongzhou Lin, Chi Jin 0001, Sanjeev Arora |
NeurIPS | 3 |
| 2024 | On the Provable Advantage of Unsupervised PretrainingabstractUnsupervised pretraining, which learns a useful representation using a large amount of unlabeled data to facilitate the learning of downstream tasks, is a critical component of modern large-scale machine learning systems. Despite its tremendous empirical success, the rigorous theoretical understanding of why unsupervised pretraining generally helps remains rather limited---most existing results are restricted to particular methods or approaches for unsupervised pretraining with specialized structural assumptions. This paper studies a generic framework,
where the unsupervised representation learning task is specified by an abstract class of latent variable models $\Phi$ and the downstream task is specified by a class of prediction functions $\Psi$. We consider a natural approach of using Maximum Likelihood Estimation (MLE) for unsupervised pretraining and Empirical Risk Minimization (ERM) for learning downstream tasks. We prove that, under a mild ``informative'' condition, our algorithm achieves an excess risk of $\\tilde{\\mathcal{O}}(\sqrt{\mathcal{C}\_\Phi/m} + \sqrt{\mathcal{C}\_\Psi/n})$ for downstream tasks, where $\mathcal{C}\_\Phi, \mathcal{C}\_\Psi$ are complexity measures of function classes $\Phi, \Psi$, and $m, n$ are the number of unlabeled and labeled data respectively. Comparing to the baseline of $\tilde{\mathcal{O}}(\sqrt{\mathcal{C}\_{\Phi \circ \Psi}/n})$ achieved by performing supervised learning using only the labeled data, our result rigorously shows the benefit of unsupervised pretraining when $m \gg n$ and $\mathcal{C}\_{\Phi\circ \Psi} > \mathcal{C}\_\Psi$. This paper further shows that our generic framework covers a wide range of approaches for unsupervised pretraining, including factor models, Gaussian mixture models, and contrastive learning. Jiawei Ge 0003, Shange Tang, Jianqing Fan, Chi Jin 0001 |
ICLR | 2 |
| 2024 | Maximum Likelihood Estimation is All You Need for Well-Specified Covariate ShiftabstractA key challenge of modern machine learning systems is to achieve Out-of-Distribution (OOD) generalization---generalizing to target data whose distribution differs from that of source data. Despite its significant importance, the fundamental question of ``what are the most effective algorithms for OOD generalization'' remains open even under the standard setting of covariate shift.
This paper addresses this fundamental question by proving that, surprisingly, classical Maximum Likelihood Estimation (MLE) purely using source data (without any modification) achieves the *minimax* optimality for covariate shift under the *well-specified* setting. That is, *no* algorithm performs better than MLE in this setting (up to a constant factor), justifying MLE is all you need.
Our result holds for a very rich class of parametric models, and does not require any boundedness condition on the density ratio. We illustrate the wide applicability of our framework by instantiating it to three concrete examples---linear regression, logistic regression, and phase retrieval. This paper further complement the study by proving that, under the *misspecified setting*, MLE is no longer the optimal choice, whereas Maximum Weighted Likelihood Estimator (MWLE) emerges as minimax optimal in certain scenarios. Jiawei Ge 0003, Shange Tang, Jianqing Fan, Cong Ma 0001, Chi Jin 0001 |
ICLR | 2 |