VLDB 2026 Research / reviewers in the wild / expert
Nan Ding 0002
dblp:68/3975-2
· DBLP profile ↗
29ranked-venue papers
13as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 11 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Graders Should Cheat: Privileged Information Enables Expert-Level Automated EvaluationsabstractAuto-evaluating language models (LMs), i.e., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and reduce the cost associated with it.But this presents a paradox: how can we trust the grader LM, which is presumably weaker than the candidate LM, to assess problems that are beyond the frontier of the capabilities of either model or both?For instance, today's LMs struggle on graduate-level physics and Olympiad-level math, making them unreliable graders in these domains.We show that providing privileged information -such as ground-truth solutions or problem-specific guidelines -improves automated evaluations on such frontier problems.This approach offers two key advantages.First, it expands the range of problems where LMs graders apply.Specifically, weaker models can now rate the predictions of stronger models.Second, privileged information can be used to devise easier variations of challenging problems which improves the separability of different LMs on tasks where their performance is generally low.With this approach, general-purpose LM graders match the state of the art performance on RewardBench, surpassing almost all the specially-tuned models.LM graders also outperform individual human raters on Vibe-Eval, and approach human expert graders on Olympiad-level math problems. Jin Peng Zhou, Sébastien M. R. Arnold, Nan Ding 0002, Kilian Q. Weinberger, Nan Hua, Fei Sha |
EMNLP | 3 |
| 2024 | CausalLM is not optimal for in-context learningabstractRecent empirical evidence indicates that transformer based in-context learning performs better when using a prefix language model (prefixLM), in which in-context samples can all attend to each other, compared to causal language models (causalLM), which use auto-regressive attention that prohibits in-context samples to attend to future samples. While this result is intuitive, it is not understood from a theoretical perspective. In this paper we take a theoretical approach and analyze the convergence behavior of prefixLM and causalLM under a certain parameter construction. Our analysis shows that both LM types converge to their stationary points at a linear rate, but that while prefixLM converges to the optimal solution of linear regression, causalLM convergence dynamics follows that of an online gradient descent algorithm, which is not guaranteed to be optimal even as the number of samples grows infinitely. We supplement our theoretical claims with empirical experiments over synthetic and real tasks and using various types of transformers. Our experiments verify that causalLM consistently underperforms prefixLM in all settings. Nan Ding 0002, Tomer Levinboim, Sebastian Goodman, Radu Soricut |
ICLR | 1 |
| 2023 | Improving Robust Generalization by Direct PAC-Bayesian Bound MinimizationabstractRecent research in robust optimization has shown an overfitting-like phenomenon in which models trained against adversarial attacks exhibit higher robustness on the training set compared to the test set. Although previous work provided theoretical explanations for this phenomenon using a robust PAC-Bayesian bound over the adversarial test error, related algorithmic derivations are at best only loosely connected to this bound, which implies that there is still a gap between their empirical success and our understanding of adversarial robustness theory. To close this gap, in this paper we consider a different form of the robust PAC-Bayesian bound and directly minimize it with respect to the model posterior. The derivation of the optimal solution connects PAC-Bayesian learning to the geometry of the robust loss surface through a Trace of Hessian (TrH) regularizer that measures the surface flatness. In practice, we restrict the TrH regularizer to the top layer only, which results in an analytical solution to the bound whose computational cost does not depend on the network depth. Finally, we evaluate our TrH regularization approach over CIFAR-10/100 and ImageNet using Vision Transformers (ViT) and compare against baseline adversarial robustness algorithms. Experimental results show that TrH regularization leads to improved ViT robustness that either matches or surpasses previous state-of-the-art approaches while at the same time requires less memory and computational cost. Zifan Wang 0001, Nan Ding 0002, Tomer Levinboim, Xi Chen 0071, Radu Soricut |
CVPR | 2 |
| 2023 | PaLI: A Jointly-Scaled Multilingual Language-Image Model
Xi Chen 0071, Xiao Wang 0038, Soravit Changpinyo, A. J. Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov 0003, Joan Puigcerver, Nan Ding 0002, Keran Rong, Hassan Akbari, Linting Xue, Ashish V. Thapliyal, Weicheng Kuo |
ICLR | 13 |
| 2022 | PACTran: PAC-Bayesian Metrics for Estimating the Transferability of Pretrained Models to Classification Tasks
Nan Ding 0002, Xi Chen 0071, Tomer Levinboim, Soravit Changpinyo, Radu Soricut |
ECCV (34) | 1 |
| 2022 | All You May Need for VQA are Image CaptionsabstractSoravit Changpinyo, Doron Kukliansy, Idan Szpektor, Xi Chen, Nan Ding, Radu Soricut. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Soravit Changpinyo, Doron Kukliansky, Idan Szpektor, Xi Chen 0071, Nan Ding 0002, Radu Soricut |
NAACL-HLT | 5 |
| 2021 | Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual ConceptsabstractThe availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pretraining. However, these datasets are often collected with overrestrictive requirements inherited from their original target tasks (e.g., image caption generation), which limit the resulting dataset scale and diversity. We take a step further in pushing the limits of vision-and-language pretraining data by relaxing the data collection pipeline used in Conceptual Captions 3M (CC3M) [54] and introduce the Conceptual 12M (CC12M), a dataset with 12 million image-text pairs specifically meant to be used for visionand-language pre-training. We perform an analysis of this dataset and benchmark its effectiveness against CC3M on multiple downstream tasks with an emphasis on long-tail visual recognition. Our results clearly illustrate the benefit of scaling up pre-training data for vision-and-language tasks, as indicated by the new state-of-the-art results on both the nocaps and Conceptual Captions benchmarks.1 Soravit Changpinyo, Piyush Sharma, Nan Ding 0002, Radu Soricut |
CVPR | 3 |
| 2021 | Do Transformer Modifications Transfer Across Implementations and Applications?abstractSharan Narang, Hyung Won Chung, Yi Tay, Liam Fedus, Thibault Fevry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhenzhong Lan, Yanqi Zhou, Wei Li, Nan Ding, Jake Marcus, Adam Roberts, Colin Raffel. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Sharan Narang, Hyung Won Chung, Yi Tay, William Fedus, Thibault Févry, Michael Matena, Karishma Malkan, Noah Fiedel, Noam Shazeer, Zhen-Zhong Lan, Yanqi Zhou, Wei Li 0133, Nan Ding 0002, Jake Marcus, Adam Roberts, Colin Raffel |
EMNLP (1) | 13 |
| 2021 | Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-LearningabstractDespite recent advances in its theoretical understanding, there still remains a significant gap in the ability of existing PAC-Bayesian theories on meta-learning to explain performance improvements in the few-shot learning setting, where the number of training examples in the target tasks is severely limited. This gap originates from an assumption in the existing theories which supposes that the number of training examples in the observed tasks and the number of training examples in the target tasks follow the same distribution, an assumption that rarely holds in practice. By relaxing this assumption, we develop two PAC-Bayesian bounds tailored for the few-shot learning setting and show that two existing meta-learning algorithms (MAML and Reptile) can be derived from our bounds, thereby bridging the gap between practice and PAC-Bayesian theories. Furthermore, we derive a new computationally-efficient PACMAML algorithm, and show it outperforms existing meta-learning algorithms on several few-shot benchmark datasets. Nan Ding 0002, Xi Chen 0071, Tomer Levinboim, Sebastian Goodman, Radu Soricut |
NeurIPS | 1 |
| 2020 | TeaForN: Teacher-Forcing with N-gramsabstractSequence generation models trained with teacher-forcing suffer from issues related to exposure bias and lack of differentiability across timesteps.Our proposed method, Teacher-Forcing with N-grams (TeaForN), addresses both these problems directly, through the use of a stack of N decoders trained to decode along a secondary time axis that allows modelparameter updates based on N prediction steps.TeaForN can be used with a wide class of decoder architectures and requires minimal modifications from a standard teacher-forcing setup.Empirically, we show that TeaForN boosts generation quality on one Machine Translation benchmark, WMT 2014 English-French, and two News Summarization benchmarks, CNN/Dailymail and Gigaword. Sebastian Goodman, Nan Ding 0002, Radu Soricut |
EMNLP (1) | 2 |
| 2018 | Conceptual Captions: A Cleaned, Hypernymed, Image Alt-text Dataset For Automatic Image CaptioningabstractWe present a new dataset of image caption annotations, Conceptual Captions, which contains an order of magnitude more images than the MS-COCO dataset (Lin et al., 2014) and represents a wider variety of both images and image caption styles.We achieve this by extracting and filtering image caption annotations from billions of webpages.We also present quantitative evaluations of a number of image captioning models and show that a model architecture based on Inception-ResNet-v2 (Szegedy et al., 2016) for image-feature extraction and Transformer (Vaswani et al., 2017) for sequence modeling achieves the best performance when trained on the Conceptual Captions dataset. Piyush Sharma, Nan Ding 0002, Sebastian Goodman, Radu Soricut |
ACL (1) | 2 |
| 2018 | SHAPED: Shared-Private Encoder-Decoder for Text Style AdaptationabstractYe Zhang, Nan Ding, Radu Soricut. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Nan Ding 0002, Radu Soricut |
NAACL-HLT | 2 |
| 2017 | Cold-Start Reinforcement Learning with Softmax Policy GradientabstractPolicy-gradient approaches to reinforcement learning have two common and undesirable overhead procedures, namely warm-start training and sample variance reduction. In this paper, we describe a reinforcement learning method based on a softmax value function that requires neither of these procedures. Our method combines the advantages of policy-gradient methods with the efficiency and simplicity of maximum-likelihood approaches. We apply this new cold-start reinforcement learning method in training sequence generation models for structured output prediction problems. Empirical evidence validates this method on automatic summarization and image captioning tasks. Nan Ding 0002, Radu Soricut |
NIPS | 1 |
| 2016 | Stochastic Gradient MCMC with Stale GradientsabstractStochastic gradient MCMC (SG-MCMC) has played an important role in large-scale Bayesian learning, with well-developed theoretical convergence properties. In such applications of SG-MCMC, it is becoming increasingly popular to employ distributed systems, where stochastic gradients are computed based on some outdated parameters, yielding what are termed stale gradients. While stale gradients could be directly used in SG-MCMC, their impact on convergence properties has not been well studied. In this paper we develop theory to show that while the bias and MSE of an SG-MCMC algorithm depend on the staleness of stochastic gradients, its estimation variance (relative to the expected estimate, based on a prescribed number of samples) is independent of it. In a simple Bayesian distributed system with SG-MCMC, where stale gradients are computed asynchronously by a set of workers, our theory indicates a linear speedup on the decrease of estimation variance w.r.t. the number of workers. Experiments on synthetic data and deep neural networks validate our theory, demonstrating the effectiveness and scalability of SG-MCMC with stale gradients. Changyou Chen, Nan Ding 0002, Chunyuan Li, Yizhe Zhang 0002, Lawrence Carin |
NIPS | 2 |
| 2015 | Probabilistic Label Relation Graphs with Ising ModelsabstractWe consider classification problems in which the label space has structure. A common example is hierarchical label spaces, corresponding to the case where one label subsumes another (e.g., animal subsumes dog). But labels can also be mutually exclusive (e.g., dog vs cat) or unrelated (e.g., furry, carnivore). To jointly model hierarchy and exclusion relations, the notion of a HEX (hierarchy and exclusion) graph was introduced in [8]. This combined a conditional random field (CRF) with a deep neural network (DNN), resulting in state of the art results when applied to visual object classification problems where the training labels were drawn from different levels of the ImageNet hierarchy (e.g., an image might be labeled with the basic level category "dog", rather than the more specific label "husky"). In this paper, we extend the HEX model to allow for soft or probabilistic relations between labels, which is useful when there is uncertainty about the relationship between two labels (e.g., an antelope is "sort of" furry, but not to the same degree as a grizzly bear). We call our new model pHEX, for probabilistic HEX. We show that the pHEX graph can be converted to an Ising model, which allows us to use existing off-the-shelf inference methods (in contrast to the HEX method, which needed specialized inference algorithms). Experimental results show significant improvements in a number of large-scale visual object classification tasks, outperforming the previous HEX model. Nan Ding 0002, Jia Deng 0001, Kevin Murphy 0002, Hartmut Neven |
ICCV | 1 |
| 2015 | On the Convergence of Stochastic Gradient MCMC Algorithms with High-Order IntegratorsabstractRecent advances in Bayesian learning with large-scale data have witnessed emergence of stochastic gradient MCMC algorithms (SG-MCMC), such as stochastic gradient Langevin dynamics (SGLD), stochastic gradient Hamiltonian MCMC (SGHMC), and the stochastic gradient thermostat. While finite-time convergence properties of the SGLD with a 1st-order Euler integrator have recently been studied, corresponding theory for general SG-MCMCs has not been explored. In this paper we consider general SG-MCMCs with high-order integrators, and develop theory to analyze finite-time convergence properties and their asymptotic invariant measures. Our theoretical results show faster convergence rates and more accurate invariant measures for SG-MCMCs with higher-order integrators. For example, with the proposed efficient 2nd-order symmetric splitting integrator, the mean square error (MSE) of the posterior average for the SGHMC achieves an optimal convergence rate of $L^{-4/5}$ at $L$ iterations, compared to $L^{-2/3}$ for the SGHMC and SGLD with 1st-order Euler integrators. Furthermore, convergence results of decreasing-step-size SG-MCMCs are also developed, with the same convergence rates as their fixed-step-size counterparts for a specific decreasing sequence. Experiments on both synthetic and real datasets verify our theory, and show advantages of the proposed method in two large-scale real applications. Changyou Chen, Nan Ding 0002, Lawrence Carin |
NIPS | 2 |
| 2015 | Embedding Inference for Structured Multilabel PredictionabstractA key bottleneck in structured output prediction is the need for inference during training and testing, usually requiring some form of dynamic programming. Rather than using approximate inference or tailoring a specialized inference method for a particular structure---standard responses to the scaling challenge---we propose to embed prediction constraints directly into the learned representation. By eliminating the need for explicit inference a more scalable approach to structured output prediction can be achieved, particularly at test time. We demonstrate the idea for multi-label prediction under subsumption and mutual exclusion constraints, where a relationship to maximum margin structured output prediction can be established. Experiments demonstrate that the benefits of structured output training can still be realized even after inference has been eliminated. Farzaneh Mirzazadeh, Siamak Ravanbakhsh, Nan Ding 0002, Dale Schuurmans |
NIPS | 3 |
| 2015 | Differential Topic ModelsabstractIn applications we may want to compare different document collections: they could have shared content but also different and unique aspects in particular collections. This task has been called comparative text mining or cross-collection modeling. We present a differential topic model for this application that models both topic differences and similarities. For this we use hierarchical Bayesian nonparametric models. Moreover, we found it was important to properly model power-law phenomena in topic-word distributions and thus we used the full Pitman-Yor process rather than just a Dirichlet process. Furthermore, we propose the transformed Pitman-Yor process (TPYP) to incorporate prior knowledge such as vocabulary variations in different collections into the model. To deal with the non-conjugate issue between model prior and likelihood in the TPYP, we thus propose an efficient sampling algorithm using a data augmentation technique based on the multinomial theorem. Experimental results show the model discovers interesting aspects of different collections. We also show the proposed MCMC based algorithm achieves a dramatically reduced test perplexity compared to some existing topic models. Finally, we show our model outperforms the state-of-the-art for document classification/ideology prediction on a number of text collections. Changyou Chen, Wray L. Buntine, Nan Ding 0002, Lexing Xie, Lan Du 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | Large-Scale Object Classification Using Label Relation Graphs
Jia Deng 0001, Nan Ding 0002, Yangqing Jia, Andrea Frome, Kevin Murphy 0002, Samy Bengio, Hartmut Neven, Hartwig Adam |
ECCV (1) | 2 |
| 2014 | Bayesian Sampling Using Stochastic Gradient Thermostats
Nan Ding 0002, Youhan Fang, Ryan Babbush, Changyou Chen, Robert D. Skeel, Hartmut Neven |
NIPS | 1 |
| 2012 | Dependent Hierarchical Normalized Random Measures for Dynamic Topic Modeling
Changyou Chen, Nan Ding 0002, Wray L. Buntine |
ICML | 2 |
| 2012 | Robust Classification with Adiabatic Quantum Optimization
Vasil S. Denchev, Nan Ding 0002, S. V. N. Vishwanathan, Hartmut Neven |
ICML | 2 |
| 2011 | t-divergence Based Approximate InferenceabstractApproximate inference is an important technique for dealing with large, intractable graphical models based on the exponential family of distributions. We extend the idea of approximate inference to the t-exponential family by defining a new t-divergence. This divergence measure is obtained via convex duality between the log-partition function of the t-exponential family and a new t-entropy. We illustrate our approach on the Bayes Point Machine with a Student's t-prior. Nan Ding 0002, S. V. N. Vishwanathan, Yuan Qi 0001 |
NIPS | 1 |
| 2010 | Variational nonparametric Bayesian Hidden Markov ModelabstractThe Hidden Markov Model (HMM) has been widely used in many applications such as speech recognition. A common challenge for applying the classical HMM is to determine the structure of the hidden state space. Based on the Dirichlet Process, a nonparametric Bayesian Hidden Markov Model is proposed, which allows an infinite number of hidden states and uses an infinite number of Gaussian components to support continuous observations. An efficient variational inference method is also proposed and applied on the model. Our experiments demonstrate that the variational Bayesian inference on the new model can discover the HMM hidden structure for both synthetic data and real-world applications. Nan Ding 0002, Zhijian Ou |
ICASSP | 1 |
| 2010 | t-logistic regressionabstractWe extend logistic regression by using t-exponential families which were introduced recently in statistical physics. This gives rise to a regularized risk minimization problem with a non-convex loss function. An efficient block coordinate descent optimization scheme can be derived for estimating the parameters. Because of the nature of the loss function, our algorithm is tolerant to label noise. Furthermore, unlike other algorithms which employ non-convex loss functions, our algorithm is fairly robust to the choice of initial values. We verify both these observations empirically on a number of synthetic and real datasets. Nan Ding 0002, S. V. N. Vishwanathan |
NIPS | 1 |
| 2008 | A Bayesian view on the polynomial distribution model in estimation of distribution algorithmsabstractEstimation of distribution algorithms(EDA) are a class of recently-developed evolutionary algorithms in which the probabilistic model are used to explicitly characterize the distribution of the population and to generate new individuals. The polynomial distribution is applied by discrete EDAs and continuous EDAs based on discretization of the domain such as histogram-based EDA. We can unify those kinds of EDA from their distribution and call them PolyEDA. In this paper, we theoretically analyze PolyEDA from a Bayesian analysis view. Our analysis is based on the assumption that the prior distribution of the parameters satisfies a Dirichlet Distribution, because under this assumption the formulation can be analytically solved. Furthermore, we notice that the prior distribution is always overlooked by previous algorithms, so we follow this way and propose some strategies to improve the PolyEDA. The experimental results show that these new strategies can help the polynomial model based estimation of distribution algorithms achieve better convergence and diversity. Nan Ding 0002, Shude Zhou, Zengqi Sun |
IEEE Congress on Evolutionary Computation | 1 |
| 2008 | Marginal probability distribution estimation in characteristic space of covariance-matrixabstractMarginal probability distribution has been widely used as the probabilistic model in EDAs because of its simplicity and efficiency. However, the obvious shortcoming of the kind of EDAs lies in its incapability of taking the correlation between variables into account. This paper tries to solve the problem from the point view of space transformation. As we know, it seems a default rule that the probabilistic model is usually constructed directly from the selected samples in the space defined by the problem. In the algorithm CM-MEDA, instead, we first transform the sampled data from the initial coordinate space into the characteristic space of covariance-matrix and then the marginal probabilistic model is constructed in the new space. We find that the marginal probabilistic model in the new space can capture the variable linkages in the initial space quite well. The relationship of CM-MEDA with Covariance-Matrix estimation and principal component analysis is also analyzed in this paper. We implement CM-MEDA in continuous domain based on both Gaussian and histogram models. The experimental results verify the effectiveness of our idea. Nan Ding 0002, Shude Zhou, Zengqi Sun |
IEEE Congress on Evolutionary Computation | 1 |
| 2008 | Histogram-Based Estimation of Distribution Algorithm: A Competent Method for Continuous Optimization
Nan Ding 0002, Shude Zhou, Zengqi Sun |
J. Comput. Sci. Technol. | 1 |
| 2007 | Reducing computational complexity of estimating multivariate histogram-based probabilistic modelabstractIn continuous domain, how to efficiently learn the complex probabilistic graphical model is a bottleneck problem for estimation of distribution algorithms (EDAs). The predominant researches focus on Gaussian probabilistic model instead of histogram distribution model because of its comparative superiority in the computational complexity. In this paper, however, we find that using the histogram model does not necessarily bring into exponential computational complexity. Based on the fact many bins are zero-height, we propose a novel method that can learn the multivariate dependency histogram based probabilistic graphical model with acceptable polynomial computational complexity. Several strategies previously used in the HEDA are combined into the new algorithm to improve the convergence and diversity. Experiments showed the superior performance of the new algorithm on several continuous problems compared with UMDAc, IDEA-G and sur-shr-HEDA. Nan Ding 0002, Shude Zhou, Zengqi Sun |
IEEE Congress on Evolutionary Computation | 1 |