Lexin Li

dblp:07/236 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Probabilistic and Bayesian machine learning · 53% Kernel, tree and ensemble methods · 22% Language models and text generation · 12%
Theoretical computer science
2 papers
Algorithmic game theory and mechanism design · 95% Mathematical optimization · 5%
Interdisciplinary, comprehensive, and emerging computing
6 papers
Computational science and engineering · 42% Bioinformatics and computational biology · 39% Computational finance and economics · 14%
Databases, data mining, and information retrieval
2 papers
Data mining · 100%

Topics — the 30 heaviest of 33, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
hallucination mitigation
0.912025
Incentivizing Truthful Language Models via Peer Elicitation Games · NeurIPS 2025
Algorithmic game theory and mechanism design
mechanism design
0.912025
Incentivizing Truthful Language Models via Peer Elicitation Games · NeurIPS 2025
Algorithmic game theory and mechanism design › mechanism design
truthful mechanism
0.912025
Incentivizing Truthful Language Models via Peer Elicitation Games · NeurIPS 2025
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.812024
Functional Directed Acyclic Graphs · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning
causal inference
0.812024
Functional Directed Acyclic Graphs · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
conditional independence
0.812024
Functional Directed Acyclic Graphs · J. Mach. Learn. Res. 2024
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel learning
0.812024
Post-Regularization Confidence Bands for Ordinary Differential Equations · J. Mach. Learn. Res. 2024
Machine learning › Kernel, tree and ensemble methods
kernel methods
0.812024
Post-Regularization Confidence Bands for Ordinary Differential Equations · J. Mach. Learn. Res. 2024
Computational science and engineering › differential equations
ordinary differential equations
0.812024
Post-Regularization Confidence Bands for Ordinary Differential Equations · J. Mach. Learn. Res. 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › density estimation
conditional density estimation
0.512021
Double Generative Adversarial Networks for Conditional Independence Testing · J. Mach. Learn. Res. 2021
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › conditional independence
conditional independence testing
0.512021
Double Generative Adversarial Networks for Conditional Independence Testing · J. Mach. Learn. Res. 2021
Machine learning › Generative modeling
generative adversarial network
0.512021
Double Generative Adversarial Networks for Conditional Independence Testing · J. Mach. Learn. Res. 2021
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis › tensor factorization
boolean tensor factorization
0.412020
Learning from Binary Multiway Data: Probabilistic Tensor Decomposition and its Statistical Optimality · J. Mach. Learn. Res. 2020
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis › tensor factorization
probabilistic tensor decomposition
0.412020
Learning from Binary Multiway Data: Probabilistic Tensor Decomposition and its Statistical Optimality · J. Mach. Learn. Res. 2020
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis
tensor completion
0.412020
Learning from Binary Multiway Data: Probabilistic Tensor Decomposition and its Statistical Optimality · J. Mach. Learn. Res. 2020
Data mining › multidimensional data analysis › multiway data analysis › tensor analysis
tensor factorization
0.412020
Learning from Binary Multiway Data: Probabilistic Tensor Decomposition and its Statistical Optimality · J. Mach. Learn. Res. 2020
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
regression
0.312017
STORE: Sparse Tensor Response Regression and Neuroimaging Analysis · J. Mach. Learn. Res. 2017
Machine learning › Probabilistic and Bayesian machine learning
statistical inference
0.112021
Double Generative Adversarial Networks for Conditional Independence Testing · J. Mach. Learn. Res. 2021
Machine learning › Learning theory › hypothesis testing
type-i error control
0.112021
Double Generative Adversarial Networks for Conditional Independence Testing · J. Mach. Learn. Res. 2021
Bioinformatics and computational biology › statistical genetics › genetic association study
association mapping
0.112012
Nonlinear dimension reduction with Wright-Fisher kernel for genotype aggregation and association mapping · Bioinform. 2012
Bioinformatics and computational biology
genomics
0.112012
Nonlinear dimension reduction with Wright-Fisher kernel for genotype aggregation and association mapping · Bioinform. 2012
Bioinformatics and computational biology
kernel methods
0.112012
Nonlinear dimension reduction with Wright-Fisher kernel for genotype aggregation and association mapping · Bioinform. 2012
Computational finance and economics › online advertising
multi-touch attribution
0.112011
Data-driven multi-touch attribution models · KDD 2011
Computational finance and economics
online advertising
0.112011
Data-driven multi-touch attribution models · KDD 2011
Medical and health informatics › neuroimaging
neuroimaging analysis
0.112017
STORE: Sparse Tensor Response Regression and Neuroimaging Analysis · J. Mach. Learn. Res. 2017
Mathematical optimization
nonconvex optimization
0.112017
STORE: Sparse Tensor Response Regression and Neuroimaging Analysis · J. Mach. Learn. Res. 2017
Bioinformatics and computational biology
cancer genomics
0.112006
Survival prediction of diffuse large-B-cell lymphoma based on both clinical and gene expression information · Bioinform. 2006
Bioinformatics and computational biology › survival analysis
survival prediction
0.112006
Survival prediction of diffuse large-B-cell lymphoma based on both clinical and gene expression information · Bioinform. 2006
Bioinformatics and computational biology › survival analysis › survival prediction
gene expression-based survival prediction
0.012004
Dimension reduction methods for microarrays with application to censored survival data · Bioinform. 2004
Bioinformatics and computational biology
survival analysis
0.012004
Dimension reduction methods for microarrays with application to censored survival data · Bioinform. 2004

Methods — techniques the papers use, named apart from their topics

online learning · 1.7mutual information · 1.7game theory · 1.7post-regularization · 1.5localized kernel learning · 1.5debiasing · 1.5partial correlation operator · 0.8conditional covariance operator · 0.8PC algorithm · 0.8sparse tensor decomposition · 0.6low-rank regularization · 0.6alternating updating · 0.6cross-fitting · 0.5multilinear bernoulli model · 0.4minimax optimality · 0.4alternating optimization · 0.4data-driven attribution modeling · 0.2wright-fisher kernel · 0.1
YearPublicationVenuePosition
2025 Incentivizing Truthful Language Models via Peer Elicitation Games
abstract
Large Language Models (LLMs) have demonstrated strong generative capabilities but remain prone to inconsistencies and hallucinations. We introduce Peer Elicitation Games (PEG), a training-free, game-theoretic framework for aligning LLMs through a peer elicitation mechanism involving a generator and multiple discriminators instantiated from distinct base models. Discriminators interact in a peer evaluation setting, where utilities are computed using a determinant-based mutual information score that provably incentivizes truthful reporting without requiring ground-truth labels. We establish theoretical guarantees showing that each agent, via online learning, achieves sublinear regret in the sense their cumulative performance approaches that of the best fixed truthful strategy in hindsight. Moreover, we prove last-iterate convergence to a truthful Nash equilibrium, ensuring that the actual policies used by agents converge to stable and truthful behavior over time. Empirical evaluations across multiple benchmarks demonstrate significant improvements in factual accuracy. These results position PEG as a practical approach for eliciting truthful behavior from LLMs without supervision or fine-tuning.
Baiting Chen, Jiale Han 0002, Lexin Li, Xiaowu Dai
NeurIPS4
2024 Post-Regularization Confidence Bands for Ordinary Differential Equations
abstract
Ordinary differential equation (ODE) is an important tool to study a system of biological and physical processes. A central question in ODE modeling is to infer the significance of individual regulatory effect of one signal variable on another. However, building confidence band for ODE with unknown regulatory relations is challenging, and it remains largely an open question. In this article, we construct the post-regularization confidence band for the individual regulatory function in ODE with unknown functionals and noisy data observations. Our proposal is the first of its kind, and is built on two novel ingredients. The first is a new localized kernel learning approach that combines reproducing kernel learning with local Taylor approximation, and the second is a new de-biasing method that tackles infinite-dimensional functionals and additional measurement errors. We show that the constructed confidence band has the desired asymptotic coverage probability, and the recovered regulatory network approaches the truth with probability tending to one. We establish the theoretical properties when the number of variables in the system can be either smaller or larger than the number of sampling time points, and we study the regime-switching phenomenon. We demonstrate the efficacy of the proposed method through both simulations and illustrations with two data applications.
Xiaowu Dai, Lexin Li
J. Mach. Learn. Res.2
2024 Functional Directed Acyclic Graphs
abstract
In this article, we introduce a new method to estimate a directed acyclic graph (DAG) from multivariate functional data. We build on the notion of faithfulness that relates a DAG with a set of conditional independences among the random functions. We develop two linear operators, the conditional covariance operator and the partial correlation operator, to characterize and evaluate the conditional independence. Based on these operators, we adapt and extend the PC-algorithm to estimate the functional directed graph, so that the computation time depends on the sparsity rather than the full size of the graph. We study the asymptotic properties of the two operators, derive their uniform convergence rates, and establish the uniform consistency of the estimated graph, all of which are obtained while allowing the graph size to diverge to infinity with the sample size. We demonstrate the efficacy of our method through both simulations and an application to a time-course proteomic dataset.
Kuang-Yao Lee, Lexin Li
J. Mach. Learn. Res.2
2023 Optimizing Pessimism in Dynamic Treatment Regimes: A Bayesian Learning Approach
abstract
In this article, we propose a novel pessimism-based Bayesian learning method for optimal dynamic treatment regimes in the offline setting. When the coverage condition does not hold, which is common for offline data, the existing solutions would produce sub-optimal policies. The pessimism principle addresses this issue by discouraging recommendation of actions that are less explored conditioning on the state. However, nearly all pessimism-based methods rely on a key hyper-parameter that quantifies the degree of pessimism, and the performance of the methods can be highly sensitive to the choice of this parameter. We propose to integrate the pessimism principle with Thompson sampling and Bayesian machine learning for optimizing the degree of pessimism. We derive a credible set whose boundary uniformly lower bounds the optimal Q-function, and thus we do not require additional tuning of the degree of pessimism. We develop a general Bayesian learning method that works with a range of models, from Bayesian linear basis model to Bayesian neural network model. We develop the computational algorithm based on variational inference, which is highly efficient and scalable. We establish the theoretical guarantees of the proposed method, and show empirically that it outperforms the existing state-of-the-art solutions through both simulations and a real data example.
Yunzhe Zhou, Zhengling Qi, Chengchun Shi, Lexin Li
AISTATS4
2023 A generalized method for dynamic noise inference in modeling sequential decision-making
Chengchun Shi, Lexin Li, Anne Gabrielle Eva Collins
CogSci3
2021 Double Generative Adversarial Networks for Conditional Independence Testing
abstract
In this article, we study the problem of high-dimensional conditional independence testing, a key building block in statistics and machine learning. We propose an inferential procedure based on double generative adversarial networks (GANs). Specifically, we first introduce a double GANs framework to learn two generators of the conditional distributions. We then integrate the two generators to construct a test statistic, which takes the form of the maximum of generalized covariance measures of multiple transformation functions. We also employ data-splitting and cross-fitting to minimize the conditions on the generators to achieve the desired asymptotic properties, and employ multiplier bootstrap to obtain the corresponding p-value. We show that the constructed test statistic is doubly robust, and the resulting test both controls type-I error and has the power approaching one asymptotically. Also notably, we establish those theoretical guarantees under much weaker and practically more feasible conditions compared to the existing tests, and our proposal gives a concrete example of how to utilize some state-of-the-art deep learning tools, such as GANs, to help address a classical but challenging statistical problem. We demonstrate the efficacy of our test through both simulations and an application to an anti-cancer drug dataset. A Python implementation of the proposed procedure is available at https://github.com/tianlinxu312/dgcit.
Chengchun Shi, Tianlin Xu, Wicher Bergsma, Lexin Li
J. Mach. Learn. Res.4
2020 Learning from Binary Multiway Data: Probabilistic Tensor Decomposition and its Statistical Optimality
abstract
We consider the problem of decomposing a higher-order tensor with binary entries. Such data problems arise frequently in applications such as neuroimaging, recommendation system, topic modeling, and sensor network localization. We propose a multilinear Bernoulli model, develop a rank-constrained likelihood-based estimation method, and obtain the theoretical accuracy guarantees. In contrast to continuous-valued problems, the binary tensor problem exhibits an interesting phase transition phenomenon according to the signal-to-noise ratio. The error bound for the parameter tensor estimation is established, and we show that the obtained rate is minimax optimal under the considered model. Furthermore, we develop an alternating optimization algorithm with convergence guarantees. The efficacy of our approach is demonstrated through both simulations and analyses of multiple data sets on the tasks of tensor completion and clustering.
Miaoyan Wang, Lexin Li
J. Mach. Learn. Res.2
2019 Spatially Adaptive Varying Correlation Analysis for Multimodal Neuroimaging Data
abstract
In this paper, we study a central problem in multimodal neuroimaging analysis, i.e., identification of significantly correlated brain regions between multiple imaging modalities. We propose a spatially varying correlation model and the associated inference procedure, which improves substantially over the common alternative solutions of voxel-wise and region-wise analysis. Compared with voxel-wise analysis, our method aggregates voxels with similar correlations into regions, takes into account spatial continuity of correlations at nearby voxels, and enjoys a much higher detection power. Compared with region-wise analysis, our method does not rely on any pre-specified brain region map, but instead finds homogenous correlation regions adaptively given the data. We applied our method to a multimodal positron emission tomography study, and found brain regions with significant correlation between tau and glucose metabolism that voxel-wise or region-wise analysis failed to identify. Our findings conform and lend additional support to prior hypotheses about how the two pathological proteins of Alzheimer's disease, tau and amyloid, interact with glucose metabolism in the aging human brain.
Lexin Li, Samuel N. Lockhart, Jenna Adams, William J. Jagust
IEEE Trans. Medical Imaging1
2017 Scalable Object Detection Using Deep but Lightweight CNN with Features Fusion
Qiaosong Chen, Shangsheng Feng, Pei Xu 0008, Lexin Li, Jin Wang 0006, Xin Deng 0003
ICIG (1)4
2017 STORE: Sparse Tensor Response Regression and Neuroimaging Analysis
abstract
Motivated by applications in neuroimaging analysis, we propose a new regression model, Sparse TensOr REsponse regression (STORE), with a tensor response and a vector predictor. STORE embeds two key sparse structures: element-wise sparsity and low-rankness. It can handle both a non-symmetric and a symmetric tensor response, and thus is applicable to both structural and functional neuroimaging data. We formulate the parameter estimation as a non-convex optimization problem, and develop an efficient alternating updating algorithm. We establish a non- asymptotic estimation error bound for the actual estimator obtained from the proposed algorithm. This error bound reveals an interesting interaction between the computational efficiency and the statistical rate of convergence. When the distribution of the error tensor is Gaussian, we further obtain a fast estimation error rate which allows the tensor dimension to grow exponentially with the sample size. We illustrate the efficacy of our model through intensive simulations and an analysis of the Autism spectrum disorder neuroimaging data.
Will Wei Sun, Lexin Li
J. Mach. Learn. Res.2
2016 Sparse Multi-Response Tensor Regression for Alzheimer's Disease Study With Multivariate Clinical Assessments
abstract
Alzheimer's disease (AD) is a progressive and irreversible neurodegenerative disorder that has recently seen serious increase in the number of affected subjects. In the last decade, neuroimaging has been shown to be a useful tool to understand AD and its prodromal stage, amnestic mild cognitive impairment (MCI). The majority of AD/MCI studies have focused on disease diagnosis, by formulating the problem as classification with a binary outcome of AD/MCI or healthy controls. There have recently emerged studies that associate image scans with continuous clinical scores that are expected to contain richer information than a binary outcome. However, very few studies aim at modeling multiple clinical scores simultaneously, even though it is commonly conceived that multivariate outcomes provide correlated and complementary information about the disease pathology. In this article, we propose a sparse multi-response tensor regression method to model multiple outcomes jointly as well as to model multiple voxels of an image jointly. The proposed method is particularly useful to both infer clinical scores and thus disease diagnosis, and to identify brain subregions that are highly relevant to the disease outcomes. We conducted experiments on the Alzheimer's Disease Neuroimaging Initiative (ADNI) dataset, and showed that the proposed method enhances the performance and clearly outperforms the competing solutions.
Heung-Il Suk, Dinggang Shen, Lexin Li
IEEE Trans. Medical Imaging4
2012 Nonlinear dimension reduction with Wright-Fisher kernel for genotype aggregation and association mapping
abstract
MOTIVATION: Association tests based on next-generation sequencing data are often under-powered due to the presence of rare variants and large amount of neutral or protective variants. A successful strategy is to aggregate genetic information within meaningful single-nucleotide polymorphism (SNP) sets, e.g. genes or pathways, and test association on SNP sets. Many existing methods for group-wise tests require specific assumptions about the direction of individual SNP effects and/or perform poorly in the presence of interactions. RESULTS: We propose a joint association test strategy based on two key components: a nonlinear supervised dimension reduction approach for effective SNP information aggregation and a novel kernel specially designed for qualitative genotype data. The new test demonstrates superior performance in identifying causal genes over existing methods across a large variety of disease models simulated from sequence data of real genes. In general, the proposed method provides an association test strategy that can (i) detect both rare and common causal variants, (ii) deal with both additive and interaction effect, (iii) handle both quantitative traits and disease dichotomies and (iv) incorporate non-genetic covariates. In addition, the new kernel can potentially boost the power of the entire family of kernel-based methods for genetic data analysis. AVAILABILITY: The method is implemented in MATLAB. Source code is available upon request. CONTACT: [email protected].
Lexin Li, Hua Zhou 0001
Bioinform.2
2011 Data-driven multi-touch attribution models
abstract
In digital advertising, attribution is the problem of assigning credit to one or more advertisements for driving the user to the desirable actions such as making a purchase. Rather than giving all the credit to the last ad a user sees, multi-touch attribution allows more than one ads to get the credit based on their corresponding contributions. Multi-touch attribution is one of the most important problems in digital advertising, especially when multiple media channels, such as search, display, social, mobile and video are involved. Due to the lack of statistical framework and a viable modeling approach, true data-driven methodology does not exist today in the industry. While predictive modeling has been thoroughly researched in recent years in the digital advertising domain, the attribution problem focuses more on accurate and stable interpretation of the influence of each user interaction to the final user decision rather than just user classification. Traditional classification models fail to achieve those goals.
Xuhui Shao, Lexin Li
KDD2
2006 Survival prediction of diffuse large-B-cell lymphoma based on both clinical and gene expression information
abstract
MOTIVATION: It is important to predict the outcome of patients with diffuse large-B-cell lymphoma after chemotherapy, since the survival rate after treatment of this common lymphoma disease is <50%. Both clinically based outcome predictors and the gene expression-based molecular factors have been proposed independently in disease prognosis. However combining the high-dimensional genomic data and the clinically relevant information to predict disease outcome is challenging. RESULTS: We describe an integrated clinicogenomic modeling approach that combines gene expression profiles and the clinically based International Prognostic Index (IPI) for personalized prediction in disease outcome. Dimension reduction methods are proposed to produce linear combinations of gene expressions, while taking into account clinical IPI information. The extracted summary measures capture all the regression information of the censored survival phenotype given both genomic and clinical data, and are employed as covariates in the subsequent survival model formulation. A case study of diffuse large-B-cell lymphoma data, as well as Monte Carlo simulations, both demonstrate that the proposed integrative modeling improves the prediction accuracy, delivering predictions more accurate than those achieved by using either clinical data or molecular predictors alone.
Lexin Li
Bioinform.1
2004 Dimension reduction methods for microarrays with application to censored survival data
abstract
MOTIVATION: Recent research has shown that gene expression profiles can potentially be used for predicting various clinical phenotypes, such as tumor class, drug response and survival time. While there has been extensive studies on tumor classification, there has been less emphasis on other phenotypic features, in particular, patient survival time or time to cancer recurrence, which are subject to right censoring. We consider in this paper an analysis of censored survival time based on microarray gene expression profiles. RESULTS: We propose a dimension reduction strategy, which combines principal components analysis and sliced inverse regression, to identify linear combinations of genes, that both account for the variability in the gene expression levels and preserve the phenotypic information. The extracted gene combinations are then employed as covariates in a predictive survival model formulation. We apply the proposed method to a large diffuse large-B-cell lymphoma dataset, which consists of 240 patients and 7399 genes, and build a Cox proportional hazards model based on the derived gene expression components. The proposed method is shown to provide a good predictive performance for patient survival, as demonstrated by both the significant survival difference between the predicted risk groups and the receiver operator characteristics analysis. AVAILABILITY: R programs are available upon request from the authors. SUPPLEMENTARY INFORMATION: http://dna.ucdavis.edu/~hli/bioinfo-surv-supp.pdf.
Lexin Li, Hongzhe Li
Bioinform.1