Chenyang Tao

dblp:170/6702 · DBLP profile ↗
← Back
27ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-2155-8180ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Photovoltaic Energy-Harvesting Chip Featuring Self-Adaptive-Monitoring-Time MPPT
Chenyang Tao, Changgui Yang, Kai Huang 0001, Bo Zhao 0003
ISCAS3
2024 MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMs
abstract
Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Salman Avestimehr. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Amir Salman Avestimehr
ACL (1)4
2024 CroMo-Mixup: Augmenting Cross-Model Representations for Continual Self-Supervised Learning
Erum Mushtaq, Duygu Nur Yaldiz, Yavuz Faruk Bakman, Jie Ding 0002, Chenyang Tao, Dimitrios Dimitriadis, Amir Salman Avestimehr
ECCV (80)5
2024 DiNADO: Norm-Disentangled Neurally-Decomposed Oracles for Controlling Language Models
abstract
NeurAlly-Decomposed Oracle (NADO) is a powerful approach for controllable generation with large language models. It is designed to avoid catastrophic forgetting while achieving guaranteed convergence to an entropy-maximized closed-form optimal solution with reasonable modeling capacity. Despite the success, several challenges arise when apply NADO to a wide range of scenarios. Vanilla NADO suffers from gradient vanishing for low-probability control signals and is highly reliant on a regularization to satisfy the stochastic version of Bellman equation. In addition, the vanilla implementation of NADO introduces a few additional transformer layers, suffering from a limited capacity especially compared to other finetune-based model adaptation methods like LoRA. In this paper, we propose a improved version of the NADO algorithm, namely DiNADO (norm-Disentangled NeurAlly-Decomposed Oracles), which improves the performance of the NADO algorithm through disentangling the step-wise global norm over the approximated oracle $R$-value for all potential next-tokens, allowing DiNADO to be combined with finetuning methods like LoRA. We discuss in depth how DiNADO achieves better capacity, stability and flexibility with both empirical and theoretical results. Experiments on formality control in machine translation and the lexically constrained generation task CommonGen demonstrates the significance of the improvements.
Sidi Lu, Wenbo Zhao 0006, Chenyang Tao, Arpit Gupta, Shanchan Wu, Tagyoung Chung, Nanyun Peng 0001
ICML3
2024 Predicting Uncertainty of Generative LLMs with MARS: Meaning-Aware Response Scoring
abstract
Generative Large Language Models (LLMs) have recently been widely utilized for their unprecedented capabil-ities across many tasks. Considering their use in high-stakes environments and for mission-critical applications, the fact that LLMs often can generate inaccurate or misleading results can be potentially harmful, which motivates us to study the correctness of generative LLM outputs. Uncertainty Estimation (UE) in generative LLMs is a developing area, with state-of-the-art probability-based techniques frequently using length-normalized scoring. As an alternative to length-normalized scoring in UE, in this work, we propose Meaning-Aware Response Scoring (MARS). The key idea of MARS is to consider the semantic contribution of each token of the generated sequence to the context of the question during UE. Through extensive experiments on three question-answering datasets across five pretrained LLMs, we show that utilizing MARS during UE results in a universal and significant improvement in UE performance.
Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Amir Salman Avestimehr, Chenyang Tao, Dimitrios Dimitriadis
ISIT5
2023 Unsupervised Melody-to-Lyrics Generation
abstract
Yufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Gunnar Sigurdsson, Chenyang Tao, Wenbo Zhao, Tagyoung Chung, Jing Huang, Nanyun Peng. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yufei Tian, Anjali Narayan-Chen, Shereen Oraby, Alessandra Cervone, Gunnar A. Sigurdsson, Chenyang Tao, Wenbo Zhao 0006, Tagyoung Chung, Jing Huang 0020, Nanyun Peng 0001
ACL (1)6
2023 Calibration and Uncertainty in Neural Time-to-Event Modeling
abstract
Models for predicting the time of a future event are crucial for risk assessment, across a diverse range of applications. Existing time-to-event (survival) models have focused primarily on preserving pairwise ordering of estimated event times (i.e., relative risk). We propose neural time-to-event models that account for calibration and uncertainty while predicting accurate absolute event times. Specifically, an adversarial nonparametric model is introduced for estimating matched time-to-event distributions for probabilistically concentrated and accurate predictions. We also consider replacing the discriminator of the adversarial nonparametric model with a survival-function matching estimator that accounts for model calibration. The proposed estimator can be used as a means of estimating and comparing conditional survival distributions while accounting for the predictive uncertainty of probabilistic models. Extensive experiments show that the distribution matching methods outperform existing approaches in terms of both calibration and concentration of time-to-event distributions.
Paidamoyo Chapfuwa, Chenyang Tao, Chunyuan Li, Karen Chandross, Michael J. Pencina, Lawrence Carin, Ricardo Henao
IEEE Trans. Neural Networks Learn. Syst.2
2022 Gradient Importance Learning for Incomplete Observations
Qitong Gao, Dong Wang 0037, Joshua D. Amason, Siyang Yuan, Chenyang Tao, Ricardo Henao, Majda Hadziahmetovic, Lawrence Carin, Miroslav Pajic
ICLR5
2022 Tight Mutual Information Estimation With Contrastive Fenchel-Legendre Optimization
abstract
Successful applications of InfoNCE (Information Noise-Contrastive Estimation) and its variants have popularized the use of contrastive variational mutual information (MI) estimators in machine learning . While featuring superior stability, these estimators crucially depend on costly large-batch training, and they sacrifice bound tightness for variance reduction. To overcome these limitations, we revisit the mathematics of popular variational MI bounds from the lens of unnormalized statistical modeling and convex optimization. Our investigation yields a new unified theoretical framework encompassing popular variational MI bounds, and leads to a novel, simple, and powerful contrastive MI estimator we name FLO. Theoretically, we show that the FLO estimator is tight, and it converges under stochastic gradient descent. Empirically, the proposed FLO estimator overcomes the limitations of its predecessors and learns more efficiently. The utility of FLO is verified using extensive benchmarks, and we further inspire the community with novel applications in meta-learning. Our presentation underscores the foundational importance of variational MI estimation in data-efficient learning.
Junya Chen, Dong Wang 0037, Yuewei Yang, Xinwei Deng, Lawrence Carin, Chenyang Tao
NeurIPS9
2021 Variational Disentanglement for Rare Event Modeling
abstract
Combining the increasing availability and abundance of healthcare data and the current advances in machine learning methods have created renewed opportunities to improve clinical decision support systems. However, in healthcare risk prediction applications, the proportion of cases with the condition (label) of interest is often very low relative to the available sample size. Though very prevalent in healthcare, such imbalanced classification settings are also common and challenging in many other scenarios. So motivated, we propose a variational disentanglement approach to semi-parametrically learn from rare events in heavily imbalanced classification problems. Specifically, we leverage the imposed extreme-distribution behavior on a latent space to extract information from low-prevalence events, and develop a robust prediction arm that joins the merits of the generalized additive model and isotonic neural nets. Results on synthetic studies and diverse real-world datasets, including mortality prediction on a COVID-19 cohort, demonstrate that the proposed approach outperforms existing alternatives.
Zidi Xiu, Chenyang Tao, Michael Gao, Connor Davis, Benjamin Goldstein 0001, Ricardo Henao
AAAI2
2021 Counterfactual Representation Learning with Balancing Weights
abstract
A key to causal inference with observational data is achieving balance in predictive features associated with each treatment type. Recent literature has explored representation learning to achieve this goal. In this work, we discuss the pitfalls of these strategies – such as a steep trade-off between achieving balance and predictive power – and present a remedy via the integration of balancing weights in causal learning. Specifically, we theoretically link balance to the quality of propensity estimation, emphasize the importance of identifying a proper target population, and elaborate on the complementary roles of feature balancing and weight adjustments. Using these concepts, we then develop an algorithm for flexible, scalable and accurate estimation of causal effects. Finally, we show how the learned weighted representations may serve to facilitate alternative causal learning procedures with appealing statistical features. We conduct an extensive set of experiments on both synthetic examples and standard benchmarks, and report encouraging results relative to state-of-the-art baselines.
Serge Assaad, Shuxi Zeng, Chenyang Tao, Shounak Datta, Nikhil Mehta 0002, Ricardo Henao, Lawrence Carin
AISTATS3
2021 APo-VAE: Text Generation in Hyperbolic Space
abstract
Shuyang Dai, Zhe Gan, Yu Cheng, Chenyang Tao, Lawrence Carin, Jingjing Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Shuyang Dai, Zhe Gan, Yu Cheng 0001, Chenyang Tao, Lawrence Carin, Jingjing Liu 0001
NAACL-HLT4
2021 Multi-Grained Knowledge Distillation for Named Entity Recognition
abstract
Xuan Zhou, Xiao Zhang, Chenyang Tao, Junya Chen, Bing Xu, Wei Wang, Jing Xiao. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Chenyang Tao, Junya Chen, Jing Xiao 0006
NAACL-HLT3
2021 Supercharging Imbalanced Data Learning With Energy-based Contrastive Representation Transfer
abstract
Dealing with severe class imbalance poses a major challenge for many real-world applications, especially when the accurate classification and generalization of minority classes are of primary interest.In computer vision and NLP, learning from datasets with long-tail behavior is a recurring theme, especially for naturally occurring labels. Existing solutions mostly appeal to sampling or weighting adjustments to alleviate the extreme imbalance, or impose inductive bias to prioritize generalizable associations. Here we take a novel perspective to promote sample efficiency and model generalization based on the invariance principles of causality. Our contribution posits a meta-distributional scenario, where the causal generating mechanism for label-conditional features is invariant across different labels. Such causal assumption enables efficient knowledge transfer from the dominant classes to their under-represented counterparts, even if their feature distributions show apparent disparities. This allows us to leverage a causal data augmentation procedure to enlarge the representation of minority classes. Our development is orthogonal to the existing imbalanced data learning techniques thus can be seamlessly integrated. The proposed approach is validated on an extensive set of synthetic and real-world tasks against state-of-the-art solutions.
Junya Chen, Zidi Xiu, Benjamin Goldstein 0001, Ricardo Henao, Lawrence Carin, Chenyang Tao
NeurIPS6
2020 Sequence Generation with Optimal-Transport-Enhanced Reinforcement Learning
abstract
Reinforcement learning (RL) has been widely used to aid training in language generation. This is achieved by enhancing standard maximum likelihood objectives with user-specified reward functions that encourage global semantic consistency. We propose a principled approach to address the difficulties associated with RL-based solutions, namely, high-variance gradients, uninformative rewards and brittle training. By leveraging the optimal transport distance, we introduce a regularizer that significantly alleviates the above issues. Our formulation emphasizes the preservation of semantic features, enabling end-to-end training instead of ad-hoc fine-tuning, and when combined with RL, it controls the exploration space for more efficient model updates. To validate the effectiveness of the proposed solution, we perform a comprehensive evaluation covering a wide variety of NLP tasks: machine translation, abstractive text summarization and image caption, with consistent improvements over competing solutions.
Liqun Chen 0001, Ke Bai 0001, Chenyang Tao, Yizhe Zhang 0002, Guoyin Wang 0002, Wenlin Wang, Ricardo Henao, Lawrence Carin
AAAI3
2020 Advancing weakly supervised cross-domain alignment with optimal transport
Siyang Yuan, Ke Bai 0001, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Chunyuan Li, Guoyin Wang 0002, Ricardo Henao, Lawrence Carin
BMVC5
2020 Improving Text Generation with Student-Forcing Optimal Transport
abstract
Jianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu, Yuhchen Lin, Liqun Chen, Yizhe Zhang, Chenyang Tao, Ruiyi Zhang, Wenlin Wang, Dinghan Shen, Qian Yang, Lawrence Carin. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020.
Jianqiao Li, Chunyuan Li, Guoyin Wang 0002, Hao Fu 0002, Yuh-Chen Lin, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Ruiyi Zhang 0002, Wenlin Wang, Dinghan Shen, Qian Yang 0003, Lawrence Carin
EMNLP (1)8
2020 Reconsidering Generative Objectives For Counterfactual Reasoning
abstract
There has been recent interest in exploring generative goals for counterfactual reasoning, such as individualized treatment effect (ITE) estimation. However, existing solutions often fail to address issues that are unique to causal inference, such as covariate balancing and (infeasible) counterfactual validation. As a step towards more flexible, scalable and accurate ITE estimation, we present a novel generative Bayesian estimation framework that integrates representation learning, adversarial matching and causal estimation. By appealing to the Robinson decomposition, we derive a reformulated variational bound that explicitly targets the causal effect estimation rather than specific predictive goals. Our procedure acknowledges the uncertainties in representation and solves a Fenchel mini-max game to resolve the representation imbalance for better counterfactual generalization, justified by new theory. Further, the latent variable formulation employed enables robustness to unobservable latent confounders, extending the scope of its applicability. The utility of the proposed solution is demonstrated via an extensive set of tests against competing solutions, both under various simulation setups and to real-world datasets, with encouraging results reported.
Danni Lu, Chenyang Tao, Junya Chen, Lawrence Carin
NeurIPS2
2019 Improving Textual Network Embedding with Global Attention via Optimal Transport
abstract
Liqun Chen, Guoyin Wang, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang, Wenlin Wang, Yizhe Zhang, Lawrence Carin. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019.
Liqun Chen 0001, Guoyin Wang 0002, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang 0001, Wenlin Wang, Yizhe Zhang 0002, Lawrence Carin
ACL (1)3
2019 Improving Sequence-to-Sequence Learning via Optimal Transport
Liqun Chen 0001, Yizhe Zhang 0002, Ruiyi Zhang 0002, Chenyang Tao, Zhe Gan, Bai Li 0001, Dinghan Shen, Changyou Chen, Lawrence Carin
ICLR (Poster)4
2019 Variational Annealing of GANs: A Langevin Perspective
abstract
The generative adversarial network (GAN) has received considerable attention recently as a model for data synthesis, without an explicit specification of a likelihood function. There has been commensurate interest in leveraging likelihood estimates to improve GAN training. To enrich the understanding of this fast-growing yet almost exclusively heuristic-driven subject, we elucidate the theoretical roots of some of the empirical attempts to stabilize and improve GAN training with the introduction of likelihoods. We highlight new insights from variational theory of diffusion processes to derive a likelihood-based regularizing scheme for GAN training, and present a novel approach to train GANs with an unnormalized distribution instead of empirical samples. To substantiate our claims, we provide experimental evidence on how our theoretically-inspired new algorithms improve upon current practice.
Chenyang Tao, Shuyang Dai, Liqun Chen 0001, Ke Bai 0001, Junya Chen, Chang Liu 0030, Ruiyi Zhang 0002, Georgiy V. Bobashev, Lawrence Carin
ICML1
2019 On Fenchel Mini-Max Learning
abstract
Inference, estimation, sampling and likelihood evaluation are four primary goals of probabilistic modeling. Practical considerations often force modeling approaches to make compromises between these objectives. We present a novel probabilistic learning framework, called Fenchel Mini-Max Learning (FML), that accommodates all four desiderata in a flexible and scalable manner. Our derivation is rooted in classical maximum likelihood estimation, and it overcomes a longstanding challenge that prevents unbiased estimation of unnormalized statistical models. By reformulating MLE as a mini-max game, FML enjoys an unbiased training objective that (i) does not explicitly involve the intractable normalizing constant and (ii) is directly amendable to stochastic gradient descent optimization. To demonstrate the utility of the proposed approach, we consider learning unnormalized statistical models, nonparametric density estimation and training generative models, with encouraging empirical results presented.
Chenyang Tao, Liqun Chen 0001, Shuyang Dai, Junya Chen, Ke Bai 0001, Dong Wang 0037, Jianfeng Feng, Wenlian Lu, Georgiy V. Bobashev, Lawrence Carin
NeurIPS1
2019 Improving Textual Network Learning with Variational Homophilic Embeddings
abstract
The performance of many network learning applications crucially hinges on the success of network embedding algorithms, which aim to encode rich network information into low-dimensional vertex-based vector representations. This paper considers a novel variational formulation of network embeddings, with special focus on textual networks. Different from most existing methods that optimize a discriminative objective, we introduce Variational Homophilic Embedding (VHE), a fully generative model that learns network embeddings by modeling the semantic (textual) information with a variational autoencoder, while accounting for the structural (topology) information through a novel homophilic prior design. Homophilic vertex embeddings encourage similar embedding vectors for related (connected) vertices. The VHE encourages better generalization for downstream tasks, robustness to incomplete observations, and the ability to generalize to unseen vertices. Extensive experiments on real-world networks, for multiple tasks, demonstrate that the proposed method achieves consistently superior performance relative to competing state-of-the-art approaches.
Wenlin Wang, Chenyang Tao, Zhe Gan, Guoyin Wang 0002, Liqun Chen 0001, Xinyuan Zhang 0001, Ruiyi Zhang 0002, Qian Yang 0003, Ricardo Henao, Lawrence Carin
NeurIPS2
2018 Adversarial Time-to-Event Modeling
abstract
Modern health data science applications leverage abundant molecular and electronic health data, providing opportunities for machine learning to build statistical models to support clinical practice. Time-to-event analysis, also called survival analysis, stands as one of the most representative examples of such statistical models. We present a deep-network-based approach that leverages adversarial learning to address a key challenge in modern time-to-event modeling: nonparametric estimation of event-time distributions. We also introduce a principled cost function to exploit information from censored events (events that occur subsequent to the observation window). Unlike most time-to-event models, we focus on the estimation of time-to-event distributions, rather than time ordering. We validate our model on both benchmark and real datasets, demonstrating that the proposed formulation yields significant performance gains relative to a parametric alternative, which we also propose.
Paidamoyo Chapfuwa, Chenyang Tao, Chunyuan Li, Courtney Page, Benjamin Goldstein 0001, Lawrence Carin, Ricardo Henao
ICML2
2018 Variational Inference and Model Selection with Generalized Evidence Bounds
abstract
Recent advances on the scalability and flexibility of variational inference have made it successful at unravelling hidden patterns in complex data. In this work we propose a new variational bound formulation, yielding an estimator that extends beyond the conventional variational bound. It naturally subsumes the importance-weighted and Renyi bounds as special cases, and it is provably sharper than these counterparts. We also present an improved estimator for variational learning, and advocate a novel high signal-to-variance ratio update rule for the variational parameters. We discuss model-selection issues associated with existing evidence-lower-bound-based variational inference procedures, and show how to leverage the flexibility of our new formulation to address them. Empirical evidence is provided to validate our claims.
Liqun Chen 0001, Chenyang Tao, Ruiyi Zhang 0002, Ricardo Henao, Lawrence Carin
ICML2
2018 Chi-square Generative Adversarial Network
abstract
To assess the difference between real and synthetic data, Generative Adversarial Networks (GANs) are trained using a distribution discrepancy measure. Three widely employed measures are information-theoretic divergences, integral probability metrics, and Hilbert space discrepancy metrics. We elucidate the theoretical connections between these three popular GAN training criteria and propose a novel procedure, called $\chi^2$ (Chi-square) GAN, that is conceptually simple, stable at training and resistant to mode collapse. Our procedure naturally generalizes to address the problem of simultaneous matching of multiple distributions. Further, we propose a resampling strategy that significantly improves sample quality, by repurposing the trained critic function via an importance weighting mechanism. Experiments show that the proposed procedure improves stability and convergence, and yields state-of-art results on a wide range of generative modeling tasks.
Chenyang Tao, Liqun Chen 0001, Ricardo Henao, Jianfeng Feng, Lawrence Carin
ICML1
2018 Adversarial Text Generation via Feature-Mover's Distance
abstract
Generative adversarial networks (GANs) have achieved significant success in generating real-valued data. However, the discrete nature of text hinders the application of GAN to text-generation tasks. Instead of using the standard GAN objective, we propose to improve text-generation GAN via a novel approach inspired by optimal transport. Specifically, we consider matching the latent feature distributions of real and synthetic sentences using a novel metric, termed the feature-mover's distance (FMD). This formulation leads to a highly discriminative critic and easy-to-optimize objective, overcoming the mode-collapsing and brittle-training problems in existing methods. Extensive experiments are conducted on a variety of tasks to evaluate the proposed model empirically, including unconditional text generation, style transfer from non-parallel text, and unsupervised cipher cracking. The proposed model yields superior performance, demonstrating wide applicability and effectiveness.
Liqun Chen 0001, Shuyang Dai, Chenyang Tao, Zhe Gan, Dinghan Shen, Yizhe Zhang 0002, Guoyin Wang 0002, Ruiyi Zhang 0002, Lawrence Carin
NeurIPS3