EDBT 2026 Demo / reviewers in the wild / expert
Ricardo Henao
dblp:27/3207
· DBLP profile ↗
89ranked-venue papers
7as first author
41since 2021 · last 2026
0000-0003-4980-845XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 6 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SIDEKICK: A Semantically Integrated Resource for Drug Effects, Indications, and ContraindicationsabstractPharmacovigilance and clinical decision support systems utilize structured drug safety data to guide medical practice. However, existing datasets frequently depend on terminologies such as MedDRA, which limits their semantic reasoning capabilities and their interoperability with Semantic Web ontologies and knowledge graphs. To address this gap, we developed SIDEKICK, a knowledge graph that standardizes drug indications, contraindications, and adverse reactions from FDA Structured Product Labels. We developed and used a workflow based on Large Language Model (LLM) extraction and Graph-Retrieval Augmented Generation (Graph RAG) for ontology mapping. We processed over 50,000 drug labels and mapped terms to the Human Phenotype Ontology (HPO), the MONDO Disease Ontology, and RxNorm. Our semantically integrated resource outperforms the SIDER and ONSIDES databases when applied to the task of drug repurposing by side effect similarity. We serialized the dataset as a Resource Description Framework (RDF) graph and employed the Semanticscience Integrated Ontology (SIO) as upper level ontology to further improve interoperability. Consequently, SIDEKICK enables automated safety surveillance and phenotype-based similarity analysis for drug repurposing. Mohammad Ashhad, Olga Mashkova, Ricardo Henao, Robert Hoehndorf |
ESWC (2) | 3 |
| 2025 | Cross-Modal Imputation and Uncertainty Estimation for Spatial TranscriptomicsabstractHigh-resolution spatial transcriptomics (ST) technologies can capture gene expression at the cellular level along with spatial information, but are limited in the number of genes that can be profiled. In contrast, single-cell RNA sequencing (SC) provides more comprehensive gene expression profiles but lacks spatial context. To bridge these gaps, existing methods typically focus on single-modality prediction tasks, leveraging complementary information from the other modality. Here, we propose an attention-based cross-modal framework that simultaneously imputes gene expression for ST and recovers spatial locations for SC, while also providing uncertainty estimates for the expression of the imputed genes. Our approach was evaluated on three real-world datasets, where it consistently outperformed state-of-the-art methods in spatial gene profile imputation. Moreover, our framework enhances latent embedding integration between the two modalities, resulting in more accurate spatial position estimates. Ricardo Henao |
AISTATS | 2 |
| 2025 | Learning Subjective Label Distributions via Sociocultural DescriptorsabstractSubjectivity in NLP tasks, e.g., toxicity classification, has emerged as a critical challenge precipitated by the increased deployment of NLP systems in content-sensitive domains.Conventional approaches aggregate annotator judgements (labels), ignoring minority perspectives and overlooking the influence of the sociocultural context behind such annotations.We propose a framework 1 where subjectivity in binary labels is modeled as an empirical distribution accounting for the variation in annotators through human values extracted from sociocultural descriptors using a language model.The framework also allows for downstream tasks such as population and sociocultural grouplevel majority label prediction.Experiments on three toxicity datasets covering human-chatbot conversations and social media posts annotated with diverse annotator pools demonstrate that our approach yields well-calibrated toxicity distribution predictions across binary toxicity labels, which are further used for majority label prediction across cultural subgroups, improving over existing methods. Mohammed Fayiz Parappan, Ricardo Henao |
EMNLP | 2 |
| 2025 | Learning Survival Distributions with the Asymmetric Laplace DistributionabstractProbabilistic survival analysis models seek to estimate the distribution of the future occurrence (time) of an event given a set of covariates. In recent years, these models have preferred nonparametric specifications that avoid directly estimating survival distributions via discretization. Specifically, they estimate the probability of an individual event at fixed times or the time of an event at fixed probabilities (quantiles), using supervised learning. Borrowing ideas from the quantile regression literature, we propose a parametric survival analysis method based on the Asymmetric Laplace Distribution (ALD). This distribution allows for closed-form calculation of popular event summaries such as mean, median, mode, variation, and quantiles. The model is optimized by maximum likelihood to learn, at the individual level, the parameters (location, scale, and asymmetry) of the ALD distribution. Extensive results on synthetic and real-world data demonstrate that the proposed method outperforms parametric and nonparametric approaches in terms of accuracy, discrimination and calibration. Deming Sheng, Ricardo Henao |
ICML | 2 |
| 2025 | On Understanding Attention-Based In-Context Learning for Categorical DataabstractIn-context learning based on attention models is examined for data with categorical outcomes, with inference in such models viewed from the perspective of functional gradient descent (GD). We develop a network composed of attention blocks, with each block employing a self-attention layer followed by a cross-attention layer, with associated skip connections. This model can exactly perform multi-step functional GD inference for in-context inference with categorical observations. We perform a theoretical analysis of this setup, generalizing many prior assumptions in this line of work, including the class of attention mechanisms for which it is appropriate. We demonstrate the framework empirically on synthetic data, image classification and language generation. Aaron T. Wang, William Convertino, Ricardo Henao, Lawrence Carin |
ICML | 4 |
| 2025 | Learning to Substitute Words with Model-based Score RankingabstractHongye Liu, Ricardo Henao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hongye Liu, Ricardo Henao |
NAACL (Long Papers) | 2 |
| 2025 | Coupling Generative Modeling and an Autoencoder with the Causal BridgeabstractWe consider inferring the causal effect of a treatment (intervention) on an outcome of interest in situations where there is potentially an unobserved confounder influencing both the treatment and the outcome. This is achievable by assuming access to two separate sets of control (proxy) measurements associated with treatment and outcomes, which are used to estimate treatment effects through a function termed the *causal bridge* (CB). We present a new theoretical perspective, associated assumptions for when estimating treatment effects with the CB is feasible, and a bound on the average error of the treatment effect when the CB assumptions are violated. From this new perspective, we then demonstrate how coupling the CB with an autoencoder architecture allows for the sharing of statistical strength between observed quantities (proxies, treatment, and outcomes), thus improving the quality of the CB estimates. Experiments on synthetic and real-world data demonstrate the effectiveness of the proposed approach relative to state-of-the-art methodology for causal inference with proxy measurements. Ruolin Meng, Ming-Yu Chung, Dhanajit Brahma, Ricardo Henao, Lawrence Carin |
NeurIPS | 4 |
| 2025 | Exploring trade-offs in equitable stroke risk prediction with parity-constrained and race-free models
Matthew Engelhard, Daniel Wojdyla, Michael J. Pencina, Ricardo Henao |
Artif. Intell. Medicine | 5 |
| 2024 | Adaptive Discretization for Event PredicTion (ADEPT)abstractRecently developed survival analysis methods improve upon existing approaches by predicting the probability of event occurrence in each of a number pre-specified (discrete) time intervals. By avoiding placing strong parametric assumptions on the event density, this approach tends to improve prediction performance, particularly when data are plentiful. However, in clinical settings with limited available data, it is often preferable to judiciously partition the event time space into a limited number of intervals well suited to the prediction task at hand. In this work, we develop Adaptive Discretization for Event PredicTion (ADEPT) to learn from data a set of cut points defining such a partition. We show that in two simulated datasets, we are able to recover intervals that match the underlying generative model. We then demonstrate improved prediction performance on three real-world observational datasets, including a large, newly harmonized stroke risk prediction dataset. Finally, we argue that our approach facilitates clinical decision-making by suggesting time intervals that are most appropriate for each task, in the sense that they facilitate more accurate risk prediction. Jimmy Hickey, Ricardo Henao, Daniel Wojdyla, Michael J. Pencina, Matthew Engelhard |
AISTATS | 2 |
| 2024 | Contrastive Learning for Clinical Outcome Prediction with Partial Data SourcesabstractThe use of machine learning models to predict clinical outcomes from (longitudinal) electronic health record (EHR) data is becoming increasingly popular due to advances in deep architectures, representation learning, and the growing availability of large EHR datasets. Existing models generally assume access to the same data sources during both training and inference stages. However, this assumption is often challenged by the fact that real-world clinical datasets originate from various data sources (with distinct sets of covariates), which though can be available for training (in a research or retrospective setting), are more realistically only partially available (a subset of such sets) for inference when deployed. So motivated, we introduce Contrastive Learning for clinical Outcome Prediction with Partial data Sources (CLOPPS), that trains encoders to capture information across different data sources and then leverages them to build classifiers restricting access to a single data source. This approach can be used with existing cross-sectional or longitudinal outcome classification models. We present experiments on two real-world datasets demonstrating that CLOPPS consistently outperforms strong baselines in several practical scenarios. Jonathan Wilson, Benjamin Goldstein 0001, Ricardo Henao |
ICML | 4 |
| 2024 | Translating ethical and quality principles for the effective, safe and fair development, deployment and use of artificial intelligence technologies in healthcareabstractOBJECTIVE: The complexity and rapid pace of development of algorithmic technologies pose challenges for their regulation and oversight in healthcare settings. We sought to improve our institution's approach to evaluation and governance of algorithmic technologies used in clinical care and operations by creating an Implementation Guide that standardizes evaluation criteria so that local oversight is performed in an objective fashion. MATERIALS AND METHODS: Building on a framework that applies key ethical and quality principles (clinical value and safety, fairness and equity, usability and adoption, transparency and accountability, and regulatory compliance), we created concrete guidelines for evaluating algorithmic technologies at our institution. RESULTS: An Implementation Guide articulates evaluation criteria used during review of algorithmic technologies and details what evidence supports the implementation of ethical and quality principles for trustworthy health AI. Application of the processes described in the Implementation Guide can lead to algorithms that are safer as well as more effective, fair, and equitable upon implementation, as illustrated through 4 examples of technologies at different phases of the algorithmic lifecycle that underwent evaluation at our academic medical center. DISCUSSION: By providing clear descriptions/definitions of evaluation criteria and embedding them within standardized processes, we streamlined oversight processes and educated communities using and developing algorithmic technologies within our institution. CONCLUSIONS: We developed a scalable, adaptable framework for translating principles into evaluation criteria and specific requirements that support trustworthy implementation of algorithmic technologies in patient care and healthcare operations. Nicoleta J. Economou-Zavlanos, Sophia Bessias, Michael P. Cary, Armando Bedoya, Benjamin Goldstein 0001, John Eric Jelovsek, Cara O'Brien, Nancy Walden, Matthew Elmore, Amanda B. Parrish, Scott Elengold, Kay Lytle, Suresh Balu, Michael E. Lipkin, Afreen Idris Shariff, Michael Gao, David Leverenz, Ricardo Henao, David Y. Ming, David M. Gallagher, Michael J. Pencina, Eric G. Poon |
J. Am. Medical Informatics Assoc. | 18 |
| 2024 | Trans-Balance: Reducing demographic disparity for prediction models in the presence of class imbalance
Chuan Hong, Molei Liu, Daniel Wojdyla, Jimmy Hickey, Michael J. Pencina, Ricardo Henao |
J. Biomed. Informatics | 6 |
| 2024 | A conditional multi-label model to improve prediction of a rare outcome: An illustration predicting autism diagnosis
Wei A. Huang, Matthew Engelhard, Marika Coffman, Elliot D. Hill, Qin Weng, Abby Scheer, Gary Maslow, Ricardo Henao, Geraldine Dawson, Benjamin Goldstein 0001 |
J. Biomed. Informatics | 8 |
| 2024 | Text Feature Adversarial Learning for Text Generation With Knowledge Transfer From GPT2abstractText generation is a key component of many natural language tasks. Motivated by the success of generative adversarial networks (GANs) for image generation, many text-specific GANs have been proposed. However, due to the discrete nature of text, these text GANs often use reinforcement learning (RL) or continuous relaxations to calculate gradients during learning, leading to high-variance or biased estimation. Furthermore, the existing text GANs often suffer from mode collapse (i.e., they have limited generative diversity). To tackle these problems, we propose a new text GAN model named text feature GAN (TFGAN), where adversarial learning is performed in a continuous text feature space. In the adversarial game, GPT2 provides the "true" features, while the generator of TFGAN learns from them. TFGAN is trained by maximum likelihood estimation on text space and adversarial learning on text feature space, effectively combining them into a single objective, while alleviating mode collapse. TFGAN achieves appealing performance in text generation tasks, and it can also be used as a flexible framework for learning text representations. Hao Zhang 0050, Yulai Cong, Zhengjue Wang, Miaoyun Zhao, Liqun Chen 0001, Shijing Si, Ricardo Henao, Lawrence Carin |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | Few-Shot Composition Learning for Image Retrieval with Prompt TuningabstractWe study the problem of composition learning for image retrieval, for which we learn to retrieve target images with search queries in the form of a composition of a reference image and a modification text that describes desired modifications of the image. Existing models of composition learning for image retrieval are generally built with large-scale datasets, demanding extensive training samples, i.e., query-target pairs, as supervision, which restricts their application for the scenario of few-shot learning with only few query-target pairs available. Recently, prompt tuning with frozen pretrained language models has shown remarkable performance when the amount of training data is limited. Inspired by this, we propose a prompt tuning mechanism with the pretrained CLIP model for the task of few-shot composition learning for image retrieval. Specifically, we regard the representation of the reference image as a trainable visual prompt, prefixed to the embedding of the text sequence. One challenge is to efficiently train visual prompt with few-shot samples. To deal with this issue, we further propose a self-upervised auxiliary task via ensuring that the reference image can retrieve itself when no modification information is given from the text, which facilitates training for the visual prompt, while not requiring additional annotations for query-target pairs. Experiments on multiple benchmarks show that our proposed model can yield superior performance when trained with only few query-target pairs. Junda Wu, Rui Wang 0088, Handong Zhao, Ruiyi Zhang 0002, Chaochao Lu, Shuai Li 0010, Ricardo Henao |
AAAI | 7 |
| 2023 | Estimating Total Correlation with Mutual Information EstimatorsabstractTotal correlation (TC) is a fundamental concept in information theory that measures statistical dependency among multiple random variables. Recently, TC has shown noticeable effectiveness as a regularizer in many learning tasks, where the correlation among multiple latent embeddings requires to be jointly minimized or maximized. However, calculating precise TC values is challenging, especially when the closed-form distributions of embedding variables are unknown. In this paper, we introduce a unified framework to estimate total correlation values with sample-based mutual information (MI) estimators. More specifically, we discover a relation between TC and MI and propose two types of calculation paths (tree-like and line-like) to decompose TC into MI terms. With each MI term being bounded, the TC values can be successfully estimated. Further, we provide theoretical analyses concerning the statistical consistency of the proposed TC estimators. Experiments are presented on both synthetic and real-world scenarios, where our estimators demonstrate effectiveness in all TC estimation, minimization, and maximization tasks. Ke Bai 0001, Pengyu Cheng, Weituo Hao, Ricardo Henao, Larry Carin |
AISTATS | 4 |
| 2023 | Toward Fairness in Text Generation via Mutual Information Minimization based on Importance SamplingabstractPretrained language models (PLMs), such as GPT- 2, have achieved remarkable empirical performance in text generation tasks. However, pre- trained on large-scale natural language corpora, the generated text from PLMs may exhibit social bias against disadvantaged demographic groups. To improve the fairness of PLMs in text generation, we propose to minimize the mutual information between the semantics in the generated text sentences and their demographic polarity, i.e., the demographic group to which the sentence is referring. In this way, the mentioning of a demographic group (e.g., male or female) is encouraged to be independent from how it is described in the generated text, thus effectively alleviating the so cial bias. Moreover, we propose to efficiently estimate the upper bound of the above mutual information via importance sampling, leveraging a natural language corpus. We also propose a distillation mechanism that preserves the language modeling ability of the PLMs after debiasing. Empirical results on real-world benchmarks demonstrate that the proposed method yields superior performance in term of both fairness and language modeling ability. Rui Wang 0088, Pengyu Cheng, Ricardo Henao |
AISTATS | 3 |
| 2023 | An Effective Meaningful Way to Evaluate Survival ModelsabstractOne straightforward metric to evaluate a survival prediction model is based on the Mean Absolute Error (MAE) – the average of the absolute difference between the time predicted by the model and the true event time, over all subjects. Unfortunately, this is challenging because, in practice, the test set includes (right) censored individuals, meaning we do not know when a censored individual actually experienced the event. In this paper, we explore various metrics to estimate MAE for survival datasets that include (many) censored individuals. Moreover, we introduce a novel and effective approach for generating realistic semi-synthetic survival datasets to facilitate the evaluation of metrics. Our findings, based on the analysis of the semi-synthetic datasets, reveal that our proposed metric (MAE using pseudo-observations) is able to rank models accurately based on their performance, and often closely matches the true MAE – in particular, is better than several alternative methods. Shiang Qi, Neeraj Kumar 0002, Mahtab Farrokh, Weijie Sun 0004, Li-Hao Kuan, Rajesh Ranganath, Ricardo Henao, Russell Greiner |
ICML | 7 |
| 2023 | Neural Insights for Digital Marketing Content DesignabstractIn digital marketing, experimenting with new website content is one of the key levers to improve customer engagement. However, creating successful marketing content is a manual and time-consuming process that lacks clear guiding principles. This paper seeks to close the loop between content creation and online experimentation by offering marketers AI-driven actionable insights based on historical data to improve their creative process. We present a neural-network-based system that scores and extracts insights from a marketing content design. Namely, a multimodal neural network predicts the attractiveness of marketing contents, and a post-hoc attribution method generates actionable insights for marketers to improve their content in specific marketing locations. Our insights not only point out the advantages and drawbacks of a given current content, but also provide design recommendations based on historical data. We show that our scoring model and insights work well both quantitatively and qualitatively. Fanjie Kong, Yuan Li 0031, Houssam Nassif, Tanner Fiez, Ricardo Henao, Shreya Chakrabarti |
KDD | 5 |
| 2023 | Mitigating Test-Time Bias for Fair Image RetrievalabstractWe address the challenge of generating fair and unbiased image retrieval results given neutral textual queries (with no explicit gender or race connotations), while maintaining the utility (performance) of the underlying vision-language (VL) model. Previous methods aim to disentangle learned representations of images and text queries from gender and racial characteristics. However, we show these are inadequate at alleviating bias for the desired equal representation result, as there usually exists test-time bias in the target retrieval set. So motivated, we introduce a straightforward technique, Post-hoc Bias Mitigation (PBM), that post-processes the outputs from the pre-trained vision-language model. We evaluate our algorithm on real-world image search datasets, Occupation 1 and 2, as well as two large-scale image-text datasets, MS-COCO and Flickr30k. Our approach achieves the lowest bias, compared with various existing bias-mitigation methods, in text-based image retrieval result while maintaining satisfactory retrieval performance. The source code is publicly available at \url{https://github.com/timqqt/Fair_Text_based_Image_Retrieval}. Fanjie Kong, Weituo Hao, Ricardo Henao |
NeurIPS | 4 |
| 2023 | InfoPrompt: Information-Theoretic Soft Prompt Tuning for Natural Language UnderstandingabstractSoft prompt tuning achieves superior performances across a wide range of few-shot tasks. However, the performances of prompt tuning can be highly sensitive to the initialization of the prompts. We have also empirically observed that conventional prompt tuning methods cannot encode and learn sufficient task-relevant information from prompt tokens. In this work, we develop an information-theoretic framework that formulates soft prompt tuning as maximizing the mutual information between prompts and other model parameters (or encoded representations). This novel view helps us to develop a more efficient, accurate and robust soft prompt tuning method, InfoPrompt. With this framework, we develop two novel mutual information based loss functions, to (i) explore proper prompt initialization for the downstream tasks and learn sufficient task-relevant information from prompt tokens and (ii) encourage the output representation from the pretrained language model to be more aware of the task-relevant information captured in the learnt prompts. Extensive experiments validate that InfoPrompt can significantly accelerate the convergence of the prompt tuning and outperform traditional prompt tuning methods. Finally, we provide a formal theoretical result to show that a gradient descent type algorithm can be used to train our mutual information loss. Junda Wu, Tong Yu 0001, Rui Wang 0088, Zhao Song 0002, Ruiyi Zhang 0002, Handong Zhao, Chaochao Lu, Shuai Li 0010, Ricardo Henao |
NeurIPS | 9 |
| 2023 | Pushing the Efficiency Limit Using Structured Sparse ConvolutionsabstractWeight pruning is among the most popular approaches for compressing deep convolutional neural networks. Recent work suggests that in a randomly initialized deep neural network, there exist sparse subnetworks that achieve performance comparable to the original network. Unfortunately, finding these subnetworks involves iterative stages of training and pruning, which can be computationally expensive. We propose Structured Sparse Convolution (SSC), that leverages the inherent structure in images to reduce the parameters in the convolutional filter. This leads to improved efficiency of convolutional architectures compared to existing methods that perform pruning at initialization. We show that SSC is a generalization of commonly used layers (depthwise, groupwise and pointwise convolution) in "efficient architectures." Extensive experiments on well-known CNN models and datasets show the effectiveness of the proposed method. Architectures based on SSC achieve state-of-the-art performance compared to baselines on CIFAR10, CIFAR-100, Tiny-ImageNet, and ImageNet classification benchmarks. Our source code is publicly available at https://github.com/vkvermaa/SSC. Vinay Kumar Verma, Nikhil Mehta 0002, Shijing Si, Ricardo Henao, Lawrence Carin |
WACV | 4 |
| 2023 | Enhancing early autism prediction based on electronic records using clinical narratives
Junya Chen, Matthew Engelhard, Ricardo Henao, Samuel Berchuck, Brian Eichner, Eliana M. Perrin, Guillermo Sapiro, Geraldine Dawson |
J. Biomed. Informatics | 3 |
| 2023 | Calibration and Uncertainty in Neural Time-to-Event ModelingabstractModels for predicting the time of a future event are crucial for risk assessment, across a diverse range of applications. Existing time-to-event (survival) models have focused primarily on preserving pairwise ordering of estimated event times (i.e., relative risk). We propose neural time-to-event models that account for calibration and uncertainty while predicting accurate absolute event times. Specifically, an adversarial nonparametric model is introduced for estimating matched time-to-event distributions for probabilistically concentrated and accurate predictions. We also consider replacing the discriminator of the adversarial nonparametric model with a survival-function matching estimator that accounts for model calibration. The proposed estimator can be used as a means of estimating and comparing conditional survival distributions while accounting for the predictive uncertainty of probabilistic models. Extensive experiments show that the distribution matching methods outperform existing approaches in terms of both calibration and concentration of time-to-event distributions. Paidamoyo Chapfuwa, Chenyang Tao, Chunyuan Li, Karen Chandross, Michael J. Pencina, Lawrence Carin, Ricardo Henao |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2023 | Learning Hierarchical Document Graphs From Multilevel Sentence RelationsabstractOrganizing the implicit topology of a document as a graph, and further performing feature extraction via the graph convolutional network (GCN), has proven effective in document analysis. However, existing document graphs are often restricted to expressing single-level relations, which are predefined and independent of downstream learning. A set of learnable hierarchical graphs are built to explore multilevel sentence relations, assisted by a hierarchical probabilistic topic model. Based on these graphs, multiple parallel GCNs are used to extract multilevel semantic features, which are aggregated by an attention mechanism for different document-comprehension tasks. Equipped with variational inference, the graph construction and GCN are learned jointly, allowing the graphs to evolve dynamically to better match the downstream task. The effectiveness and efficiency of the proposed multilevel sentence relation graph convolutional network (MuserGCN) is demonstrated via experiments on document classification, abstractive summarization, and matching. Hao Zhang 0050, Chaojie Wang 0001, Zhengjue Wang, Zhibin Duan, Bo Chen 0001, Mingyuan Zhou, Ricardo Henao, Lawrence Carin |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2022 | Few-Shot Class-Incremental Learning for Named Entity RecognitionabstractRui Wang, Tong Yu, Handong Zhao, Sungchul Kim, Subrata Mitra, Ruiyi Zhang, Ricardo Henao. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Rui Wang 0088, Tong Yu 0001, Handong Zhao, Sungchul Kim, Subrata Mitra, Ruiyi Zhang 0002, Ricardo Henao |
ACL (1) | 7 |
| 2022 | Disentangling Whether from When in a Neural Mixture Cure Model for Failure Time DataabstractThe mixture cure model allows failure probability to be estimated separately from failure timing in settings wherein failure never occurs in a subset of the population. In this paper, we draw on insights from representation learning and causal inference to develop a neural network based mixture cure model that is free of distributional assumptions, yielding improved prediction of failure timing, yet still effectively disentangles information about failure timing from information about failure probability. Our approach also mitigates effects of selection biases in the observation of failure and censoring times on estimation of the failure density and censoring density, respectively. Results suggest this approach could be applied to distinguish factors predicting failure occurrence versus timing and mitigate biases in real-world observational datasets. Matthew Engelhard, Ricardo Henao |
AISTATS | 2 |
| 2022 | Efficient Classification of Very Large Images with Tiny ObjectsabstractAn increasing number of applications in computer vision, specially, in medical imaging and remote sensing, become challenging when the goal is to classify very large images with tiny informative objects. Specifically, these classification tasks face two key challenges: i) the size of the input image is usually in the order of mega- or giga-pixels, however, existing deep architectures do not easily operate on such big images due to memory constraints, consequently, we seek a memory-efficient method to process these images; and ii) only a very small fraction of the input images are informative of the label of interest, resulting in low region of interest (ROI) to image ratio. However, most of the current convolutional neural networks (CNNs) are designed for image classification datasets that have relatively large ROIs and small image sizes (sub-megapixel). Existing approaches have addressed these two challenges in isolation. We present an end-to-end CNN model termed Zoom-In network that leverages hierarchical attention sampling for classification of large images with tiny objects using a single GPU. We evaluate our method on four large-image histopathology, road-scene and satellite imaging datasets, and one gigapixel pathology dataset. Experimental results show that our model achieves higher accuracy than existing methods while requiring less memory resources. Fanjie Kong, Ricardo Henao |
CVPR | 2 |
| 2022 | Open World Classification with Adaptive Negative SamplesabstractOpen world classification is a task in natural language processing with key practical relevance and impact.Since the open or unknown category data only manifests in the inference phase, finding a model with a suitable decision boundary accommodating for the identification of known classes and discrimination of the open category is challenging.The performance of existing models is limited by the lack of effective open category data during the training stage or the lack of a good mechanism to learn appropriate decision boundaries.We propose an approach based on adaptive negative samples (ANS) designed to generate effective synthetic open category samples in the training stage and without requiring any prior knowledge or external datasets.Empirically, we find a significant advantage in using auxiliary one-versus-rest binary classifiers, which effectively utilize the generated negative samples and avoid the complex threshold-seeking stage in previous works.Extensive experiments on three benchmark datasets show that ANS achieves significant improvements over stateof-the-art methods. Ke Bai 0001, Guoyin Wang 0002, Jiwei Li 0001, Puyang Xu, Ricardo Henao, Lawrence Carin |
EMNLP | 7 |
| 2022 | Wasserstein Cross-Lingual Alignment For Named Entity RecognitionabstractSupervised training of Named Entity Recognition (NER) models generally require large amounts of annotations, which are hardly available for less widely used (low resource) languages, e.g., Armenian and Dutch. Therefore, it will be desirable if we could leverage knowledge extracted from a high resource language (source), e.g., English, so that NER models for the low resource languages (target) could be trained more efficiently with less cost associated with annotations. In this paper, we study cross-lingual alignment for NER, an approach for transferring knowledge from high-to low-resource languages, via the alignment of token embeddings between different languages. Specifically, we propose to align by minimizing the Wasserstein distance between the contextualized token embeddings from source and target languages. Experimental results show that our method yields improved performance over existing works for cross-lingual alignment in NER tasks. Rui Wang 0088, Ricardo Henao |
ICASSP | 2 |
| 2022 | Gradient Importance Learning for Incomplete Observations
Qitong Gao, Dong Wang 0037, Joshua D. Amason, Siyang Yuan, Chenyang Tao, Ricardo Henao, Majda Hadziahmetovic, Lawrence Carin, Miroslav Pajic |
ICLR | 6 |
| 2022 | Capturing actionable dynamics with structured latent ordinary differential equationsabstractEnd-to-end learning of dynamical systems with black-box models, such as neural ordinary differential equations (ODEs), provides a flexible framework for learning dynamics from data without prescribing a mathematical model for the dynamics. Unfortunately, this flexibility comes at the cost of understanding the dynamical system, for which ODEs are used ubiquitously. Further, experimental data are collected under various conditions (inputs), such as treatments, or grouped in some way, such as part of sub-populations. Understanding the effects of these system inputs on system outputs is crucial to have any meaningful model of a dynamical system. To that end, we propose a structured latent ODE model that explicitly captures system input variations within its latent representation. Building on a static latent variable specification, our model learns (independent) stochastic factors of variation for each input to the system, thus separating the effects of the system inputs in the latent space. This approach provides actionable modeling through the controlled generation of time-series data for novel input combinations (or perturbations). Additionally, we propose a flexible approach for quantifying uncertainties, leveraging a quantile regression formulation. Results on challenging biological datasets show consistent improvements over competitive baselines in the controlled generation of observational data and inference of biologically meaningful system inputs. Paidamoyo Chapfuwa, Sherri Rose, Lawrence Carin, Edward Meeds, Ricardo Henao |
UAI | 5 |
| 2022 | TAMC: A deep-learning approach to predict motif-centric transcriptional factor binding activity based on ATAC-seq profileabstractDetermining transcriptional factor binding sites (TFBSs) is critical for understanding the molecular mechanisms regulating gene expression in different biological conditions. Biological assays designed to directly mapping TFBSs require large sample size and intensive resources. As an alternative, ATAC-seq assay is simple to conduct and provides genomic cleavage profiles that contain rich information for imputing TFBSs indirectly. Previous footprint-based tools are inheritably limited by the accuracy of their bias correction algorithms and the efficiency of their feature extraction models. Here we introduce TAMC (Transcriptional factor binding prediction from ATAC-seq profile at Motif-predicted binding sites using Convolutional neural networks), a deep-learning approach for predicting motif-centric TF binding activity from paired-end ATAC-seq data. TAMC does not require bias correction during signal processing. By leveraging a one-dimensional convolutional neural network (1D-CNN) model, TAMC make predictions based on both footprint and non-footprint features at binding sites for each TF and outperforms existing footprinting tools in TFBS prediction particularly for ATAC-seq data with limited sequencing depth. Ricardo Henao |
PLoS Comput. Biol. | 2 |
| 2021 | Variational Disentanglement for Rare Event ModelingabstractCombining the increasing availability and abundance of healthcare data and the current advances in machine learning methods have created renewed opportunities to improve clinical decision support systems. However, in healthcare risk prediction applications, the proportion of cases with the condition (label) of interest is often very low relative to the available sample size. Though very prevalent in healthcare, such imbalanced classification settings are also common and challenging in many other scenarios. So motivated, we propose a variational disentanglement approach to semi-parametrically learn from rare events in heavily imbalanced classification problems. Specifically, we leverage the imposed extreme-distribution behavior on a latent space to extract information from low-prevalence events, and develop a robust prediction arm that joins the merits of the generalized additive model and isotonic neural nets. Results on synthetic studies and diverse real-world datasets, including mortality prediction on a COVID-19 cohort, demonstrate that the proposed approach outperforms existing alternatives. Zidi Xiu, Chenyang Tao, Michael Gao, Connor Davis, Benjamin Goldstein 0001, Ricardo Henao |
AAAI | 6 |
| 2021 | Counterfactual Representation Learning with Balancing WeightsabstractA key to causal inference with observational data is achieving balance in predictive features associated with each treatment type. Recent literature has explored representation learning to achieve this goal. In this work, we discuss the pitfalls of these strategies – such as a steep trade-off between achieving balance and predictive power – and present a remedy via the integration of balancing weights in causal learning. Specifically, we theoretically link balance to the quality of propensity estimation, emphasize the importance of identifying a proper target population, and elaborate on the complementary roles of feature balancing and weight adjustments. Using these concepts, we then develop an algorithm for flexible, scalable and accurate estimation of causal effects. Finally, we show how the learned weighted representations may serve to facilitate alternative causal learning procedures with appealing statistical features. We conduct an extensive set of experiments on both synthetic examples and standard benchmarks, and report encouraging results relative to state-of-the-art baselines. Serge Assaad, Shuxi Zeng, Chenyang Tao, Shounak Datta, Nikhil Mehta 0002, Ricardo Henao, Lawrence Carin |
AISTATS | 6 |
| 2021 | Wasserstein Contrastive Representation DistillationabstractThe primary goal of knowledge distillation (KD) is to encapsulate the information of a model learned from a teacher network into a student network, with the latter being more compact than the former. Existing work, e.g., using Kullback-Leibler divergence for distillation, may fail to capture important structural knowledge in the teacher network and often lacks the ability for feature generalization, particularly in situations when teacher and student are built to address different classification tasks. We propose Wasserstein Contrastive Representation Distillation (WCoRD), which leverages both primal and dual forms of Wasserstein distance for KD. The dual form is used for global knowledge transfer, yielding a contrastive learning objective that maximizes the lower bound of mutual information between the teacher and the student networks. The primal form is used for local contrastive knowledge transfer within a mini-batch, effectively matching the distributions of features between the teacher and the student networks. Experiments demonstrate that the proposed WCoRD method outperforms state-of-the-art approaches on privileged information distillation, model compression and cross-modal transfer. Liqun Chen 0001, Dong Wang 0037, Zhe Gan, Jingjing Liu 0001, Ricardo Henao, Lawrence Carin |
CVPR | 5 |
| 2021 | Unsupervised Paraphrasing Consistency Training for Low Resource Named Entity RecognitionabstractUnsupervised consistency training is a way of semi-supervised learning that encourages consistency in model predictions between the original and augmented data.For Named Entity Recognition (NER), existing approaches augment the input sequence with token replacement, assuming annotations on the replaced positions unchanged.In this paper, we explore the use of paraphrasing as a more principled data augmentation scheme for NER unsupervised consistency training.Specifically, we convert Conditional Random Field (CRF) into a multi-label classification module and encourage consistency on the entity appearance between the original and paraphrased sequences.Experiments show that our method is especially effective when annotations are limited. Rui Wang 0088, Ricardo Henao |
EMNLP (1) | 2 |
| 2021 | SpanPredict: Extraction of Predictive Document Spans with Neural AttentionabstractVivek Subramanian, Matthew Engelhard, Sam Berchuck, Liqun Chen, Ricardo Henao, Lawrence Carin. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Vivek Subramanian, Matthew Engelhard, Samuel Berchuck, Liqun Chen 0001, Ricardo Henao, Lawrence Carin |
NAACL-HLT | 5 |
| 2021 | Supercharging Imbalanced Data Learning With Energy-based Contrastive Representation TransferabstractDealing with severe class imbalance poses a major challenge for many real-world applications, especially when the accurate classification and generalization of minority classes are of primary interest.In computer vision and NLP, learning from datasets with long-tail behavior is a recurring theme, especially for naturally occurring labels. Existing solutions mostly appeal to sampling or weighting adjustments to alleviate the extreme imbalance, or impose inductive bias to prioritize generalizable associations. Here we take a novel perspective to promote sample efficiency and model generalization based on the invariance principles of causality. Our contribution posits a meta-distributional scenario, where the causal generating mechanism for label-conditional features is invariant across different labels. Such causal assumption enables efficient knowledge transfer from the dominant classes to their under-represented counterparts, even if their feature distributions show apparent disparities. This allows us to leverage a causal data augmentation procedure to enlarge the representation of minority classes. Our development is orthogonal to the existing imbalanced data learning techniques thus can be seamlessly integrated. The proposed approach is validated on an extensive set of synthetic and real-world tasks against state-of-the-art solutions. Junya Chen, Zidi Xiu, Benjamin Goldstein 0001, Ricardo Henao, Lawrence Carin, Chenyang Tao |
NeurIPS | 4 |
| 2021 | Weakly supervised instance learning for thyroid malignancy prediction from whole slide cytopathology images
David Dov, Shahar Z. Kovalsky, Serge Assaad, Jonathan Cohen 0004, Danielle Range, Avani A. Pendse, Ricardo Henao, Lawrence Carin |
Medical Image Anal. | 7 |
| 2021 | Machine-learning-based multiple abnormality prediction with large-scale chest computed tomography volumes
Rachel Lea Draelos, David Dov, Maciej A. Mazurowski, Joseph Y. Lo, Ricardo Henao, Geoffrey D. Rubin, Lawrence Carin |
Medical Image Anal. | 5 |
| 2020 | Sequence Generation with Optimal-Transport-Enhanced Reinforcement LearningabstractReinforcement learning (RL) has been widely used to aid training in language generation. This is achieved by enhancing standard maximum likelihood objectives with user-specified reward functions that encourage global semantic consistency. We propose a principled approach to address the difficulties associated with RL-based solutions, namely, high-variance gradients, uninformative rewards and brittle training. By leveraging the optimal transport distance, we introduce a regularizer that significantly alleviates the above issues. Our formulation emphasizes the preservation of semantic features, enabling end-to-end training instead of ad-hoc fine-tuning, and when combined with RL, it controls the exploration space for more efficient model updates. To validate the effectiveness of the proposed solution, we perform a comprehensive evaluation covering a wide variety of NLP tasks: machine translation, abstractive text summarization and image caption, with consistent improvements over competing solutions. Liqun Chen 0001, Ke Bai 0001, Chenyang Tao, Yizhe Zhang 0002, Guoyin Wang 0002, Wenlin Wang, Ricardo Henao, Lawrence Carin |
AAAI | 7 |
| 2020 | Advancing weakly supervised cross-domain alignment with optimal transport
Siyang Yuan, Ke Bai 0001, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Chunyuan Li, Guoyin Wang 0002, Ricardo Henao, Lawrence Carin |
BMVC | 8 |
| 2020 | Learning Autoencoders with Relational RegularizationabstractWe propose a new algorithmic framework for learning autoencoders of data distributions. In this framework, we minimize the discrepancy between the model distribution and the target one, with relational regularization on learnable latent prior. This regularization penalizes the fused Gromov-Wasserstein (FGW) distance between the latent prior and its corresponding posterior, which allows us to learn a structured prior distribution associated with the generative model in a flexible way. Moreover, it helps us co-train multiple autoencoders even if they are with heterogeneous architectures and incomparable latent spaces. We implement the framework with two scalable algorithms, making it applicable for both probabilistic and deterministic autoencoders. Our relational regularized autoencoder (RAE) outperforms existing methods, e.g., variational autoencoder, Wasserstein autoencoder, and their variants, on generating images. Additionally, our relational co-training strategy of autoencoders achieves encouraging results in both synthesis and real-world multi-view learning tasks. Hongteng Xu, Dixin Luo, Ricardo Henao, Svati Shah, Lawrence Carin |
ICML | 3 |
| 2019 | Communication-Efficient Stochastic Gradient MCMC for Neural NetworksabstractLearning probability distributions on the weights of neural networks has recently proven beneficial in many applications. Bayesian methods such as Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) offer an elegant framework to reason about model uncertainty in neural networks. However, these advantages usually come with a high computational cost. We propose accelerating SG-MCMC under the masterworker framework: workers asynchronously and in parallel share responsibility for gradient computations, while the master collects the final samples. To reduce communication overhead, two protocols (downpour and elastic) are developed to allow periodic interaction between the master and workers. We provide a theoretical analysis on the finite-time estimation consistency of posterior expectations, and establish connections to sample thinning. Our experiments on various neural networks demonstrate that the proposed algorithms can greatly reduce training time while achieving comparable (or better) test accuracy/log-likelihood levels, relative to traditional SG-MCMC. When applied to reinforcement learning, it naturally provides exploration for asynchronous policy optimization, with encouraging performance improvement. Chunyuan Li, Changyou Chen, Yunchen Pu, Ricardo Henao, Lawrence Carin |
AAAI | 4 |
| 2019 | Kernel-Based Approaches for Sequence Modeling: Connections to Neural MethodsabstractWe investigate time-dependent data analysis from the perspective of recurrent kernel machines, from which models with hidden units and gated memory cells arise naturally. By considering dynamic gating of the memory cell, a model closely related to the long short-term memory (LSTM) recurrent neural network is derived. Extending this setup to $n$-gram filters, the convolutional neural network (CNN), Gated CNN, and recurrent additive network (RAN) are also recovered as special cases. Our analysis provides a new perspective on the LSTM, while also extending it to $n$-gram convolutional filters. Experiments are performed on natural language processing tasks and on analysis of local field potentials (neuroscience). We demonstrate that the variants we derive from kernels perform on par or even better than traditional neural methods. For the neuroscience application, the new models demonstrate significant improvements relative to the prior state of the art. Kevin J. Liang, Guoyin Wang 0002, Yitong Li 0001, Ricardo Henao, Lawrence Carin |
NeurIPS | 4 |
| 2019 | Improving Textual Network Learning with Variational Homophilic EmbeddingsabstractThe performance of many network learning applications crucially hinges on the success of network embedding algorithms, which aim to encode rich network information into low-dimensional vertex-based vector representations. This paper considers a novel variational formulation of network embeddings, with special focus on textual networks. Different from most existing methods that optimize a discriminative objective, we introduce Variational Homophilic Embedding (VHE), a fully generative model that learns network embeddings by modeling the semantic (textual) information with a variational autoencoder, while accounting for the structural (topology) information through a novel homophilic prior design. Homophilic vertex embeddings encourage similar embedding vectors for related (connected) vertices. The VHE encourages better generalization for downstream tasks, robustness to incomplete observations, and the ability to generalize to unseen vertices. Extensive experiments on real-world networks, for multiple tasks, demonstrate that the proposed method achieves consistently superior performance relative to competing state-of-the-art approaches. Wenlin Wang, Chenyang Tao, Zhe Gan, Guoyin Wang 0002, Liqun Chen 0001, Xinyuan Zhang 0001, Ruiyi Zhang 0002, Qian Yang 0003, Ricardo Henao, Lawrence Carin |
NeurIPS | 9 |
| 2018 | Deconvolutional Latent-Variable Model for Text Sequence MatchingabstractA latent-variable model is introduced for text matching, inferring sentence representations by jointly optimizing generative and discriminative objectives. To alleviate typical optimization challenges in latent-variable models for text, we employ deconvolutional networks as the sequence decoder (generator), providing learned latent codes with more semantic information and better generalization. Our model, trained in an unsupervised manner, yields stronger empirical predictive performance than a decoder based on Long Short-Term Memory (LSTM), with less parameters and considerably faster training. Further, we apply it to text sequence-matching problems. The proposed model significantly outperforms several strong sentence-encoding baselines, especially in the semi-supervised setting. Dinghan Shen, Yizhe Zhang 0002, Ricardo Henao, Qinliang Su, Lawrence Carin |
AAAI | 3 |
| 2018 | NASH: Toward End-to-End Neural Architecture for Generative Semantic HashingabstractDinghan Shen, Qinliang Su, Paidamoyo Chapfuwa, Wenlin Wang, Guoyin Wang, Ricardo Henao, Lawrence Carin. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Dinghan Shen, Qinliang Su, Paidamoyo Chapfuwa, Wenlin Wang, Guoyin Wang 0002, Ricardo Henao, Lawrence Carin |
ACL (1) | 6 |
| 2018 | Baseline Needs More Love: On Simple Word-Embedding-Based Models and Associated Pooling MechanismsabstractDinghan Shen, Guoyin Wang, Wenlin Wang, Martin Renqiang Min, Qinliang Su, Yizhe Zhang, Chunyuan Li, Ricardo Henao, Lawrence Carin. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Dinghan Shen, Guoyin Wang 0002, Wenlin Wang, Martin Renqiang Min, Qinliang Su, Yizhe Zhang 0002, Chunyuan Li, Ricardo Henao, Lawrence Carin |
ACL (1) | 8 |
| 2018 | Joint Embedding of Words and Labels for Text ClassificationabstractGuoyin Wang, Chunyuan Li, Wenlin Wang, Yizhe Zhang, Dinghan Shen, Xinyuan Zhang, Ricardo Henao, Lawrence Carin. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Guoyin Wang 0002, Chunyuan Li, Wenlin Wang, Yizhe Zhang 0002, Dinghan Shen, Xinyuan Zhang 0001, Ricardo Henao, Lawrence Carin |
ACL (1) | 7 |
| 2018 | The Duke Health Data Science Internship Program: Integrating the Educational Mission into Real-World Research
Shelley A. Rusincovitch, Lisa Wruck, Ricardo Henao, Larisa Rodgers, Allison Dunning, Peter Merrill, Hillary Mulder, Robert Overton, Matthew Phelan, Erich Huang, Lawrence Carin, Michael J. Pencina |
AMIA | 3 |
| 2018 | Improved Semantic-Aware Network Embedding with Fine-Grained Word AlignmentabstractNetwork embeddings, which learn lowdimensional representations for each vertex in a large-scale network, have received considerable attention in recent years.For a wide range of applications, vertices in a network are typically accompanied by rich textual information such as user profiles, paper abstracts, etc.We propose to incorporate semantic features into network embeddings by matching important words between text sequences for all pairs of vertices.We introduce a word-by-word alignment framework that measures the compatibility of embeddings between word pairs, and then adaptively accumulates these alignment features with a simple yet effective aggregation function.In experiments, we evaluate the proposed framework on three real-world benchmarks for downstream tasks, including link prediction and multi-label vertex classification.Results demonstrate that our model outperforms state-of-the-art network embedding methods by a large margin. Dinghan Shen, Xinyuan Zhang 0001, Ricardo Henao, Lawrence Carin |
EMNLP | 3 |
| 2018 | Adversarial Time-to-Event ModelingabstractModern health data science applications leverage abundant molecular and electronic health data, providing opportunities for machine learning to build statistical models to support clinical practice. Time-to-event analysis, also called survival analysis, stands as one of the most representative examples of such statistical models. We present a deep-network-based approach that leverages adversarial learning to address a key challenge in modern time-to-event modeling: nonparametric estimation of event-time distributions. We also introduce a principled cost function to exploit information from censored events (events that occur subsequent to the observation window). Unlike most time-to-event models, we focus on the estimation of time-to-event distributions, rather than time ordering. We validate our model on both benchmark and real datasets, demonstrating that the proposed formulation yields significant performance gains relative to a parametric alternative, which we also propose. Paidamoyo Chapfuwa, Chenyang Tao, Chunyuan Li, Courtney Page, Benjamin Goldstein 0001, Lawrence Carin, Ricardo Henao |
ICML | 7 |
| 2018 | Variational Inference and Model Selection with Generalized Evidence BoundsabstractRecent advances on the scalability and flexibility of variational inference have made it successful at unravelling hidden patterns in complex data. In this work we propose a new variational bound formulation, yielding an estimator that extends beyond the conventional variational bound. It naturally subsumes the importance-weighted and Renyi bounds as special cases, and it is provably sharper than these counterparts. We also present an improved estimator for variational learning, and advocate a novel high signal-to-variance ratio update rule for the variational parameters. We discuss model-selection issues associated with existing evidence-lower-bound-based variational inference procedures, and show how to leverage the flexibility of our new formulation to address them. Empirical evidence is provided to validate our claims. Liqun Chen 0001, Chenyang Tao, Ruiyi Zhang 0002, Ricardo Henao, Lawrence Carin |
ICML | 4 |
| 2018 | JointGAN: Multi-Domain Joint Distribution Learning with Generative Adversarial NetsabstractA new generative adversarial network is developed for joint distribution matching.Distinct from most existing approaches, that only learn conditional distributions, the proposed model aims to learn a joint distribution of multiple random variables (domains). This is achieved by learning to sample from conditional distributions between the domains, while simultaneously learning to sample from the marginals of each individual domain.The proposed framework consists of multiple generators and a single softmax-based critic, all jointly trained via adversarial learning.From a simple noise source, the proposed framework allows synthesis of draws from the marginals, conditional draws given observations from a subset of random variables, or complete draws from the full joint distribution. Most examples considered are for joint analysis of two domains, with examples for three domains also presented. Yunchen Pu, Shuyang Dai, Zhe Gan, Weiyao Wang 0002, Guoyin Wang 0002, Yizhe Zhang 0002, Ricardo Henao, Lawrence Carin |
ICML | 7 |
| 2018 | Chi-square Generative Adversarial NetworkabstractTo assess the difference between real and synthetic data, Generative Adversarial Networks (GANs) are trained using a distribution discrepancy measure. Three widely employed measures are information-theoretic divergences, integral probability metrics, and Hilbert space discrepancy metrics. We elucidate the theoretical connections between these three popular GAN training criteria and propose a novel procedure, called $\chi^2$ (Chi-square) GAN, that is conceptually simple, stable at training and resistant to mode collapse. Our procedure naturally generalizes to address the problem of simultaneous matching of multiple distributions. Further, we propose a resampling strategy that significantly improves sample quality, by repurposing the trained critic function via an importance weighting mechanism. Experiments show that the proposed procedure improves stability and convergence, and yields state-of-art results on a wide range of generative modeling tasks. Chenyang Tao, Liqun Chen 0001, Ricardo Henao, Jianfeng Feng, Lawrence Carin |
ICML | 3 |
| 2017 | Guiding Principles for the Duke Connected Care Predictive Modeling Pilot
Eugenie Komives, Shelley A. Rusincovitch, John Paat, Lawrence Carin, Daniel Costello, Michael Gao, Bradley G. Hammill, Ricardo Henao, Nigel B. Neely, Ursula Rogers, Devdutta Sangvai, Mary Schilder, Erich Huang |
AMIA | 8 |
| 2017 | Rationale and Design for the Duke Connected Care Predictive Modeling Pilot with a Medicare Shared Savings Program Population
Shelley A. Rusincovitch, Ricardo Henao, Michael Gao, Lawrence Carin, Ursula Rogers, Nigel B. Neely, Mary Schilder, Daniel Costello, Eugenie Komives, Erich Huang |
AMIA | 2 |
| 2017 | Learning Generic Sentence Representations Using Convolutional Neural NetworksabstractWe propose a new encoder-decoder approach to learn distributed sentence representations that are applicable to multiple purposes.The model is learned by using a convolutional neural network as an encoder to map an input sentence into a continuous vector, and using a long short-term memory recurrent neural network as a decoder.Several tasks are considered, including sentence reconstruction and future sentence prediction.Further, a hierarchical encoderdecoder model is proposed to encode a sentence to predict multiple future sentences.By training our models on a large collection of novels, we obtain a highly generic convolutional sentence encoder that performs well in practice.Experimental results on several benchmark datasets, and across a broad range of applications, demonstrate the superiority of the proposed model over competing methods. Zhe Gan, Yunchen Pu, Ricardo Henao, Chunyuan Li, Xiaodong He 0001, Lawrence Carin |
EMNLP | 3 |
| 2017 | Stochastic Gradient Monomial Gamma SamplerabstractScaling Markov Chain Monte Carlo (MCMC) to estimate posterior distributions from large datasets has been made possible as a result of advances in stochastic gradient techniques. Despite their success, mixing performance of existing methods when sampling from multimodal distributions can be less efficient with insufficient Monte Carlo samples; this is evidenced by slow convergence and insufficient exploration of posterior distributions. We propose a generalized framework to improve the sampling efficiency of stochastic gradient MCMC, by leveraging a generalized kinetics that delivers superior stationary mixing, especially in multimodal distributions, and propose several techniques to overcome the practical issues. We show that the proposed approach is better at exploring a complicated multimodal posterior distribution, and demonstrate improvements over other stochastic gradient MCMC methods on various applications. Yizhe Zhang 0002, Changyou Chen, Zhe Gan, Ricardo Henao, Lawrence Carin |
ICML | 4 |
| 2017 | Adversarial Feature Matching for Text GenerationabstractThe Generative Adversarial Network (GAN) has achieved great success in generating realistic (real-valued) synthetic data. However, convergence issues and difficulties dealing with discrete data hinder the applicability of GAN to text. We propose a framework for generating realistic text via adversarial training. We employ a long short-term memory network as generator, and a convolutional network as discriminator. Instead of using the standard objective of GAN, we propose matching the high-dimensional latent feature distributions of real and synthetic sentences, via a kernelized discrepancy metric. This eases adversarial training by alleviating the mode-collapsing problem. Our experiments show superior performance in quantitative evaluation, and demonstrate that our model can generate realistic-looking sentences. Yizhe Zhang 0002, Zhe Gan, Kai Fan 0002, Zhi Chen 0009, Ricardo Henao, Dinghan Shen, Lawrence Carin |
ICML | 5 |
| 2017 | ALICE: Towards Understanding Adversarial Learning for Joint Distribution MatchingabstractWe investigate the non-identifiability issues associated with bidirectional adversarial training for joint distribution matching. Within a framework of conditional entropy, we propose both adversarial and non-adversarial approaches to learn desirable matched joint distributions for unsupervised and supervised tasks. We unify a broad family of adversarial models as joint distribution matching problems. Our approach stabilizes learning of unsupervised bidirectional adversarial learning methods. Further, we introduce an extension for semi-supervised learning tasks. Theoretical results are validated in synthetic data and real-world applications. Chunyuan Li, Hao Liu 0015, Changyou Chen, Yunchen Pu, Liqun Chen 0001, Ricardo Henao, Lawrence Carin |
NIPS | 6 |
| 2017 | VAE Learning via Stein Variational Gradient DescentabstractA new method for learning variational autoencoders (VAEs) is developed, based on Stein variational gradient descent. A key advantage of this approach is that one need not make parametric assumptions about the form of the encoder distribution. Performance is further enhanced by integrating the proposed encoder with importance sampling. Excellent performance is demonstrated across multiple unsupervised and semi-supervised problems, including semi-supervised analysis of the ImageNet data, demonstrating the scalability of the model to large datasets. Yunchen Pu, Zhe Gan, Ricardo Henao, Chunyuan Li, Shaobo Han, Lawrence Carin |
NIPS | 3 |
| 2017 | Adversarial Symmetric Variational AutoencoderabstractA new form of variational autoencoder (VAE) is developed, in which the joint distribution of data and codes is considered in two (symmetric) forms: (i) from observed data fed through the encoder to yield codes, and (ii) from latent codes drawn from a simple prior and propagated through the decoder to manifest data. Lower bounds are learned for marginal log-likelihood fits observed data and latent codes. When learning with the variational bound, one seeks to minimize the symmetric Kullback-Leibler divergence of joint density functions from (i) and (ii), while simultaneously seeking to maximize the two marginal log-likelihoods. To facilitate learning, a new form of adversarial training is developed. An extensive set of experiments is performed, in which we demonstrate state-of-the-art data reconstruction and generation on several image benchmarks datasets. Yunchen Pu, Weiyao Wang 0002, Ricardo Henao, Liqun Chen 0001, Zhe Gan, Chunyuan Li, Lawrence Carin |
NIPS | 3 |
| 2017 | Deconvolutional Paragraph Representation LearningabstractLearning latent representations from long text sequences is an important first step in many natural language processing applications. Recurrent Neural Networks (RNNs) have become a cornerstone for this challenging task. However, the quality of sentences during RNN-based decoding (reconstruction) decreases with the length of the text. We propose a sequence-to-sequence, purely convolutional and deconvolutional autoencoding framework that is free of the above issue, while also being computationally efficient. The proposed method is simple, easy to implement and can be leveraged as a building block for many applications. We show empirically that compared to RNNs, our framework is better at reconstructing and correcting long paragraphs. Quantitative evaluation on semi-supervised text classification and summarization tasks demonstrate the potential for better utilization of long unlabeled text data. Yizhe Zhang 0002, Dinghan Shen, Guoyin Wang 0002, Zhe Gan, Ricardo Henao, Lawrence Carin |
NIPS | 5 |
| 2016 | Learning a Hybrid Architecture for Sequence Regression and AnnotationabstractWhen learning a hidden Markov model (HMM), sequential observations can often be complemented by real-valued summary response variables generated from the path of hidden states. Such settings arise in numerous domains, including many applications in biology, like motif discovery and genome annotation. In this paper, we present a flexible framework for jointly modeling both latent sequence features and the functional mapping that relates the summary response variables to the hidden state sequence. The algorithm is compatible with a rich set of mapping functions. Results show that the availability of additional continuous response variables can simultaneously improve the annotation of the sequential observations and yield good prediction performance in both synthetic data and real-world datasets. Yizhe Zhang 0002, Ricardo Henao, Lawrence Carin, Jianling Zhong, Alexander J. Hartemink |
AAAI | 2 |
| 2016 | Learning Sigmoid Belief Networks via Monte Carlo Expectation MaximizationabstractBelief networks are commonly used generative models of data, but require expensive posterior estimation to train and test the model. Learning typically proceeds by posterior sampling, variational approximations, or recognition networks, combined with stochastic optimization. We propose using an online Monte Carlo expectation-maximization (MCEM) algorithm to learn the maximum a posteriori (MAP) estimator of the generative model or optimize the variational lower bound of a recognition network. The E-step in this algorithm requires posterior samples, which are already generated in current learning schema. For the M-step, we augment with Polya-Gamma (PG) random variables to give an analytic updating scheme. We show relationships to standard learning approaches by deriving stochastic gradient ascent in the MCEM framework. We apply the proposed methods to both binary and count data. Experimental results show that MCEM improves the convergence speed and often improves hold-out performance over existing learning methods. Our approach is readily generalized to other recognition networks. Zhao Song 0001, Ricardo Henao, David E. Carlson, Lawrence Carin |
AISTATS | 2 |
| 2016 | Triply Stochastic Variational Inference for Non-linear Beta Process Factor AnalysisabstractWe propose a non-linear extension to factor analysis with beta process priors for improved data representation ability. This non-linear Beta Process Factor Analysis (nBPFA) allows data to be represented as a non-linear transformation of a standard sparse factor decomposition. We develop a scalable variational inference framework, which builds upon the ideas of the variational auto-encoder, by allowing latent variables of the model to be sparse. Our framework can be readily used for real-valued, binary and count data. We show theoretically and with experiments that our training scheme, with additive or multiplicative noise on observations, improves performance and prevents overfitting. We benchmark our algorithms on image, text and collaborative filtering datasets. We demonstrate faster convergence rates and competitive performance compared to standard gradient-based approaches. Kai Fan 0002, Yizhe Zhang 0002, Ricardo Henao, Katherine A. Heller |
ICDM | 3 |
| 2016 | Dynamic Poisson Factor AnalysisabstractWe introduce a novel dynamic model for discrete time-series data, in which the temporal sampling may be nonuniform. The model is specified by constructing a hierarchy of Poisson factor analysis blocks, one for the transitions between latent states and the other for the emissions between latent states and observations. Latent variables are binary and linked to Poisson factor analysis via Bernoulli-Poisson specifications. The model is derived for count data but can be readily modified for binary observations. We derive efficient inference via Markov chain Monte Carlo, that scales with the number of non-zeros in the data and latent binary states, yielding significant acceleration compared to related models. Experimental results on benchmark data show the proposed model achieves state-of-the-art predictive performance. Additional experiments on microbiome data demonstrate applicability of the proposed model to interesting problems in computational biology where interpretability is of utmost importance. Yizhe Zhang 0002, Lawrence David, Ricardo Henao, Lawrence Carin |
ICDM | 4 |
| 2016 | Bayesian Dictionary Learning with Gaussian Processes and Sigmoid Belief Networks
Yizhe Zhang 0002, Ricardo Henao, Chunyuan Li, Lawrence Carin |
IJCAI | 2 |
| 2016 | Variational Autoencoder for Deep Learning of Images, Labels and CaptionsabstractA novel variational autoencoder is developed to model images, as well as associated labels or captions. The Deep Generative Deconvolutional Network (DGDN) is used as a decoder of the latent image features, and a deep Convolutional Neural Network (CNN) is used as an image encoder; the CNN is used to approximate a distribution for the latent DGDN features/code. The latent code is also linked to generative models for labels (Bayesian support vector machine) or captions (recurrent neural network). When predicting a label/caption for a new image at test, averaging is performed across the distribution of latent codes; this is computationally efficient as a consequence of the learned CNN-based encoder. Since the framework is capable of modeling the image in the presence/absence of associated labels/captions, a new semi-supervised setting is manifested for CNN learning with images; the framework even allows unsupervised CNN learning, based on images alone. Yunchen Pu, Zhe Gan, Ricardo Henao, Xin Yuan 0002, Chunyuan Li, Andrew Stevens 0005, Lawrence Carin |
NIPS | 3 |
| 2016 | Towards Unifying Hamiltonian Monte Carlo and Slice SamplingabstractWe unify slice sampling and Hamiltonian Monte Carlo (HMC) sampling, demonstrating their connection via the Hamiltonian-Jacobi equation from Hamiltonian mechanics. This insight enables extension of HMC and slice sampling to a broader family of samplers, called Monomial Gamma Samplers (MGS). We provide a theoretical analysis of the mixing performance of such samplers, proving that in the limit of a single parameter, the MGS draws decorrelated samples from the desired target distribution. We further show that as this parameter tends toward this limit, performance gains are achieved at a cost of increasing numerical difficulty and some practical convergence issues. Our theoretical results are validated with synthetic data and real-world applications. Yizhe Zhang 0002, Xiangyu Wang 0006, Changyou Chen, Ricardo Henao, Kai Fan 0002, Lawrence Carin |
NIPS | 4 |
| 2016 | Laplacian Hamiltonian Monte Carlo
Yizhe Zhang 0002, Changyou Chen, Ricardo Henao, Lawrence Carin |
ECML/PKDD (1) | 3 |
| 2016 | Electronic Health Record Analysis via Deep Poisson Factor ModelsabstractElectronic Health Record (EHR) phenotyping utilizes patient data captured through normal medical practice, to identify features that may represent computational medical phenotypes. These features may be used to identify at-risk patients and improve prediction of patient morbidity and mortality. We present a novel deep multi-modality architecture for EHR analysis (applicable to joint analysis of multiple forms of EHR data), based on Poisson Factor Analysis (PFA) modules. Each modality, composed of observed counts, is represented as a Poisson distribution, parameterized in terms of hidden binary units. Information from different modalities is shared via a deep hierarchy of common hidden units. Activation of these binary units occurs with probability characterized as Bernoulli- Poisson link functions, instead of more traditional logistic link functions. In addition, we demonstrate that PFA modules can be adapted to discriminative modalities. To compute model parameters, we derive efficient Markov Chain Monte Carlo (MCMC) inference that scales efficiently, with significant computational gains when compared to related models based on logistic link functions. To explore the utility of these models, we apply them to a subset of patients from the Duke-Durham patient cohort. We identified a cohort of over 16,000 patients with Type 2 Diabetes Mellitus (T2DM) based on diagnosis codes and laboratory tests out of our patient population of over 240,000. Examining the common hidden units uniting the PFA modules, we identify patient features that represent medical concepts. Experiments indicate that our learned features are better able to predict mortality and morbidity than clinical features identified previously in a large-scale clinical trial. Ricardo Henao, James Lu, Joseph E. Lucas, Jeffrey M. Ferranti, Lawrence Carin |
J. Mach. Learn. Res. | 1 |
| 2015 | Learning Deep Sigmoid Belief Networks with Data AugmentationabstractDeep directed generative models are developed. The multi-layered model is designed by stacking sigmoid belief networks, with sparsity-encouraging priors placed on the model parameters. Learning and inference of layer-wise model parameters are implemented in a Bayesian setting. By exploring the idea of data augmentation and introducing auxiliary Polya-Gamma variables, simple and efficient Gibbs sampling and mean-field variational Bayes (VB) inference are implemented. To address large-scale datasets, an online version of VB is also developed. Experimental results are presented for three publicly available datasets: MNIST, Caltech 101 Silhouettes and OCR letters. Zhe Gan, Ricardo Henao, David E. Carlson, Lawrence Carin |
AISTATS | 2 |
| 2015 | Scalable Deep Poisson Factor Analysis for Topic ModelingabstractA new framework for topic modeling is developed, based on deep graphical models, where interactions between topics are inferred through deep latent binary hierarchies. The proposed multi-layer model employs a deep sigmoid belief network or restricted Boltzmann machine, the bottom binary layer of which selects topics for use in a Poisson factor analysis model. Under this setting, topics live on the bottom layer of the model, while the deep specification serves as a flexible prior for revealing topic structure. Scalable inference algorithms are derived by applying Bayesian conditional density filtering algorithm, in addition to extending recently proposed work on stochastic gradient thermostats. Experimental results on several corpora show that the proposed approach readily handles very large collections of text documents, infers structured topic representations, and obtains superior test perplexities when compared with related models. Zhe Gan, Changyou Chen, Ricardo Henao, David E. Carlson, Lawrence Carin |
ICML | 3 |
| 2015 | A Multitask Point Process Predictive ModelabstractPoint process data are commonly observed in fields like healthcare and social science. Designing predictive models for such event streams is an under-explored problem, due to often scarce training data. In this work we propose a multitask point process model, leveraging information from all tasks via a hierarchical Gaussian process (GP). Nonparametric learning functions implemented by a GP, which map from past events to future rates, allow analysis of flexible arrival patterns. To facilitate efficient inference, we propose a sparse construction for this hierarchical model, and derive a variational Bayes method for learning and inference. Experimental results are shown on both synthetic data and an application on real electronic health records. Wenzhao Lian, Ricardo Henao, Vinayak A. Rao, Joseph E. Lucas, Lawrence Carin |
ICML | 2 |
| 2015 | Non-Gaussian Discriminative Factor Models via the Max-Margin Rank-LikelihoodabstractWe consider the problem of discriminative factor analysis for data that are in general non-Gaussian. A Bayesian model based on the ranks of the data is proposed. We first introduce a max-margin version of the rank-likelihood. A discriminative factor model is then developed, integrating the new max-margin rank-likelihood and (linear) Bayesian support vector machines, which are also built on the max-margin principle. The discriminative factor model is further extended to the nonlinear case through mixtures of local linear classifiers, via Dirichlet processes. Fully local conjugacy of the model yields efficient inference with both Markov Chain Monte Carlo and variational Bayes approaches. Extensive experiments on benchmark and real data demonstrate superior performance of the proposed model and its potential for applications in computational biology. Xin Yuan 0002, Ricardo Henao, Ephraim Tsalik, Raymond Langley, Lawrence Carin |
ICML | 2 |
| 2015 | Deep Temporal Sigmoid Belief Networks for Sequence ModelingabstractDeep dynamic generative models are developed to learn sequential dependencies in time-series data. The multi-layered model is designed by constructing a hierarchy of temporal sigmoid belief networks (TSBNs), defined as a sequential stack of sigmoid belief networks (SBNs). Each SBN has a contextual hidden state, inherited from the previous SBNs in the sequence, and is used to regulate its hidden bias. Scalable learning and inference algorithms are derived by introducing a recognition model that yields fast sampling from the variational posterior. This recognition model is trained jointly with the generative model, by maximizing its variational lower bound on the log-likelihood. Experimental results on bouncing balls, polyphonic music, motion capture, and text streams show that the proposed approach achieves state-of-the-art predictive performance, and has the capacity to synthesize various sequences. Zhe Gan, Chunyuan Li, Ricardo Henao, David E. Carlson, Lawrence Carin |
NIPS | 3 |
| 2015 | Deep Poisson Factor ModelingabstractWe propose a new deep architecture for topic modeling, based on Poisson Factor Analysis (PFA) modules. The model is composed of a Poisson distribution to model observed vectors of counts, as well as a deep hierarchy of hidden binary units. Rather than using logistic functions to characterize the probability that a latent binary unit is on, we employ a Bernoulli-Poisson link, which allows PFA modules to be used repeatedly in the deep architecture. We also describe an approach to build discriminative topic models, by adapting PFA modules. We derive efficient inference via MCMC and stochastic variational methods, that scale with the number of non-zeros in the data and binary units, yielding significant efficiency, relative to models based on logistic links. Experiments on several corpora demonstrate the advantages of our model when compared to related deep models. Ricardo Henao, Zhe Gan, James Lu, Lawrence Carin |
NIPS | 1 |
| 2015 | Large-Scale Bayesian Multi-Label Learning via Topic-Based Label EmbeddingsabstractWe present a scalable Bayesian multi-label learning model based on learning low-dimensional label embeddings. Our model assumes that each label vector is generated as a weighted combination of a set of topics (each topic being a distribution over labels), where the combination weights (i.e., the embeddings) for each label vector are conditioned on the observed feature vector. This construction, coupled with a Bernoulli-Poisson link function for each label of the binary label vector, leads to a model with a computational cost that scales in the number of positive labels in the label matrix. This makes the model particularly appealing for real-world multi-label learning problems where the label matrix is usually very massive but highly sparse. Using a data-augmentation strategy leads to full local conjugacy in our model, facilitating simple and very efficient Gibbs sampling, as well as an Expectation Maximization algorithm for inference. Also, predicting the label vector at test time does not require doing an inference for the label embeddings and can be done in closed form. We report results on several benchmark data sets, comparing our model with various state-of-the art methods. Piyush Rai, Changwei Hu, Ricardo Henao, Lawrence Carin |
NIPS | 3 |
| 2014 | Bayesian Nonlinear Support Vector Machines and Discriminative Factor Modeling
Ricardo Henao, Xin Yuan 0002, Lawrence Carin |
NIPS | 1 |
| 2013 | Patient Clustering with Uncoded Text in Electronic Medical Records
Ricardo Henao, Jared Murray, Geoffrey S. Ginsburg, Lawrence Carin, Joseph E. Lucas |
AMIA | 1 |
| 2013 | A flexible statistical model for alignment of label-free proteomics data - incorporating ion mobility and product ion informationabstractBACKGROUND: The goal of many proteomics experiments is to determine the abundance of proteins in biological samples, and the variation thereof in various physiological conditions. High-throughput quantitative proteomics, specifically label-free LC-MS/MS, allows rapid measurement of thousands of proteins, enabling large-scale studies of various biological systems. Prior to analyzing these information-rich datasets, raw data must undergo several computational processing steps. We present a method to address one of the essential steps in proteomics data processing--the matching of peptide measurements across samples. RESULTS: We describe a novel method for label-free proteomics data alignment with the ability to incorporate previously unused aspects of the data, particularly ion mobility drift times and product ion information. We compare the results of our alignment method to PEPPeR and OpenMS, and compare alignment accuracy achieved by different versions of our method utilizing various data characteristics. Our method results in increased match recall rates and similar or improved mismatch rates compared to PEPPeR and OpenMS feature-based alignment. We also show that the inclusion of drift time and product ion information results in higher recall rates and more confident matches, without increases in error rates. CONCLUSIONS: Based on the results presented here, we argue that the incorporation of ion mobility drift time and product ion information are worthy pursuits. Alignment methods should be flexible enough to utilize all available data, particularly with recent advancements in experimental separation methods. Ashlee M. Benjamin, J. Will Thompson, Erik J. Soderblom, Scott J. Geromanos, Ricardo Henao, Virginia B. Kraus, M. Arthur Moseley, Joseph E. Lucas |
BMC Bioinform. | 5 |
| 2012 | Predictive active set selection methods for Gaussian processes
Ricardo Henao, Ole Winther |
Neurocomputing | 1 |
| 2011 | Sparse Linear Identifiable Multivariate Modeling
Ricardo Henao, Ole Winther |
J. Mach. Learn. Res. | 1 |
| 2009 | Bayesian Sparse Factor Models and DAGs Inference and ComparisonabstractIn this paper we present a novel approach to learn directed acyclic graphs (DAG) and factor models within the same framework while also allowing for model comparison between them. For this purpose, we exploit the connection between factor models and DAGs to propose Bayesian hierarchies based on spike and slab priors to promote sparsity, heavy-tailed priors to ensure identifiability and predictive densities to perform the model comparison. We require identifiability to be able to produce variable orderings leading to valid DAGs and sparsity to learn the structures. The effectiveness of our approach is demonstrated through extensive experiments on artificial and biological data showing that our approach outperform a number of state of the art methods. Ricardo Henao, Ole Winther |
NIPS | 1 |
| 2006 | Probabilistic Kernel Principal Component Analysis Through Time
Mauricio A. Álvarez, Ricardo Henao |
ICONIP (1) | 2 |