EDBT 2026 Demo / reviewers in the wild / expert
Lawrence Carin
dblp:23/2168
· DBLP profile ↗
365ranked-venue papers
4as first author
37since 2021 · last 2025
0000-0001-6277-7948ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 249 · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 110 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 44 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 11 · 2 since 2021Theory of computation · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LangMark: A Multilingual Dataset for Automatic Post-EditingabstractDiego Velazquez, Mikaela Grace, Konstantinos Karageorgos, Lawrence Carin, Aaron Schliem, Dimitrios Zaikis, Roger Wechsler. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Diego Velazquez, Mikaela Grace, Konstantinos Karageorgos, Lawrence Carin, Aaron Schliem, Dimitrios Zaikis, Roger Wechsler |
ACL (1) | 4 |
| 2025 | Graph Transformers Dream of Electric FlowabstractWe show theoretically and empirically that the linear Transformer, when applied to graph data, can implement algorithms that solve canonical problems such as electric flow and eigenvector decomposition. The Transformer has access to information on the input graph only via the graph's incidence matrix. We present explicit weight configurations for implementing each algorithm, and we bound the constructed Transformers' errors by the errors of the underlying algorithms. Our theoretical findings are corroborated by experiments on synthetic data. Additionally, on a real-world molecular regression task, we observe that the linear Transformer is capable of learning a more effective positional encoding than the default one based on Laplacian eigenvectors. Our work is an initial step towards elucidating the inner-workings of the Transformer for graph data. Code is available at https://github.com/chengxiang/LinearGraphTransformer Lawrence Carin, Suvrit Sra |
ICLR | 2 |
| 2025 | On Understanding Attention-Based In-Context Learning for Categorical DataabstractIn-context learning based on attention models is examined for data with categorical outcomes, with inference in such models viewed from the perspective of functional gradient descent (GD). We develop a network composed of attention blocks, with each block employing a self-attention layer followed by a cross-attention layer, with associated skip connections. This model can exactly perform multi-step functional GD inference for in-context inference with categorical observations. We perform a theoretical analysis of this setup, generalizing many prior assumptions in this line of work, including the class of attention mechanisms for which it is appropriate. We demonstrate the framework empirically on synthetic data, image classification and language generation. Aaron T. Wang, William Convertino, Ricardo Henao, Lawrence Carin |
ICML | 5 |
| 2025 | Coupling Generative Modeling and an Autoencoder with the Causal BridgeabstractWe consider inferring the causal effect of a treatment (intervention) on an outcome of interest in situations where there is potentially an unobserved confounder influencing both the treatment and the outcome. This is achievable by assuming access to two separate sets of control (proxy) measurements associated with treatment and outcomes, which are used to estimate treatment effects through a function termed the *causal bridge* (CB). We present a new theoretical perspective, associated assumptions for when estimating treatment effects with the CB is feasible, and a bound on the average error of the treatment effect when the CB assumptions are violated. From this new perspective, we then demonstrate how coupling the CB with an autoencoder architecture allows for the sharing of statistical strength between observed quantities (proxies, treatment, and outcomes), thus improving the quality of the CB estimates. Experiments on synthetic and real-world data demonstrate the effectiveness of the proposed approach relative to state-of-the-art methodology for causal inference with proxy measurements. Ruolin Meng, Ming-Yu Chung, Dhanajit Brahma, Ricardo Henao, Lawrence Carin |
NeurIPS | 5 |
| 2025 | From Softmax to Score: Transformers Can Effectively Implement In-Context Denoising StepsabstractTransformers have emerged as powerful meta-learners, with growing evidence that they implement learning algorithms within their forward pass. We study this phenomenon in the context of denoising, presenting a unified framework that shows Transformers can implement (a) manifold denoising via Laplacian flows, (b) score-based denoising from diffusion models, and (c) a generalized form of anisotropic diffusion denoising. Our theory establishes exact equivalence between Transformer attention updates and these algorithms. Empirically, we validate these findings on image denoising tasks, showing that even simple Transformers can perform robust denoising both with and without context. These results illustrate the Transformer’s flexibility as a denoising meta-learner. Code available at https://github.com/paulrosu11/Transformers_are_Diffusion_Denoisers. Paul Rosu, Lawrence Carin |
NeurIPS | 2 |
| 2024 | Meta-Learned Attribute Self-Interaction Network for Continual and Generalized Zero-Shot LearningabstractZero-shot learning (ZSL) is a promising approach to generalizing a model to categories unseen during training by leveraging class attributes, but challenges remain. Recently, methods using generative models to combat bias towards classes seen during training have pushed state of the art, but these generative models can be slow or computationally expensive to train. Also, these generative models assume that the attribute vector of each unseen class is available a priori at training, which is not always practical. Additionally, while many previous ZSL methods assume a one-time adaptation to unseen classes, in reality, the world is always changing, necessitating a constant adjustment of deployed models. Models unprepared to handle a sequential stream of data are likely to experience catastrophic forgetting. We propose a Meta-learned Attribute self-Interaction Network (MAIN) for continual ZSL. By pairing attribute self-interaction trained using meta-learning with inverse regularization of the attribute encoder, we are able to outperform state-of-the-art results without leveraging the unseen class attributes while also being able to train our models substantially faster (> 100×) than expensive generative-based approaches. We demonstrate this with experiments on five standard ZSL datasets (CUB, aPY, AWA1, AWA2, and SUN) in the generalized zero-shot learning and continual (fixed/dynamic) zero-shot learning settings. Extensive ablations and analyses demonstrate the efficacy of various components proposed. Vinay Kumar Verma, Nikhil Mehta 0002, Kevin J. Liang, Aakansha Mishra, Lawrence Carin |
WACV | 5 |
| 2024 | Text Feature Adversarial Learning for Text Generation With Knowledge Transfer From GPT2abstractText generation is a key component of many natural language tasks. Motivated by the success of generative adversarial networks (GANs) for image generation, many text-specific GANs have been proposed. However, due to the discrete nature of text, these text GANs often use reinforcement learning (RL) or continuous relaxations to calculate gradients during learning, leading to high-variance or biased estimation. Furthermore, the existing text GANs often suffer from mode collapse (i.e., they have limited generative diversity). To tackle these problems, we propose a new text GAN model named text feature GAN (TFGAN), where adversarial learning is performed in a continuous text feature space. In the adversarial game, GPT2 provides the "true" features, while the generator of TFGAN learns from them. TFGAN is trained by maximum likelihood estimation on text space and adversarial learning on text feature space, effectively combining them into a single objective, while alleviating mode collapse. TFGAN achieves appealing performance in text generation tasks, and it can also be used as a flexible framework for learning text representations. Hao Zhang 0050, Yulai Cong, Zhengjue Wang, Miaoyun Zhao, Liqun Chen 0001, Shijing Si, Ricardo Henao, Lawrence Carin |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2023 | Pushing the Efficiency Limit Using Structured Sparse ConvolutionsabstractWeight pruning is among the most popular approaches for compressing deep convolutional neural networks. Recent work suggests that in a randomly initialized deep neural network, there exist sparse subnetworks that achieve performance comparable to the original network. Unfortunately, finding these subnetworks involves iterative stages of training and pruning, which can be computationally expensive. We propose Structured Sparse Convolution (SSC), that leverages the inherent structure in images to reduce the parameters in the convolutional filter. This leads to improved efficiency of convolutional architectures compared to existing methods that perform pruning at initialization. We show that SSC is a generalization of commonly used layers (depthwise, groupwise and pointwise convolution) in "efficient architectures." Extensive experiments on well-known CNN models and datasets show the effectiveness of the proposed method. Architectures based on SSC achieve state-of-the-art performance compared to baselines on CIFAR10, CIFAR-100, Tiny-ImageNet, and ImageNet classification benchmarks. Our source code is publicly available at https://github.com/vkvermaa/SSC. Vinay Kumar Verma, Nikhil Mehta 0002, Shijing Si, Ricardo Henao, Lawrence Carin |
WACV | 5 |
| 2023 | Differentiable Hierarchical Optimal Transport for Robust Multi-View LearningabstractTraditional multi-view learning methods often rely on two assumptions: ( i) the samples in different views are well-aligned, and ( ii) their representations obey the same distribution in a latent space. Unfortunately, these two assumptions may be questionable in practice, which limits the application of multi-view learning. In this work, we propose a differentiable hierarchical optimal transport (DHOT) method to mitigate the dependency of multi-view learning on these two assumptions. Given arbitrary two views of unaligned multi-view data, the DHOT method calculates the sliced Wasserstein distance between their latent distributions. Based on these sliced Wasserstein distances, the DHOT method further calculates the entropic optimal transport across different views and explicitly indicates the clustering structure of the views. Accordingly, the entropic optimal transport, together with the underlying sliced Wasserstein distances, leads to a hierarchical optimal transport distance defined for unaligned multi-view data, which works as the objective function of multi-view learning and leads to a bi-level optimization task. Moreover, our DHOT method treats the entropic optimal transport as a differentiable operator of model parameters. It considers the gradient of the entropic optimal transport in the backpropagation step and thus helps improve the descent direction for the model in the training phase. We demonstrate the superiority of our bi-level optimization strategy by comparing it to the traditional alternating optimization strategy. The DHOT method is applicable for both unsupervised and semi-supervised learning. Experimental results show that our DHOT method is at least comparable to state-of-the-art multi-view learning methods on both synthetic and real-world tasks, especially for challenging scenarios with unaligned multi-view data. Dixin Luo, Hongteng Xu, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Representing Graphs via Gromov-Wasserstein FactorizationabstractGraph representation is a challenging and significant problem for many real-world applications. In this work, we propose a novel paradigm called "Gromov-Wasserstein Factorization" (GWF) to learn graph representations in a flexible and interpretable way. Given a set of graphs, whose correspondence between nodes is unknown and whose sizes can be different, our GWF model reconstructs each graph by a weighted combination of some "graph factors" under a pseudo-metric called Gromov-Wasserstein (GW) discrepancy. This model leads to a new nonlinear factorization mechanism of the graphs. The graph factors are shared by all the graphs, which represent the typical patterns of the graphs' structures. The weights associated with each graph indicate the graph factors' contributions to the graph's reconstruction, which lead to a permutation-invariant graph representation. We learn the graph factors of the GWF model and the weights of the graphs jointly by minimizing the overall reconstruction error. When learning the model, we reparametrize the graph factors and the weights to unconstrained model parameters and simplify the backpropagation of gradient with the help of the envelope theorem. For the GW discrepancy (the critical training step), we consider two algorithms to compute it, which correspond to the proximal point algorithm (PPA) and Bregman alternating direction method of multipliers (BADMM), respectively. Furthermore, we propose some extensions of the GWF model, including (i) combining with a graph neural network and learning graph representations in an auto-encoding manner, (ii) representing the graphs with node attributes, and (iii) working as a regularizer for semi-supervised graph classification. Experiments on various datasets demonstrate that our GWF model is comparable to the state-of-the-art methods. The graph representations derived by it perform well in graph clustering and classification tasks. Hongteng Xu, Jiachang Liu 0001, Dixin Luo, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Calibration and Uncertainty in Neural Time-to-Event ModelingabstractModels for predicting the time of a future event are crucial for risk assessment, across a diverse range of applications. Existing time-to-event (survival) models have focused primarily on preserving pairwise ordering of estimated event times (i.e., relative risk). We propose neural time-to-event models that account for calibration and uncertainty while predicting accurate absolute event times. Specifically, an adversarial nonparametric model is introduced for estimating matched time-to-event distributions for probabilistically concentrated and accurate predictions. We also consider replacing the discriminator of the adversarial nonparametric model with a survival-function matching estimator that accounts for model calibration. The proposed estimator can be used as a means of estimating and comparing conditional survival distributions while accounting for the predictive uncertainty of probabilistic models. Extensive experiments show that the distribution matching methods outperform existing approaches in terms of both calibration and concentration of time-to-event distributions. Paidamoyo Chapfuwa, Chenyang Tao, Chunyuan Li, Karen Chandross, Michael J. Pencina, Lawrence Carin, Ricardo Henao |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2023 | Learning Hierarchical Document Graphs From Multilevel Sentence RelationsabstractOrganizing the implicit topology of a document as a graph, and further performing feature extraction via the graph convolutional network (GCN), has proven effective in document analysis. However, existing document graphs are often restricted to expressing single-level relations, which are predefined and independent of downstream learning. A set of learnable hierarchical graphs are built to explore multilevel sentence relations, assisted by a hierarchical probabilistic topic model. Based on these graphs, multiple parallel GCNs are used to extract multilevel semantic features, which are aggregated by an attention mechanism for different document-comprehension tasks. Equipped with variational inference, the graph construction and GCN are learned jointly, allowing the graphs to evolve dynamically to better match the downstream task. The effectiveness and efficiency of the proposed multilevel sentence relation graph convolutional network (MuserGCN) is demonstrated via experiments on document classification, abstractive summarization, and matching. Hao Zhang 0050, Chaojie Wang 0001, Zhengjue Wang, Zhibin Duan, Bo Chen 0001, Mingyuan Zhou, Ricardo Henao, Lawrence Carin |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2022 | Improving Downstream Task Performance by Treating Numbers as EntitiesabstractNumbers are essential components of text, like any other word tokens, from which natural language processing (NLP) models are built and deployed. Though numbers are typically not accounted for distinctly in most NLP tasks, there is still an underlying amount of numeracy already exhibited by NLP models. For instance, in named entity recognition (NER), numbers are not treated as an entity with distinct tags. In this work, we attempt to tap the potential of state-of-the-art language models and transfer their ability to boost performance in related downstream tasks dealing with numbers. Our proposed classification of numbers into entities helps NLP models perform well on several tasks, including a handcrafted Fill-In-The-Blank (FITB) task and on question answering, using joint embeddings, outperforming the BERT and RoBERTa baseline classification. Dhanasekar Sundararaman, Vivek Subramanian, Guoyin Wang 0002, Liyan Xu, Lawrence Carin |
CIKM | 5 |
| 2022 | Open World Classification with Adaptive Negative SamplesabstractOpen world classification is a task in natural language processing with key practical relevance and impact.Since the open or unknown category data only manifests in the inference phase, finding a model with a suitable decision boundary accommodating for the identification of known classes and discrimination of the open category is challenging.The performance of existing models is limited by the lack of effective open category data during the training stage or the lack of a good mechanism to learn appropriate decision boundaries.We propose an approach based on adaptive negative samples (ANS) designed to generate effective synthetic open category samples in the training stage and without requiring any prior knowledge or external datasets.Empirically, we find a significant advantage in using auxiliary one-versus-rest binary classifiers, which effectively utilize the generated negative samples and avoid the complex threshold-seeking stage in previous works.Extensive experiments on three benchmark datasets show that ANS achieves significant improvements over stateof-the-art methods. Ke Bai 0001, Guoyin Wang 0002, Jiwei Li 0001, Puyang Xu, Ricardo Henao, Lawrence Carin |
EMNLP | 8 |
| 2022 | Gradient Importance Learning for Incomplete Observations
Qitong Gao, Dong Wang 0037, Joshua D. Amason, Siyang Yuan, Chenyang Tao, Ricardo Henao, Majda Hadziahmetovic, Lawrence Carin, Miroslav Pajic |
ICLR | 8 |
| 2022 | Tight Mutual Information Estimation With Contrastive Fenchel-Legendre OptimizationabstractSuccessful applications of InfoNCE (Information Noise-Contrastive Estimation) and its variants have popularized the use of contrastive variational mutual information (MI) estimators in machine learning . While featuring superior stability, these estimators crucially depend on costly large-batch training, and they sacrifice bound tightness for variance reduction. To overcome these limitations, we revisit the mathematics of popular variational MI bounds from the lens of unnormalized statistical modeling and convex optimization. Our investigation yields a new unified theoretical framework encompassing popular variational MI bounds, and leads to a novel, simple, and powerful contrastive MI estimator we name FLO. Theoretically, we show that the FLO estimator is tight, and it converges under stochastic gradient descent. Empirically, the proposed FLO estimator overcomes the limitations of its predecessors and learns more efficiently. The utility of FLO is verified using extensive benchmarks, and we further inspire the community with novel applications in meta-learning. Our presentation underscores the foundational importance of variational MI estimation in data-efficient learning. Junya Chen, Dong Wang 0037, Yuewei Yang, Xinwei Deng, Lawrence Carin, Chenyang Tao |
NeurIPS | 7 |
| 2022 | Capturing actionable dynamics with structured latent ordinary differential equationsabstractEnd-to-end learning of dynamical systems with black-box models, such as neural ordinary differential equations (ODEs), provides a flexible framework for learning dynamics from data without prescribing a mathematical model for the dynamics. Unfortunately, this flexibility comes at the cost of understanding the dynamical system, for which ODEs are used ubiquitously. Further, experimental data are collected under various conditions (inputs), such as treatments, or grouped in some way, such as part of sub-populations. Understanding the effects of these system inputs on system outputs is crucial to have any meaningful model of a dynamical system. To that end, we propose a structured latent ODE model that explicitly captures system input variations within its latent representation. Building on a static latent variable specification, our model learns (independent) stochastic factors of variation for each input to the system, thus separating the effects of the system inputs in the latent space. This approach provides actionable modeling through the controlled generation of time-series data for novel input combinations (or perturbations). Additionally, we propose a flexible approach for quantifying uncertainties, leveraging a quantile regression formulation. Results on challenging biological datasets show consistent improvements over competitive baselines in the controlled generation of observational data and inference of biologically meaningful system inputs. Paidamoyo Chapfuwa, Sherri Rose, Lawrence Carin, Edward Meeds, Ricardo Henao |
UAI | 3 |
| 2022 | Learning to Weight Filter Groups for Robust ClassificationabstractIn many real-world tasks, a canonical “big data” problem is created by combining data from several individual groups or domains. Because test data will likely come from a new group of data, we want to utilize the grouped structure of our training data to enforce generalization between groups of data, not just individual samples. This can be viewed as a multiple-domain generalization problem. Specifically, the goal is to encourage generalization between previously seen labeled source data from multiple domains and unlabeled target domain data. To address this challenge, we introduce Domain-Specific Filter Group (DSFG), where each training domain has a unique filter group and each test data point is predicted by a weighted sum over the outputs of different domain filters. A separate neural network learns to estimate the appropriate filter group weights through a meta-learning strategy. Empirically, experiments on three benchmark datasets demonstrate improved performance compared to current state-of-the-art approaches. Siyang Yuan, Yitong Li 0001, Dong Wang 0037, Ke Bai 0001, Lawrence Carin, David E. Carlson |
WACV | 5 |
| 2022 | Explainable multiple abnormality classification of chest CT volumes
Rachel Lea Draelos, Lawrence Carin |
Artif. Intell. Medicine | 2 |
| 2022 | elBERto: Self-supervised commonsense learning for question answering
Xunlin Zhan, Yuan Li 0032, Xiaodan Liang, Zhiting Hu, Lawrence Carin |
Knowl. Based Syst. | 6 |
| 2021 | GO Hessian for Expectation-Based Objectives
Yulai Cong, Miaoyun Zhao, Jianqiao Li, Junya Chen, Lawrence Carin |
AAAI | 5 |
| 2021 | Learning Graphons via Structured Gromov-Wasserstein BarycentersabstractWe propose a novel and principled method to learn a nonparametric graph model called graphon, which is defined in an infinite-dimensional space and represents arbitrary-size graphs. Based on the weak regularity lemma from the theory of graphons, we leverage a step function to approximate a graphon. We show that the cut distance of graphons can be relaxed to the Gromov-Wasserstein distance of their step functions. Accordingly, given a set of graphs generated by an underlying graphon, we learn the corresponding step function as the Gromov-Wasserstein barycenter of the given graphs. Furthermore, we develop several enhancements and extensions of the basic algorithm, e.g., the smoothed Gromov-Wasserstein barycenter for guaranteeing the continuity of the learned graphons and the mixed Gromov-Wasserstein barycenters for learning multiple structured graphons. The proposed approach overcomes drawbacks of prior state-of-the-art methods, and outperforms them on both synthetic and real-world data. The code is available at https://github.com/HongtengXu/SGWB-Graphon. Hongteng Xu, Dixin Luo, Lawrence Carin, Hongyuan Zha |
AAAI | 3 |
| 2021 | Counterfactual Representation Learning with Balancing WeightsabstractA key to causal inference with observational data is achieving balance in predictive features associated with each treatment type. Recent literature has explored representation learning to achieve this goal. In this work, we discuss the pitfalls of these strategies – such as a steep trade-off between achieving balance and predictive power – and present a remedy via the integration of balancing weights in causal learning. Specifically, we theoretically link balance to the quality of propensity estimation, emphasize the importance of identifying a proper target population, and elaborate on the complementary roles of feature balancing and weight adjustments. Using these concepts, we then develop an algorithm for flexible, scalable and accurate estimation of causal effects. Finally, we show how the learned weighted representations may serve to facilitate alternative causal learning procedures with appealing statistical features. We conduct an extensive set of experiments on both synthetic examples and standard benchmarks, and report encouraging results relative to state-of-the-art baselines. Serge Assaad, Shuxi Zeng, Chenyang Tao, Shounak Datta, Nikhil Mehta 0002, Ricardo Henao, Lawrence Carin |
AISTATS | 8 |
| 2021 | Continual Learning using a Bayesian Nonparametric Dictionary of Weight FactorsabstractNaively trained neural networks tend to experience catastrophic forgetting in sequential task settings, where data from previous tasks are unavailable. A number of methods, using various model expansion strategies, have been proposed recently as possible solutions. However, determining how much to expand the model is left to the practitioner, and often a constant schedule is chosen for simplicity, regardless of how complex the incoming task is. Instead, we propose a principled Bayesian nonparametric approach based on the Indian Buffet Process (IBP) prior, letting the data determine how much to expand the model complexity. We pair this with a factorization of the neural network’s weight matrices. Such an approach allows us to scale the number of factors of each weight matrix to the complexity of the task, while the IBP prior encourages sparse weight factor selection and factor reuse, promoting positive knowledge transfer between tasks. We demonstrate the effectiveness of our method on a number of continual learning benchmarks and analyze how weight factors are allocated and reused throughout the training. Nikhil Mehta 0002, Kevin J. Liang, Vinay Kumar Verma, Lawrence Carin |
AISTATS | 4 |
| 2021 | Wasserstein Contrastive Representation DistillationabstractThe primary goal of knowledge distillation (KD) is to encapsulate the information of a model learned from a teacher network into a student network, with the latter being more compact than the former. Existing work, e.g., using Kullback-Leibler divergence for distillation, may fail to capture important structural knowledge in the teacher network and often lacks the ability for feature generalization, particularly in situations when teacher and student are built to address different classification tasks. We propose Wasserstein Contrastive Representation Distillation (WCoRD), which leverages both primal and dual forms of Wasserstein distance for KD. The dual form is used for global knowledge transfer, yielding a contrastive learning objective that maximizes the lower bound of mutual information between the teacher and the student networks. The primal form is used for local contrastive knowledge transfer within a mini-batch, effectively matching the distributions of features between the teacher and the student networks. Experiments demonstrate that the proposed WCoRD method outperforms state-of-the-art approaches on privileged information distillation, model compression and cross-modal transfer. Liqun Chen 0001, Dong Wang 0037, Zhe Gan, Jingjing Liu 0001, Ricardo Henao, Lawrence Carin |
CVPR | 6 |
| 2021 | Efficient Feature Transformations for Discriminative and Generative Continual LearningabstractAs neural networks are increasingly being applied to real-world applications, mechanisms to address distributional shift and sequential task learning without forgetting are critical. Methods incorporating network expansion have shown promise by naturally adding model capacity for learning new tasks while simultaneously avoiding catastrophic forgetting. However, the growth in the number of additional parameters of many of these types of methods can be computationally expensive at larger scales, at times prohibitively so. Instead, we propose a simple task-specific feature map transformation strategy for continual learning, which we call Efficient Feature Transformations (EFTs). These EFTs provide powerful flexibility for learning new tasks, achieved with minimal parameters added to the base architecture. We further propose a feature distance maximization strategy, which significantly improves task prediction in class incremental settings, without needing expensive generative models. We demonstrate the efficacy and efficiency of our method with an extensive set of experiments in discriminative (CIFAR-100 and ImageNet-1K) and generative (LSUN, CUB-200, Cats) sequences of tasks. Even with low single-digit parameter growth rates, EFTs can outperform many other continual learning methods in a wide range of settings. Vinay Kumar Verma, Kevin J. Liang, Nikhil Mehta 0002, Piyush Rai, Lawrence Carin |
CVPR | 5 |
| 2021 | FairFil: Contrastive Neural Debiasing Method for Pretrained Text Encoders
Pengyu Cheng, Weituo Hao, Siyang Yuan, Shijing Si, Lawrence Carin |
ICLR | 5 |
| 2021 | MixKD: Towards Efficient Distillation of Large-scale Language Models
Kevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou 0001, Weizhu Chen, Changyou Chen, Lawrence Carin |
ICLR | 7 |
| 2021 | Improving Zero-Shot Voice Style Transfer via Disentangled Representation Learning
Siyang Yuan, Pengyu Cheng, Ruiyi Zhang 0002, Weituo Hao, Zhe Gan, Lawrence Carin |
ICLR | 6 |
| 2021 | FLOP: Federated Learning on Medical Datasets using Partial NetworksabstractThe outbreak of COVID-19 Disease due to the novel coronavirus has caused a shortage of medical resources. To aid and accelerate the diagnosis process, automatic diagnosis of COVID-19 via deep learning models has recently been explored by researchers across the world. While different data-driven deep learning models have been developed to mitigate the diagnosis of COVID-19, the data itself is still scarce due to patient privacy concerns. Federated Learning (FL) is a natural solution because it allows different organizations to cooperatively learn an effective deep learning model without sharing raw data. However, recent studies show that FL still lacks privacy protection and may cause data leakage. We investigate this challenging problem by proposing a simple yet effective algorithm, named Federated Learning on Medical Datasets using Partial Networks (FLOP), that shares only a partial model between the server and clients. Extensive experiments on benchmark data and real-world healthcare tasks show that our approach achieves comparable or better performance while reducing the privacy and security risks. Of particular interest, we conduct experiments on the COVID-19 dataset and find that our FLOP algorithm can allow different hospitals to collaboratively and effectively train a partially shared model without sharing local patients' data. Qian Yang 0003, Weituo Hao, Gregory Spell, Lawrence Carin |
KDD | 5 |
| 2021 | APo-VAE: Text Generation in Hyperbolic SpaceabstractShuyang Dai, Zhe Gan, Yu Cheng, Chenyang Tao, Lawrence Carin, Jingjing Liu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Shuyang Dai, Zhe Gan, Yu Cheng 0001, Chenyang Tao, Lawrence Carin, Jingjing Liu 0001 |
NAACL-HLT | 5 |
| 2021 | SpanPredict: Extraction of Predictive Document Spans with Neural AttentionabstractVivek Subramanian, Matthew Engelhard, Sam Berchuck, Liqun Chen, Ricardo Henao, Lawrence Carin. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Vivek Subramanian, Matthew Engelhard, Samuel Berchuck, Liqun Chen 0001, Ricardo Henao, Lawrence Carin |
NAACL-HLT | 6 |
| 2021 | Supercharging Imbalanced Data Learning With Energy-based Contrastive Representation TransferabstractDealing with severe class imbalance poses a major challenge for many real-world applications, especially when the accurate classification and generalization of minority classes are of primary interest.In computer vision and NLP, learning from datasets with long-tail behavior is a recurring theme, especially for naturally occurring labels. Existing solutions mostly appeal to sampling or weighting adjustments to alleviate the extreme imbalance, or impose inductive bias to prioritize generalizable associations. Here we take a novel perspective to promote sample efficiency and model generalization based on the invariance principles of causality. Our contribution posits a meta-distributional scenario, where the causal generating mechanism for label-conditional features is invariant across different labels. Such causal assumption enables efficient knowledge transfer from the dominant classes to their under-represented counterparts, even if their feature distributions show apparent disparities. This allows us to leverage a causal data augmentation procedure to enlarge the representation of minority classes. Our development is orthogonal to the existing imbalanced data learning techniques thus can be seamlessly integrated. The proposed approach is validated on an extensive set of synthetic and real-world tasks against state-of-the-art solutions. Junya Chen, Zidi Xiu, Benjamin Goldstein 0001, Ricardo Henao, Lawrence Carin, Chenyang Tao |
NeurIPS | 5 |
| 2021 | CAM-GAN: Continual Adaptation Modules for Generative Adversarial NetworksabstractWe present a continual learning approach for generative adversarial networks (GANs), by designing and leveraging parameter-efficient feature map transformations. Our approach is based on learning a set of global and task-specific parameters. The global parameters are fixed across tasks whereas the task-specific parameters act as local adapters for each task, and help in efficiently obtaining task-specific feature maps. Moreover, we propose an element-wise addition of residual bias in the transformed feature space, which further helps stabilize GAN training in such settings. Our approach also leverages task similarities based on the Fisher information matrix. Leveraging this knowledge from previous tasks significantly improves the model performance. In addition, the similarity measure also helps reduce the parameter growth in continual adaptation and helps to learn a compact model. In contrast to the recent approaches for continually-learned GANs, the proposed approach provides a memory-efficient way to perform effective continual data generation. Through extensive experiments on challenging and diverse datasets, we show that the feature-map-transformation approach outperforms state-of-the-art methods for continually-learned GANs, with substantially fewer parameters. The proposed method generates high-quality samples that can also improve the generative-replay-based continual learning for discriminative tasks. Sakshi Varshney, Vinay Kumar Verma, P. K. Srijith, Lawrence Carin, Piyush Rai |
NeurIPS | 4 |
| 2021 | Zero-Shot Recognition via Optimal TransportabstractWe propose an optimal transport (OT) framework for generalized zero-shot learning (GZSL), seeking to distinguish samples for both seen and unseen classes, with the assist of auxiliary attributes. The discrepancy between features and attributes is minimized by solving an optimal transport problem. Specifically, we build a conditional generative model to generate features from seen-class attributes, and establish an optimal transport between the distribution of the generated features and that of the real features. The generative model and the optimal transport are optimized iteratively with an attribute-based regularizer, that further enhances the discriminative power of the generated features. A classifier is learned based on the features generated for both the seen and unseen classes. In addition to generalized zero-shot learning, our framework is also applicable to standard and transductive ZSL problems. Experiments show that our optimal transport-based method outperforms state-of-the-art methods on several benchmark datasets. Wenlin Wang, Hongteng Xu, Guoyin Wang 0002, Wenqi Wang 0001, Lawrence Carin |
WACV | 5 |
| 2021 | Weakly supervised instance learning for thyroid malignancy prediction from whole slide cytopathology images
David Dov, Shahar Z. Kovalsky, Serge Assaad, Jonathan Cohen 0004, Danielle Range, Avani A. Pendse, Ricardo Henao, Lawrence Carin |
Medical Image Anal. | 8 |
| 2021 | Machine-learning-based multiple abnormality prediction with large-scale chest computed tomography volumes
Rachel Lea Draelos, David Dov, Maciej A. Mazurowski, Joseph Y. Lo, Ricardo Henao, Geoffrey D. Rubin, Lawrence Carin |
Medical Image Anal. | 7 |
| 2020 | Sequence Generation with Optimal-Transport-Enhanced Reinforcement LearningabstractReinforcement learning (RL) has been widely used to aid training in language generation. This is achieved by enhancing standard maximum likelihood objectives with user-specified reward functions that encourage global semantic consistency. We propose a principled approach to address the difficulties associated with RL-based solutions, namely, high-variance gradients, uninformative rewards and brittle training. By leveraging the optimal transport distance, we introduce a regularizer that significantly alleviates the above issues. Our formulation emphasizes the preservation of semantic features, enabling end-to-end training instead of ad-hoc fine-tuning, and when combined with RL, it controls the exploration space for more efficient model updates. To validate the effectiveness of the proposed solution, we perform a comprehensive evaluation covering a wide variety of NLP tasks: machine translation, abstractive text summarization and image caption, with consistent improvements over competing solutions. Liqun Chen 0001, Ke Bai 0001, Chenyang Tao, Yizhe Zhang 0002, Guoyin Wang 0002, Wenlin Wang, Ricardo Henao, Lawrence Carin |
AAAI | 8 |
| 2020 | Dynamic Embedding on Textual Networks via a Gaussian ProcessabstractTextual network embedding aims to learn low-dimensional representations of text-annotated nodes in a graph. Prior work in this area has typically focused on fixed graph structures; however, real-world networks are often dynamic. We address this challenge with a novel end-to-end node-embedding model, called Dynamic Embedding for Textual Networks with a Gaussian Process (DetGP). After training, DetGP can be applied efficiently to dynamic graphs without re-training or backpropagation. The learned representation of each node is a combination of textual and structural embeddings. Because the structure is allowed to be dynamic, our method uses the Gaussian process to take advantage of its non-parametric properties. To use both local and global graph structures, diffusion is used to model multiple hops between neighbors. The relative importance of global versus local structure for the embeddings is learned automatically. With the non-parametric nature of the Gaussian process, updating the embeddings for a changed graph structure requires only a forward pass through the learned model. Considering link prediction and node classification, experiments demonstrate the empirical effectiveness of our method compared to baseline approaches. We further show that DetGP can be straightforwardly and efficiently applied to dynamic textual networks. Pengyu Cheng, Yitong Li 0001, Xinyuan Zhang 0001, Liqun Chen 0001, David E. Carlson, Lawrence Carin |
AAAI | 6 |
| 2020 | Complementary Auxiliary Classifiers for Label-Conditional Text GenerationabstractLearning to generate text with a given label is a challenging task because natural language sentences are highly variable and ambiguous. It renders difficulties in trade-off between sentence quality and label fidelity. In this paper, we present CARA to alleviate the issue, where two auxiliary classifiers work simultaneously to ensure that (1) the encoder learns disentangled features and (2) the generator produces label-related sentences. Two practical techniques are further proposed to improve the performance, including annealing the learning signal from the auxiliary classifier, and enhancing the encoder with pre-trained language models. To establish a comprehensive benchmark fostering future research, we consider a suite of four datasets, and systematically reproduce three representative methods. CARA shows consistent improvement over the previous methods on the task of label-conditional text generation, and achieves state-of-the-art on the task of attribute transfer. Yuan Li 0032, Chunyuan Li, Yizhe Zhang 0002, Xiujun Li, Guoqing Zheng, Lawrence Carin, Jianfeng Gao 0001 |
AAAI | 6 |
| 2020 | Graph-Driven Generative Models for Heterogeneous Multi-Task LearningabstractWe propose a novel graph-driven generative model, that unifies multiple heterogeneous learning tasks into the same framework. The proposed model is based on the fact that heterogeneous learning tasks, which correspond to different generative processes, often rely on data with a shared graph structure. Accordingly, our model combines a graph convolutional network (GCN) with multiple variational autoencoders, thus embedding the nodes of the graph (i.e., samples for the tasks) in a uniform manner, while specializing their organization and usage to different tasks. With a focus on healthcare applications (tasks), including clinical topic modeling, procedure recommendation and admission-type prediction, we demonstrate that our method successfully leverages information across different tasks, boosting performance in all tasks and outperforming existing state-of-the-art approaches. Wenlin Wang, Hongteng Xu, Zhe Gan, Bai Li 0001, Guoyin Wang 0002, Liqun Chen 0001, Qian Yang 0003, Wenqi Wang 0001, Lawrence Carin |
AAAI | 9 |
| 2020 | Bridging Maximum Likelihood and Adversarial Learning via α-DivergenceabstractMaximum likelihood (ML) and adversarial learning are two popular approaches for training generative models, and from many perspectives these techniques are complementary. ML learning encourages the capture of all data modes, and it is typically characterized by stable training. However, ML learning tends to distribute probability mass diffusely over the data space, e.g., yielding blurry synthetic images. Adversarial learning is well known to synthesize highly realistic natural images, despite practical challenges like mode dropping and delicate training. We propose an α-Bridge to unify the advantages of ML and adversarial learning, enabling the smooth transfer from one to the other via the α-divergence. We reveal that generalizations of the α-Bridge are closely related to approaches developed recently to regularize adversarial learning, providing insights into that prior work, and further understanding of why the α-Bridge performs well in practice. Miaoyun Zhao, Yulai Cong, Shuyang Dai, Lawrence Carin |
AAAI | 4 |
| 2020 | Contrastively Smoothed Class Alignment for Unsupervised Domain Adaptation
Shuyang Dai, Yu Cheng 0001, Yizhe Zhang 0002, Zhe Gan, Jingjing Liu 0001, Lawrence Carin |
ACCV (4) | 6 |
| 2020 | Improving Disentangled Text Representation Learning with Information-Theoretic GuidanceabstractPengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon, Yizhe Zhang, Yitong Li, Lawrence Carin. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Pengyu Cheng, Martin Renqiang Min, Dinghan Shen, Christopher Malon, Yizhe Zhang 0002, Yitong Li 0001, Lawrence Carin |
ACL | 7 |
| 2020 | Improving Adversarial Text Generation by Modeling the Distant FutureabstractAuto-regressive text generation models usually focus on local fluency, and may cause inconsistent semantic meaning in long text generation.Further, automatically generating words with similar semantics is challenging, and hand-crafted linguistic rules are difficult to apply.We consider a text planning scheme and present a model-based imitation-learning approach to alleviate the aforementioned issues.Specifically, we propose a novel guider network to focus on the generative process over a longer horizon, which can assist next-word prediction and provide intermediate rewards for generator optimization.Extensive experiments demonstrate that the proposed method leads to improved performance. Ruiyi Zhang 0002, Changyou Chen, Zhe Gan, Wenlin Wang, Dinghan Shen, Guoyin Wang 0002, Lawrence Carin |
ACL | 8 |
| 2020 | Nested-Wasserstein Self-Imitation Learning for Sequence GenerationabstractReinforcement learning (RL) has been widely studied for improving sequence-generation models. However, the conventional rewards used for RL training typically cannot capture sufficient semantic information and therefore render model bias. Further, the sparse and delayed rewards make RL exploration inefficient. To alleviate these issues, we propose the concept of nested-Wasserstein distance for distributional semantic matching. To further exploit it, a novel nested-Wasserstein self-imitation learning framework is developed, encouraging the model to exploit historical high-rewarded sequences for enhanced exploration and better semantic matching. Our solution can be understood as approximately executing proximal policy optimization with Wasserstein trust-regions. Experiments on a variety of unconditional and conditional sequence-generation tasks demonstrate the proposed approach consistently leads to improved performance. Ruiyi Zhang 0002, Changyou Chen, Zhe Gan, Wenlin Wang, Lawrence Carin |
AISTATS | 6 |
| 2020 | Stochastic Particle-Optimization Sampling and the Non-Asymptotic Convergence TheoryabstractParticle-optimization-based sampling (POS) is a recently developed effective sampling technique that interactively updates a set of particles. A representative algorithm is the Stein variational gradient descent (SVGD). We prove, under certain conditions, SVGD experiences a theoretical pitfall, {\it i.e.}, particles tend to collapse. As a remedy, we generalize POS to a stochastic setting by injecting random noise into particle updates, thus termed stochastic particle-optimization sampling (SPOS). Notably, for the first time, we develop non-asymptotic convergence theory for the SPOS framework (related to SVGD), characterizing algorithm convergence in terms of the 1-Wasserstein distance w.r.t. the numbers of particles and iterations. Somewhat surprisingly, with the same number of updates (not too large) for each particle, our theory suggests adopting more particles does not necessarily lead to a better approximation of a target distribution, due to limited computational budget and numerical errors. This phenomenon is also observed in SVGD and verified via a synthetic experiment. Extensive experimental results verify our theory and demonstrate the effectiveness of our proposed framework. Ruiyi Zhang 0002, Lawrence Carin, Changyou Chen |
AISTATS | 3 |
| 2020 | Adaptation Across Extreme Variations using Unlabeled Bridges
Shuyang Dai, Kihyuk Sohn, Yi-Hsuan Tsai, Lawrence Carin, Manmohan Krishna Chandraker |
BMVC | 4 |
| 2020 | Object Detection as a Positive-Unlabeled Problem
Yuewei Yang, Kevin J. Liang, Lawrence Carin |
BMVC | 3 |
| 2020 | Advancing weakly supervised cross-domain alignment with optimal transport
Siyang Yuan, Ke Bai 0001, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Chunyuan Li, Guoyin Wang 0002, Ricardo Henao, Lawrence Carin |
BMVC | 9 |
| 2020 | Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-TrainingabstractLearning to navigate in a visual environment following natural-language instructions is a challenging task, because the multimodal inputs to the agent are highly variable, and the training data on a new task is often limited. In this paper, we present the first pre-training and fine-tuning paradigm for vision-and-language navigation (VLN) tasks. By training on a large amount of image-text-action triplets in a self-supervised learning manner, the pre-trained model provides generic representations of visual environments and language instructions. It can be easily used as a drop-in for existing VLN frameworks, leading to the proposed agent PREVALENT. It learns more effectively in new tasks and generalizes better in a previously unseen environment. The performance is validated on three VLN tasks. On the Room-to-Room benchmark, our model improves the state-of-the-art from 47\% to 51\% on success rate weighted by path length. Further, the learned representation is transferable to other VLN tasks. On two recent tasks, vision-and-dialog navigation and ``Help, Anna!'', the proposed PREVALENT leads to significant improvement over existing methods, achieving a new state of the art. Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, Jianfeng Gao 0001 |
CVPR | 4 |
| 2020 | Enhancing Cross-Task Black-Box Transferability of Adversarial Examples With Dispersion ReductionabstractNeural networks are known to be vulnerable to carefully crafted adversarial examples, and these malicious samples often transfer, i.e., they remain adversarial even against other models. Although significant effort has been devoted to the transferability across models, surprisingly little attention has been paid to cross-task transferability, which represents the real-world cybercriminal's situation, where an ensemble of different defense/detection mechanisms need to be evaded all at once. We investigate the transferability of adversarial examples across a wide range of real-world computer vision tasks, including image classification, object detection, semantic segmentation, explicit content detection, and text detection. Our proposed attack minimizes the “dispersion” of the internal feature map, overcoming the limitations of existing attacks, that require task-specific loss functions and/or probing a target model. We conduct evaluation on open-source detection and segmentation models, as well as four different computer vision tasks provided by Google Cloud Vision (GCV) APIs. We demonstrate that our approach outperforms existing attacks by degrading performance of multiple CV tasks by a large margin with only modest perturbations. Yantao Lu, Yunhan Jia, Bai Li 0001, Weiheng Chai, Lawrence Carin, Senem Velipasalar |
CVPR | 6 |
| 2020 | Improving Text Generation with Student-Forcing Optimal TransportabstractJianqiao Li, Chunyuan Li, Guoyin Wang, Hao Fu, Yuhchen Lin, Liqun Chen, Yizhe Zhang, Chenyang Tao, Ruiyi Zhang, Wenlin Wang, Dinghan Shen, Qian Yang, Lawrence Carin. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Jianqiao Li, Chunyuan Li, Guoyin Wang 0002, Hao Fu 0002, Yuh-Chen Lin, Liqun Chen 0001, Yizhe Zhang 0002, Chenyang Tao, Ruiyi Zhang 0002, Wenlin Wang, Dinghan Shen, Qian Yang 0003, Lawrence Carin |
EMNLP (1) | 13 |
| 2020 | An Embedding Model for Estimating Legislative Preferences from the Frequency and Sentiment of TweetsabstractLegislator preferences are typically represented as measures of general ideology estimated from roll call votes on legislation, potentially masking important nuances in legislators’ political attitudes. In this paper we introduce a method of measuring more specific legislator attitudes using an alternative expression of preferences: tweeting. Specifically, we present an embedding-based model for predicting the frequency and sentiment of legislator tweets. To illustrate our method, we model legislators’ attitudes towards President Donald Trump as vector embeddings that interact with embeddings for Trump himself constructed using a neural network from the text of his daily tweets. We demonstrate the predictive performance of our model on tweets authored by members of the U.S. House and Senate related to the president from November 2016 to February 2018. We further assess the quality of our learned representations for legislators by comparing to traditional measures of legislator preferences. Gregory Spell, Brian Guay, Sunshine Hillygus, Lawrence Carin |
EMNLP (1) | 4 |
| 2020 | Methods for Numeracy-Preserving Word EmbeddingsabstractDhanasekar Sundararaman, Shijing Si, Vivek Subramanian, Guoyin Wang, Devamanyu Hazarika, Lawrence Carin. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Dhanasekar Sundararaman, Shijing Si, Vivek Subramanian, Guoyin Wang 0002, Devamanyu Hazarika, Lawrence Carin |
EMNLP (1) | 6 |
| 2020 | Transferable Perturbations of Deep Feature Distributions
Nathan Inkawhich, Kevin J. Liang, Lawrence Carin, Yiran Chen 0001 |
ICLR | 3 |
| 2020 | RaCT: Toward Amortized Ranking-Critical Training For Collaborative Filtering
Sam Lobel, Chunyuan Li, Jianfeng Gao 0001, Lawrence Carin |
ICLR | 4 |
| 2020 | Graph Optimal Transport for Cross-Domain AlignmentabstractCross-domain alignment between two sets of entities (e.g., objects in an image, words in a sentence) is fundamental to both computer vision and natural language processing. Existing methods mainly focus on designing advanced attention mechanisms to simulate soft alignment, where no training signals are provided to explicitly encourage alignment. Plus, the learned attention matrices are often dense and difficult to interpret. We propose Graph Optimal Transport (GOT), a principled framework that builds upon recent advances in Optimal Transport (OT). In GOT, cross-domain alignment is formulated as a graph matching problem, by representing entities as a dynamically-constructed graph. Two types of OT distances are considered: (i) Wasserstein distance (WD) for node (entity) matching; and (ii) Gromov-Wasserstein distance (GWD) for edge (structure) matching. Both WD and GWD can be incorporated into existing neural network models, effectively acting as a drop-in regularizer. The inferred transport plan also yields sparse and self-normalized alignment, enhancing the interpretability of the learned model. Experiments show consistent outperformance of GOT over baselines across a wide range of tasks, including image-text retrieval, visual question answering, image captioning, machine translation, and text summarization. Liqun Chen 0001, Zhe Gan, Yu Cheng 0001, Lawrence Carin, Jingjing Liu 0001 |
ICML | 5 |
| 2020 | CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationabstractMutual information (MI) minimization has gained considerable interests in various machine learning tasks. However, estimating and minimizing MI in high-dimensional spaces remains a challenging problem, especially when only samples, rather than distribution forms, are accessible. Previous works mainly focus on MI lower bound approximation, which is not applicable to MI minimization problems. In this paper, we propose a novel Contrastive Log-ratio Upper Bound (CLUB) of mutual information. We provide a theoretical analysis of the properties of CLUB and its variational approximation. Based on this upper bound, we introduce a MI minimization training scheme and further accelerate it with a negative sampling strategy. Simulation studies on Gaussian distributions show the reliable estimation ability of CLUB. Real-world MI minimization experiments, including domain adaptation and information bottleneck, demonstrate the effectiveness of the proposed method. The code is at https://github.com/Linear95/CLUB. Pengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu 0001, Zhe Gan, Lawrence Carin |
ICML | 6 |
| 2020 | Learning Autoencoders with Relational RegularizationabstractWe propose a new algorithmic framework for learning autoencoders of data distributions. In this framework, we minimize the discrepancy between the model distribution and the target one, with relational regularization on learnable latent prior. This regularization penalizes the fused Gromov-Wasserstein (FGW) distance between the latent prior and its corresponding posterior, which allows us to learn a structured prior distribution associated with the generative model in a flexible way. Moreover, it helps us co-train multiple autoencoders even if they are with heterogeneous architectures and incomparable latent spaces. We implement the framework with two scalable algorithms, making it applicable for both probabilistic and deterministic autoencoders. Our relational regularized autoencoder (RAE) outperforms existing methods, e.g., variational autoencoder, Wasserstein autoencoder, and their variants, on generating images. Additionally, our relational co-training strategy of autoencoders achieves encouraging results in both synthesis and real-world multi-view learning tasks. Hongteng Xu, Dixin Luo, Ricardo Henao, Svati Shah, Lawrence Carin |
ICML | 5 |
| 2020 | On Leveraging Pretrained GANs for Generation with Limited DataabstractRecent work has shown generative adversarial networks (GANs) can generate highly realistic images, that are often indistinguishable (by humans) from real images. Most images so generated are not contained in the training dataset, suggesting potential for augmenting training sets with GAN-generated data. While this scenario is of particular relevance when there are limited data available, there is still the issue of training the GAN itself based on that limited data. To facilitate this, we leverage existing GAN models pretrained on large-scale datasets (like ImageNet) to introduce additional knowledge (which may not exist within the limited data), following the concept of transfer learning. Demonstrated by natural-image generation, we reveal that low-level filters (those close to observations) of both the generator and discriminator of pretrained GANs can be transferred to facilitate generation in a perceptually-distinct target domain with limited training data. To further adapt the transferred filters to the target domain, we propose adaptive filter modulation (AdaFM). An extensive set of experiments is presented to demonstrate the effectiveness of the proposed techniques on generation with limited data. Miaoyun Zhao, Yulai Cong, Lawrence Carin |
ICML | 3 |
| 2020 | AutoSync: Learning to Synchronize for Data-Parallel Distributed Deep LearningabstractSynchronization is a key step in data-parallel distributed machine learning (ML). Different synchronization systems and strategies perform differently, and to achieve optimal parallel training throughput requires synchronization strategies that adapt to model structures and cluster configurations. Existing synchronization systems often only consider a single or a few synchronization aspects, and the burden of deciding the right synchronization strategy is then placed on the ML practitioners, who may lack the required expertise. In this paper, we develop a model- and resource-dependent representation for synchronization, which unifies multiple synchronization aspects ranging from architecture, message partitioning, placement scheme, to communication topology. Based on this representation, we build an end-to-end pipeline, AutoSync, to automatically optimize synchronization strategies given model structures and resource specifications, lowering the bar for data-parallel distributed ML. By learning from low-shot data collected in only 200 trial runs, AutoSync can discover synchronization strategies up to 1.6x better than manually optimized ones. We develop transfer-learning mechanisms to further reduce the auto-optimization cost -- the simulators can transfer among similar model architectures, among similar cluster configurations, or both. We also present a dataset that contains over 10000 synchronization strategies and run-time pairs on a diverse set of models and cluster specifications. Hao Zhang 0025, Yuan Li 0032, Zhijie Deng, Xiaodan Liang, Lawrence Carin, Eric P. Xing |
NeurIPS | 5 |
| 2020 | GAN Memory with No ForgettingabstractAs a fundamental issue in lifelong learning, catastrophic forgetting is directly caused by inaccessible historical data; accordingly, if the data (information) were memorized perfectly, no forgetting should be expected. Motivated by that, we propose a GAN memory for lifelong learning, which is capable of remembering a stream of datasets via generative processes, with \emph{no} forgetting. Our GAN memory is based on recognizing that one can modulate the ``style'' of a GAN model to form perceptually-distant targeted generation. Accordingly, we propose to do sequential style modulations atop a well-behaved base GAN model, to form sequential targeted generative models, while simultaneously benefiting from the transferred base knowledge. The GAN memory -- that is motivated by lifelong learning -- is therefore itself manifested by a form of lifelong learning, via forward transfer and modulation of information from prior tasks. Experiments demonstrate the superiority of our method over existing approaches and its effectiveness in alleviating catastrophic forgetting for lifelong classification problems. Code is available at \url{https://github.com/MiaoyunZhao/GANmemory_LifelongLearning}. Yulai Cong, Miaoyun Zhao, Jianqiao Li, Lawrence Carin |
NeurIPS | 5 |
| 2020 | Perturbing Across the Feature Hierarchy to Improve Standard and Strict Blackbox Attack TransferabilityabstractWe consider the blackbox transfer-based targeted adversarial attack threat model in the realm of deep neural network (DNN) image classifiers. Rather than focusing on crossing decision boundaries at the output layer of the source model, our method perturbs representations throughout the extracted feature hierarchy to resemble other classes. We design a flexible attack framework that allows for multi-layer perturbations and demonstrates state-of-the-art targeted transfer performance between ImageNet DNNs. We also show the superiority of our feature space methods under a relaxation of the common assumption that the source and target models are trained on the same dataset and label space, in some instances achieving a $10\times$ increase in targeted success rate relative to other blackbox transfer methods. Finally, we analyze why the proposed methods outperform existing attack strategies and show an extension of the method in the case when limited queries to the blackbox model are allowed. Nathan Inkawhich, Kevin J. Liang, Binghui Wang, Matthew Inkawhich, Lawrence Carin, Yiran Chen 0001 |
NeurIPS | 5 |
| 2020 | Reconsidering Generative Objectives For Counterfactual ReasoningabstractThere has been recent interest in exploring generative goals for counterfactual reasoning, such as individualized treatment effect (ITE) estimation. However, existing solutions often fail to address issues that are unique to causal inference, such as covariate balancing and (infeasible) counterfactual validation. As a step towards more flexible, scalable and accurate ITE estimation, we present a novel generative Bayesian estimation framework that integrates representation learning, adversarial matching and causal estimation. By appealing to the Robinson decomposition, we derive a reformulated variational bound that explicitly targets the causal effect estimation rather than specific predictive goals. Our procedure acknowledges the uncertainties in representation and solves a Fenchel mini-max game to resolve the representation imbalance for better counterfactual generalization, justified by new theory. Further, the latent variable formulation employed enables robustness to unobservable latent confounders, extending the scope of its applicability. The utility of the proposed solution is demonstrated via an extensive set of tests against competing solutions, both under various simulation setups and to real-world datasets, with encouraging results reported. Danni Lu, Chenyang Tao, Junya Chen, Lawrence Carin |
NeurIPS | 6 |
| 2020 | Calibrating CNNs for Lifelong LearningabstractWe present an approach for lifelong/continual learning of convolutional neural networks (CNN) that does not suffer from the problem of catastrophic forgetting when moving from one task to the other. We show that the activation maps generated by the CNN trained on the old task can be calibrated using very few calibration parameters, to become relevant to the new task. Based on this, we calibrate the activation maps produced by each network layer using spatial and channel-wise calibration modules and train only these calibration parameters for each new task in order to perform lifelong learning. Our calibration modules introduce significantly less computation and parameters as compared to the approaches that dynamically expand the network. Our approach is immune to catastrophic forgetting since we store the task-adaptive calibration parameters, which contain all the task-specific knowledge and is exclusive to each task. Further, our approach does not require storing data samples from the old tasks, which is done by many replay based methods. We perform extensive experiments on multiple benchmark datasets (SVHN, CIFAR, ImageNet, and MS-Celeb), all of which show substantial improvements over state-of-the-art methods (e.g., a 29% absolute increase in accuracy on CIFAR-100 with 10 classes at a time). On large-scale datasets, our approach yields 23.8% and 9.7% absolute increase in accuracy on ImageNet-100 and MS-Celeb-10K datasets, respectively, by employing very few (0.51% and 0.35% of model parameters) task-adaptive calibration parameters. Pravendra Singh, Vinay Kumar Verma, Pratik Mazumder, Lawrence Carin, Piyush Rai |
NeurIPS | 4 |
| 2019 | Communication-Efficient Stochastic Gradient MCMC for Neural NetworksabstractLearning probability distributions on the weights of neural networks has recently proven beneficial in many applications. Bayesian methods such as Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) offer an elegant framework to reason about model uncertainty in neural networks. However, these advantages usually come with a high computational cost. We propose accelerating SG-MCMC under the masterworker framework: workers asynchronously and in parallel share responsibility for gradient computations, while the master collects the final samples. To reduce communication overhead, two protocols (downpour and elastic) are developed to allow periodic interaction between the master and workers. We provide a theoretical analysis on the finite-time estimation consistency of posterior expectations, and establish connections to sample thinning. Our experiments on various neural networks demonstrate that the proposed algorithms can greatly reduce training time while achieving comparable (or better) test accuracy/log-likelihood levels, relative to traditional SG-MCMC. When applied to reinforcement learning, it naturally provides exploration for asynchronous policy optimization, with encouraging performance improvement. Chunyuan Li, Changyou Chen, Yunchen Pu, Ricardo Henao, Lawrence Carin |
AAAI | 5 |
| 2019 | Improving Textual Network Embedding with Global Attention via Optimal TransportabstractLiqun Chen, Guoyin Wang, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang, Wenlin Wang, Yizhe Zhang, Lawrence Carin. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Liqun Chen 0001, Guoyin Wang 0002, Chenyang Tao, Dinghan Shen, Pengyu Cheng, Xinyuan Zhang 0001, Wenlin Wang, Yizhe Zhang 0002, Lawrence Carin |
ACL (1) | 9 |
| 2019 | Learning Compressed Sentence Representations for On-Device Text ProcessingabstractDinghan Shen, Pengyu Cheng, Dhanasekar Sundararaman, Xinyuan Zhang, Qian Yang, Meng Tang, Asli Celikyilmaz, Lawrence Carin. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Dinghan Shen, Pengyu Cheng, Dhanasekar Sundararaman, Xinyuan Zhang 0001, Qian Yang 0003, Asli Celikyilmaz, Lawrence Carin |
ACL (1) | 8 |
| 2019 | Towards Generating Long and Coherent Text with Multi-Level Latent Variable ModelsabstractVariational autoencoders (VAEs) have received much attention recently as an end-toend architecture for text generation with latent variables.However, previous works typically focus on synthesizing relatively short sentences (up to 20 words), and the posterior collapse issue has been widely identified in text-VAEs.In this paper, we propose to leverage several multi-level structures to learn a VAE model for generating long, and coherent text.In particular, a hierarchy of stochastic layers between the encoder and decoder networks is employed to abstract more informative and semantic-rich latent codes.Besides, we utilize a multi-level decoder structure to capture the coherent long-term structure inherent in long-form texts, by generating intermediate sentence representations as highlevel plan vectors.Extensive experimental results demonstrate that the proposed multi-level VAE model produces more coherent and less repetitive long text compared to baselines as well as can mitigate the posterior-collapse issue. Dinghan Shen, Asli Celikyilmaz, Yizhe Zhang 0002, Liqun Chen 0001, Xin Wang 0061, Jianfeng Gao 0001, Lawrence Carin |
ACL (1) | 7 |
| 2019 | Syntax-Infused Variational Autoencoder for Text GenerationabstractWe present a syntax-infused variational autoencoder (SIVAE), that integrates sentences with their syntactic trees to improve the grammar of generated sentences.Distinct from existing VAE-based text generative models, SIVAE contains two separate latent spaces, for sentences and syntactic trees.The evidence lower bound objective is redesigned correspondingly, by optimizing a joint distribution that accommodates two encoders and two decoders.SIVAE works with long shortterm memory architectures to simultaneously generate sentences and syntactic trees.Two versions of SIVAE are proposed: one captures the dependencies between the latent variables through a conditional prior network, and the other treats the latent variables independently such that syntactically-controlled sentence generation can be performed.Experimental results demonstrate the generative superiority of SIVAE on both reconstruction and targeted syntactic evaluations.Finally, we show that the proposed models can be used for unsupervised paraphrasing given different syntactic tree templates. Xinyuan Zhang 0001, Yi Yang 0038, Siyang Yuan, Dinghan Shen, Lawrence Carin |
ACL (1) | 5 |
| 2019 | Adversarial Learning of a Sampler Based on an Unnormalized DistributionabstractFundamental aspects of adversarial learning are investigated, with learning based on samples from the target distribution (conventional GAN setup). With insights so garnered, adversarial learning is extended to the case for which one has access to an unnormalized form $u(x)$ of the target density function, but no samples. Further, new concepts in GAN regularization are developed, based on learning from samples or from $u(x)$. The proposed method is compared to alternative approaches, with encouraging results demonstrated across a range of applications, including deep soft Q-learning. Chunyuan Li, Ke Bai 0001, Jianqiao Li, Guoyin Wang 0002, Changyou Chen, Lawrence Carin |
AISTATS | 6 |
| 2019 | On Connecting Stochastic Gradient MCMC and Differential PrivacyabstractConcerns related to data security and confidentiality have been raised when applying machine learning to real-world applications. Differential privacy provides a principled and rigorous privacy guarantee for machine learning models. While it is common to inject noise to design a model satisfying a required differential-privacy property, it is generally hard to balance the trade-off between privacy and utility. We show that stochastic gradient Markov chain Monte Carlo (SG-MCMC) – a class of scalable Bayesian posterior sampling algorithms – satisfies strong differential privacy, when carefully chosen stepsizes are employed. We develop theory on the performance of the proposed differentially-private SG-MCMC method. We conduct experiments to support our analysis, and show that a standard SG-MCMC sampler with minor modification can reach state-of-the-art performance in terms of both privacy and utility on Bayesian learning. Bai Li 0001, Changyou Chen, Hao Liu 0015, Lawrence Carin |
AISTATS | 4 |
| 2019 | Scalable Thompson Sampling via Optimal TransportabstractThompson sampling (TS) is a class of algorithms for sequential decision-making, which requires maintaining a posterior distribution over a reward model. However, calculating exact posterior distributions is intractable for all but the simplest models. Consequently, how to computationally-efficiently approximate a posterior distribution is a crucial problem for scalable TS with complex models, such as neural networks. In this paper, we use distribution optimization techniques to approximate the posterior distribution, solved via Wasserstein gradient flows. Based on the framework, a principled particle-optimization algorithm is developed for TS to approximate the posterior efficiently. Our approach is scalable and does not make explicit distribution assumptions on posterior approximations. Extensive experiments on both synthetic data and large-scale real data demonstrate the superior performance of the proposed methods. Ruiyi Zhang 0002, Changyou Chen, Tong Yu 0001, Lawrence Carin |
AISTATS | 6 |
| 2019 | StoryGAN: A Sequential Conditional GAN for Story VisualizationabstractIn this work, we propose a new task called Story Visualization. Given a multi-sentence paragraph, the story is visualized by generating a sequence of images, one for each sentence. In contrast to video generation, story visualization focuses less on the continuity in generated images (frames), but more on the global consistency across dynamic scenes and characters -- a challenge that has not been addressed by any single-image or video generation methods. Therefore, we propose a new story-to-image-sequence generation model, StoryGAN, based on the sequential conditional GAN framework. Our model is unique in that it consists of a deep Context Encoder that dynamically tracks the story flow, and two discriminators at the story and image levels, to enhance the image quality and the consistency of the generated sequences. To evaluate the model, we modified existing datasets to create the CLEVR-SV and Pororo-SV datasets. Empirically, StoryGAN outperformed state-of-the-art models in image quality, contextual consistency metrics, and human evaluation. Yitong Li 0001, Zhe Gan, Yelong Shen, Jingjing Liu 0001, Yu Cheng 0001, Yuexin Wu, Lawrence Carin, David E. Carlson, Jianfeng Gao 0001 |
CVPR | 7 |
| 2019 | An End-to-End Generative Architecture for Paraphrase GenerationabstractQian Yang, Zhouyuan Huo, Dinghan Shen, Yong Cheng, Wenlin Wang, Guoyin Wang, Lawrence Carin. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Qian Yang 0003, Zhouyuan Huo, Dinghan Shen, Yong Cheng 0003, Wenlin Wang, Guoyin Wang 0002, Lawrence Carin |
EMNLP/IJCNLP (1) | 7 |
| 2019 | Improving Sequence-to-Sequence Learning via Optimal Transport
Liqun Chen 0001, Yizhe Zhang 0002, Ruiyi Zhang 0002, Chenyang Tao, Zhe Gan, Bai Li 0001, Dinghan Shen, Changyou Chen, Lawrence Carin |
ICLR (Poster) | 10 |
| 2019 | GO Gradient for Expectation-Based Objectives
Yulai Cong, Miaoyun Zhao, Ke Bai 0001, Lawrence Carin |
ICLR (Poster) | 4 |
| 2019 | Stochastic Blockmodels meet Graph Neural NetworksabstractStochastic blockmodels (SBM) and their variants, $e.g.$, mixed-membership and overlapping stochastic blockmodels, are latent variable based generative models for graphs. They have proven to be successful for various tasks, such as discovering the community structure and link prediction on graph-structured data. Recently, graph neural networks, $e.g.$, graph convolutional networks, have also emerged as a promising approach to learn powerful representations (embeddings) for the nodes in the graph, by exploiting graph properties such as locality and invariance. In this work, we unify these two directions by developing a sparse variational autoencoder for graphs, that retains the interpretability of SBMs, while also enjoying the excellent predictive performance of graph neural nets. Moreover, our framework is accompanied by a fast recognition model that enables fast inference of the node embeddings (which are of independent interest for inference in SBM and its variants). Although we develop this framework for a particular type of SBM, namely the overlapping stochastic blockmodel, the proposed framework can be adapted readily for other types of SBMs. Experimental results on several benchmarks demonstrate encouraging results on link prediction while learning an interpretable latent structure that can be used for community discovery. Nikhil Mehta 0002, Lawrence Carin, Piyush Rai |
ICML | 2 |
| 2019 | Revisiting the Softmax Bellman Operator: New Benefits and New PerspectiveabstractThe impact of softmax on the value function itself in reinforcement learning (RL) is often viewed as problematic because it leads to sub-optimal value (or Q) functions and interferes with the contraction properties of the Bellman operator. Surprisingly, despite these concerns, and independent of its effect on exploration, the softmax Bellman operator when combined with Deep Q-learning, leads to Q-functions with superior policies in practice, even outperforming its double Q-learning counterpart. To better understand how and why this occurs, we revisit theoretical properties of the softmax Bellman operator, and prove that (i) it converges to the standard Bellman operator exponentially fast in the inverse temperature parameter, and (ii) the distance of its Q function from the optimal one can be bounded. These alone do not explain its superior performance, so we also show that the softmax operator can reduce the overestimation error, which may give some insight into why a sub-optimal operator leads to better performance in the presence of value function approximation. A comparison among different Bellman operators is then presented, showing the trade-offs when selecting them. Zhao Song 0001, Ronald Parr, Lawrence Carin |
ICML | 3 |
| 2019 | Variational Annealing of GANs: A Langevin PerspectiveabstractThe generative adversarial network (GAN) has received considerable attention recently as a model for data synthesis, without an explicit specification of a likelihood function. There has been commensurate interest in leveraging likelihood estimates to improve GAN training. To enrich the understanding of this fast-growing yet almost exclusively heuristic-driven subject, we elucidate the theoretical roots of some of the empirical attempts to stabilize and improve GAN training with the introduction of likelihoods. We highlight new insights from variational theory of diffusion processes to derive a likelihood-based regularizing scheme for GAN training, and present a novel approach to train GANs with an unnormalized distribution instead of empirical samples. To substantiate our claims, we provide experimental evidence on how our theoretically-inspired new algorithms improve upon current practice. Chenyang Tao, Shuyang Dai, Liqun Chen 0001, Ke Bai 0001, Junya Chen, Chang Liu 0030, Ruiyi Zhang 0002, Georgiy V. Bobashev, Lawrence Carin |
ICML | 9 |
| 2019 | Gromov-Wasserstein Learning for Graph Matching and Node EmbeddingabstractA novel Gromov-Wasserstein learning framework is proposed to jointly match (align) graphs and learn embedding vectors for the associated graph nodes. Using Gromov-Wasserstein discrepancy, we measure the dissimilarity between two graphs and find their correspondence, according to the learned optimal transport. The node embeddings associated with the two graphs are learned under the guidance of the optimal transport, the distance of which not only reflects the topological structure of each graph but also yields the correspondence across the graphs. These two learning steps are mutually-beneficial, and are unified here by minimizing the Gromov-Wasserstein discrepancy with structural regularizers. This framework leads to an optimization problem that is solved by a proximal point method. We apply the proposed method to matching problems in real-world networks, and demonstrate its superior performance compared to alternative approaches. Hongteng Xu, Dixin Luo, Hongyuan Zha, Lawrence Carin |
ICML | 4 |
| 2019 | Certified Adversarial Robustness with Additive NoiseabstractThe existence of adversarial data examples has drawn significant attention in the deep-learning community; such data are seemingly minimally perturbed relative to the original data, but lead to very different outputs from a deep-learning algorithm. Although a significant body of work on developing defense models has been developed, most such models are heuristic and are often vulnerable to adaptive attacks. Defensive methods that provide theoretical robustness guarantees have been studied intensively, yet most fail to obtain non-trivial robustness when a large-scale model and data are present. To address these limitations, we introduce a framework that is scalable and provides certified bounds on the norm of the input manipulation for constructing adversarial examples. We establish a connection between robustness against adversarial perturbation and additive random noise, and propose a training strategy that can significantly improve the certified bounds. Our evaluation on MNIST, CIFAR-10 and ImageNet suggests that our method is scalable to complicated models and large data sets, while providing competitive robustness to state-of-the-art provable defense methods. Bai Li 0001, Changyou Chen, Wenlin Wang, Lawrence Carin |
NeurIPS | 4 |
| 2019 | Kernel-Based Approaches for Sequence Modeling: Connections to Neural MethodsabstractWe investigate time-dependent data analysis from the perspective of recurrent kernel machines, from which models with hidden units and gated memory cells arise naturally. By considering dynamic gating of the memory cell, a model closely related to the long short-term memory (LSTM) recurrent neural network is derived. Extending this setup to $n$-gram filters, the convolutional neural network (CNN), Gated CNN, and recurrent additive network (RAN) are also recovered as special cases. Our analysis provides a new perspective on the LSTM, while also extending it to $n$-gram convolutional filters. Experiments are performed on natural language processing tasks and on analysis of local field potentials (neuroscience). We demonstrate that the variants we derive from kernels perform on par or even better than traditional neural methods. For the neuroscience application, the new models demonstrate significant improvements relative to the prior state of the art. Kevin J. Liang, Guoyin Wang 0002, Yitong Li 0001, Ricardo Henao, Lawrence Carin |
NeurIPS | 5 |
| 2019 | On Fenchel Mini-Max LearningabstractInference, estimation, sampling and likelihood evaluation are four primary goals of probabilistic modeling. Practical considerations often force modeling approaches to make compromises between these objectives. We present a novel probabilistic learning framework, called Fenchel Mini-Max Learning (FML), that accommodates all four desiderata in a flexible and scalable manner. Our derivation is rooted in classical maximum likelihood estimation, and it overcomes a longstanding challenge that prevents unbiased estimation of unnormalized statistical models. By reformulating MLE as a mini-max game, FML enjoys an unbiased training objective that (i) does not explicitly involve the intractable normalizing constant and (ii) is directly amendable to stochastic gradient descent optimization. To demonstrate the utility of the proposed approach, we consider learning unnormalized statistical models, nonparametric density estimation and training generative models, with encouraging empirical results presented. Chenyang Tao, Liqun Chen 0001, Shuyang Dai, Junya Chen, Ke Bai 0001, Dong Wang 0037, Jianfeng Feng, Wenlian Lu, Georgiy V. Bobashev, Lawrence Carin |
NeurIPS | 10 |
| 2019 | Improving Textual Network Learning with Variational Homophilic EmbeddingsabstractThe performance of many network learning applications crucially hinges on the success of network embedding algorithms, which aim to encode rich network information into low-dimensional vertex-based vector representations. This paper considers a novel variational formulation of network embeddings, with special focus on textual networks. Different from most existing methods that optimize a discriminative objective, we introduce Variational Homophilic Embedding (VHE), a fully generative model that learns network embeddings by modeling the semantic (textual) information with a variational autoencoder, while accounting for the structural (topology) information through a novel homophilic prior design. Homophilic vertex embeddings encourage similar embedding vectors for related (connected) vertices. The VHE encourages better generalization for downstream tasks, robustness to incomplete observations, and the ability to generalize to unseen vertices. Extensive experiments on real-world networks, for multiple tasks, demonstrate that the proposed method achieves consistently superior performance relative to competing state-of-the-art approaches. Wenlin Wang, Chenyang Tao, Zhe Gan, Guoyin Wang 0002, Liqun Chen 0001, Xinyuan Zhang 0001, Ruiyi Zhang 0002, Qian Yang 0003, Ricardo Henao, Lawrence Carin |
NeurIPS | 10 |
| 2019 | Scalable Gromov-Wasserstein Learning for Graph Partitioning and MatchingabstractWe propose a scalable Gromov-Wasserstein learning (S-GWL) method and establish a novel and theoretically-supported paradigm for large-scale graph analysis. The proposed method is based on the fact that Gromov-Wasserstein discrepancy is a pseudometric on graphs. Given two graphs, the optimal transport associated with their Gromov-Wasserstein discrepancy provides the correspondence between their nodes and achieves graph matching. When one of the graphs is a predefined graph with isolated but self-connected nodes ($i.e.$, disconnected graph), the optimal transport indicates the clustering structure of the other graph and achieves graph partitioning. Further, we extend our method to multi-graph partitioning and matching by learning a Gromov-Wasserstein barycenter graph for multiple observed graphs. Our method combines a recursive $K$-partition mechanism with a warm-start proximal gradient algorithm, whose time complexity is $\mathcal{O}(K(E+V)\log_K V)$ for graphs with $V$ nodes and $E$ edges. To our knowledge, our method is the first attempt to make Gromov-Wasserstein discrepancy applicable to large-scale graph analysis and unify graph partitioning and matching into the same framework. It outperforms state-of-the-art graph partitioning and matching methods, achieving a trade-off between accuracy and efficiency. Hongteng Xu, Dixin Luo, Lawrence Carin |
NeurIPS | 3 |
| 2019 | Ouroboros: On Accelerating Training of Transformer-Based Language ModelsabstractLanguage models are essential for natural language processing (NLP) tasks, such as machine translation and text summarization. Remarkable performance has been demonstrated recently across many NLP domains via a Transformer-based language model with over a billion parameters, verifying the benefits of model size. Model parallelism is required if a model is too large to fit in a single computing device. Current methods for model parallelism either suffer from backward locking in backpropagation or are not applicable to language models. We propose the first model-parallel algorithm that speeds the training of Transformer-based language models. We also prove that our proposed algorithm is guaranteed to converge to critical points for non-convex problems. Extensive experiments on Transformer and Transformer-XL language models demonstrate that the proposed algorithm obtains a much faster speedup beyond data parallelism, with comparable or better accuracy. Code to reproduce experiments is to be found at \url{https://github.com/LaraQianYang/Ouroboros}. Qian Yang 0003, Zhouyuan Huo, Wenlin Wang, Heng Huang 0001, Lawrence Carin |
NeurIPS | 5 |
| 2019 | A convergence analysis for a class of practical variance-reduction stochastic gradient MCMC
Changyou Chen, Wenlin Wang, Yizhe Zhang 0002, Qinliang Su, Lawrence Carin |
Sci. China Inf. Sci. | 5 |
| 2018 | Video Generation From TextabstractGenerating videos from text has proven to be a significant challenge for existing generative models. We tackle this problem by training a conditional generative model to extract both static and dynamic information from text. This is manifested in a hybrid framework, employing a Variational Autoencoder (VAE) and a Generative Adversarial Network (GAN). The static features, called "gist," are used to sketch text-conditioned background color and object layout structure. Dynamic features are considered by transforming input text into an image filter. To obtain a large amount of data for training the deep-learning model, we develop a method to automatically create a matched text-video corpus from publicly available online videos. Experimental results show that the proposed framework generates plausible and diverse short-duration smooth videos, while accurately reflecting the input text information. It significantly outperforms baseline models that directly adapt text-to-image generation procedures to produce videos. Performance is evaluated both visually and by adapting the inception score used to evaluate image generation in GANs. Yitong Li 0001, Martin Renqiang Min, Dinghan Shen, David E. Carlson, Lawrence Carin |
AAAI | 5 |
| 2018 | Adaptive Feature Abstraction for Translating Video to TextabstractPrevious models for video captioning often use the output from a specific layer of a Convolutional Neural Network (CNN) as video features. However, the variable context-dependent semantics in the video may make it more appropriate to adaptively select features from the multiple CNN layers. We propose a new approach for generating adaptive spatiotemporal representations of videos for the captioning task. A novel attention mechanism is developed, that adaptively and sequentially focuses on different layers of CNN features (levels of feature "abstraction"), as well as local spatiotemporal regions of the feature maps at each layer. The proposed approach is evaluated on three benchmark datasets: YouTube2Text, M-VAD and MSR-VTT. Along with visualizing the results and how the model works, these experiments quantitatively demonstrate the effectiveness of the proposed adaptive spatiotemporal feature abstraction for translating videos to sentences with rich semantics. Yunchen Pu, Martin Renqiang Min, Zhe Gan, Lawrence Carin |
AAAI | 4 |
| 2018 | Deconvolutional Latent-Variable Model for Text Sequence MatchingabstractA latent-variable model is introduced for text matching, inferring sentence representations by jointly optimizing generative and discriminative objectives. To alleviate typical optimization challenges in latent-variable models for text, we employ deconvolutional networks as the sequence decoder (generator), providing learned latent codes with more semantic information and better generalization. Our model, trained in an unsupervised manner, yields stronger empirical predictive performance than a decoder based on Long Short-Term Memory (LSTM), with less parameters and considerably faster training. Further, we apply it to text sequence-matching problems. The proposed model significantly outperforms several strong sentence-encoding baselines, especially in the semi-supervised setting. Dinghan Shen, Yizhe Zhang 0002, Ricardo Henao, Qinliang Su, Lawrence Carin |
AAAI | 5 |
| 2018 | Zero-Shot Learning via Class-Conditioned Deep Generative ModelsabstractWe present a deep generative model for Zero-Shot Learning (ZSL). Unlike most existing methods for this problem, that represent each class as a point (via a semantic embedding), we represent each seen/unseen class using a class-specific latent-space distribution, conditioned on class attributes. We use these latent-space distributions as a prior for a supervised variational autoencoder (VAE), which also facilitates learning highly discriminative feature representations for the inputs. The entire framework is learned end-to-end using only the seen-class training data. At test time, the label for an unseen-class test input is the class that maximizes the VAE lower bound. We further extend the model to a (i) semi-supervised/transductive setting by leveraging unlabeled unseen-class data via an unsupervised learning module, and (ii) few-shot learning where we also have a small number of labeled inputs from the unseen classes. We compare our model with several state-of-the-art methods through a comprehensive set of experiments on a variety of benchmark data sets. Wenlin Wang, Yunchen Pu, Vinay Kumar Verma, Kai Fan 0002, Yizhe Zhang 0002, Changyou Chen, Piyush Rai, Lawrence Carin |
AAAI | 8 |
| 2018 | NASH: Toward End-to-End Neural Architecture for Generative Semantic HashingabstractDinghan Shen, Qinliang Su, Paidamoyo Chapfuwa, Wenlin Wang, Guoyin Wang, Ricardo Henao, Lawrence Carin. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Dinghan Shen, Qinliang Su, Paidamoyo Chapfuwa, Wenlin Wang, Guoyin Wang 0002, Ricardo Henao, Lawrence Carin |
ACL (1) | 7 |
| 2018 | Baseline Needs More Love: On Simple Word-Embedding-Based Models and Associated Pooling MechanismsabstractDinghan Shen, Guoyin Wang, Wenlin Wang, Martin Renqiang Min, Qinliang Su, Yizhe Zhang, Chunyuan Li, Ricardo Henao, Lawrence Carin. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Dinghan Shen, Guoyin Wang 0002, Wenlin Wang, Martin Renqiang Min, Qinliang Su, Yizhe Zhang 0002, Chunyuan Li, Ricardo Henao, Lawrence Carin |
ACL (1) | 9 |
| 2018 | Joint Embedding of Words and Labels for Text ClassificationabstractGuoyin Wang, Chunyuan Li, Wenlin Wang, Yizhe Zhang, Dinghan Shen, Xinyuan Zhang, Ricardo Henao, Lawrence Carin. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Guoyin Wang 0002, Chunyuan Li, Wenlin Wang, Yizhe Zhang 0002, Dinghan Shen, Xinyuan Zhang 0001, Ricardo Henao, Lawrence Carin |
ACL (1) | 8 |
| 2018 | Symmetric Variational Autoencoder and Connections to Adversarial LearningabstractA new form of the variational autoencoder (VAE) is proposed, based on the symmetric Kullback- Leibler divergence. It is demonstrated that learn- ing of the resulting symmetric VAE (sVAE) has close connections to previously developed adversarial-learning methods. This relationship helps unify the previously distinct techniques of VAE and adversarially learning, and provides insights that allow us to ameliorate shortcomings with some previously developed adversarial methods. In addition to an analysis that motivates and explains the sVAE, an extensive set of experiments validate the utility of the approach. Liqun Chen 0001, Shuyang Dai, Yunchen Pu, Erjin Zhou, Chunyuan Li, Qinliang Su, Changyou Chen, Lawrence Carin |
AISTATS | 8 |
| 2018 | Topic Compositional Neural Language ModelabstractWe propose a Topic Compositional Neural Language Model (TCNLM), a novel method designed to simultaneously capture both the global semantic meaning and the local word-ordering structure in a document. The TCNLM learns the global semantic coherence of a document via a neural topic model, and the probability of each learned latent topic is further used to build a Mixture-of-Experts (MoE) language model, where each expert (corresponding to one topic) is a recurrent neural network (RNN) that accounts for learning the local structure of a word sequence. In order to train the MoE model efficiently, a matrix factorization method is applied, by extending each weight matrix of the RNN to be an ensemble of topic-dependent weight matrices. The degree to which each member of the ensemble is used is tied to the document-dependent probability of the corresponding topics. Experimental results on several corpora show that the proposed approach outperforms both a pure RNN-based model and other topic-guided language models. Further, our model yields sensible topics, and also has the capacity to generate meaningful sentences conditioned on given topics. Wenlin Wang, Zhe Gan, Wenqi Wang 0001, Dinghan Shen, Jiaji Huang, Wei Ping, Sanjeev Satheesh, Lawrence Carin |
AISTATS | 8 |
| 2018 | Benefits from Superposed Hawkes ProcessesabstractThe superposition of temporal point processes has been studied for many years, although the usefulness of such models for practical applications has not be fully developed. We investigate superposed Hawkes process as an important class of such models, with properties studied in the framework of least squares estimation. The superposition of Hawkes processes is demonstrated to be beneficial for tightening the upper bound of excess risk under certain conditions, and we show the feasibility of the benefit in typical situations. The usefulness of superposed Hawkes processes is verified on synthetic data, and its potential to solve the cold-start problem of recommendation systems is demonstrated on real-world data. Hongteng Xu, Dixin Luo, Xu Chen 0017, Lawrence Carin |
AISTATS | 4 |
| 2018 | Learning Structural Weight Uncertainty for Sequential Decision-MakingabstractLearning probability distributions on the weights of neural networks (NNs) has recently proven beneficial in many applications. Bayesian methods, such as Stein variational gradient descent (SVGD), offer an elegant framework to reason about NN model uncertainty. However, by assuming independent Gaussian priors for the individual NN weights (as often applied), SVGD does not impose prior knowledge that there is often structural information (dependence) among weights. We propose efficient posterior learning of structural weight uncertainty, within an SVGD framework, by employing matrix variate Gaussian priors on NN parameters. We further investigate the learned structural uncertainty in sequential decision-making problems, including contextual bandits and reinforcement learning. Experiments on several synthetic and real datasets indicate the superiority of our model, compared with state-of-the-art methods. Ruiyi Zhang 0002, Chunyuan Li, Changyou Chen, Lawrence Carin |
AISTATS | 4 |
| 2018 | The Duke Health Data Science Internship Program: Integrating the Educational Mission into Real-World Research
Shelley A. Rusincovitch, Lisa Wruck, Ricardo Henao, Larisa Rodgers, Allison Dunning, Peter Merrill, Hillary Mulder, Robert Overton, Matthew Phelan, Erich Huang, Lawrence Carin, Michael J. Pencina |
AMIA | 11 |
| 2018 | Nonlocal Low-Rank Tensor Factor Analysis for Image RestorationabstractLow-rank signal modeling has been widely leveraged to capture non-local correlation in image processing applications. We propose a new method that employs low-rank tensor factor analysis for tensors generated by grouped image patches. The low-rank tensors are fed into the alternative direction multiplier method (ADMM) to further improve image reconstruction. The motivating application is compressive sensing (CS), and a deep convolutional architecture is adopted to approximate the expensive matrix inversion in CS applications. An iterative algorithm based on this low-rank tensor factorization strategy, called NLR-TFA, is presented in detail. Experimental results on noiseless and noisy CS measurements demonstrate the superiority of the proposed approach, especially at low CS sampling rates. Xinyuan Zhang 0001, Xin Yuan 0002, Lawrence Carin |
CVPR | 3 |
| 2018 | Learning Context-Aware Convolutional Filters for Text ProcessingabstractConvolutional neural networks (CNNs) have recently emerged as a popular building block for natural language processing (NLP).Despite their success, most existing CNN models employed in NLP share the same learned (and static) set of filters for all input sentences.In this paper, we consider an approach of using a small meta network to learn contextaware convolutional filters for text processing.The role of meta network is to abstract the contextual information of a sentence or document into a set of input-aware filters.We further generalize this framework to model sentence pairs, where a bidirectional filter generation mechanism is introduced to encapsulate co-dependent sentence representations.In our benchmarks on four different tasks, including ontology classification, sentiment analysis, answer sentence selection, and paraphrase identification, our proposed model, a modified CNN with context-aware filters, consistently outperforms the standard CNN and attentionbased CNN baselines.By visualizing the learned context-aware filters, we further validate and rationalize the effectiveness of proposed framework. Dinghan Shen, Martin Renqiang Min, Yitong Li 0001, Lawrence Carin |
EMNLP | 4 |
| 2018 | Improved Semantic-Aware Network Embedding with Fine-Grained Word AlignmentabstractNetwork embeddings, which learn lowdimensional representations for each vertex in a large-scale network, have received considerable attention in recent years.For a wide range of applications, vertices in a network are typically accompanied by rich textual information such as user profiles, paper abstracts, etc.We propose to incorporate semantic features into network embeddings by matching important words between text sequences for all pairs of vertices.We introduce a word-by-word alignment framework that measures the compatibility of embeddings between word pairs, and then adaptively accumulates these alignment features with a simple yet effective aggregation function.In experiments, we evaluate the proposed framework on three real-world benchmarks for downstream tasks, including link prediction and multi-label vertex classification.Results demonstrate that our model outperforms state-of-the-art network embedding methods by a large margin. Dinghan Shen, Xinyuan Zhang 0001, Ricardo Henao, Lawrence Carin |
EMNLP | 4 |
| 2018 | Adversarial Time-to-Event ModelingabstractModern health data science applications leverage abundant molecular and electronic health data, providing opportunities for machine learning to build statistical models to support clinical practice. Time-to-event analysis, also called survival analysis, stands as one of the most representative examples of such statistical models. We present a deep-network-based approach that leverages adversarial learning to address a key challenge in modern time-to-event modeling: nonparametric estimation of event-time distributions. We also introduce a principled cost function to exploit information from censored events (events that occur subsequent to the observation window). Unlike most time-to-event models, we focus on the estimation of time-to-event distributions, rather than time ordering. We validate our model on both benchmark and real datasets, demonstrating that the proposed formulation yields significant performance gains relative to a parametric alternative, which we also propose. Paidamoyo Chapfuwa, Chenyang Tao, Chunyuan Li, Courtney Page, Benjamin Goldstein 0001, Lawrence Carin, Ricardo Henao |
ICML | 6 |
| 2018 | Continuous-Time Flows for Efficient Inference and Density EstimationabstractTwo fundamental problems in unsupervised learning are efficient inference for latent-variable models and robust density estimation based on large amounts of unlabeled data. Algorithms for the two tasks, such as normalizing flows and generative adversarial networks (GANs), are often developed independently. In this paper, we propose the concept of continuous-time flows (CTFs), a family of diffusion-based methods that are able to asymptotically approach a target distribution. Distinct from normalizing flows and GANs, CTFs can be adopted to achieve the above two goals in one framework, with theoretical guarantees. Our framework includes distilling knowledge from a CTF for efficient inference, and learning an explicit energy-based distribution with CTFs for density estimation. Both tasks rely on a new technique for distribution matching within amortized learning. Experiments on various tasks demonstrate promising performance of the proposed CTF framework, compared to related techniques. Changyou Chen, Chunyuan Li, Liquan Chen, Wenlin Wang, Yunchen Pu, Lawrence Carin |
ICML | 6 |
| 2018 | Variational Inference and Model Selection with Generalized Evidence BoundsabstractRecent advances on the scalability and flexibility of variational inference have made it successful at unravelling hidden patterns in complex data. In this work we propose a new variational bound formulation, yielding an estimator that extends beyond the conventional variational bound. It naturally subsumes the importance-weighted and Renyi bounds as special cases, and it is provably sharper than these counterparts. We also present an improved estimator for variational learning, and advocate a novel high signal-to-variance ratio update rule for the variational parameters. We discuss model-selection issues associated with existing evidence-lower-bound-based variational inference procedures, and show how to leverage the flexibility of our new formulation to address them. Empirical evidence is provided to validate our claims. Liqun Chen 0001, Chenyang Tao, Ruiyi Zhang 0002, Ricardo Henao, Lawrence Carin |
ICML | 5 |
| 2018 | JointGAN: Multi-Domain Joint Distribution Learning with Generative Adversarial NetsabstractA new generative adversarial network is developed for joint distribution matching.Distinct from most existing approaches, that only learn conditional distributions, the proposed model aims to learn a joint distribution of multiple random variables (domains). This is achieved by learning to sample from conditional distributions between the domains, while simultaneously learning to sample from the marginals of each individual domain.The proposed framework consists of multiple generators and a single softmax-based critic, all jointly trained via adversarial learning.From a simple noise source, the proposed framework allows synthesis of draws from the marginals, conditional draws given observations from a subset of random variables, or complete draws from the full joint distribution. Most examples considered are for joint analysis of two domains, with examples for three domains also presented. Yunchen Pu, Shuyang Dai, Zhe Gan, Weiyao Wang 0002, Guoyin Wang 0002, Yizhe Zhang 0002, Ricardo Henao, Lawrence Carin |
ICML | 8 |
| 2018 | Chi-square Generative Adversarial NetworkabstractTo assess the difference between real and synthetic data, Generative Adversarial Networks (GANs) are trained using a distribution discrepancy measure. Three widely employed measures are information-theoretic divergences, integral probability metrics, and Hilbert space discrepancy metrics. We elucidate the theoretical connections between these three popular GAN training criteria and propose a novel procedure, called $\chi^2$ (Chi-square) GAN, that is conceptually simple, stable at training and resistant to mode collapse. Our procedure naturally generalizes to address the problem of simultaneous matching of multiple distributions. Further, we propose a resampling strategy that significantly improves sample quality, by repurposing the trained critic function via an importance weighting mechanism. Experiments show that the proposed procedure improves stability and convergence, and yields state-of-art results on a wide range of generative modeling tasks. Chenyang Tao, Liqun Chen 0001, Ricardo Henao, Jianfeng Feng, Lawrence Carin |
ICML | 5 |
| 2018 | Learning Registered Point Processes from Idiosyncratic ObservationsabstractA parametric point process model is developed, with modeling based on the assumption that sequential observations often share latent phenomena, while also possessing idiosyncratic effects. An alternating optimization method is proposed to learn a “registered” point process that accounts for shared structure, as well as “warping” functions that characterize idiosyncratic aspects of each observed sequence. Under reasonable constraints, in each iteration we update the sample-specific warping functions by solving a set of constrained nonlinear programming problems in parallel, and update the model by maximum likelihood estimation. The justifiability, complexity and robustness of the proposed method are investigated in detail, and the influence of sequence stitching on the learning results is examined empirically. Experiments on both synthetic and real-world data demonstrate that the method yields explainable point process models, achieving encouraging results compared to state-of-the-art methods. Hongteng Xu, Lawrence Carin, Hongyuan Zha |
ICML | 2 |
| 2018 | Policy Optimization as Wasserstein Gradient FlowsabstractPolicy optimization is a core component of reinforcement learning (RL), and most existing RL methods directly optimize parameters of a policy based on maximizing the expected total reward, or its surrogate. Though often achieving encouraging empirical success, its correspondence to policy-distribution optimization has been unclear mathematically. We place policy optimization into the space of probability measures, and interpret it as Wasserstein gradient flows. On the probability-measure space, under specified circumstances, policy optimization becomes convex in terms of distribution optimization. To make optimization feasible, we develop efficient algorithms by numerically solving the corresponding discrete gradient flows. Our technique is applicable to several RL settings, and is related to many state-of-the-art policy-optimization algorithms. Specifically, we define gradient flows on both the parameter-distribution space and policy-distribution space, leading to what we term indirect-policy and direct-policy learning frameworks, respectively. Extensive experiments verify the effectiveness of our framework, often obtaining better performance compared to related algorithms. Ruiyi Zhang 0002, Changyou Chen, Chunyuan Li, Lawrence Carin |
ICML | 4 |
| 2018 | Online Continuous-Time Tensor Factorization Based on Pairwise Interactive Point ProcessesabstractA continuous-time tensor factorization method is developed for event sequences containing multiple "modalities." Each data element is a point in a tensor, whose dimensions are associated with the discrete alphabet of the modalities. Each tensor data element has an associated time of occurence and a feature vector. We model such data based on pairwise interactive point processes, and the proposed framework connects pairwise tensor factorization with a feature-embedded point process. The model accounts for interactions within each modality, interactions across different modalities, and continuous-time dynamics of the interactions. Model learning is formulated as a convex optimization problem, based on online alternating direction method of multipliers. Compared to existing state-of-the-art methods, our approach captures the latent structure of the tensor and its evolution over time, obtaining superior results on real-world datasets. Hongteng Xu, Dixin Luo, Lawrence Carin |
IJCAI | 3 |
| 2018 | Adversarial Text Generation via Feature-Mover's DistanceabstractGenerative adversarial networks (GANs) have achieved significant success in generating real-valued data. However, the discrete nature of text hinders the application of GAN to text-generation tasks. Instead of using the standard GAN objective, we propose to improve text-generation GAN via a novel approach inspired by optimal transport. Specifically, we consider matching the latent feature distributions of real and synthetic sentences using a novel metric, termed the feature-mover's distance (FMD). This formulation leads to a highly discriminative critic and easy-to-optimize objective, overcoming the mode-collapsing and brittle-training problems in existing methods. Extensive experiments are conducted on a variety of tasks to evaluate the proposed model empirically, including unconditional text generation, style transfer from non-parallel text, and unsupervised cipher cracking. The proposed model yields superior performance, demonstrating wide applicability and effectiveness. Liqun Chen 0001, Shuyang Dai, Chenyang Tao, Zhe Gan, Dinghan Shen, Yizhe Zhang 0002, Guoyin Wang 0002, Ruiyi Zhang 0002, Lawrence Carin |
NeurIPS | 10 |
| 2018 | Distilled Wasserstein Learning for Word Embedding and Topic ModelingabstractWe propose a novel Wasserstein method with a distillation mechanism, yielding joint learning of word embeddings and topics. The proposed method is based on the fact that the Euclidean distance between word embeddings may be employed as the underlying distance in the Wasserstein topic model. The word distributions of topics, their optimal transport to the word distributions of documents, and the embeddings of words are learned in a unified framework. When learning the topic model, we leverage a distilled ground-distance matrix to update the topic distributions and smoothly calculate the corresponding optimal transports. Such a strategy provides the updating of word embeddings with robust guidance, improving algorithm convergence. As an application, we focus on patient admission records, in which the proposed method embeds the codes of diseases and procedures and learns the topics of admissions, obtaining superior performance on clinically-meaningful disease network construction, mortality prediction as a function of admission codes, and procedure recommendation. Hongteng Xu, Wenlin Wang, Wei Liu 0005, Lawrence Carin |
NeurIPS | 4 |
| 2018 | Diffusion Maps for Textual Network EmbeddingabstractTextual network embedding leverages rich text information associated with the network to learn low-dimensional vectorial representations of vertices. Rather than using typical natural language processing (NLP) approaches, recent research exploits the relationship of texts on the same edge to graphically embed text. However, these models neglect to measure the complete level of connectivity between any two texts in the graph. We present diffusion maps for textual network embedding (DMTE), integrating global structural information of the graph to capture the semantic relatedness between texts, with a diffusion-convolution operation applied on the text inputs. In addition, a new objective function is designed to efficiently preserve the high-order proximity using the graph diffusion. Experimental results show that the proposed approach outperforms state-of-the-art methods on the vertex-classification and link-prediction tasks. Xinyuan Zhang 0001, Yitong Li 0001, Dinghan Shen, Lawrence Carin |
NeurIPS | 4 |
| 2017 | Unsupervised Learning with Truncated Gaussian Graphical ModelsabstractGaussian graphical models (GGMs) are widely used for statistical modeling, because of ease of inference and the ubiquitous use of the normal distribution in practical approximations. However, they are also known for their limited modeling abilities, due to the Gaussian assumption. In this paper, we introduce a novel variant of GGMs, which relaxes the Gaussian restriction and yet admits efficient inference. Specifically, we impose a bipartite structure on the GGM and govern the hidden variables by truncated normal distributions. The nonlinearity of the model is revealed by its connection to rectified linear unit (ReLU) neural networks. Meanwhile, thanks to the bipartite structure and appealing properties of truncated normals, we are able to train the models efficiently using contrastive divergence. We consider three output constructs, accounting for real-valued, binary and count data. We further extend the model to deep constructions and show that deep models can be used for unsupervised pre-training of rectifier neural networks. Extensive experimental results are provided to validate the proposed models and demonstrate their superiority over competing models. Qinliang Su, Xuejun Liao, Chunyuan Li, Zhe Gan, Lawrence Carin |
AAAI | 5 |
| 2017 | Scalable Bayesian Learning of Recurrent Neural Networks for Language ModelingabstractRecurrent neural networks (RNNs) have shown promising performance for language modeling.However, traditional training of RNNs using back-propagation through time often suffers from overfitting.One reason for this is that stochastic optimization (used for large training sets) does not provide good estimates of model uncertainty.This paper leverages recent advances in stochastic gradient Markov Chain Monte Carlo (also appropriate for large training sets) to learn weight uncertainty in RNNs.It yields a principled Bayesian learning algorithm, adding gradient noise during training (enhancing exploration of the model-parameter space) and model averaging when testing.Extensive experiments on various RNN models and across a broad range of applications demonstrate the superiority of the proposed approach relative to stochastic optimization. Zhe Gan, Chunyuan Li, Changyou Chen, Yunchen Pu, Qinliang Su, Lawrence Carin |
ACL (1) | 6 |
| 2017 | Tensor-Dictionary Learning with Deep Kruskal-Factor AnalysisabstractA multi-way factor analysis model is introduced for tensor-variate data of any order. Each data item is represented as a (sparse) sum of Kruskal decompositions, a Kruskal- factor analysis (KFA). KFA is nonparametric and can infer both the tensor-rank of each dictionary atom and the number of dictionary atoms. The model is adapted for online learning, which allows dictionary learning on large data sets. After KFA is introduced, the model is extended to a deep convolutional tensor-factor analysis, supervised by a Bayesian SVM. The experiments section demonstrates the improvement of KFA over vectorized approaches (e.g., BPFA), tensor decompositions, and convolutional neural networks (CNN) in multi-way denoising, blind inpainting, and image classification. The improvement in PSNR for the inpainting results over other methods exceeds 1dB in several cases and we achieve state of the art results on Caltech101 image classification. Andrew Stevens 0005, Yunchen Pu, Yannan Sun, Gregory Spell, Lawrence Carin |
AISTATS | 5 |
| 2017 | Learning Structured Weight Uncertainty in Bayesian Neural NetworksabstractDeep neural networks (DNNs) are increasingly popular in modern machine learning. Bayesian learning affords the opportunity to quantify posterior uncertainty on DNN model parameters. Most existing work adopts independent Gaussian priors on the model weights, ignoring possible structural information. In this paper, we consider the matrix variate Gaussian (MVG) distribution to model structured correlations within the weights of a DNN. To make posterior inference feasible, a reparametrization is proposed for the MVG prior, simplifying the complex MVG-based model to an equivalent yet simpler model with independent Gaussian priors on the transformed weights. Consequently, we develop a scalable Bayesian online inference algorithm by adopting the recently proposed probabilistic backpropagation framework. Experiments on several synthetic and real datasets indicate the superiority of our model, achieving competitive performance in terms of model likelihood and predictive root mean square error. Importantly, it also yields faster convergence speed compared to related Bayesian DNN models. Shengyang Sun, Changyou Chen, Lawrence Carin |
AISTATS | 3 |
| 2017 | Guiding Principles for the Duke Connected Care Predictive Modeling Pilot
Eugenie Komives, Shelley A. Rusincovitch, John Paat, Lawrence Carin, Daniel Costello, Michael Gao, Bradley G. Hammill, Ricardo Henao, Nigel B. Neely, Ursula Rogers, Devdutta Sangvai, Mary Schilder, Erich Huang |
AMIA | 4 |
| 2017 | Rationale and Design for the Duke Connected Care Predictive Modeling Pilot with a Medicare Shared Savings Program Population
Shelley A. Rusincovitch, Ricardo Henao, Michael Gao, Lawrence Carin, Ursula Rogers, Nigel B. Neely, Mary Schilder, Daniel Costello, Eugenie Komives, Erich Huang |
AMIA | 4 |
| 2017 | Semantic Compositional Networks for Visual CaptioningabstractA Semantic Compositional Network (SCN) is developed for image captioning, in which semantic concepts (i.e., tags) are detected from the image, and the probability of each tag is used to compose the parameters in a long short-term memory (LSTM) network. The SCN extends each weight matrix of the LSTM to an ensemble of tag-dependent weight matrices. The degree to which each member of the ensemble is used to generate an image caption is tied to the image-dependent probability of the corresponding tag. In addition to captioning images, we also extend the SCN to generate captions for video clips. We qualitatively analyze semantic composition in SCNs, and quantitatively evaluate the algorithm on three benchmark datasets: COCO, Flickr30k, and Youtube2Text. Experimental results show that the proposed method significantly outperforms prior state-of-the-art approaches, across multiple evaluation metrics. Zhe Gan, Chuang Gan 0001, Xiaodong He 0001, Yunchen Pu, Kenneth Tran, Jianfeng Gao 0001, Lawrence Carin, Li Deng 0001 |
CVPR | 7 |
| 2017 | Learning Generic Sentence Representations Using Convolutional Neural NetworksabstractWe propose a new encoder-decoder approach to learn distributed sentence representations that are applicable to multiple purposes.The model is learned by using a convolutional neural network as an encoder to map an input sentence into a continuous vector, and using a long short-term memory recurrent neural network as a decoder.Several tasks are considered, including sentence reconstruction and future sentence prediction.Further, a hierarchical encoderdecoder model is proposed to encode a sentence to predict multiple future sentences.By training our models on a large collection of novels, we obtain a highly generic convolutional sentence encoder that performs well in practice.Experimental results on several benchmark datasets, and across a broad range of applications, demonstrate the superiority of the proposed model over competing methods. Zhe Gan, Yunchen Pu, Ricardo Henao, Chunyuan Li, Xiaodong He 0001, Lawrence Carin |
EMNLP | 6 |
| 2017 | Deep Generative Models for Relational Data with Side InformationabstractWe present a probabilistic framework for overlapping community discovery and link prediction for relational data, given as a graph. The proposed framework has: (1) a deep architecture which enables us to infer multiple layers of latent features/communities for each node, providing superior link prediction performance on more complex networks and better interpretability of the latent features; and (2) a regression model which allows directly conditioning the node latent features on the side information available in form of node attributes. Our framework handles both (1) and (2) via a clean, unified model, which enjoys full local conjugacy via data augmentation, and facilitates efficient inference via closed form Gibbs sampling. Moreover, inference cost scales in the number of edges which is attractive for massive but sparse networks. Our framework is also easily extendable to model weighted networks with count-valued edges. We compare with various state-of-the-art methods and report results, both quantitative and qualitative, on several benchmark data sets. Changwei Hu, Piyush Rai, Lawrence Carin |
ICML | 3 |
| 2017 | Stochastic Gradient Monomial Gamma SamplerabstractScaling Markov Chain Monte Carlo (MCMC) to estimate posterior distributions from large datasets has been made possible as a result of advances in stochastic gradient techniques. Despite their success, mixing performance of existing methods when sampling from multimodal distributions can be less efficient with insufficient Monte Carlo samples; this is evidenced by slow convergence and insufficient exploration of posterior distributions. We propose a generalized framework to improve the sampling efficiency of stochastic gradient MCMC, by leveraging a generalized kinetics that delivers superior stationary mixing, especially in multimodal distributions, and propose several techniques to overcome the practical issues. We show that the proposed approach is better at exploring a complicated multimodal posterior distribution, and demonstrate improvements over other stochastic gradient MCMC methods on various applications. Yizhe Zhang 0002, Changyou Chen, Zhe Gan, Ricardo Henao, Lawrence Carin |
ICML | 5 |
| 2017 | Adversarial Feature Matching for Text GenerationabstractThe Generative Adversarial Network (GAN) has achieved great success in generating realistic (real-valued) synthetic data. However, convergence issues and difficulties dealing with discrete data hinder the applicability of GAN to text. We propose a framework for generating realistic text via adversarial training. We employ a long short-term memory network as generator, and a convolutional network as discriminator. Instead of using the standard objective of GAN, we propose matching the high-dimensional latent feature distributions of real and synthetic sentences, via a kernelized discrepancy metric. This eases adversarial training by alleviating the mode-collapsing problem. Our experiments show superior performance in quantitative evaluation, and demonstrate that our model can generate realistic-looking sentences. Yizhe Zhang 0002, Zhe Gan, Kai Fan 0002, Zhi Chen 0009, Ricardo Henao, Dinghan Shen, Lawrence Carin |
ICML | 7 |
| 2017 | Evaluating U.S. Electoral Representation with a Joint Statistical Model of Congressional Roll-Calls, Legislative Text, and Voter Registration DataabstractExtensive information on 3 million randomly sampled United States citizens is used to construct a statistical model of constituent preferences for each U.S. congressional district. This model is linked to the legislative voting record of the legislator from each district, yielding an integrated model for constituency data, legislative roll-call votes, and the text of the legislation. The model is used to examine the extent to which legislators' voting records are aligned with constituent preferences, and the implications of that alignment (or lack thereof) on subsequent election outcomes. The analysis is based on a Bayesian formalism, with fast inference via a stochastic variational Bayesian analysis. Zhengming Xing, Sunshine Hillygus, Lawrence Carin |
KDD | 3 |
| 2017 | An inner-loop free solution to inverse problems using deep neural networksabstractWe propose a new method that uses deep learning techniques to accelerate the popular alternating direction method of multipliers (ADMM) solution for inverse problems. The ADMM updates consist of a proximity operator, a least squares regression that includes a big matrix inversion, and an explicit solution for updating the dual variables. Typically, inner loops are required to solve the first two sub-minimization problems due to the intractability of the prior and the matrix inversion. To avoid such drawbacks or limitations, we propose an inner-loop free update rule with two pre-trained deep convolutional architectures. More specifically, we learn a conditional denoising auto-encoder which imposes an implicit data-dependent prior/regularization on ground-truth in the first sub-minimization problem. This design follows an empirical Bayesian strategy, leading to so-called amortized inference. For matrix inversion in the second sub-problem, we learn a convolutional neural network to approximate the matrix inversion, i.e., the inverse mapping is learned by feeding the input through the learned forward network. Note that training this neural network does not require ground-truth or measurements, i.e., data-independent. Extensive experiments on both synthetic data and real datasets demonstrate the efficiency and accuracy of the proposed method compared with the conventional ADMM solution using inner loops for solving inverse problems. Kai Fan 0002, Lawrence Carin, Katherine A. Heller |
NIPS | 3 |
| 2017 | Cross-Spectral Factor AnalysisabstractIn neuropsychiatric disorders such as schizophrenia or depression, there is often a disruption in the way that regions of the brain synchronize with one another. To facilitate understanding of network-level synchronization between brain regions, we introduce a novel model of multisite low-frequency neural recordings, such as local field potentials (LFPs) and electroencephalograms (EEGs). The proposed model, named Cross-Spectral Factor Analysis (CSFA), breaks the observed signal into factors defined by unique spatio-spectral properties. These properties are granted to the factors via a Gaussian process formulation in a multiple kernel learning framework. In this way, the LFP signals can be mapped to a lower dimensional space in a way that retains information of relevance to neuroscientists. Critically, the factors are interpretable. The proposed approach empirically allows similar performance in classifying mouse genotype and behavioral context when compared to commonly used approaches that lack the interpretability of CSFA. We also introduce a semi-supervised approach, termed discriminative CSFA (dCSFA). CSFA and dCSFA provide useful tools for understanding neural dynamics, particularly by aiding in the design of causal follow-up experiments. Neil Gallagher, Kyle R. Ulrich, Austin Talbot, Kafui Dzirasa, Lawrence Carin, David E. Carlson |
NIPS | 5 |
| 2017 | Triangle Generative Adversarial NetworksabstractA Triangle Generative Adversarial Network ($\Delta$-GAN) is developed for semi-supervised cross-domain joint distribution matching, where the training data consists of samples from each domain, and supervision of domain correspondence is provided by only a few paired samples. $\Delta$-GAN consists of four neural networks, two generators and two discriminators. The generators are designed to learn the two-way conditional distributions between the two domains, while the discriminators implicitly define a ternary discriminative function, which is trained to distinguish real data pairs and two kinds of fake data pairs. The generators and discriminators are trained together using adversarial learning. Under mild assumptions, in theory the joint distributions characterized by the two generators concentrate to the data distribution. In experiments, three different kinds of domain pairs are considered, image-label, image-image and image-attribute pairs. Experiments on semi-supervised image classification, image-to-image translation and attribute-based image generation demonstrate the superiority of the proposed approach. Zhe Gan, Liqun Chen 0001, Weiyao Wang 0002, Yunchen Pu, Yizhe Zhang 0002, Hao Liu 0015, Chunyuan Li, Lawrence Carin |
NIPS | 8 |
| 2017 | ALICE: Towards Understanding Adversarial Learning for Joint Distribution MatchingabstractWe investigate the non-identifiability issues associated with bidirectional adversarial training for joint distribution matching. Within a framework of conditional entropy, we propose both adversarial and non-adversarial approaches to learn desirable matched joint distributions for unsupervised and supervised tasks. We unify a broad family of adversarial models as joint distribution matching problems. Our approach stabilizes learning of unsupervised bidirectional adversarial learning methods. Further, we introduce an extension for semi-supervised learning tasks. Theoretical results are validated in synthetic data and real-world applications. Chunyuan Li, Hao Liu 0015, Changyou Chen, Yunchen Pu, Liqun Chen 0001, Ricardo Henao, Lawrence Carin |
NIPS | 7 |
| 2017 | Targeting EEG/LFP Synchrony with Neural NetsabstractWe consider the analysis of Electroencephalography (EEG) and Local Field Potential (LFP) datasets, which are “big” in terms of the size of recorded data but rarely have sufficient labels required to train complex models (e.g., conventional deep learning methods). Furthermore, in many scientific applications, the goal is to be able to understand the underlying features related to the classification, which prohibits the blind application of deep networks. This motivates the development of a new model based on {\em parameterized} convolutional filters guided by previous neuroscience research; the filters learn relevant frequency bands while targeting synchrony, which are frequency-specific power and phase correlations between electrodes. This results in a highly expressive convolutional neural network with only a few hundred parameters, applicable to smaller datasets. The proposed approach is demonstrated to yield competitive (often state-of-the-art) predictive performance during our empirical tests while yielding interpretable features. Furthermore, a Gaussian process adapter is developed to combine analysis over distinct electrode layouts, allowing the joint processing of multiple datasets to address overfitting and improve generalizability. Finally, it is demonstrated that the proposed framework effectively tracks neural dynamics on children in a clinical trial on Autism Spectrum Disorder. Yitong Li 0001, Michael Murias, Samantha Major, Geraldine Dawson, Kafui Dzirasa, Lawrence Carin, David E. Carlson |
NIPS | 6 |
| 2017 | VAE Learning via Stein Variational Gradient DescentabstractA new method for learning variational autoencoders (VAEs) is developed, based on Stein variational gradient descent. A key advantage of this approach is that one need not make parametric assumptions about the form of the encoder distribution. Performance is further enhanced by integrating the proposed encoder with importance sampling. Excellent performance is demonstrated across multiple unsupervised and semi-supervised problems, including semi-supervised analysis of the ImageNet data, demonstrating the scalability of the model to large datasets. Yunchen Pu, Zhe Gan, Ricardo Henao, Chunyuan Li, Shaobo Han, Lawrence Carin |
NIPS | 6 |
| 2017 | Adversarial Symmetric Variational AutoencoderabstractA new form of variational autoencoder (VAE) is developed, in which the joint distribution of data and codes is considered in two (symmetric) forms: (i) from observed data fed through the encoder to yield codes, and (ii) from latent codes drawn from a simple prior and propagated through the decoder to manifest data. Lower bounds are learned for marginal log-likelihood fits observed data and latent codes. When learning with the variational bound, one seeks to minimize the symmetric Kullback-Leibler divergence of joint density functions from (i) and (ii), while simultaneously seeking to maximize the two marginal log-likelihoods. To facilitate learning, a new form of adversarial training is developed. An extensive set of experiments is performed, in which we demonstrate state-of-the-art data reconstruction and generation on several image benchmarks datasets. Yunchen Pu, Weiyao Wang 0002, Ricardo Henao, Liqun Chen 0001, Zhe Gan, Chunyuan Li, Lawrence Carin |
NIPS | 7 |
| 2017 | Scalable Model Selection for Belief NetworksabstractWe propose a scalable algorithm for model selection in sigmoid belief networks (SBNs), based on the factorized asymptotic Bayesian (FAB) framework. We derive the corresponding generalized factorized information criterion (gFIC) for the SBN, which is proven to be statistically consistent with the marginal log-likelihood. To capture the dependencies within hidden variables in SBNs, a recognition network is employed to model the variational distribution. The resulting algorithm, which we call FABIA, can simultaneously execute both model selection and inference by maximizing the lower bound of gFIC. On both synthetic and real data, our experiments suggest that FABIA, when compared to state-of-the-art algorithms for learning SBNs, $(i)$ produces a more concise model, thus enabling faster testing; $(ii)$ improves predictive performance; $(iii)$ accelerates convergence; and $(iv)$ prevents overfitting. Zhao Song 0001, Yusuke Muraoka, Ryohei Fujimaki, Lawrence Carin |
NIPS | 4 |
| 2017 | A Probabilistic Framework for Nonlinearities in Stochastic Neural NetworksabstractWe present a probabilistic framework for nonlinearities, based on doubly truncated Gaussian distributions. By setting the truncation points appropriately, we are able to generate various types of nonlinearities within a unified framework, including sigmoid, tanh and ReLU, the most commonly used nonlinearities in neural networks. The framework readily integrates into existing stochastic neural networks (with hidden units characterized as random variables), allowing one for the first time to learn the nonlinearities alongside model weights in these networks. Extensive experiments demonstrate the performance improvements brought about by the proposed framework when integrated with the restricted Boltzmann machine (RBM), temporal RBM and the truncated Gaussian graphical model (TGGM). Qinliang Su, Xuejun Liao, Lawrence Carin |
NIPS | 3 |
| 2017 | Deconvolutional Paragraph Representation LearningabstractLearning latent representations from long text sequences is an important first step in many natural language processing applications. Recurrent Neural Networks (RNNs) have become a cornerstone for this challenging task. However, the quality of sentences during RNN-based decoding (reconstruction) decreases with the length of the text. We propose a sequence-to-sequence, purely convolutional and deconvolutional autoencoding framework that is free of the above issue, while also being computationally efficient. The proposed method is simple, easy to implement and can be leveraged as a building block for many applications. We show empirically that compared to RNNs, our framework is better at reconstructing and correcting long paragraphs. Quantitative evaluation on semi-supervised text classification and summarization tasks demonstrate the potential for better utilization of long unlabeled text data. Yizhe Zhang 0002, Dinghan Shen, Guoyin Wang 0002, Zhe Gan, Ricardo Henao, Lawrence Carin |
NIPS | 6 |
| 2017 | Information-Theoretic Compressive Measurement DesignabstractAn information-theoretic projection design framework is proposed, of interest for feature design and compressive measurements. Both Gaussian and Poisson measurement models are considered. The gradient of a proposed information-theoretic metric (ITM) is derived, and a gradient-descent algorithm is applied in design; connections are made to the information bottleneck. The fundamental solution structure of such design is revealed in the case of a Gaussian measurement model and arbitrary input statistics. This new theoretical result reveals how ITM parameter settings impact the number of needed projection measurements, with this verified experimentally. The ITM achieves promising results on real data, for both signal recovery and classification. Liming Wang 0004, Minhua Chen, Miguel R. D. Rodrigues, David Wilcox, A. Robert Calderbank, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2016 | Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural NetworksabstractEffective training of deep neural networks suffers from two main issues. The first is that the parameter space of these models exhibit pathological curvature. Recent methods address this problem by using adaptive preconditioning for Stochastic Gradient Descent (SGD). These methods improve convergence by adapting to the local geometry of parameter space. A second issue is overfitting, which is typically addressed by early stopping. However, recent work has demonstrated that Bayesian model averaging mitigates this problem. The posterior can be sampled by using Stochastic Gradient Langevin Dynamics (SGLD). However, the rapidly changing curvature renders default SGLD methods inefficient. Here, we propose combining adaptive preconditioners with SGLD. In support of this idea, we give theoretical properties on asymptotic convergence and predictive risk. We also provide empirical results for Logistic Regression, Feedforward Neural Nets, and Convolutional Neural Nets, demonstrating that our preconditioned SGLD method gives state-of-the-art performance on these models. Chunyuan Li, Changyou Chen, David E. Carlson, Lawrence Carin |
AAAI | 4 |
| 2016 | High-Order Stochastic Gradient Thermostats for Bayesian Learning of Deep ModelsabstractLearning in deep models using Bayesian methods has generated significant attention recently. This is largely because of the feasibility of modern Bayesian methods to yield scalable learning and inference, while maintaining a measure of uncertainty in the model parameters. Stochastic gradient MCMC algorithms (SG-MCMC) are a family of diffusion-based sampling methods for large-scale Bayesian learning. In SG-MCMC, multivariate stochastic gradient thermostats (mSGNHT) augment each parameter of interest, with a momentum and a thermostat variable to maintain stationary distributions as target posterior distributions. As the number of variables in a continuous-time diffusion increases, its numerical approximation error becomes a practical bottleneck, so better use of a numerical integrator is desirable. To this end, we propose use of an efficient symmetric splitting integrator in mSGNHT, instead of the traditional Euler integrator. We demonstrate that the proposed scheme is more accurate, robust, and converges faster. These properties are demonstrated to be desirable in Bayesian deep learning. Extensive experiments on two canonical models and their deep extensions demonstrate that the proposed scheme improves general Bayesian posterior sampling, particularly for deep models. Chunyuan Li, Changyou Chen, Kai Fan 0002, Lawrence Carin |
AAAI | 4 |
| 2016 | Learning a Hybrid Architecture for Sequence Regression and AnnotationabstractWhen learning a hidden Markov model (HMM), sequential observations can often be complemented by real-valued summary response variables generated from the path of hidden states. Such settings arise in numerous domains, including many applications in biology, like motif discovery and genome annotation. In this paper, we present a flexible framework for jointly modeling both latent sequence features and the functional mapping that relates the summary response variables to the hidden state sequence. The algorithm is compatible with a rich set of mapping functions. Results show that the availability of additional continuous response variables can simultaneously improve the annotation of the sequential observations and yield good prediction performance in both synthetic data and real-world datasets. Yizhe Zhang 0002, Ricardo Henao, Lawrence Carin, Jianling Zhong, Alexander J. Hartemink |
AAAI | 3 |
| 2016 | Bridging the Gap between Stochastic Gradient MCMC and Stochastic OptimizationabstractStochastic gradient Markov chain Monte Carlo (SG-MCMC) methods are Bayesian analogs to popular stochastic optimization methods; however, this connection is not well studied. We explore this relationship by applying simulated annealing to an SG-MCMC algorithm. Furthermore, we extend recent SG-MCMC methods with two key components: i) adaptive preconditioners (as in ADAgrad or RMSprop), and ii) adaptive element-wise momentum weights. The zero-temperature limit gives a novel stochastic optimization method with adaptive element-wise momentum weights, while conventional optimization methods only have a shared, static momentum weight. Under certain assumptions, our theoretical analysis suggests the proposed simulated annealing approach converges close to the global optima. Experiments on several deep neural network models show state-of-the-art results compared to related stochastic optimization algorithms. Changyou Chen, David E. Carlson, Zhe Gan, Chunyuan Li, Lawrence Carin |
AISTATS | 5 |
| 2016 | Variational Gaussian Copula InferenceabstractWe utilize copulas to constitute a unified framework for constructing and optimizing variational proposals in hierarchical Bayesian models. For models with continuous and non-Gaussian hidden variables, we propose a semiparametric and automated variational Gaussian copula approach, in which the parametric Gaussian copula family is able to preserve multivariate posterior dependence, and the nonparametric transformations based on Bernstein polynomials provide ample flexibility in characterizing the univariate marginal posteriors. Shaobo Han, Xuejun Liao, David B. Dunson, Lawrence Carin |
AISTATS | 4 |
| 2016 | Non-negative Matrix Factorization for Discrete Data with Hierarchical Side-InformationabstractWe present a probabilistic framework for efficient non-negative matrix factorization of discrete (count/binary) data with side-information. The side-information is given as a multi-level structure, taxonomy, or ontology, with nodes at each level being categorical-valued observations. For example, when modeling documents with a two-level side-information (documents being at level-zero), level-one may represent (one or more) authors associated with each document and level-two may represent affiliations of each author. The model easily generalizes to more than two levels (or taxonomy/ontology of arbitrary depth). Our model can learn embeddings of entities present at each level in the data/side-information hierarchy (e.g., documents, authors, affiliations, in the previous example), with appropriate sharing of information across levels. The model also enjoys full local conjugacy, facilitating efficient Gibbs sampling for model inference. Inference cost scales in the number of non-zero entries in the data matrix, which is especially appealing for real-world massive but sparse matrices. We demonstrate the effectiveness of the model on several real-world data sets. Changwei Hu, Piyush Rai, Lawrence Carin |
AISTATS | 3 |
| 2016 | Topic-Based Embeddings for Learning from Large Knowledge GraphsabstractWe present a scalable probabilistic framework for learning from multi-relational data given in form of entity-relation-entity triplets, with a potentially massive number of entities and relations (e.g., in multi-relational networks, knowledge bases, etc.). We define each triplet via a relation-specific bilinear function of the embeddings of entities associated with it (these embeddings correspond to “topics”). To handle massive number of relations and the data sparsity problem (very few observations per relation), we also extend this model to allow sharing of parameters across relations, which leads to a substantial reduction in the number of parameters to be learned. In addition to yielding excellent predictive performance (e.g., for knowledge base completion tasks), the interpretability of our topic-based embedding framework enables easy qualitative analyses. Computational cost of our models scales in the number of positive triplets, which makes it easy to scale to massive real-world multi-relational data sets, which are usually extremely sparse. We develop simple-to-implement batch as well as online Gibbs sampling algorithms and demonstrate the effectiveness of our models on tasks such as multi-relational link-prediction, and learning from large knowledge bases. Changwei Hu, Piyush Rai, Lawrence Carin |
AISTATS | 3 |
| 2016 | Parallel Majorization Minimization with Dynamically Restricted Domains for Nonconvex OptimizationabstractWe propose an optimization framework for nonconvex problems based on majorization-minimization that is particularity well-suited for parallel computing. It reduces the optimization of a high dimensional nonconvex objective function to successive optimizations of locally tight and convex upper bounds which are additively separable into low dimensional objectives. The original problem is then broken into simpler and parallel tasks, while guaranteeing the monotonic reduction of the original objective function and convergence to a local minimum. This framework also allows one to restrict the upper bound to a local dynamic convex domain, so that the bound is better matched to the local curvature of the objective function, resulting in accelerated convergence. We test the proposed framework on a nonconvex support vector machine based on a sigmoid loss function and on nonconvex penalized logistic regression. Yan Kaganovsky, Ikenna Odinaka, David E. Carlson, Lawrence Carin |
AISTATS | 4 |
| 2016 | A Deep Generative Deconvolutional Image ModelabstractA deep generative model is developed for representation and analysis of images, based on a hierarchical convolutional dictionary-learning framework. Stochastic unpooling is employed to link consecutive layers in the model, yielding top-down image generation. A Bayesian support vector machine is linked to the top-layer features, yielding max-margin discrimination. Deep deconvolutional inference is employed when testing, to infer the latent features, and the top-layer features are connected with the max-margin classifier for discrimination tasks. The algorithm is efficiently trained via Monte Carlo expectation-maximization (MCEM), with implementation on graphical processor units (GPUs) for efficient large-scale learning, and fast testing. Excellent results are obtained on several benchmark datasets, including ImageNet, demonstrating that the proposed model achieves results that are highly competitive with similarly sized convolutional neural networks. Yunchen Pu, Xin Yuan 0002, Andrew Stevens 0005, Chunyuan Li, Lawrence Carin |
AISTATS | 5 |
| 2016 | Learning Sigmoid Belief Networks via Monte Carlo Expectation MaximizationabstractBelief networks are commonly used generative models of data, but require expensive posterior estimation to train and test the model. Learning typically proceeds by posterior sampling, variational approximations, or recognition networks, combined with stochastic optimization. We propose using an online Monte Carlo expectation-maximization (MCEM) algorithm to learn the maximum a posteriori (MAP) estimator of the generative model or optimize the variational lower bound of a recognition network. The E-step in this algorithm requires posterior samples, which are already generated in current learning schema. For the M-step, we augment with Polya-Gamma (PG) random variables to give an analytic updating scheme. We show relationships to standard learning approaches by deriving stochastic gradient ascent in the MCEM framework. We apply the proposed methods to both binary and count data. Experimental results show that MCEM improves the convergence speed and often improves hold-out performance over existing learning methods. Our approach is readily generalized to other recognition networks. Zhao Song 0001, Ricardo Henao, David E. Carlson, Lawrence Carin |
AISTATS | 4 |
| 2016 | Learning Weight Uncertainty with Stochastic Gradient MCMC for Shape ClassificationabstractLearning the representation of shape cues in 2D & 3D objects for recognition is a fundamental task in computer vision. Deep neural networks (DNNs) have shown promising performance on this task. Due to the large variability of shapes, accurate recognition relies on good estimates of model uncertainty, ignored in traditional training of DNNs, typically learned via stochastic optimization. This paper leverages recent advances in stochastic gradient Markov Chain Monte Carlo (SG-MCMC) to learn weight uncertainty in DNNs. It yields principled Bayesian interpretations for the commonly used Dropout/DropConnect techniques and incorporates them into the SG-MCMC framework. Extensive experiments on 2D & 3D shape datasets and various DNN models demonstrate the superiority of the proposed approach over stochastic optimization. Our approach yields higher recognition accuracy when used in conjunction with Dropout and Batch-Normalization. Chunyuan Li, Andrew Stevens 0005, Changyou Chen, Yunchen Pu, Zhe Gan, Lawrence Carin |
CVPR | 6 |
| 2016 | A general framework for reconstruction and classification from compressive measurements with side informationabstractWe develop a general framework for compressive linear-projection measurements with side information. Side information is an additional signal correlated with the signal of interest. We investigate the impact of side information on classification and signal recovery from low-dimensional measurements. Motivated by real applications, two special cases of the general model are studied. In the first, a joint Gaussian mixture model is manifested on the signal and side information. The second example again employs a Gaussian mixture model for the signal, with side information drawn from a mixture in the exponential family. Theoretical results on recovery and classification accuracy are derived. The presence of side information is shown to yield improved performance, both theoretically and experimentally. Liming Wang 0004, Francesco Renna, Xin Yuan 0002, Miguel R. D. Rodrigues, A. Robert Calderbank, Lawrence Carin |
ICASSP | 6 |
| 2016 | Dynamic Poisson Factor AnalysisabstractWe introduce a novel dynamic model for discrete time-series data, in which the temporal sampling may be nonuniform. The model is specified by constructing a hierarchy of Poisson factor analysis blocks, one for the transitions between latent states and the other for the emissions between latent states and observations. Latent variables are binary and linked to Poisson factor analysis via Bernoulli-Poisson specifications. The model is derived for count data but can be readily modified for binary observations. We derive efficient inference via Markov chain Monte Carlo, that scales with the number of non-zeros in the data and latent binary states, yielding significant acceleration compared to related models. Experimental results on benchmark data show the proposed model achieves state-of-the-art predictive performance. Additional experiments on microbiome data demonstrate applicability of the proposed model to interesting problems in computational biology where interpretability is of utmost importance. Yizhe Zhang 0002, Lawrence David, Ricardo Henao, Lawrence Carin |
ICDM | 5 |
| 2016 | Factored Temporal Sigmoid Belief Networks for Sequence LearningabstractDeep conditional generative models are developed to simultaneously learn the temporal dependencies of multiple sequences. The model is designed by introducing a three-way weight tensor to capture the multiplicative interactions between side information and sequences. The proposed model builds on the Temporal Sigmoid Belief Network (TSBN), a sequential stack of Sigmoid Belief Networks (SBNs). The transition matrices are further factored to reduce the number of parameters and improve generalization. When side information is not available, a general framework for semi-supervised learning based on the proposed model is constituted, allowing robust sequence classification. Experimental results show that the proposed approach achieves state-of-the-art predictive and classification performance on sequential data, and has the capacity to synthesize sequences, with controlled style transitioning and blending. Jiaming Song, Zhe Gan, Lawrence Carin |
ICML | 3 |
| 2016 | Nonlinear Statistical Learning with Truncated Gaussian Graphical ModelsabstractWe introduce the truncated Gaussian graphical model (TGGM) as a novel framework for designing statistical models for nonlinear learning. A TGGM is a Gaussian graphical model (GGM) with a subset of variables truncated to be nonnegative. The truncated variables are assumed latent and integrated out to induce a marginal model. We show that the variables in the marginal model are non-Gaussian distributed and their expected relations are nonlinear. We use expectation-maximization to break the inference of the nonlinear model into a sequence of TGGM inference problems, each of which is efficiently solved by using the properties and numerical methods of multivariate Gaussian distributions. We use the TGGM to design models for nonlinear regression and classification, with the performances of these models demonstrated on extensive benchmark datasets and compared to state-of-the-art competing results. Qinliang Su, Xuejun Liao, Changyou Chen, Lawrence Carin |
ICML | 4 |
| 2016 | Bayesian Dictionary Learning with Gaussian Processes and Sigmoid Belief Networks
Yizhe Zhang 0002, Ricardo Henao, Chunyuan Li, Lawrence Carin |
IJCAI | 4 |
| 2016 | Stochastic Gradient MCMC with Stale GradientsabstractStochastic gradient MCMC (SG-MCMC) has played an important role in large-scale Bayesian learning, with well-developed theoretical convergence properties. In such applications of SG-MCMC, it is becoming increasingly popular to employ distributed systems, where stochastic gradients are computed based on some outdated parameters, yielding what are termed stale gradients. While stale gradients could be directly used in SG-MCMC, their impact on convergence properties has not been well studied. In this paper we develop theory to show that while the bias and MSE of an SG-MCMC algorithm depend on the staleness of stochastic gradients, its estimation variance (relative to the expected estimate, based on a prescribed number of samples) is independent of it. In a simple Bayesian distributed system with SG-MCMC, where stale gradients are computed asynchronously by a set of workers, our theory indicates a linear speedup on the decrease of estimation variance w.r.t. the number of workers. Experiments on synthetic data and deep neural networks validate our theory, demonstrating the effectiveness and scalability of SG-MCMC with stale gradients. Changyou Chen, Nan Ding 0002, Chunyuan Li, Yizhe Zhang 0002, Lawrence Carin |
NIPS | 5 |
| 2016 | Variational Autoencoder for Deep Learning of Images, Labels and CaptionsabstractA novel variational autoencoder is developed to model images, as well as associated labels or captions. The Deep Generative Deconvolutional Network (DGDN) is used as a decoder of the latent image features, and a deep Convolutional Neural Network (CNN) is used as an image encoder; the CNN is used to approximate a distribution for the latent DGDN features/code. The latent code is also linked to generative models for labels (Bayesian support vector machine) or captions (recurrent neural network). When predicting a label/caption for a new image at test, averaging is performed across the distribution of latent codes; this is computationally efficient as a consequence of the learned CNN-based encoder. Since the framework is capable of modeling the image in the presence/absence of associated labels/captions, a new semi-supervised setting is manifested for CNN learning with images; the framework even allows unsupervised CNN learning, based on images alone. Yunchen Pu, Zhe Gan, Ricardo Henao, Xin Yuan 0002, Chunyuan Li, Andrew Stevens 0005, Lawrence Carin |
NIPS | 7 |
| 2016 | Linear Feature Encoding for Reinforcement LearningabstractFeature construction is of vital importance in reinforcement learning, as the quality of a value function or policy is largely determined by the corresponding features. The recent successes of deep reinforcement learning (RL) only increase the importance of understanding feature construction. Typical deep RL approaches use a linear output layer, which means that deep RL can be interpreted as a feature construction/encoding network followed by linear value function approximation. This paper develops and evaluates a theory of linear feature encoding. We extend theoretical results on feature quality for linear value function approximation from the uncontrolled case to the controlled case. We then develop a supervised linear feature encoding method that is motivated by insights from linear value function approximation theory, as well as empirical successes from deep RL. The resulting encoder is a surprisingly effective method for linear value function approximation using raw images as inputs. Zhao Song 0001, Ronald Parr, Xuejun Liao, Lawrence Carin |
NIPS | 4 |
| 2016 | Towards Unifying Hamiltonian Monte Carlo and Slice SamplingabstractWe unify slice sampling and Hamiltonian Monte Carlo (HMC) sampling, demonstrating their connection via the Hamiltonian-Jacobi equation from Hamiltonian mechanics. This insight enables extension of HMC and slice sampling to a broader family of samplers, called Monomial Gamma Samplers (MGS). We provide a theoretical analysis of the mixing performance of such samplers, proving that in the limit of a single parameter, the MGS draws decorrelated samples from the desired target distribution. We further show that as this parameter tends toward this limit, performance gains are achieved at a cost of increasing numerical difficulty and some practical convergence issues. Our theoretical results are validated with synthetic data and real-world applications. Yizhe Zhang 0002, Xiangyu Wang 0006, Changyou Chen, Ricardo Henao, Kai Fan 0002, Lawrence Carin |
NIPS | 6 |
| 2016 | Deep Metric Learning with Data Summarization
Wenlin Wang, Changyou Chen, Piyush Rai, Lawrence Carin |
ECML/PKDD (1) | 5 |
| 2016 | Laplacian Hamiltonian Monte Carlo
Yizhe Zhang 0002, Changyou Chen, Ricardo Henao, Lawrence Carin |
ECML/PKDD (1) | 4 |
| 2016 | Electronic Health Record Analysis via Deep Poisson Factor ModelsabstractElectronic Health Record (EHR) phenotyping utilizes patient data captured through normal medical practice, to identify features that may represent computational medical phenotypes. These features may be used to identify at-risk patients and improve prediction of patient morbidity and mortality. We present a novel deep multi-modality architecture for EHR analysis (applicable to joint analysis of multiple forms of EHR data), based on Poisson Factor Analysis (PFA) modules. Each modality, composed of observed counts, is represented as a Poisson distribution, parameterized in terms of hidden binary units. Information from different modalities is shared via a deep hierarchy of common hidden units. Activation of these binary units occurs with probability characterized as Bernoulli- Poisson link functions, instead of more traditional logistic link functions. In addition, we demonstrate that PFA modules can be adapted to discriminative modalities. To compute model parameters, we derive efficient Markov Chain Monte Carlo (MCMC) inference that scales efficiently, with significant computational gains when compared to related models based on logistic link functions. To explore the utility of these models, we apply them to a subset of patients from the Duke-Durham patient cohort. We identified a cohort of over 16,000 patients with Type 2 Diabetes Mellitus (T2DM) based on diagnosis codes and laboratory tests out of our patient population of over 240,000. Examining the common hidden units uniting the PFA modules, we identify patient features that represent medical concepts. Experiments indicate that our learned features are better able to predict mortality and morbidity than clinical features identified previously in a large-scale clinical trial. Ricardo Henao, James Lu, Joseph E. Lucas, Jeffrey M. Ferranti, Lawrence Carin |
J. Mach. Learn. Res. | 5 |
| 2016 | Classification and Reconstruction of High-Dimensional Signals From Low-Dimensional Features in the Presence of Side InformationabstractThis paper offers a characterization of fundamental limits on the classification and reconstruction of high-dimensional signals from low-dimensional features, in the presence of side information. We consider a scenario where a decoder has access both to linear features of the signal of interest and to linear features of the side information signal; while the side information may be in a compressed form, the objective is recovery or classification of the primary signal, not the side information. The signal of interest and the side information are each assumed to have (distinct) latent discrete labels; conditioned on these two labels, the signal of interest and side information are drawn from a multivariate Gaussian distribution that correlates the two. With joint probabilities on the latent labels, the overall signal-(side information) representation is defined by a Gaussian mixture model. By considering bounds to the misclassification probability associated with the recovery of the underlying signal label, and bounds to the reconstruction error associated with the recovery of the signal of interest itself, we then provide sharp sufficient and/or necessary conditions for these quantities to approach zero when the covariance matrices of the Gaussians are nearly low rank. These conditions, which are reminiscent of the well-known Slepian-Wolf and Wyner-Ziv conditions, are the function of the number of linear features extracted from signal of interest, the number of linear features extracted from the side information signal, and the geometry of these signals and their interplay. Moreover, on assuming that the signal of interest and the side information obey such an approximately low-rank model, we derive the expansions of the reconstruction error as a function of the deviation from an exactly low-rank model; such expansions also allow the identification of operational regimes, where the impact of side information on signal reconstruction is most relevant. Our framework, which offers a principled mechanism to integrate side information in high-dimensional data problems, is also tested in the context of imaging applications. In particular, we report state-of-theart results in compressive hyperspectral imaging applications, where the accompanying side information is a conventional digital photograph. Francesco Renna, Liming Wang 0004, Xin Yuan 0002, Jianbo Yang, Galen Reeves, A. Robert Calderbank, Lawrence Carin, Miguel R. D. Rodrigues |
IEEE Trans. Inf. Theory | 7 |
| 2015 | Integrating Features and Similarities: Flexible Models for Heterogeneous Multiview DataabstractWe present a probabilistic framework for learning with heterogeneous multiview data where some views are given as ordinal, binary, or real-valued feature matrices, and some views as similarity matrices. Our framework has the following distinguishing aspects: (i) a unified latent factor model for integrating information from diverse feature (ordinal, binary, real) and similarity based views, and predicting the missing data in each view, leveraging view correlations; (ii) seamless adaptation to binary/multiclass classification where data consists of multiple feature and/or similarity-based views; and (iii) an efficient, variational inference algorithm which is especially flexible in modeling the views with ordinal-valued data (by learning the cutpoints for the ordinal data), and extends naturally to streaming data settings. Our framework subsumes methods such as multiview learning and multiple kernel learning as special cases. We demonstrate the effectiveness of our framework on several real-world and benchmarks datasets. Wenzhao Lian, Piyush Rai, Esther Salazar, Lawrence Carin |
AAAI | 4 |
| 2015 | Leveraging Features and Networks for Probabilistic Tensor DecompositionabstractWe present a probabilistic model for tensor decomposition where one or more tensor modes may have side-information about the mode entities in form of their features and/or their adjacency network. We consider a Bayesian approach based on the Canonical PARAFAC (CP) decomposition and enrich this single-layer decomposition approach with a two-layer decomposition. The second layer fits a factor model for each layer-one factor matrix and models the factor matrix via the mode entities' features and/or the network between the mode entities. The second-layer decomposition of each factor matrix also learns a binary latent representation for the entities of that mode, which can be useful in its own right. Our model can handle both continuous as well as binary tensor observations. Another appealing aspect of our model is the simplicity of the model inference, with easy-to-sample Gibbs updates. We demonstrate the results of our model on several benchmarks datasets, consisting of both real and binary tensors. Piyush Rai, Yingjian Wang 0004, Lawrence Carin |
AAAI | 3 |
| 2015 | Cross-Modal Similarity Learning via Pairs, Preferences, and Active SupervisionabstractWe present a probabilistic framework for learning pairwise similarities between objects belonging to different modalities, such as drugs and proteins, or text and images. Our framework is based on learning a binary code based representation for objects in each modality, and has the following key properties: (i) it can leverage both pairwise as well as easy-to-obtain relative preference based cross-modal constraints, (ii) the probabilistic framework naturally allows querying for the most useful/informative constraints, facilitating an active learning setting (existing methods for cross-modal similarity learning do not have such a mechanism), and (iii) the binary code length is learned from the data. We demonstrate the effectiveness of the proposed approach on two problems that require computing pairwise similarities between cross-modal object pairs: cross-modal link prediction in bipartite graphs, and hashing based cross-modal similarity search. Yi Zhen, Piyush Rai, Hongyuan Zha, Lawrence Carin |
AAAI | 4 |
| 2015 | Stochastic Spectral Descent for Restricted Boltzmann MachinesabstractRestricted Boltzmann Machines (RBMs) are widely used as building blocks for deep learning models. Learning typically proceeds by using stochastic gradient descent, and the gradients are estimated with sampling methods. However, the gradient estimation is a computational bottleneck, so better use of the gradients will speed up the descent algorithm. To this end, we first derive upper bounds on the RBM cost function, then show that descent methods can have natural ad- vantages by operating in the L∞and Shatten-∞norm. We introduce a new method called “Stochastic Spectral Descent” that updates parameters in the normed space. Empirical results show dramatic improvements over stochastic gradient descent, and have only have a fractional increase on the per-iteration cost. David E. Carlson, Volkan Cevher, Lawrence Carin |
AISTATS | 3 |
| 2015 | Learning Deep Sigmoid Belief Networks with Data AugmentationabstractDeep directed generative models are developed. The multi-layered model is designed by stacking sigmoid belief networks, with sparsity-encouraging priors placed on the model parameters. Learning and inference of layer-wise model parameters are implemented in a Bayesian setting. By exploring the idea of data augmentation and introducing auxiliary Polya-Gamma variables, simple and efficient Gibbs sampling and mean-field variational Bayes (VB) inference are implemented. To address large-scale datasets, an online version of VB is also developed. Experimental results are presented for three publicly available datasets: MNIST, Caltech 101 Silhouettes and OCR letters. Zhe Gan, Ricardo Henao, David E. Carlson, Lawrence Carin |
AISTATS | 4 |
| 2015 | Scalable Deep Poisson Factor Analysis for Topic ModelingabstractA new framework for topic modeling is developed, based on deep graphical models, where interactions between topics are inferred through deep latent binary hierarchies. The proposed multi-layer model employs a deep sigmoid belief network or restricted Boltzmann machine, the bottom binary layer of which selects topics for use in a Poisson factor analysis model. Under this setting, topics live on the bottom layer of the model, while the deep specification serves as a flexible prior for revealing topic structure. Scalable inference algorithms are derived by applying Bayesian conditional density filtering algorithm, in addition to extending recently proposed work on stochastic gradient thermostats. Experimental results on several corpora show that the proposed approach readily handles very large collections of text documents, infers structured topic representations, and obtains superior test perplexities when compared with related models. Zhe Gan, Changyou Chen, Ricardo Henao, David E. Carlson, Lawrence Carin |
ICML | 5 |
| 2015 | A Multitask Point Process Predictive ModelabstractPoint process data are commonly observed in fields like healthcare and social science. Designing predictive models for such event streams is an under-explored problem, due to often scarce training data. In this work we propose a multitask point process model, leveraging information from all tasks via a hierarchical Gaussian process (GP). Nonparametric learning functions implemented by a GP, which map from past events to future rates, allow analysis of flexible arrival patterns. To facilitate efficient inference, we propose a sparse construction for this hierarchical model, and derive a variational Bayes method for learning and inference. Experimental results are shown on both synthetic data and an application on real electronic health records. Wenzhao Lian, Ricardo Henao, Vinayak A. Rao, Joseph E. Lucas, Lawrence Carin |
ICML | 5 |
| 2015 | Non-Gaussian Discriminative Factor Models via the Max-Margin Rank-LikelihoodabstractWe consider the problem of discriminative factor analysis for data that are in general non-Gaussian. A Bayesian model based on the ranks of the data is proposed. We first introduce a max-margin version of the rank-likelihood. A discriminative factor model is then developed, integrating the new max-margin rank-likelihood and (linear) Bayesian support vector machines, which are also built on the max-margin principle. The discriminative factor model is further extended to the nonlinear case through mixtures of local linear classifiers, via Dirichlet processes. Fully local conjugacy of the model yields efficient inference with both Markov Chain Monte Carlo and variational Bayes approaches. Extensive experiments on benchmark and real data demonstrate superior performance of the proposed model and its potential for applications in computational biology. Xin Yuan 0002, Ricardo Henao, Ephraim Tsalik, Raymond Langley, Lawrence Carin |
ICML | 5 |
| 2015 | Stick-Breaking Policy Learning in Dec-POMDPs
Miao Liu 0001, Christopher Amato, Xuejun Liao, Lawrence Carin, Jonathan P. How |
IJCAI | 4 |
| 2015 | Scalable Probabilistic Tensor Factorization for Binary and Count Data
Piyush Rai, Changwei Hu, Matthew Harding, Lawrence Carin |
IJCAI | 4 |
| 2015 | Classification and reconstruction of compressed GMM signals with side informationabstractThis paper offers a characterization of performance limits for classification and reconstruction of high-dimensional signals from noisy compressive measurements, in the presence of side information. We assume the signal of interest and the side information signal are drawn from a correlated mixture of distributions/components, where each component associated with a specific class label follows a Gaussian mixture model (GMM). We provide sharp sufficient and/or necessary conditions for the phase transition of the misclassification probability and the reconstruction error in the low-noise regime. These conditions, which are reminiscent of the well-known Slepian-Wolf and Wyner-Ziv conditions, are a function of the number of measurements taken from the signal of interest, the number of measurements taken from the side information signal, and the geometry of these signals and their interplay. Francesco Renna, Liming Wang 0004, Xin Yuan 0002, Jianbo Yang, Galen Reeves, A. Robert Calderbank, Lawrence Carin, Miguel R. D. Rodrigues |
ISIT | 7 |
| 2015 | A concentration-of-measure inequality for multiple-measurement modelsabstractClassical compressive sensing typically assumes a single measurement, and theoretical analysis often relies on corresponding concentration-of-measure results. There are many real-world applications involving multiple compressive measurements, from which the underlying signals may be estimated. In this paper, we establish a new concentration-of-measure inequality for a block-diagonal structured random compressive sensing matrix with Rademacher-ensembles. We discuss applications of this newly-derived inequality to two appealing compressive multiple-measurement models: for Gaussian and Poisson systems. In particular, Johnson-Lindenstrauss-type results and a compressed-domain classification result are derived for a Gaussian multiple-measurement model. We also propose, as another contribution, theoretical performance guarantees for signal recovery for multi-measurement Poisson systems, via the inequality. Liming Wang 0004, Jiaji Huang, Xin Yuan 0002, Volkan Cevher, Miguel R. D. Rodrigues, A. Robert Calderbank, Lawrence Carin |
ISIT | 7 |
| 2015 | Preconditioned Spectral Descent for Deep LearningabstractDeep learning presents notorious computational challenges. These challenges include, but are not limited to, the non-convexity of learning objectives and estimating the quantities needed for optimization algorithms, such as gradients. While we do not address the non-convexity, we present an optimization solution that ex- ploits the so far unused “geometry” in the objective function in order to best make use of the estimated gradients. Previous work attempted similar goals with preconditioned methods in the Euclidean space, such as L-BFGS, RMSprop, and ADA-grad. In stark contrast, our approach combines a non-Euclidean gradient method with preconditioning. We provide evidence that this combination more accurately captures the geometry of the objective function compared to prior work. We theoretically formalize our arguments and derive novel preconditioned non-Euclidean algorithms. The results are promising in both computational time and quality when applied to Restricted Boltzmann Machines, Feedforward Neural Nets, and Convolutional Neural Nets. David E. Carlson, Edo Collins, Ya-Ping Hsieh, Lawrence Carin, Volkan Cevher |
NIPS | 4 |
| 2015 | On the Convergence of Stochastic Gradient MCMC Algorithms with High-Order IntegratorsabstractRecent advances in Bayesian learning with large-scale data have witnessed emergence of stochastic gradient MCMC algorithms (SG-MCMC), such as stochastic gradient Langevin dynamics (SGLD), stochastic gradient Hamiltonian MCMC (SGHMC), and the stochastic gradient thermostat. While finite-time convergence properties of the SGLD with a 1st-order Euler integrator have recently been studied, corresponding theory for general SG-MCMCs has not been explored. In this paper we consider general SG-MCMCs with high-order integrators, and develop theory to analyze finite-time convergence properties and their asymptotic invariant measures. Our theoretical results show faster convergence rates and more accurate invariant measures for SG-MCMCs with higher-order integrators. For example, with the proposed efficient 2nd-order symmetric splitting integrator, the mean square error (MSE) of the posterior average for the SGHMC achieves an optimal convergence rate of $L^{-4/5}$ at $L$ iterations, compared to $L^{-2/3}$ for the SGHMC and SGLD with 1st-order Euler integrators. Furthermore, convergence results of decreasing-step-size SG-MCMCs are also developed, with the same convergence rates as their fixed-step-size counterparts for a specific decreasing sequence. Experiments on both synthetic and real datasets verify our theory, and show advantages of the proposed method in two large-scale real applications. Changyou Chen, Nan Ding 0002, Lawrence Carin |
NIPS | 3 |
| 2015 | Deep Temporal Sigmoid Belief Networks for Sequence ModelingabstractDeep dynamic generative models are developed to learn sequential dependencies in time-series data. The multi-layered model is designed by constructing a hierarchy of temporal sigmoid belief networks (TSBNs), defined as a sequential stack of sigmoid belief networks (SBNs). Each SBN has a contextual hidden state, inherited from the previous SBNs in the sequence, and is used to regulate its hidden bias. Scalable learning and inference algorithms are derived by introducing a recognition model that yields fast sampling from the variational posterior. This recognition model is trained jointly with the generative model, by maximizing its variational lower bound on the log-likelihood. Experimental results on bouncing balls, polyphonic music, motion capture, and text streams show that the proposed approach achieves state-of-the-art predictive performance, and has the capacity to synthesize various sequences. Zhe Gan, Chunyuan Li, Ricardo Henao, David E. Carlson, Lawrence Carin |
NIPS | 5 |
| 2015 | Deep Poisson Factor ModelingabstractWe propose a new deep architecture for topic modeling, based on Poisson Factor Analysis (PFA) modules. The model is composed of a Poisson distribution to model observed vectors of counts, as well as a deep hierarchy of hidden binary units. Rather than using logistic functions to characterize the probability that a latent binary unit is on, we employ a Bernoulli-Poisson link, which allows PFA modules to be used repeatedly in the deep architecture. We also describe an approach to build discriminative topic models, by adapting PFA modules. We derive efficient inference via MCMC and stochastic variational methods, that scale with the number of non-zeros in the data and binary units, yielding significant efficiency, relative to models based on logistic links. Experiments on several corpora demonstrate the advantages of our model when compared to related deep models. Ricardo Henao, Zhe Gan, James Lu, Lawrence Carin |
NIPS | 4 |
| 2015 | Large-Scale Bayesian Multi-Label Learning via Topic-Based Label EmbeddingsabstractWe present a scalable Bayesian multi-label learning model based on learning low-dimensional label embeddings. Our model assumes that each label vector is generated as a weighted combination of a set of topics (each topic being a distribution over labels), where the combination weights (i.e., the embeddings) for each label vector are conditioned on the observed feature vector. This construction, coupled with a Bernoulli-Poisson link function for each label of the binary label vector, leads to a model with a computational cost that scales in the number of positive labels in the label matrix. This makes the model particularly appealing for real-world multi-label learning problems where the label matrix is usually very massive but highly sparse. Using a data-augmentation strategy leads to full local conjugacy in our model, facilitating simple and very efficient Gibbs sampling, as well as an Expectation Maximization algorithm for inference. Also, predicting the label vector at test time does not require doing an inference for the label embeddings and can be done in closed form. We report results on several benchmark data sets, comparing our model with various state-of-the art methods. Piyush Rai, Changwei Hu, Ricardo Henao, Lawrence Carin |
NIPS | 4 |
| 2015 | GP Kernels for Cross-Spectrum AnalysisabstractMulti-output Gaussian processes provide a convenient framework for multi-task problems. An illustrative and motivating example of a multi-task problem is multi-region electrophysiological time-series data, where experimentalists are interested in both power and phase coherence between channels. Recently, Wilson and Adams (2013) proposed the spectral mixture (SM) kernel to model the spectral density of a single task in a Gaussian process framework. In this paper, we develop a novel covariance kernel for multiple outputs, called the cross-spectral mixture (CSM) kernel. This new, flexible kernel represents both the power and phase relationship between multiple observation channels. We demonstrate the expressive capabilities of the CSM kernel through implementation of a Bayesian hidden Markov model, where the emission distribution is a multi-output Gaussian process with a CSM covariance kernel. Results are presented for measured multi-region electrophysiological data. Kyle R. Ulrich, David E. Carlson, Kafui Dzirasa, Lawrence Carin |
NIPS | 4 |
| 2015 | Scalable Bayesian Non-negative Tensor Factorization for Massive Count Data
Changwei Hu, Piyush Rai, Changyou Chen, Matthew Harding, Lawrence Carin |
ECML/PKDD (2) | 5 |
| 2015 | Zero-Truncated Poisson Tensor Factorization for Massive Binary Tensors
Changwei Hu, Piyush Rai, Lawrence Carin |
UAI | 3 |
| 2015 | A Bayesian Nonparametric Approach to Image Super-ResolutionabstractSuper-resolution methods form high-resolution images from low-resolution images. In this paper, we develop a new Bayesian nonparametric model for super-resolution. Our method uses a beta-Bernoulli process to learn a set of recurring visual patterns, called dictionary elements, from the data. Because it is nonparametric, the number of elements found is also determined from the data. We test the results on both benchmark and natural images, comparing with several other models from the research literature. We perform large-scale human evaluation experiments to assess the visual quality of the results. In a first implementation, we use Gibbs sampling to approximate the posterior. However, this algorithm is not feasible for large-scale data. To circumvent this, we then develop an online variational Bayes (VB) algorithm. This algorithm finds high quality dictionaries in a fraction of the time needed by the Gibbs sampler. Gungor Polatkan, Mingyuan Zhou, Lawrence Carin, David M. Blei, Ingrid Daubechies |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Negative Binomial Process Count and Mixture ModelingabstractThe seemingly disjoint problems of count and mixture modeling are united under the negative binomial (NB) process. A gamma process is employed to model the rate measure of a Poisson process, whose normalization provides a random probability measure for mixture modeling and whose marginalization leads to an NB process for count modeling. A draw from the NB process consists of a Poisson distributed finite number of distinct atoms, each of which is associated with a logarithmic distributed number of data samples. We reveal relationships between various count- and mixture-modeling distributions and construct a Poisson-logarithmic bivariate distribution that connects the NB and Chinese restaurant table distributions. Fundamental properties of the models are developed, and we derive efficient Bayesian inference. It is shown that with augmentation and normalization, the NB process and gamma-NB process can be reduced to the Dirichlet process and hierarchical Dirichlet process, respectively. These relationships highlight theoretical, structural, and computational advantages of the NB process. A variety of NB processes, including the beta-geometric, beta-NB, marked-beta-NB, marked-gamma-NB and zero-inflated-NB processes, with distinct sharing mechanisms, are also constructed. These models are applied to topic modeling, with connections made to existing algorithms under Poisson factor analysis. Example results show the importance of inferring both the NB dispersion and probability parameters. Mingyuan Zhou, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Alternating Minimization Algorithm with Automatic Relevance Determination for Transmission Tomography under Poisson NoiseabstractWe propose a globally convergent alternating minimization (AM) algorithm for image reconstruction in transmission tomography, which extends automatic relevance determination (ARD) to Poisson noise models with Beer's law. The algorithm promotes solutions that are sparse in the pixel/voxel--difference domain by introducing additional latent variables, one for each pixel/voxel, and then learning these variables from the data using a hierarchical Bayesian model. Importantly, the proposed AM algorithm is free of any tuning parameters with image quality comparable to standard penalized likelihood methods. Our algorithm exploits optimization transfer principles which reduce the problem into parallel one-dimensional optimization tasks (one for each pixel/voxel), making the algorithm feasible for large-scale problems. This approach considerably reduces the computational bottleneck of ARD associated with the posterior variances. Positivity constraints inherent in transmission tomography problems are also enforced. We demonstrate the performance of the proposed algorithm for x-ray computed tomography using synthetic and real-world datasets. The algorithm is shown to have much better performance than prior ARD algorithms based on approximate Gaussian noise models, even for high photon flux. Sample code is available from http://www.yan-kaganovsky.com/\#!code/c24bp. Yan Kaganovsky, Shaobo Han, Soysal Degirmenci, David G. Politte, David J. Brady, Joseph A. O'Sullivan, Lawrence Carin |
SIAM J. Imaging Sci. | 7 |
| 2015 | Signal Recovery and System Calibration from Multiple Compressive Poisson MeasurementsabstractThe measurement matrix employed in compressive sensing typically cannot be known precisely a priori and must be estimated via calibration. One may take multiple compressive measurements, from which the measurement matrix and underlying signals may be estimated jointly. This is of interest as well when the measurement matrix may change as a function of the details of what is measured. This problem has been considered recently for Gaussian measurement noise, and here we develop this idea with application to Poisson systems. A collaborative maximum likelihood algorithm and alternating proximal gradient algorithm are proposed, and associated theoretical performance guarantees are established based on newly derived concentration-of-measure results. A Bayesian model is then introduced, to improve flexibility and generality. Connections between the maximum likelihood methods and the Bayesian model are developed, and example results are presented for a real compressive X-ray imaging system. Liming Wang 0004, Jiaji Huang, Xin Yuan 0002, Kalyani Krishnamurthy, Joel A. Greenberg, Volkan Cevher, Miguel R. D. Rodrigues, David J. Brady, A. Robert Calderbank, Lawrence Carin |
SIAM J. Imaging Sci. | 10 |
| 2015 | Multivariate time-series analysis and diffusion maps
Wenzhao Lian, Ronen Talmon, Hitten Zaveri, Lawrence Carin, Ronald R. Coifman |
Signal Process. | 4 |
| 2015 | Compressive Sensing by Learning a Gaussian Mixture Model From MeasurementsabstractCompressive sensing of signals drawn from a Gaussian mixture model (GMM) admits closed-form minimum mean squared error reconstruction from incomplete linear measurements. An accurate GMM signal model is usually not available a priori, because it is difficult to obtain training signals that match the statistics of the signals being sensed. We propose to solve that problem by learning the signal model in situ, based directly on the compressive measurements of the signals, without resorting to other signals to train a model. A key feature of our method is that the signals being sensed are treated as random variables and are integrated out in the likelihood. We derive a maximum marginal likelihood estimator (MMLE) that maximizes the likelihood of the GMM of the underlying signals given only their linear compressive measurements. We extend the MMLE to a GMM with dominantly low-rank covariance matrices, to gain computational speedup. We report extensive experimental results on image inpainting, compressive sensing of high-speed video, and compressive hyperspectral imaging (the latter two based on real compressive cameras). The results demonstrate that the proposed methods outperform state-of-the-art methods by significant margins. Jianbo Yang, Xuejun Liao, Xin Yuan 0002, Patrick Llull, David J. Brady, Guillermo Sapiro, Lawrence Carin |
IEEE Trans. Image Process. | 7 |
| 2014 | Latent Gaussian Models for Topic ModelingabstractA new approach is proposed for topic modeling, in which the latent matrix factorization employs Gaussian priors, rather than the Dirichlet-class priors widely used in such models. The use of a latent-Gaussian model permits simple and efficient approximate Bayesian posterior inference, via the Laplace approximation. On multiple datasets, the proposed approach is demonstrated to yield results as accurate as state-of-the-art approaches based on Dirichlet constructions, at a small fraction of the computation. The framework is general enough to jointly model text and binary data, here demonstrated to produce accurate and fast results for joint analysis of voting rolls and the associated legislative text. Further, it is demonstrated how the technique may be scaled up to massive data, with encouraging performance relative to alternative methods. Changwei Hu, Eunsu Ryu, David E. Carlson, Yingjian Wang 0004, Lawrence Carin |
AISTATS | 5 |
| 2014 | Low-Cost Compressive Sensing for Color Video and DepthabstractA simple and inexpensive (low-power and low-bandwidth) modification is made to a conventional off-the-shelf color video camera, from which we recover multiple color frames for each of the original measured frames, and each of the recovered frames can be focused at a different depth. The recovery of multiple frames for each measured frame is made possible via high-speed coding, manifested via translation of a single coded aperture, the inexpensive translation is constituted by mounting the binary code on a piezoelectric device. To simultaneously recover depth information, a liquid lens is modulated at high speed, via a variable voltage. Consequently, during the aforementioned coding process, the liquid lens allows the camera to sweep the focus through multiple depths. In addition to designing and implementing the camera, fast recovery is achieved by an anytime algorithm exploiting the group-sparsity of wavelet/DCT coefficients. Xin Yuan 0002, Patrick Llull, Xuejun Liao, Jianbo Yang, David J. Brady, Guillermo Sapiro, Lawrence Carin |
CVPR | 7 |
| 2014 | Multi-shot Imaging: Joint Alignment, Deblurring, and Resolution-EnhancementabstractThe capture of multiple images is a simple way to increase the chance of capturing a good photo with a light-weight hand-held camera, for which the camera-shake blur is typically a nuisance problem. The naive approach of selecting the single best captured photo as output does not take full advantage of all the observations. Conventional multi-image blind deblurring methods can take all observations as input but usually require the multiple images are well aligned. However, the multiple blurry images captured in the presence of camera shake are rarely free from mis-alignment. Registering multiple blurry images is a challenging task due to the presence of blur while deblurring of multiple blurry images requires accurate alignment, leading to an intrinsically coupled problem. In this paper, we propose a blind multi-image restoration method which can achieve joint alignment, non-uniform deblurring, together with resolution enhancement from multiple low-quality images. Experiments on several real-world images with comparison to some previous methods validate the effectiveness of the proposed method. Lawrence Carin |
CVPR | 2 |
| 2014 | Modeling Correlated Arrival Events with Latent Semi-Markov ProcessesabstractThe analysis and characterization of correlated point process data has wide applications, ranging from biomedical research to network analysis. In this work, we model such data as generated by a latent collection of continuous-time binary semi-Markov processes, corresponding to external events appearing and disappearing. A continuous-time modeling framework is more appropriate for multichannel point process data than a binning approach requiring time discretization, and we show connections between our model and recent ideas from the discrete-time literature. We describe an efficient MCMC algorithm for posterior inference, and apply our ideas to both synthetic data and a real-world biometrics application. Wenzhao Lian, Vinayak A. Rao, Brian Eriksson, Lawrence Carin |
ICML | 4 |
| 2014 | Scalable Bayesian Low-Rank Decomposition of Incomplete Multiway TensorsabstractWe present a scalable Bayesian framework for low-rank decomposition of multiway tensor data with missing observations. The key issue of pre-specifying the rank of the decomposition is sidestepped in a principled manner using a multiplicative gamma process prior. Both continuous and binary data can be analyzed under the framework, in a coherent way using fully conjugate Bayesian analysis. In particular, the analysis in the non-conjugate binary case is facilitated via the use of the Pólya-Gamma sampling strategy which elicits closed-form Gibbs sampling updates. The resulting samplers are efficient and enable us to apply our framework to large-scale problems, with time-complexity that is linear in the number of observed entries in the tensor. This is especially attractive in analyzing very large but sparsely observed tensors with very few known entries. Moreover, our method admits easy extension to the supervised setting where entities in one or more tensor modes have labels. Our method outperforms several state-of-the-art tensor decomposition methods on various synthetic and benchmark real-world datasets. Piyush Rai, Yingjian Wang 0004, Shengbo Guo, Gary Chen, David B. Dunson, Lawrence Carin |
ICML | 6 |
| 2014 | Nonlinear Information-Theoretic Compressive Measurement DesignabstractWe investigate design of general nonlinear functions for mapping high-dimensional data into a lower-dimensional (compressive) space. The nonlinear measurements are assumed contaminated by additive Gaussian noise. Depending on the application, we are either interested in recovering the high-dimensional data from the nonlinear compressive measurements, or performing classification directly based on these measurements. The latter case corresponds to classification based on nonlinearly constituted and noisy features. The nonlinear measurement functions are designed based on constrained mutual-information optimization. New analytic results are developed for the gradient of mutual information in this setting, for arbitrary input-signal statistics. We make connections to kernel-based methods, such as the support vector machine. Encouraging results are presented on multiple datasets, for both signal recovery and classification. The nonlinear approach is shown to be particularly valuable in high-noise scenarios. Liming Wang 0004, Abolfazl Razi, Miguel R. D. Rodrigues, A. Robert Calderbank, Lawrence Carin |
ICML | 5 |
| 2014 | On the relations of LFPs & Neural Spike Trains
David E. Carlson, Jana Schaich Borg, Kafui Dzirasa, Lawrence Carin |
NIPS | 4 |
| 2014 | Dynamic Rank Factor Model for Text Streams
Shaobo Han, Esther Salazar, Lawrence Carin |
NIPS | 4 |
| 2014 | Bayesian Nonlinear Support Vector Machines and Discriminative Factor Modeling
Ricardo Henao, Xin Yuan 0002, Lawrence Carin |
NIPS | 3 |
| 2014 | Analysis of Brain States from Multi-Region LFP Time-Series
Kyle R. Ulrich, David E. Carlson, Wenzhao Lian, Jana Schaich Borg, Kafui Dzirasa, Lawrence Carin |
NIPS | 6 |
| 2014 | Compressive Sensing of Signals from a GMM with Sparse Precision Matrices
Jianbo Yang, Xuejun Liao, Minhua Chen, Lawrence Carin |
NIPS | 4 |
| 2014 | Bayesian joint analysis of heterogeneous genomics dataabstractSUMMARY: A non-parametric Bayesian factor model is proposed for joint analysis of multi-platform genomics data. The approach is based on factorizing the latent space (feature space) into a shared component and a data-specific component with the dimensionality of these components (spaces) inferred via a beta-Bernoulli process. The proposed approach is demonstrated by jointly analyzing gene expression/copy number variations and gene expression/methylation data for ovarian cancer patients, showing that the proposed model can potentially uncover key drivers related to cancer. AVAILABILITY AND IMPLEMENTATION: The source code for this model is written in MATLAB and has been made publicly available at https://sites.google.com/site/jointgenomics/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Priyadip Ray, Lingling Zheng, Joseph E. Lucas, Lawrence Carin |
Bioinform. | 4 |
| 2014 | Generalized Alternating Projection for Weighted-퓁2, 1 Minimization with Applications to Model-Based Compressive SensingabstractWe consider the group basis pursuit problem, which extends basis pursuit by replacing the $\ell_{1}$ norm with a weighted-$\ell_{2,1}$ norm. We provide an anytime algorithm, called generalized alternating projection (GAP), to solve this problem. The GAP algorithm extends classical alternating projection to the case in which projections are performed between convex sets that undergo a systematic sequence of changes. We prove that, under a set of group-restricted isometry property (group-RIP) conditions, the reconstruction error of GAP monotonically converges to zero. Thus the algorithm can be interrupted at any time to return a valid solution and can be resumed subsequently to improve the solution. This anytime convergence property saves iterations on retracting and correcting mistakes, which, along with an effective acceleration scheme, makes GAP converge fast. Moreover, the per-iteration computation is inexpensive, consisting of sorting of a linear array followed by groupwise thresholding and linear transform of vectors for which fast algorithms often exist. We evaluate the algorithmic performance through extensive experiments in which GAP is compared to other state-of-the-art algorithms and applied to compressive sensing of natural images and video. Xuejun Liao, Hui Li 0068, Lawrence Carin |
SIAM J. Imaging Sci. | 3 |
| 2014 | Video Compressive Sensing Using Gaussian Mixture ModelsabstractA Gaussian mixture model (GMM)-based algorithm is proposed for video reconstruction from temporally compressed video measurements. The GMM is used to model spatio-temporal video patches, and the reconstruction can be efficiently computed based on analytic expressions. The GMM-based inversion method benefits from online adaptive learning and parallel computation. We demonstrate the efficacy of the proposed inversion method with videos reconstructed from simulated compressive video measurements, and from a real compressive video camera. We also use the GMM as a tool to investigate adaptive video compressive sensing, i.e., adaptive rate of temporal compression. Jianbo Yang, Xin Yuan 0002, Xuejun Liao, Patrick Llull, David J. Brady, Guillermo Sapiro, Lawrence Carin |
IEEE Trans. Image Process. | 7 |
| 2014 | A Bregman Matrix and the Gradient of Mutual Information for Vector Poisson and Gaussian ChannelsabstractA generalization of Bregman divergence is developed and utilized to unify vector Poisson and Gaussian channel models, from the perspective of the gradient of mutual information. The gradient is with respect to the measurement matrix in a compressive-sensing setting, and mutual information is considered for signal recovery and classification. Existing gradient-of-mutual-information results for scalar Poisson models are recovered as special cases, as are known results for the vector Gaussian model. The Bregman-divergence generalization yields a Bregman matrix, and this matrix induces numerous matrix-valued metrics. The metrics associated with the Bregman matrix are detailed, as are its other properties. The Bregman matrix is also utilized to connect the relative entropy and mismatched minimum mean squared error. Two applications are considered: 1) compressive sensing with a Poisson measurement model and 2) compressive topic modeling for analysis of a document corpora (word-count data). In both of these settings, we use the developed theory to optimize the compressive measurement matrix, for signal recovery and classification. Liming Wang 0004, David E. Carlson, Miguel R. D. Rodrigues, A. Robert Calderbank, Lawrence Carin |
IEEE Trans. Inf. Theory | 5 |
| 2013 | Patient Clustering with Uncoded Text in Electronic Medical Records
Ricardo Henao, Jared Murray, Geoffrey S. Ginsburg, Lawrence Carin, Joseph E. Lucas |
AMIA | 4 |
| 2013 | Test-size Reduction for Concept Estimation
Divyanshu Vats, Christoph Studer, Andrew S. Lan, Lawrence Carin, Richard G. Baraniuk |
EDM | 4 |
| 2013 | Compressive sensing for incoherent imaging systems with optical constraintsabstractWe consider the problem of linear projection design for incoherent optical imaging systems. We propose a computationally efficient method to obtain effective measurement kernels that satisfy the physical constraints imposed by an optical system, starting first from arbitrary kernels, including those that satisfy a less demanding power constraint. Performance is measured in terms of mutual information between the source input and the projection measurement, as well as reconstruction error for real world images. A clear improvement in the quality of image reconstructions is shown with respect to both random and adaptive projection designs in the literature. Francesco Renna, Miguel R. D. Rodrigues, Minhua Chen, A. Robert Calderbank, Lawrence Carin |
ICASSP | 5 |
| 2013 | Gaussian mixture model for video compressive sensingabstractA Gaussian Mixture Model (GMM)-based algorithm is proposed for video reconstruction from temporal compressed measurements. The GMM is used to model spatio-temporal video patches, and the reconstruction can be efficiently computed based on analytic expressions. The developed GMM reconstruction method benefits from online adaptive learning and parallel computation. We demonstrate the efficacy of the proposed GMM with videos reconstructed from simulated compressive video measurements and from a real compressive video camera. Jianbo Yang, Xin Yuan 0002, Xuejun Liao, Patrick Llull, Guillermo Sapiro, David J. Brady, Lawrence Carin |
ICIP | 7 |
| 2013 | Adaptive temporal compressive sensing for videoabstractThis paper introduces the concept of adaptive temporal compressive sensing (CS) for video. We propose a CS algorithm to adapt the compression ratio based on the scene's temporal complexity, computed from the compressed data, without compromising the quality of the reconstructed video. The temporal adaptivity is manifested by manipulating the integration time of the camera, opening the possibility to realtime implementation. The proposed algorithm is a generalized temporal CS approach that can be incorporated with a diverse set of existing hardware systems. Xin Yuan 0002, Jianbo Yang, Patrick Llull, Xuejun Liao, Guillermo Sapiro, David J. Brady, Lawrence Carin |
ICIP | 7 |
| 2013 | Exploring the Mind: Integrating Questionnaires and fMRIabstractA new model is developed for joint analysis of ordered, categorical, real and count data. The ordered and categorical data are answers to questionnaires, the (word) count data correspond to the text questions from the questionnaires, and the real data correspond to fMRI responses for each subject. The Bayesian model employs the von Mises distribution in a novel manner to infer sparse graphical models jointly across people, questions, fMRI stimuli and brain region, with this integrated within a new matrix factorization based on latent binary features. The model is compared with simpler alternatives on two real datasets. We also demonstrate the ability to predict the response of the brain to visual stimuli (as measured by fMRI), based on knowledge of how the associated person answered classical questionnaires. Esther Salazar, Ryan Bogdan, Adam Gorka, Ahmad Hariri, Lawrence Carin |
ICML (2) | 5 |
| 2013 | Online Expectation Maximization for Reinforcement Learning in POMDPs
Miao Liu 0001, Xuejun Liao, Lawrence Carin |
IJCAI | 3 |
| 2013 | Generalized Bregman divergence and gradient of mutual information for vector Poisson channelsabstractWe investigate connections between information-theoretic and estimation-theoretic quantities in vector Poisson channel models. In particular, we generalize the gradient of mutual information with respect to key system parameters from the scalar to the vector Poisson channel model. We also propose, as another contribution, a generalization of the classical Bregman divergence that offers a means to encapsulate under a unifying framework the gradient of mutual information results for scalar and vector Poisson and Gaussian channel models. The so-called generalized Bregman divergence is also shown to exhibit various properties akin to the properties of the classical version. The vector Poisson channel model is drawing considerable attention in view of its application in various domains: as an example, the availability of the gradient of mutual information can be used in conjunction with gradient descent methods to effect compressive-sensing projection designs in emerging X-ray and document classification applications. Liming Wang 0004, Miguel R. D. Rodrigues, Lawrence Carin |
ISIT | 3 |
| 2013 | Dynamic Clustering via Asymptotics of the Dependent Dirichlet Process MixtureabstractThis paper presents a novel algorithm, based upon the dependent Dirichlet process mixture model (DDPMM), for clustering batch-sequential data containing an unknown number of evolving clusters. The algorithm is derived via a low-variance asymptotic analysis of the Gibbs sampling algorithm for the DDPMM, and provides a hard clustering with convergence guarantees similar to those of the k-means algorithm. Empirical results from a synthetic test with moving Gaussian clusters and a test with real ADS-B aircraft trajectory data demonstrate that the algorithm requires orders of magnitude less computational time than contemporary probabilistic and hard clustering algorithms, while providing higher accuracy on the examined datasets. Trevor Campbell, Miao Liu 0001, Brian Kulis, Jonathan P. How, Lawrence Carin |
NIPS | 5 |
| 2013 | Real-Time Inference for a Gamma Process Model of Neural SpikingabstractWith simultaneous measurements from ever increasing populations of neurons, there is a growing need for sophisticated tools to recover signals from individual neurons. In electrophysiology experiments, this classically proceeds in a two-step process: (i) threshold the waveforms to detect putative spikes and (ii) cluster the waveforms into single units (neurons). We extend previous Bayesian nonparamet- ric models of neural spiking to jointly detect and cluster neurons using a Gamma process model. Importantly, we develop an online approximate inference scheme enabling real-time analysis, with performance exceeding the previous state-of-the- art. Via exploratory data analysis—using data with partial ground truth as well as two novel data sets—we find several features of our model collectively contribute to our improved performance including: (i) accounting for colored noise, (ii) de- tecting overlapping spikes, (iii) tracking waveform dynamics, and (iv) using mul- tiple channels. We hope to enable novel experiments simultaneously measuring many thousands of neurons and possibly adapting stimuli dynamically to probe ever deeper into the mysteries of the brain. David E. Carlson, Vinayak A. Rao, Joshua T. Vogelstein, Lawrence Carin |
NIPS | 4 |
| 2013 | Integrated Non-Factorized Variational InferenceabstractWe present a non-factorized variational method for full posterior inference in Bayesian hierarchical models, with the goal of capturing the posterior variable dependencies via efficient and possibly parallel computation. Our approach unifies the integrated nested Laplace approximation (INLA) under the variational framework. The proposed method is applicable in more challenging scenarios than typically assumed by INLA, such as Bayesian Lasso, which is characterized by the non-differentiability of the $\ell_{1}$ norm arising from independent Laplace priors. We derive an upper bound for the Kullback-Leibler divergence, which yields a fast closed-form solution via decoupled optimization. Our method is a reliable analytic alternative to Markov chain Monte Carlo (MCMC), and it results in a tighter evidence lower bound than that of mean-field variational Bayes (VB) method. Shaobo Han, Xuejun Liao, Lawrence Carin |
NIPS | 3 |
| 2013 | Designed Measurements for Vector Count DataabstractWe consider design of linear projection measurements for a vector Poisson signal model. The projections are performed on the vector Poisson rate, $X\in\mathbb{R}_+^n$, and the observed data are a vector of counts, $Y\in\mathbb{Z}_+^m$. The projection matrix is designed by maximizing mutual information between $Y$ and $X$, $I(Y;X)$. When there is a latent class label $C\in\{1,\dots,L\}$ associated with $X$, we consider the mutual information with respect to $Y$ and $C$, $I(Y;C)$. New analytic expressions for the gradient of $I(Y;X)$ and $I(Y;C)$ are presented, with gradient performed with respect to the measurement matrix. Connections are made to the more widely studied Gaussian measurement model. Example results are presented for compressive topic modeling of a document corpora (word counting), and hyperspectral compressive sensing for chemical classification (photon counting). Liming Wang 0004, David E. Carlson, Miguel R. D. Rodrigues, David Wilcox, A. Robert Calderbank, Lawrence Carin |
NIPS | 6 |
| 2013 | Deep Learning with Hierarchical Convolutional Factor AnalysisabstractUnsupervised multilayered (“deep”) models are considered for imagery. The model is represented using a hierarchical convolutional factor-analysis construction, with sparse factor loadings and scores. The computation of layer-dependent model parameters is implemented within a Bayesian setting, employing a Gibbs sampler and variational Bayesian (VB) analysis that explicitly exploit the convolutional nature of the expansion. To address large-scale and streaming data, an online version of VB is also developed. The number of dictionary elements at each layer is inferred from the data, based on a beta-Bernoulli implementation of the Indian buffet process. Example results are presented for several image-processing applications, with comparisons to related models in the literature. Bo Chen 0001, Gungor Polatkan, Guillermo Sapiro, David M. Blei, David B. Dunson, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2013 | Coded Hyperspectral Imaging and Blind Compressive SensingabstractBlind compressive sensing (CS) is considered for reconstruction of hyperspectral data imaged by a coded aperture camera. The measurements are manifested as a superposition of the coded wavelength-dependent data, with the ambient three-dimensional hyperspectral datacube mapped to a two-dimensional measurement. The hyperspectral datacube is recovered using a Bayesian implementation of blind CS. Several demonstration experiments are presented, including measurements performed using a coded aperture snapshot spectral imager (CASSI) camera. The proposed approach is capable of efficiently reconstructing large hyperspectral datacubes. Comparisons are made between the proposed algorithm and other techniques employed in compressive sensing, dictionary learning, and matrix factorization. David S. Kittle, Tsung-Han Tsai 0002, David J. Brady, Lawrence Carin |
SIAM J. Imaging Sci. | 5 |
| 2012 | How to focus the discriminative power of a dictionaryabstractThis paper is motivated by the challenge of high fidelity processing of images using a relatively small set of projection measurements. This is a problem of great interest in many sensing applications, for example where high photodetector counts are precluded by a combination of available power, form factor and expense. The emerging methods of dictionary learning and compressive sensing offer great potential for addressing this challenge. Combining these methods requires that the signals of interest be representable as a sparse combination of elements of some dictionary. This paper develops a method that aligns the discriminative power of such a dictionary with the physical limitations of the imaging system. Alignment is accomplished by designing a projection matrix that exposes and then aligns the modes of the noise with those of the dictionary. The design algorithm is obtained by modifying an algorithm for designing the pre-filter to maximize the rate and reliability of a Multiple Input Multiple Output (MIMO) communications channel. The difference is that in the communications problem a source is being matched to a channel, whereas in the imaging problem a channel, or equivalently the noise covariance, is being matched to a source. Our results shown that using the proposed communications design framework we can reduce reconstruction error between 20%, after only 20 projections of a 28 × 28 image, and 10% after 100 projections. Furthermore, we noticeably see the superior quality of the reconstructed images. William R. Carson, Miguel R. D. Rodrigues, Minhua Chen, Lawrence Carin, A. Robert Calderbank |
ICASSP | 4 |
| 2012 | Adapted statistical compressive sensing: Learning to sense gaussian mixture modelsabstractA framework for learning sensing kernels adapted to signals that follow a Gaussian mixture model (GMM) is introduced in this paper. This follows the paradigm of statistical compressive sensing (SCS), where a statistical model, a GMM in particular, replaces the standard sparsity model of classical compressive sensing (CS), leading to both theoretical and practical improvements. We show that the optimized sensing matrix outperforms random sampling matrices originally exploited both in CS and SCS. Julio Martin Duarte-Carvajalino, Guoshen Yu, Lawrence Carin, Guillermo Sapiro |
ICASSP | 3 |
| 2012 | Online Bayesian dictionary learning for large datasetsabstractThe problem of learning a data-adaptive dictionary for a very large collection of signals is addressed. This paper proposes a statistical model and associated variational Bayesian (VB) inference for simultaneously learning the dictionary and performing sparse coding of the signals. The model builds upon beta process factor analysis (BPFA), with the number of factors automatically inferred, and posterior distributions are estimated for both the dictionary and the signals. Crucially, an online learning procedure is employed, allowing scalability to very large datasets which would be beyond the capabilities of existing batch methods. State-of-the-art performance is demonstrated by experiments with large natural images containing tens of millions of pixels. Lingbo Li 0002, Jorge G. Silva, Mingyuan Zhou, Lawrence Carin |
ICASSP | 4 |
| 2012 | Communications Inspired Linear Discriminant Analysis
Minhua Chen, William R. Carson, Miguel R. D. Rodrigues, Lawrence Carin, A. Robert Calderbank |
ICML | 4 |
| 2012 | Cross-Domain Multitask Learning with Latent Probit Models
Shaobo Han, Xuejun Liao, Lawrence Carin |
ICML | 3 |
| 2012 | Inferring Latent Structure From Mixed Real and Categorical Relational Data
Esther Salazar, Lawrence Carin |
ICML | 2 |
| 2012 | Levy Measure Decompositions for the Beta and Gamma Processes
Yingjian Wang 0004, Lawrence Carin |
ICML | 2 |
| 2012 | Lognormal and Gamma Mixed Negative Binomial Regression
Mingyuan Zhou, Lingbo Li 0002, David B. Dunson, Lawrence Carin |
ICML | 4 |
| 2012 | The contextual focused topic modelabstractA nonparametric Bayesian contextual focused topic model (cFTM) is proposed. The cFTM infers a sparse ("focused") set of topics for each document, while also leveraging contextual information about the author(s) and document venue. The hierarchical beta process, coupled with a Bernoulli process, is employed to infer the focused set of topics associated with each author and venue; the same construction is also employed to infer those topics associated with a given document that are unusual (termed "random effects"), relative to topics that are inferred as probable for the associated author(s) and venue. To leverage statistical strength and infer latent interrelationships between authors and venues, the Dirichlet process is utilized to cluster authors and venues. The cFTM automatically infers the number of topics needed to represent the corpus, the number of author and venue clusters, and the probabilistic importance of the author, venue and random-effect information on word assignment for a given document. Efficient MCMC inference is presented. Example results and interpretations are presented for two real datasets, demonstrating promising performance, with comparison to other state-of-the-art methods. Xu Chen 0017, Mingyuan Zhou, Lawrence Carin |
KDD | 3 |
| 2012 | Active learning for online bayesian matrix factorizationabstractThe problem of large-scale online matrix completion is addressed via a Bayesian approach. The proposed method learns a factor analysis (FA) model for large matrices, based on a small number of observed matrix elements, and leverages the statistical model to actively select which new matrix entries/observations would be most informative if they could be acquired, to improve the model; the model inference and active learning are performed in an online setting. In the context of online learning, a greedy, fast and provably near-optimal algorithm is employed to sequentially maximize the mutual information between past and future observations, taking advantage of submodularity properties. Additionally, a simpler procedure, which directly uses the posterior parameters learned by the Bayesian approach, is shown to achieve slightly lower estimation quality, with far less computational effort. Inference is performed using a computationally efficient online variational Bayes (VB) procedure. Competitive results are obtained in a very large collaborative filtering problem, namely the Yahoo! Music ratings dataset. Jorge G. Silva, Lawrence Carin |
KDD | 2 |
| 2012 | Joint Modeling of a Matrix with Associated Text via Latent Binary FeaturesabstractA new methodology is developed for joint analysis of a matrix and accompanying documents, with the documents associated with the matrix rows/columns. The documents are modeled with a focused topic model, inferring latent binary features (topics) for each document. A new matrix decomposition is developed, with latent binary features associated with the rows/columns, and with imposition of a low-rank constraint. The matrix decomposition and topic model are coupled by sharing the latent binary feature vectors associated with each. The model is applied to roll-call data, with the associated documents defined by the legislation. State-of-the-art results are manifested for prediction of votes on a new piece of legislation, based only on the observed text legislation. The coupling of the text and legislation is also demonstrated to yield insight into the properties of the matrix decomposition for roll-call data. XianXing Zhang, Lawrence Carin |
NIPS | 2 |
| 2012 | Augment-and-Conquer Negative Binomial ProcessesabstractBy developing data augmentation methods unique to the negative binomial (NB) distribution, we unite seemingly disjoint count and mixture models under the NB process framework. We develop fundamental properties of the models and derive efficient Gibbs sampling inference. We show that the gamma-NB process can be reduced to the hierarchical Dirichlet process with normalization, highlighting its unique theoretical, structural and computational advantages. A variety of NB processes with distinct sharing mechanisms are constructed and applied to topic modeling, with connections to existing algorithms, showing the importance of inferring both the NB dispersion and probability parameters. Mingyuan Zhou, Lawrence Carin |
NIPS | 2 |
| 2012 | Nested Dictionary Learning for Hierarchical Organization of Imagery and Text
Lingbo Li 0002, XianXing Zhang, Mingyuan Zhou, Lawrence Carin |
UAI | 4 |
| 2012 | Communications-Inspired Projection Design with Application to Compressive SensingabstractWe consider the recovery of an underlying signal $\mathbf{x}\in\mathbb{C}^m$ based on projection measurements of the form $\mathbf{y}=\mathbf{M}\mathbf{x}+\mathbf{w}$, where $\mathbf{y}\in\mathbb{C}^\ell$ and $\mathbf{w}$ is measurement noise; we are interested in the case $\ell\ll m$. It is assumed that the signal model $p(\mathbf{x})$ is known and that $\mathbf{w}\sim\mathcal{CN}(\mathbf{w};\boldsymbol{0},\bf \Sigma_w)$ for known $\bf \Sigma_w$. The objective is to design a projection matrix $\mathbf{M}\in\mathbb{C}^{\ell\times m}$ to maximize key information-theoretic quantities with operational significance, including the mutual information between the signal and the projections $\mathcal{I}(\mathbf{x};\mathbf{y})$ or the Rényi entropy of the projections $\mbox{h}_\alpha \left( \mathbf{y} \right)$ (Shannon entropy is a special case). By capitalizing on explicit characterizations of the gradients of the information measures with respect to the projection matrix, where we also partially extend the well-known results of Palomar and Verdú from the mutual information to the Rényi entropy domain, we reveal the key operations carried out by the optimal projection designs: mode exposure and mode alignment. Experiments are considered for the case of compressive sensing (CS) applied to imagery. In this context, we provide a demonstration of the performance improvement possible through the application of the novel projection designs in relation to conventional ones, as well as justification for a fast online projection design method with which state-of-the-art adaptive CS signal recovery is achieved. William R. Carson, Minhua Chen, Miguel R. D. Rodrigues, A. Robert Calderbank, Lawrence Carin |
SIAM J. Imaging Sci. | 5 |
| 2012 | Dictionary Learning for Noisy and Incomplete Hyperspectral ImagesabstractWe consider analysis of noisy and incomplete hyperspectral imagery, with the objective of removing the noise and inferring the missing data. The noise statistics may be wavelength dependent, and the fraction of data missing (at random) may be substantial, including potentially entire bands, offering the potential to significantly reduce the quantity of data that need be measured. To achieve this objective, the imagery is divided into contiguous three-dimensional (3D) spatio-spectral blocks of spatial dimension much less than the image dimension. It is assumed that each such 3D block may be represented as a linear combination of dictionary elements of the same dimension, plus noise, and the dictionary elements are learned in situ based on the observed data (no a priori training). The number of dictionary elements needed for representation of any particular block is typically small relative to the block dimensions, and all the image blocks are processed jointly (“collaboratively") to infer the underlying dictionary. We address dictionary learning from a Bayesian perspective, considering two distinct means of imposing sparse dictionary usage. These models allow inference of the number of dictionary elements needed as well as the underlying wavelength-dependent noise statistics. It is demonstrated that drawing the dictionary elements from a Gaussian process prior, imposing structure on the wavelength dependence of the dictionary elements, yields significant advantages, relative to the more conventional approach of using an independent and identically distributed Gaussian prior for the dictionary elements; this advantage is particularly evident in the presence of noise. The framework is demonstrated by processing hyperspectral imagery with a significant number of voxels missing uniformly at random, with imagery at specific wavelengths missing entirely, and in the presence of substantial additive noise. Zhengming Xing, Mingyuan Zhou, Alexey Castrodad, Guillermo Sapiro, Lawrence Carin |
SIAM J. Imaging Sci. | 5 |
| 2012 | Nonparametric Bayesian Dictionary Learning for Analysis of Noisy and Incomplete ImagesabstractNonparametric Bayesian methods are considered for recovery of imagery based upon compressive, incomplete, and/or noisy measurements. A truncated beta-Bernoulli process is employed to infer an appropriate dictionary for the data under test and also for image recovery. In the context of compressive sensing, significant improvements in image recovery are manifested using learned dictionaries, relative to using standard orthonormal image expansions. The compressive-measurement projections are also optimized for the learned dictionary. Additionally, we consider simpler (incomplete) measurements, defined by measuring a subset of image pixels, uniformly selected at random. Spatial interrelationships within imagery are exploited through use of the Dirichlet and probit stick-breaking processes. Several example results are presented, with comparisons to other methods in the literature. Mingyuan Zhou, Haojun Chen, John W. Paisley, Lingbo Li 0002, Zhengming Xing, David B. Dunson, Guillermo Sapiro, Lawrence Carin |
IEEE Trans. Image Process. | 9 |
| 2011 | Non-parametric Bayesian modeling and fusion of spatio-temporal information sources
Priyadip Ray, Lawrence Carin |
FUSION | 2 |
| 2011 | Bayesian topic models for describing computer network behaviorsabstractWe consider the use of Bayesian topic models in the analysis of computer network traffic. Our approach utilizes latent Dirichlet allocation and time-varying dynamic latent Dirichlet allocation, with the goal of identifying significant co-occurrences of types of network traffic, these forming topics of user behavior. In our experiments, these topics of user behavior included: (i) web traffic, (ii) email client and instant messaging, (iii) Microsoft file access, (iv) email server, and (v) other miscellaneous traffic. Each identified behavior topic included a variety of different, but related, protocols without using any a priori knowledge of the purpose of the protocol. We believe that the techniques presented in this paper can be used to form more complex topics through the use of deep packet inspection, and that such topic models could prove useful in the identification of zero-day exploits or other network threats. Christopher Cramer, Lawrence Carin |
ICASSP | 2 |
| 2011 | Nonparametric Bayesian feature selection for multi-task learningabstractWe present a nonparametric Bayesian model for multi-task learning, with a focus on feature selection in binary classification. The model jointly identifies groups of similar tasks and selects the subset of features relevant to the tasks within each group. The model employs a Dirchlet process with a beta Bernoulli hierarchical base measure. The posterior inference is accomplished efficiently using a Gibbs sampler. Experimental results are presented on simulated as well as real data. Hui Li 0068, Xuejun Liao, Lawrence Carin |
ICASSP | 3 |
| 2011 | Joint dictionary learning and topic modeling for image clusteringabstractA new Bayesian model is proposed, integrating dictionary learning and topic modeling into a unified framework. The model is applied to cluster multiple images, and a subset of the images may be annotated. Example results are presented on the MNIST digit data and on the Microsoft MSRC multi-scene image data. These results reveal the working mechanisms of the model and demonstrate state-of-the-art performance. Lingbo Li 0002, Mingyuan Zhou, Lawrence Carin |
ICASSP | 4 |
| 2011 | Time-evolving modeling of social networksabstractA statistical framework for modeling and prediction of binary matrices is presented. The method is applied to social network analysis, specifically the database of US Supreme Court rulings. It is shown that the ruling behavior of Supreme Court judges can be accurately modeled by using a small number of latent features whose values evolve with time. The learned model facilitates the discovery of inter-relationships between judges and of the gradual evolution of their stances over time. In addition, the analysis in this paper extends previous results by considering automatic estimation of the number of latent features and other model parameters, based on a nonparametric-Bayesian approach. Inference is efficiently performed using Gibbs sampling. Jorge G. Silva, Rebecca Willett, Lawrence Carin |
ICASSP | 4 |
| 2011 | Covariate-dependent dictionary learning and sparse codingabstractA dependent hierarchical beta process (dHBP) is developed as a prior for data that may be represented in terms of a sparse set of latent features (dictionary elements), with covariate dependent feature usage. The dHBP is applicable to general covariates and data models, imposing that signals with similar covariates are likely to be manifested in terms of similar features. As an application, we consider the simultaneous sparse modeling of multiple images, with the covariate of a given image linked to its similarity to all other images (as applied in manifold learning). Efficient inference is performed using hybrid Gibbs, Metropolis-Hastings and slice sampling. Mingyuan Zhou, Hongxia Yang, Guillermo Sapiro, David B. Dunson, Lawrence Carin |
ICASSP | 5 |
| 2011 | Topic Modeling with Nonparametric Markov Tree
Haojun Chen, David B. Dunson, Lawrence Carin |
ICML | 3 |
| 2011 | The Hierarchical Beta Process for Convolutional Factor Analysis and Deep Learning
Bo Chen 0001, Gungor Polatkan, Guillermo Sapiro, David B. Dunson, Lawrence Carin |
ICML | 5 |
| 2011 | On the Integration of Topic Modeling and Dictionary Learning
Lingbo Li 0002, Mingyuan Zhou, Guillermo Sapiro, Lawrence Carin |
ICML | 4 |
| 2011 | The Infinite Regionalized Policy Representation
Miao Liu 0001, Xuejun Liao, Lawrence Carin |
ICML | 3 |
| 2011 | Variational Inference for Stick-Breaking Beta Process Priors
John W. Paisley, Lawrence Carin, David M. Blei |
ICML | 2 |
| 2011 | Tree-Structured Infinite Sparse Factor Model
XianXing Zhang, David B. Dunson, Lawrence Carin |
ICML | 3 |
| 2011 | On the Analysis of Multi-Channel Neural Spike DataabstractNonparametric Bayesian methods are developed for analysis of multi-channel spike-train data, with the feature learning and spike sorting performed jointly. The feature learning and sorting are performed simultaneously across all channels. Dictionary learning is implemented via the beta-Bernoulli process, with spike sorting performed via the dynamic hierarchical Dirichlet process (dHDP), with these two models coupled. The dHDP is augmented to eliminate refractoryperiod violations, it allows the “appearance” and “disappearance” of neurons over time, and it models smooth variation in the spike statistics. Bo Chen 0001, David E. Carlson, Lawrence Carin |
NIPS | 3 |
| 2011 | The Kernel Beta ProcessabstractA new Le ́vy process prior is proposed for an uncountable collection of covariate- dependent feature-learning measures; the model is called the kernel beta process (KBP). Available covariates are handled efficiently via the kernel construction, with covariates assumed observed with each data sample (“customer”), and latent covariates learned for each feature (“dish”). Each customer selects dishes from an infinite buffet, in a manner analogous to the beta process, with the added constraint that a customer first decides probabilistically whether to “consider” a dish, based on the distance in covariate space between the customer and dish. If a customer does consider a particular dish, that dish is then selected probabilistically as in the beta process. The beta process is recovered as a limiting case of the KBP. An efficient Gibbs sampler is developed for computations, and state-of-the-art results are presented for image processing and music analysis tasks. Yingjian Wang 0004, David B. Dunson, Lawrence Carin |
NIPS | 4 |
| 2011 | Hierarchical Topic Modeling for Analysis of Time-Evolving Personal ChoicesabstractThe nested Chinese restaurant process is extended to design a nonparametric topic-model tree for representation of human choices. Each tree branch corresponds to a type of person, and each node (topic) has a corresponding probability vector over items that may be selected. The observed data are assumed to have associated temporal covariates (corresponding to the time at which choices are made), and we wish to impose that with increasing time it is more probable that topics deeper in the tree are utilized. This structure is imposed by developing a new “change point" stick-breaking model that is coupled with a Poisson and product-of-gammas construction. To share topics across the tree nodes, topic distributions are drawn from a Dirichlet process. As a demonstration of this concept, we analyze real data on course selections of undergraduate students at Duke University, with the goal of uncovering and concisely representing structure in the curriculum and in the characteristics of the student body. XianXing Zhang, David B. Dunson, Lawrence Carin |
NIPS | 3 |
| 2011 | Logistic Stick-Breaking Process
Lan Du 0001, Lawrence Carin, David B. Dunson |
J. Mach. Learn. Res. | 3 |
| 2011 | Learning Discriminative Sparse Representations for Modeling, Source Separation, and Mapping of Hyperspectral ImageryabstractA method is presented for subpixel modeling, mapping, and classification in hyperspectral imagery using learned block-structured discriminative dictionaries, where each block is adapted and optimized to represent a material in a compact and sparse manner. The spectral pixels are modeled by linear combinations of subspaces defined by the learned dictionary atoms, allowing for linear mixture analysis. This model provides flexibility in source representation and selection, thus accounting for spectral variability, small-magnitude errors, and noise. A spatial-spectral coherence regularizer in the optimization allows pixel classification to be influenced by similar neighbors. We extend the proposed approach for cases for which there is no knowledge of the materials in the scene, unsupervised classification, and provide experiments and comparisons with simulated and real data. We also present results when the data have been significantly undersampled and then reconstructed, still retaining high-performance classification, showing the potential role of compressive sensing and sparse modeling techniques in efficient acquisition/transmission missions for hyperspectral imagery. Alexey Castrodad, Zhengming Xing, John B. Greer, Edward Bosch, Lawrence Carin, Guillermo Sapiro |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2011 | Bayesian Robust Principal Component AnalysisabstractA hierarchical Bayesian model is considered for decomposing a matrix into low-rank and sparse components, assuming the observed matrix is a superposition of the two. The matrix is assumed noisy, with unknown and possibly non-stationary noise statistics. The Bayesian framework infers an approximate representation for the noise statistics while simultaneously inferring the low-rank and sparse-outlier contributions; the model is robust to a broad range of noise levels, without having to change model hyperparameter settings. In addition, the Bayesian framework allows exploitation of additional structure in the matrix. For example, in video applications each row (or column) corresponds to a video frame, and we introduce a Markov dependency between consecutive rows in the matrix (corresponding to consecutive frames in the video). The properties of this Markov process are also inferred based on the observed matrix, while simultaneously denoising and recovering the low-rank and sparse components. We compare the Bayesian model to a state-of-the-art optimization-based implementation of robust PCA; considering several examples, we demonstrate competitive performance of the proposed model. Xinghao Ding, Lihan He, Lawrence Carin |
IEEE Trans. Image Process. | 3 |
| 2010 | Sparse linear regression with beta process priorsabstractA Bayesian approximation to finding the minimum ℓ0norm solution for an underdetermined linear system is proposed that is based on the beta process prior. The beta process linear regression (BP-LR) model finds sparse solutions to the underdetermined model y = Φx + ϵ, by modeling the vector x as an element-wise product of a non-sparse weight vector, w, and a sparse binary vector, z, that is drawn from the beta process prior. The hierarchical model is fully conjugate and therefore is amenable to fast inference methods. We demonstrate the model on a compressive sensing problem and on a correlated-feature problem, where we show the ability of the BP-LR to selectively remove the irrelevant features, while preserving the relevant groups of correlated features. Bo Chen 0001, John W. Paisley, Lawrence Carin |
ICASSP | 3 |
| 2010 | A nonparametric Bayesian model for kernel matrix completionabstractWe present a nonparametric Bayesian model for completing low-rank, positive semidefinite matrices. Given an N × N matrix with underlying rank r, and noisy measured values and missing values with a symmetric pattern, the proposed Bayesian hierarchical model nonparametrically uncovers the underlying rank from all positive semidefinite matrices, and completes the matrix by approximating the missing values. We analytically derive all posterior distributions for the fully conjugate model hierarchy and discuss variational Bayes and MCMC Gibbs sampling for inference, as well as an efficient measurement selection procedure. We present results on a toy problem, and a music recommendation problem, where we complete the kernel matrix of 2,250 pieces of music. John W. Paisley, Lawrence Carin |
ICASSP | 2 |
| 2010 | Discriminative sparse representations in hyperspectral imageryabstractRecent advances in sparse modeling and dictionary learning for discriminative applications show high potential for numerous classification tasks. In this paper, we show that highly accurate material classification from hyperspectral imagery (HSI) can be obtained with these models, even when the data is reconstructed from a very small percentage of the original image samples. The proposed supervised HSI classification is performed using a measure that accounts for both reconstruction errors and sparsity levels for sparse representations based on class-dependent learned dictionaries. Combining the dictionaries learned for the different materials, a linear mixing model is derived for sub-pixel classification. Results with real hyperspectral data cubes are shown both for urban and non-urban terrain. Alexey Castrodad, Zhengming Xing, John B. Greer, Edward Bosch, Lawrence Carin, Guillermo Sapiro |
ICIP | 5 |
| 2010 | Nonparametric image interpolation and dictionary learning using spatially-dependent Dirichlet and beta process priorsabstractWe present a Bayesian model for image interpolation and dictionary learning that uses two nonparametric priors for sparse signal representations: the beta process and the Dirichlet process. Additionally, the model uses spatial information within the image to encourage sharing of information within image subregions. We derive a hybrid MAP/Gibbs sampler, which performs Gibbs sampling for the latent indicator variables and MAP estimation for all other parameters. We present experimental results, where we show an improvement over other state-of-the-art algorithms in the low-measurement regime. John W. Paisley, Mingyuan Zhou, Guillermo Sapiro, Lawrence Carin |
ICIP | 4 |
| 2010 | A Stick-Breaking Construction of the Beta Process
John W. Paisley, Aimee K. Zaas, Christopher W. Woods, Geoffrey S. Ginsburg, Lawrence Carin |
ICML | 5 |
| 2010 | Joint Analysis of Time-Evolving Binary Matrices and Associated DocumentsabstractWe consider problems for which one has incomplete binary matrices that evolve with time (e.g., the votes of legislators on particular legislation, with each year characterized by a different such matrix). An objective of such analysis is to infer structure and inter-relationships underlying the matrices, here defined by latent features associated with each axis of the matrix. In addition, it is assumed that documents are available for the entities associated with at least one of the matrix axes. By jointly analyzing the matrices and documents, one may be used to inform the other within the analysis, and the model offers the opportunity to predict matrix values (e.g., votes) based only on an associated document (e.g., legislation). The research presented here merges two areas of machine-learning that have previously been investigated separately: incomplete-matrix analysis and topic modeling. The analysis is performed from a Bayesian perspective, with efficient inference constituted via Gibbs sampling. The framework is demonstrated by considering all voting data and available documents (legislation) during the 220-year lifetime of the United States Senate and House of Representatives. Dehong Liu, Jorge G. Silva, David B. Dunson, Lawrence Carin |
NIPS | 5 |
| 2010 | Bayesian Inference of the Number of Factors in Gene-Expression Analysis: Application to Human Virus Challenge StudiesabstractBACKGROUND: Nonparametric Bayesian techniques have been developed recently to extend the sophistication of factor models, allowing one to infer the number of appropriate factors from the observed data. We consider such techniques for sparse factor analysis, with application to gene-expression data from three virus challenge studies. Particular attention is placed on employing the Beta Process (BP), the Indian Buffet Process (IBP), and related sparseness-promoting techniques to infer a proper number of factors. The posterior density function on the model parameters is computed using Gibbs sampling and variational Bayesian (VB) analysis. RESULTS: Time-evolving gene-expression data are considered for respiratory syncytial virus (RSV), Rhino virus, and influenza, using blood samples from healthy human subjects. These data were acquired in three challenge studies, each executed after receiving institutional review board (IRB) approval from Duke University. Comparisons are made between several alternative means of per-forming nonparametric factor analysis on these data, with comparisons as well to sparse-PCA and Penalized Matrix Decomposition (PMD), closely related non-Bayesian approaches. CONCLUSIONS: Applying the Beta Process to the factor scores, or to the singular values of a pseudo-SVD construction, the proposed algorithms infer the number of factors in gene-expression data. For real data the "true" number of factors is unknown; in our simulations we consider a range of noise variances, and the proposed Bayesian models inferred the number of factors accurately relative to other methods in the literature, such as sparse-PCA and PMD. We have also identified a "pan-viral" factor of importance for each of the three viruses considered in this study. We have identified a set of genes associated with this pan-viral factor, of interest for early detection of such viruses based upon the host response, as quantified via gene-expression data. Bo Chen 0001, Minhua Chen, John W. Paisley, Aimee K. Zaas, Christopher W. Woods, Geoffrey S. Ginsburg, Alfred O. Hero III, Joseph E. Lucas, David B. Dunson, Lawrence Carin |
BMC Bioinform. | 10 |
| 2010 | Classification with Incomplete Data Using Dirichlet Process Priors
Chunping Wang 0001, Xuejun Liao, Lawrence Carin, David B. Dunson |
J. Mach. Learn. Res. | 3 |
| 2010 | Hierarchical Bayesian Modeling of Topics in Time-Stamped DocumentsabstractWe consider the problem of inferring and modeling topics in a sequence of documents with known publication dates. The documents at a given time are each characterized by a topic and the topics are drawn from a mixture model. The proposed model infers the change in the topic mixture weights as a function of time. The details of this general framework may take different forms, depending on the specifics of the model. For the examples considered here, we examine base measures based on independent multinomial-Dirichlet measures for representation of topic-dependent word counts. The form of the hierarchical model allows efficient variational Bayesian inference, of interest for large-scale problems. We demonstrate results and make comparisons to the model when the dynamic character is removed, and also compare to latent Dirichlet allocation (LDA) and Topics over Time (TOT). We consider a database of Neural Information Processing Systems papers as well as the US Presidential State of the Union addresses from 1790 to 2008. Iulian Pruteanu-Malinici, John W. Paisley, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2010 | Tree-Structured Compressive Sensing With Variational Bayesian AnalysisabstractIn compressive sensing (CS) the known structure in the transform coefficients may be leveraged to improve reconstruction accuracy. We here develop a hierarchical statistical model applicable to both wavelet and JPEG-based DCT bases, in which the tree structure in the sparseness pattern is exploited explicitly. The analysis is performed efficiently via variational Bayesian (VB) analysis, and comparisons are made with MCMC-based inference, and with many of the CS algorithms in the literature. Performance is assessed for both noise-free and noisy CS measurements, based on both JPEG-DCT and wavelet representations. Lihan He, Haojun Chen, Lawrence Carin |
IEEE Signal Process. Lett. | 3 |
| 2009 | Active learning for semi-supervised multi-task learningabstractWe present an algorithm for active learning (adaptive selection of training data) within the context of semi-supervised multi-task classifier design. The semi-supervised multi-task classifier exploits manifold information provided by the unlabeled data, while also leveraging relevant information across multiple data sets. The active-learning component defines which data would be most informative to classifier design if the associated labels are acquired. The framework is demonstrated through application to a real landmine detection problem. Hui Li 0068, Xuejun Liao, Lawrence Carin |
ICASSP | 3 |
| 2009 | Dirichlet process mixture models with multiple modalitiesabstractThe Dirichlet process can be used as a nonparametric prior for an infinite-dimensional probability mass function on the parameter space of a mixture model. The set of parameters over which it is defined is generally used for a single, parametric distribution. We extend this idea to parameter spaces that characterize multiple distributions, or modalities. In this framework, observations containing multiple, incompatible pieces of information can be mixed upon, allowing for all information to inform the final clustering result. We provide a general MCMC sampling scheme and demonstrate this framework on a Gaussian-HMM mixture model applied to synthetic and Major League Baseball data. John W. Paisley, Lawrence Carin |
ICASSP | 2 |
| 2009 | Music analysis with a Bayesian dynamic modelabstractA Bayesian dynamic model is developed to model complex sequential data, with a focus on audio signals from music. The music is represented in terms of a sequence of discrete observations, and the sequence is modeled using a hidden Markov model (HMM) with time-evolving parameters. The model imposes the belief that observations that are temporally proximate are more likely to be drawn from HMMs with similar parameters, while also allowing for “innovation” associated with abrupt changes in the music texture. Segmentation of a given musical piece is constituted via the model inference and the results are compared with other models and also to a conventional music-theoretic analysis. David B. Dunson, Scott Lindroth, Lawrence Carin |
ICASSP | 4 |
| 2009 | Multi-task classification with infinite local expertsabstractWe propose a multi-task learning (MTL) framework for non-linear classification, based on an infinite set of local experts in feature space. The usage of local experts enables sharing at the expert-level, encouraging the borrowing of information even if tasks are similar only in subregions of feature space. A kernel stick-breaking process (KSBP) prior is imposed on the underlying distribution of class labels, so that the number of experts is inferred in the posterior and thus model selection issues are avoided. The MTL is implemented by imposing a Dirichlet process (DP) prior on a layer above the task-dependent KSBPs. Chunping Wang 0001, Lawrence Carin, David B. Dunson |
ICASSP | 3 |
| 2009 | Nonparametric factor analysis with beta process priorsabstractWe propose a nonparametric extension to the factor analysis problem using a beta process prior. This beta process factor analysis (BP-FA) model allows for a dataset to be decomposed into a linear combination of a sparse set of factors, providing information on the underlying structure of the observations. As with the Dirichlet process, the beta process is a fully Bayesian conjugate prior, which allows for analytical posterior calculation and straightforward inference. We derive a variational Bayes inference algorithm and demonstrate the model on the MNIST digits and HGDP-CEPH cell line panel datasets. 1. John W. Paisley, Lawrence Carin |
ICML | 2 |
| 2009 | Learning to Explore and Exploit in POMDPsabstractA fundamental objective in reinforcement learning is the maintenance of a proper balance between exploration and exploitation. This problem becomes more challenging when the agent can only partially observe the states of its environment. In this paper we propose a dual-policy method for jointly learning the agent behavior and the balance between exploration exploitation, in partially observable environments. The method subsumes traditional exploration, in which the agent takes actions to gather information about the environment, and active learning, in which the agent queries an oracle for optimal actions (with an associated cost for employing the oracle). The form of the employed exploration is dictated by the specific problem. Theoretical guarantees are provided concerning the optimality of the balancing of exploration and exploitation. The effectiveness of the method is demonstrated by experimental results on benchmark problems. Chenghui Cai, Xuejun Liao, Lawrence Carin |
NIPS | 3 |
| 2009 | A Bayesian Model for Simultaneous Image Clustering, Annotation and Object SegmentationabstractA non-parametric Bayesian model is proposed for processing multiple images. The analysis employs image features and, when present, the words associated with accompanying annotations. The model clusters the images into classes, and each image is segmented into a set of objects, also allowing the opportunity to assign a word to each object (localized labeling). Each object is assumed to be represented as a heterogeneous mix of components, with this realized via mixture models linking image features to object types. The number of image classes, number of object types, and the characteristics of the object-feature mixture models are inferred non-parametrically. To constitute spatially contiguous objects, a new logistic stick-breaking process is developed. Inference is performed efficiently via variational Bayesian analysis, with example results presented on two image databases. Lan Du 0001, David B. Dunson, Lawrence Carin |
NIPS | 4 |
| 2009 | Non-Parametric Bayesian Dictionary Learning for Sparse Image RepresentationsabstractNon-parametric Bayesian techniques are considered for learning dictionaries for sparse image representations, with applications in denoising, inpainting and compressive sensing (CS). The beta process is employed as a prior for learning the dictionary, and this non-parametric method naturally infers an appropriate dictionary size. The Dirichlet process and a probit stick-breaking process are also considered to exploit structure within an image. The proposed method can learn a sparse dictionary in situ; training images may be exploited if available, but they are not required. Further, the noise variance need not be known, and can be non-stationary. Another virtue of the proposed method is that sequential inference can be readily employed, thereby allowing scaling to large images. Several example results are presented, using both Gibbs and variational Bayesian inference, with comparisons to other state-of-the-art approaches. Mingyuan Zhou, Haojun Chen, John W. Paisley, Guillermo Sapiro, Lawrence Carin |
NIPS | 6 |
| 2009 | Multi-task Reinforcement Learning in Partially Observable Stochastic Environments
Hui Li 0068, Xuejun Liao, Lawrence Carin |
J. Mach. Learn. Res. | 3 |
| 2009 | Semisupervised Learning of Hidden Markov Models via a Homotopy Method
Shihao Ji 0001, Layne T. Watson, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2009 | Semisupervised Multitask LearningabstractContext plays an important role when performing classification, and in this paper we examine context from two perspectives. First, the classification of items within a single task is placed within the context of distinct concurrent or previous classification tasks (multiple distinct data collections). This is referred to as multi-task learning (MTL), and is implemented here in a statistical manner, using a simplified form of the Dirichlet process. In addition, when performing many classification tasks one has simultaneous access to all unlabeled data that must be classified, and therefore there is an opportunity to place the classification of any one feature vector within the context of all unlabeled feature vectors; this is referred to as semi-supervised learning. In this paper we integrate MTL and semi-supervised learning into a single framework, thereby exploiting two forms of contextual information. Example results are presented on a "toy" example, to demonstrate the concept, and the algorithm is also applied to three real data sets. Qiuhua Liu, Xuejun Liao, Hui Li 0068, Jason R. Stack, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2009 | Migratory Logistic Regression for Learning Concept Drift Between Two Data Sets With Application to UXO SensingabstractTo achieve good generalization in supervised learning, the training and testing examples are usually required to be drawn from the same source distribution. In this paper, we propose a method to relax this requirement in the context of logistic regression. AssumingDpandDaare two sets of examples drawn from two different distributionsTandA(called concepts, borrowing a term from psychology), whereDaare fully labeled andDppartially labeled, our objective is to complete the labels ofDp. We introduce an auxiliary variable mu for each example inDato reflect its mismatch withDp. Under an appropriate constraint the mus are estimated as a byproduct, along with the classifier. We also present an active learning approach for selecting the labeled examples inDp. The proposed algorithm, calledmigratorylogisticregression, is demonstrated successfully on simulated data as well as on real measured data of interest for unexploded ordnance cleanup. Xuejun Liao, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2009 | Kernel-Matching Pursuits With Arbitrary Loss FunctionsabstractThe purpose of this research is to develop a classifier capable of state-of-the-art performance in both computational efficiency and generalization ability while allowing the algorithm designer to choose arbitrary loss functions as appropriate for a give problem domain. This is critical in applications involving heavily imbalanced, noisy, or non-Gaussian distributed data. To achieve this goal, a kernel-matching pursuit (KMP) framework is formulated where the objective is margin maximization rather than the standard error minimization. This approach enables excellent performance and computational savings in the presence of large, imbalanced training data sets and facilitates the development of two general algorithms. These algorithms support the use of arbitrary loss functions allowing the algorithm designer to control the degree to which outliers are penalized and the manner in which non-Gaussian distributed data is handled. Example loss functions are provided and algorithm performance is illustrated in two groups of experimental results. The first group demonstrates that the proposed algorithms perform equivalent to several state-of-the-art machine learning algorithms on well-published, balanced data. The second group of results illustrates superior performance by the proposed algorithms on imbalanced, non-Gaussian data achieved by employing loss functions appropriate for the data characteristics and problem domain. Jason R. Stack, Gerald J. Dobeck, Xuejun Liao, Lawrence Carin |
IEEE Trans. Neural Networks | 4 |
| 2008 | Hierarchical kernel stick-breaking process for multi-task image analysisabstractThe kernel stick-breaking process (KSBP) is employed to segment general imagery, imposing the condition that patches (small blocks of pixels) that are spatially proximate are more likely to be associated with the same cluster (segment). The number of clusters is not set a priori and is inferred from the hierarchical Bayesian model. Further, KSBP is integrated with a shared Dirichlet process prior to simultaneously model multiple images, inferring their inter-relationships. This latter application may be useful for sorting and learning relationships between multiple images. The Bayesian inference algorithm is based on a hybrid of variational Bayesian analysis and local sampling. In addition to providing details on the model and associated inference framework, example results are presented for several image-analysis problems. Chunping Wang 0001, Ivo Shterev, Lawrence Carin, David B. Dunson |
ICML | 5 |
| 2008 | Multi-task compressive sensing with Dirichlet process priorsabstractCompressive sensing (CS) is an emerging £eld that, under appropriate conditions, can signi£cantly reduce the number of measurements required for a given signal. In many applications, one is interested in multiple signals that may be measured in multiple CS-type measurements, where here each signal corresponds to a sensing "task". In this paper we propose a novel multitask compressive sensing framework based on a Bayesian formalism, where a Dirichlet process (DP) prior is employed, yielding a principled means of simultaneously inferring the appropriate sharing mechanisms as well as CS inversion for each task. A variational Bayesian (VB) inference algorithm is employed to estimate the full posterior on the model parameters. Yuting Qi, Dehong Liu, David B. Dunson, Lawrence Carin |
ICML | 4 |
| 2008 | The dynamic hierarchical Dirichlet processabstractThe dynamic hierarchical Dirichlet process (dHDP) is developed to model the time-evolving statistical properties of sequential data sets. The data collected at any time point are represented via a mixture associated with an appropriate underlying model, in the framework of HDP. The statistical properties of data collected at consecutive time points are linked via a random parameter that controls their probabilistic similarity. The sharing mechanisms of the time-evolving data are derived, and a relatively simple Markov Chain Monte Carlo sampler is developed. Experimental results are presented to demonstrate the model. David B. Dunson, Lawrence Carin |
ICML | 3 |
| 2008 | Multitask Classification by Learning the Task RelevanceabstractWe consider the problem of multitask learning (MTL), in which we simultaneously learn classifiers for multiple data sets (tasks), with sharing of intertask data as appropriate. We introduce a set of relevance parameters that control the degree to which data from other tasks are used in estimating the current task's classifier parameters. The set of relevance parameters are learned by maximizing their posterior probability, yielding an expectation-maximization (EM) algorithm. We illustrate the effectiveness of our approach through experimental results on a practical data set. Jun Fang 0001, Shihao Ji 0001, Ya Xue, Lawrence Carin |
IEEE Signal Process. Lett. | 4 |
| 2008 | An Investigation of Using the Spectral Characteristics From Ground Penetrating Radar for Landmine/Clutter DiscriminationabstractGround penetrating radar (GPR)-based discrimination of landmines from clutter is known to be challenging due to the wide variability of possible clutter (e.g., rocks, roots, and general soil heterogeneity). This paper discusses the use of GPR frequency-domain spectral features to improve the detection of weak-scattering plastic mines and to reduce the number of false alarms resulting from clutter. The motivation for this approach comes from the fact that landmine targets and clutter objects often have different shapes and/or composition, yielding different energy density spectrum (EDS) that may be exploited for their discrimination (this information is also present in time-domain data, but in the frequency domain we can remove a phase if desired and can reveal better spatial characteristics and therefore often achieve greater robustness). This paper first applies the finite-difference time-domain (FDTD) modeling technique to establish the theoretical foundation. The method to generate EDS from GPR measurements is then described. The consistency of the frequency-domain features is examined through two different GPRs that have different spatial sampling rates and frequency bandwidths. Experimental results from several test sites, based on GPR data collected over buried mines and emplaced buried clutter objects, corroborate the theoretical development and the effectiveness of the proposed spectral feature to increase the accuracy of landmine detection and discrimination. K. C. Ho 0001, Lawrence Carin, Paul D. Gader, Joseph N. Wilson |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2008 | Detection of Unexploded Ordnance via Efficient Semisupervised and Active LearningabstractSemi supervised learning and active learning are considered for unexploded ordnance (UXO) detection. Semi supervised learning algorithms are designed using both labeled and unlabeled data, where here labeled data correspond to sensor signatures for which the identity of the buried item (UXO/non-UXO) is known; for unlabeled data, one only has access to the corresponding sensor data. Active learning is used to define which unlabeled signatures would be most informative to improve the classifier design if the associated label could be acquired (where for UXO sensing, the label is acquired by excavation). A graph-based semi supervised algorithm is applied, which employs the idea of a random Markov walk on a graph, thereby exploiting knowledge of the data manifold (where the manifold is defined by both the labeled and unlabeled data). The algorithm is used to infer labels for the unlabeled data, providing a probability that a given unlabeled signature corresponds to a buried UXO. An efficient active-learning procedure is developed for this algorithm, based on a mutual information measure. In this manner, one initially performs excavation with the purpose of acquiring labels to improve the classifier, and once this active-learning phase is completed, the resulting semi supervised classifier is then applied to the remaining unlabeled signatures to quantify the probability that each such item is a UXO. Example classification results are presented for an actual UXO site, based on electromagnetic induction and magnetometer data. Performance is assessed in comparison to other semi supervised approaches, as well as to supervised algorithms. Qiuhua Liu, Xuejun Liao, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2008 | Infinite Hidden Markov Models for Unusual-Event Detection in VideoabstractWe address the problem of unusual-event detection in a video sequence. Invariant subspace analysis (ISA) is used to extract features from the video, and the time-evolving properties of these features are modeled via an infinite hidden Markov model (iHMM), which is trained using "normal"/"typical" video. The iHMM retains a full posterior density function on all model parameters, including the number of underlying HMM states. Anomalies (unusual events) are detected subsequently if a low likelihood is observed when associated sequential features are submitted to the trained iHMM. A hierarchical Dirichlet process framework is employed in the formulation of the iHMM. The evaluation of posterior distributions for the iHMM is achieved in two ways: via Markov chain Monte Carlo and using a variational Bayes formulation. Comparisons are made to modeling based on conventional maximum-likelihood-based HMMs, as well as to Dirichlet-process-based Gaussian-mixture models. Iulian Pruteanu-Malinici, Lawrence Carin |
IEEE Trans. Image Process. | 2 |
| 2007 | Point-Based Policy Iteration
Shihao Ji 0001, Ronald Parr, Hui Li 0068, Xuejun Liao, Lawrence Carin |
AAAI | 5 |
| 2007 | Wideband Array Imaging of a Target Situated in an Unknown Random MediaabstractWe propose two new methods for wideband array signal imaging for targets situated in unknown random media. First, a normalized coherent interferometric (N-CINT) imaging algorithm is developed based on coherent interferometric (CINT) imaging theory, yielding improved imaging performance with experimental data. Second, a phase-difference analysis (PDA) method is proposed to significantly reduce computation time and to improve imaging quality. The parameters in the two methods are determined adaptively by optimizing an objective function. Experiments are carried out for electromagnetic scattering using a linear antenna array, providing a demonstration of these methods. Dehong Liu, Lawrence Carin |
ICASSP (1) | 2 |
| 2007 | Learning Classifiers on a Partially Labeled Data ManifoldabstractWe present an algorithm for learning parametric classifiers on a partially labeled data manifold, based on a graph representation of the manifold. The unlabeled data are utilized by basing classifier learning on neighborhoods, formed via Markov random walks. The proposed algorithm yields superior performance on three benchmark data sets and the margin of improvements over existing semi-supervised algorithms is significant. Qiuhua Liu, Xuejun Liao, Lawrence Carin |
ICASSP (2) | 3 |
| 2007 | Multi-Aspect Target Classification and Detection via the Infinite Hidden Markov ModelabstractA new multi-aspect target detection method is presented based on the infinite hidden Markov model (iHMM). The scattering of waves from multiple targets is modeled as an iHMM with the number of underlying states treated as infinite, from which a full posterior distribution on the number of states associated with the targets is inferred and the target-dependent states are learned collectively. A set of Dirichlet processes (DPs) are used to define the rows of the HMM transition matrix and these DPs are linked and shared via a hierarchical Dirichlet process (HDP). Learning and inference for the iHMM are based on an effective Gibbs sampler. The framework is demonstrated using measured acoustic scattering data. Kai Ni 0007, Yuting Qi, Lawrence Carin |
ICASSP (2) | 3 |
| 2007 | Dirichlet Process HMM Mixture Models with Application to Music AnalysisabstractA hidden Markov mixture model is developed using a Dirichlet process (DP) prior, to represent the statistics of sequential data for which a single hidden Markov model (HMM) may not be sufficient. The DP prior has an intrinsic clustering property that encourages parameter sharing, naturally revealing the proper number of mixture components. The evaluation of posterior distributions for all model parameters is achieved via a variational Bayes formulation. We focus on exploring music similarities as an important application, highlighting the effectiveness of the HMM mixture model. Experimental results are presented from classical music clips. Yuting Qi, John W. Paisley, Lawrence Carin |
ICASSP (2) | 3 |
| 2007 | Infinite Hidden Markov Models and ISA Features for Unusual-Event Detection in VideoabstractWe address the problem of unusual-event detection in a video sequence. Invariant subspace analysis (ISA) is used to extract features from the video, and the time-evolving properties of these features are modeled via an infinite hidden Markov model (iHMM), which is trained using "normal"/"typical" video data. The iHMM automatically determines the proper number of HMM states, and it retains a full posterior density function on all model parameters. Anomalies (unusual events) are detected subsequently if a low likelihood is observed when associated sequential features are submitted to the trained iHMM. A hierarchical Dirichlet process (HDP) framework is employed in the formulation of the iHMM. The evaluation of posterior distributions for the iHMM is achieved in two ways: via MCMC and using a variational Bayes (VB) formulation. Iulian Pruteanu-Malinici, Lawrence Carin |
ICIP (5) | 2 |
| 2007 | Bayesian compressive sensing and projection optimizationabstractThis paper introduces a new problem for which machine-learning tools may make an impact. The problem considered is termed "compressive sensing", in which a real signal of dimension N is measured accurately based on K << N real measurements. This is achieved under the assumption that the underlying signal has a sparse representation in some basis (e.g., wavelets). In this paper we demonstrate how techniques developed in machine learning, specifically sparse Bayesian regression and active learning, may be leveraged to this new problem. We also point out future research directions in compressive sensing of interest to the machine-learning community. Shihao Ji 0001, Lawrence Carin |
ICML | 2 |
| 2007 | Quadratically gated mixture of experts for incomplete data classificationabstractWe introduce quadratically gated mixture of experts (QGME), a statistical model for multi-class nonlinear classification. The QGME is formulated in the setting of incomplete data, where the data values are partially observed. We show that the missing values entail joint estimation of the data manifold and the classifier, which allows adaptive imputation during classifier learning. The expectation maximization (EM) algorithm is derived for joint likelihood maximization, with adaptive imputation performed analytically in the E-step. The performance of QGME is evaluated on three benchmark data sets and the results show that the QGME yields significant improvements over competing methods. Xuejun Liao, Hui Li 0068, Lawrence Carin |
ICML | 3 |
| 2007 | Multi-task learning for sequential data via iHMMs and the nested Dirichlet processabstractA new hierarchical nonparametric Bayesian model is proposed for the problem of multitask learning (MTL) with sequential data. Sequential data are typically modeled with a hidden Markov model (HMM), for which one often must choose an appropriate model structure (number of states) before learning. Here we model sequential data from each task with an infinite hidden Markov model (iHMM), avoiding the problem of model selection. The MTL for iHMMs is implemented by imposing a nested Dirichlet process (nDP) prior on the base distributions of the iHMMs. The nDP-iHMM MTL method allows us to perform task-level clustering and data-level clustering simultaneously, with which the learning for individual iHMMs is enhanced and between-task similarities are learned. Learning and inference for the nDP-iHMM MTL are based on a Gibbs sampler. The effectiveness of the framework is demonstrated using synthetic data as well as real music data. Kai Ni 0007, Lawrence Carin, David B. Dunson |
ICML | 2 |
| 2007 | The matrix stick-breaking process for flexible multi-task learningabstractIn multi-task learning our goal is to design regression or classification models for each of the tasks and appropriately share information between tasks. A Dirichlet process (DP) prior can be used to encourage task clustering. However, the DP prior does not allow local clustering of tasks with respect to a subset of the feature vector without making independence assumptions. Motivated by this problem, we develop a new multitask-learning prior, termed the matrix stick-breaking process (MSBP), which encourages cross-task sharing of data. However, the MSBP allows separate clustering and borrowing of information for the different feature components. This is important when tasks are more closely related for certain features than for others. Bayesian inference proceeds by a Gibbs sampling algorithm and the approach is illustrated using a simulated example and a multi-national application. Ya Xue, David B. Dunson, Lawrence Carin |
ICML | 3 |
| 2007 | Semi-Supervised Multitask LearningabstractA semi-supervised multitask learning (MTL) framework is presented, in which M parameterized semi-supervised classifiers, each associated with one of M par- tially labeled data manifolds, are learned jointly under the constraint of a soft- sharing prior imposed over the parameters of the classifiers. The unlabeled data are utilized by basing classifier learning on neighborhoods, induced by a Markov random walk over a graph representation of each manifold. Experimental results on real data sets demonstrate that semi-supervised MTL yields significant im- provements in generalization performance over either semi-supervised single-task learning (STL) or supervised MTL. Qiuhua Liu, Xuejun Liao, Lawrence Carin |
NIPS | 3 |
| 2007 | Multi-Task Learning for Classification with Dirichlet Process PriorsabstractConsider the problem of learning logistic-regression models for multiple classification tasks, where the training data set for each task is not drawn from the same statistical distribution. In such a multi-task learning (MTL) scenario, it is necessary to identify groups of similar tasks that should be learned jointly. Relying on a Dirichlet process (DP) based statistical model to learn the extent of similarity between classification tasks, we develop computationally efficient algorithms for two different forms of the MTL problem. First, we consider a symmetric multi-task learning (SMTL) situation in which classifiers for multiple tasks are learned jointly using a variational Bayesian (VB) algorithm. Second, we consider an asymmetric multi-task learning (AMTL) formulation in which the posterior density function from the SMTL model parameters (from previous tasks) is used as a prior for a new task: this approach has the significant advantage of not requiring storage and use of all previous data from prior tasks. The AMTL formulation is solved with a simple Markov Chain Monte Carlo (MCMC) construction. Experimental results on two real life MTL problems indicate that the proposed algorithms: (a) automatically identify subgroups of related tasks whose training data appear to be drawn from similar distributions; and (b) are more accurate than simpler approaches such as single-task learning, pooling of data across all tasks, and simplified approximations to DP. Ya Xue, Xuejun Liao, Lawrence Carin, Balaji Krishnapuram |
J. Mach. Learn. Res. | 3 |
| 2007 | A Bivariate Gaussian Model for Unexploded Ordnance Classification with EMI DataabstractA bivariate Gaussian model is proposed for modeling spatially varying electromagnetic-induction (EMI) response of unexploded ordnance (UXO). This model is proposed for EMI sensors that do not exploit enough physics to warrant using the popular magnetic-dipole model currently commonly used. These two competing models are applied to measured EM61 sensor data at a real UXO site. UXO classification performance using the proposed bivariate Gaussian model is shown to be superior to an approach employing the magnetic-dipole model. Moreover, the bivariate Gaussian model requires no labeled training data, obviates classifier construction, and has fewer model parameters to learn. Yijun Yu 0004, Levi Kennedy, Xianyang Zhu, Lawrence Carin |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2007 | On Classification with Incomplete DataabstractWe address the incomplete-data problem in which feature vectors to be classified are missing data (features). A (supervised) logistic regression algorithm for the classification of incomplete data is developed. Single or multiple imputation for the missing data is avoided by performing analytic integration with an estimated conditional density function (conditioned on the observed data). Conditional density functions are estimated using a Gaussian mixture model (GMM), with parameter estimation performed using both Expectation-Maximization (EM) and Variational Bayesian EM (VB-EM). The proposed supervised algorithm is then extended to the semisupervised case by incorporating graph-based regularization. The semisupervised algorithm utilizes all available data-both incomplete and complete, as well as labeled and unlabeled. Experimental results of the proposed classification algorithms are shown. Xuejun Liao, Ya Xue, Lawrence Carin, Balaji Krishnapuram |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2007 | Cost-sensitive feature acquisition and classification
Shihao Ji 0001, Lawrence Carin |
Pattern Recognit. | 2 |
| 2007 | Adaptive Multimodality Sensing of LandminesabstractThe problem of adaptive multimodality sensing of landmines is considered based on electromagnetic induction (EMI) and ground-penetrating radar (GPR) sensors. Two formulations are considered based on a partially observable Markov decision process (POMDP) framework. In the first formulation, it is assumed that sufficient training data are available, and a POMDP model is designed based on physics-based features, with model selection performed via a variational Bayes analysis of several possible models. In the second approach, the training data are assumed absent or insufficient, and a lifelong-learning approach is considered, in which exploration and exploitation are integrated. We provide a detailed description of both formulations, with example results presented using measured EMI and GPR data, for buried mines and clutter Lihan He, Shihao Ji 0001, Waymond R. Scott, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2007 | Electromagnetic Target Detection in Uncertain Media: Time-Reversal and Minimum-Variance AlgorithmsabstractAn experimental study is performed on imaging targets that are situated in a highly scattering environment, employing electromagnetic time-reversal methods. A particular focus is placed on performance when the electrical properties of the background environment (medium) are uncertain. It is assumed that the (unknown) medium characteristic of the scattered fields represents one sample from an underlying random process, with this random process representing our uncertainty in the media properties associated with the scattering measurement. While the specific Green's function associated with the scattered fields is unknown, we assume access to an ensemble of Green's functions sampled from the aforementioned distribution. This ensemble of Green's functions may be used in several ways to mitigate uncertainty in the true Green's function. Specifically, when performing time-reversal imaging, we consider a Green's function as a representative of the average of the ensemble, as well as Green's functions based on a principal components analysis of the ensemble. We also develop a wideband minimum-variance beamformer with environment perturbation constraints, in which the unknown Green's function is constrained to reside in a subspace spanned by the Green's function ensemble. These algorithms are examined using electromagnetic scattering data measured in a canonical set of laboratory experiments. The qualitative performance of the different techniques is presented in the form of images, with quantitative results presented in the form of receiver operating characteristic performance Dehong Liu, Jeffrey L. Krolik, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2007 | Classification of Unexploded Ordnance Using Incomplete Multisensor Multiresolution DataabstractWe address the problem of unexploded ordnance (UXO) detection in which data to be classified is available from multiple sensor modalities and multiple resolutions. Specifically, features are extracted from measured magnetometer and electromagnetic induction data; multiple-resolution data are manifested when the sensors are separated from the buried targets of interest by different distances (e.g., different sensor-platform heights). The proposed classification algorithm explicitly emphasizes features extracted from fine-resolution imagery over those extracted from less reliable coarse-resolution data. When fine-resolution features are unavailable (due to undeployed sensors), the algorithm analytically integrates out the missing features via an estimated conditional density function, which is conditioned on the observed features (from deployed sensors). This density function exploits the statistical relationships that exist among features at different resolutions, as well as those among features from different sensors (in the multisensor case). Experimental classification results are shown for real UXO data, on which the proposed algorithm consistently achieves better classification performance than common alternative approaches. Chunping Wang 0001, Xuejun Liao, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2007 | Three-Dimensional Bayesian Inversion With Application to Subsurface SensingabstractA Bayesian formalism is considered for inverting for the parameters of a heterogeneity profile based on measured scattering data. It is shown that the typical use of regularization (e.g., Thikonov) corresponds to a maximum a posteriori point approximation to the full-posterior density function on the heterogeneity parameters, given the observed data. In the Bayesian framework considered here, the full posterior is approximated as a multidimensional Gaussian distribution. The mean of this distribution may be used as a point estimate of the heterogeneity profile, with the covariance matrix providing associated "error bars" (a measure of confidence in the inversion). In addition to providing an approximation to the full posterior of the heterogeneity profile, this formalism addresses the proper weighting to apply for inversion regularization. Specifically, an important limitation of previous regularization procedures is the need to place a weight on the importance of the regularization relative to the importance of fitting the data to the underlying model. In the Bayesian analysis outlined here, we also assign such a weight, but now the weight is treated as a random variable, with a statistical prior. The measured data are then used to determine a posterior distribution on the parameter, based on the measured data. We present here the basic Bayesian inversion framework, with several example results presented for subsurface-sensing problems Yijun Yu 0004, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2006 | Incremental Least Squares Policy Iteration for POMDPs
Hui Li 0068, Xuejun Liao, Lawrence Carin |
AAAI | 3 |
| 2006 | Homotopy-Based Semi-Supervised Hidden Markov Tree for Texture AnalysisabstractA semi-supervised hidden Markov tree (HMT) model is developed for texture analysis, incorporating both labeled and unlabeled data for training; the optimal balance between labeled and unlabeled data is estimated via the homotopy method. In traditional EM-based semi-supervised modeling, this balance is dictated by the relative size of labeled and unlabeled data, often leading to poor performance. Semi-supervised modeling may be viewed as a source allocation problem between labeled and unlabeled data, controlled by a parameter λ ∈ [0, 1], where λ = 0 and 1 correspond to the purely supervised HMT model and purely unsupervised HMT-based clustering, respectively. We consider the homotopy method to track a path of fixed points starting from λ = 0, with the optimal source allocation identified as a critical transition point where the solution is unsupported by the initial labeled data. Experimental results on real textures demonstrate the superiority of this method compared to the EM-based semi-supervised HMT training. Nilanjan Dasgupta, Shihao Ji 0001, Lawrence Carin |
ICASSP (2) | 3 |
| 2006 | A Reward-Directed Bayesian ClassifierabstractWe consider a classification problem wherein the class features are not given a priori. The classifier is responsible for selecting the features, to minimize the cost of observing features while also maximizing the classification performance. We propose a reward-directed Bayesian classifier (RDBC) to solve this problem. The RDBC features an internal state structure for preserving the feature dependence, and is formulated as a partially observable Markov decision process (POMDP). The results on a diabetes dataset show the RDBC with a moderate number of states significantly improves over the naive Bayes classifier, both in prediction accuracy and observation parsimony. It is also demonstrated that the RDBC performs better by using more states to increase its memory. Hui Li 0068, Xuejun Liao, Lawrence Carin |
ICASSP (5) | 3 |
| 2006 | Region-based value iteration for partially observable Markov decision processesabstractAn approximate region-based value iteration (RBVI) algorithm is proposed to find the optimal policy for a partially observable Markov decision process (POMDP). The proposed RBVI approximates the true polyhedral partition of the belief simplex with an ellipsoidal partition, such that the optimal value function is linear in each of the ellipsoidal regions. The position and shape of each region, as well as the gradient (alpha-vector) of the optimal value function in the region, are parameterized explicitly, and are estimated via efficient expectation maximization (EM) and variational Bayesian EM (VBEM), based on a set of selected sample belief points. The RBVI maintains a much smaller number of alpha-vectors than point-based methods and yields a more parsimonious representation that approximates the true value function in the maximum likelihood (ML) sense. The results on benchmark problems show that the proposed RBVI is comparable in performance to state-of-the-art algorithms, despite of the small number of alpha-vectors that are used. Hui Li 0068, Xuejun Liao, Lawrence Carin |
ICML | 3 |
| 2006 | Variational Bayes for Continuous Hidden Markov Models and Its Application to Active LearningabstractIn this paper, we present a varitional Bayes (VB) framework for learning continuous hidden Markov models (CHMMs), and we examine the VB framework within active learning. Unlike a maximum likelihood or maximum a posteriori training procedure, which yield a point estimate of the CHMM parameters, VB-based training yields an estimate of the full posterior of the model parameters. This is particularly important for small training sets since it gives a measure of confidence in the accuracy of the learned model. This is utilized within the context of active learning, for which we acquire labels for those feature vectors for which knowledge of the associated label would be most informative for reducing model-parameter uncertainty. Three active learning algorithms are considered in this paper: 1) query by committee (QBC), with the goal of selecting data for labeling that minimize the classification variance, 2) a maximum expected information gain method that seeks to label data with the goal of reducing the entropy of the model parameters, and 3) an error-reduction-based procedure that attempts to minimize classification error over the test data. The experimental results are presented for synthetic and measured data. We demonstrate that all of these active learning methods can significantly reduce the amount of required labeling, compared to random selection of samples for labeling. Shihao Ji 0001, Balaji Krishnapuram, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2006 | A modified SPIHT algorithm for image coding with a joint MSE and classification distortion measureabstractThe set partitioning in hierarchical trees (SPIHT) algorithm is an efficient wavelet-based progressive image-compression technique, designed to minimize the mean-squared error (MSE) between the original and decoded imagery. However, the MSE-based distortion measure is not in general well correlated with image-recognition quality, especially at low bit rates. Specifically, low-amplitude wavelet coefficients that may be important for classification are given low priority by conventional SPIHT. In this paper, we use the kernel matching pursuits (KMP) method to autonomously estimate the importance of each wavelet subband for distinguishing between different textures, with textural segmentation first performed via a hidden Markov tree. Based on subband importance determined via KMP, we scale the wavelet coefficients prior to SPIHT coding, with the goal of minimizing a Lagrangian distortion based jointly on the MSE and classification error. For comparison we consider Bayes tree-structured vector quantization (B-TSVQ), also designed to obtain a tradeoff between MSE and classification error. The performances of the original SPIHT, the modified SPIHT, and B-TSVQ are compared. Shaorong Chang, Lawrence Carin |
IEEE Trans. Image Process. | 2 |
| 2005 | A Bayesian Approach to Unsupervised Feature Selection and Density Estimation Using Expectation PropagationabstractWe propose an approximate Bayesian approach for unsupervised feature selection and density estimation, where the importance of the features for clustering is used as the measure for feature selection. Traditional maximum-likelihood (ML) model-parameter optimization schemes estimate the feature saliencies for a fixed model structure (i.e., a fixed number of clusters). In practice, the number of clusters present in the data for mixture-based modeling is unknown. In an ML framework, the number of clusters typically needs to be ascertained prior to estimating the feature saliencies. We propose a density estimation scheme that addresses model complexity (number of clusters present) and model-parameter estimation (feature saliencies) in a single optimization framework. The approximate Bayesian approach presented here, based on the expectation propagation method, obtains a full posterior distribution on the saliency of the features, along with full posterior distribution of other model parameters (including the number of clusters) that represent the underlying statistics of the data. The performance of the algorithm, is analyzed based on its ability to identify the features salient for clustering the multivariate data. Shaorong Chang, Nilanjan Dasgupta, Lawrence Carin |
CVPR (2) | 3 |
| 2005 | Logistic regression with an auxiliary data sourceabstractTo achieve good generalization in supervised learning, the training and testing examples are usually required to be drawn from the same source distribution. In this paper we propose a method to relax this requirement in the context of logistic regression. Assuming Dp and Da are two sets of examples drawn from two mismatched distributions, where Da are fully labeled and Dp partially labeled, our objective is to complete the labels of Dp. We introduce an auxiliary variable μ for each example in Da to reflect its mismatch with Dp. Under an appropriate constraint the μ's are estimated as a byproduct, along with the classifier. We also present an active learning approach for selecting the labeled examples in Dp. The proposed algorithm, called "Migratory-Logit" or M-Logit, is demonstrated successfully on simulated as well as real data sets. Xuejun Liao, Ya Xue, Lawrence Carin |
ICML | 3 |
| 2005 | Incomplete-data classification using logistic regressionabstractA logistic regression classification algorithm is developed for problems in which the feature vectors may be missing data (features). Single or multiple imputation for the missing data is avoided by performing analytic integration with an estimated conditional density function (conditioned on the non-missing data). Conditional density functions are estimated using a Gaussian mixture model (GMM), with parameter estimation performed using both expectation maximization (EM) and Variational Bayesian EM (VB-EM). Using widely available real data, we demonstrate the general advantage of the VB-EM GMM estimation for handling incomplete data, vis-à-vis the EM algorithm. Moreover, it is demonstrated that the approach proposed here is generally superior to standard imputation procedures. Xuejun Liao, Ya Xue, Lawrence Carin |
ICML | 4 |
| 2005 | Radial Basis Function Network for Multi-task LearningabstractWe extend radial basis function (RBF) networks to the scenario in which multiple correlated tasks are learned simultaneously, and present the cor- responding learning algorithms. We develop the algorithms for learn- ing the network structure, in either a supervised or unsupervised manner. Training data may also be actively selected to improve the network’s gen- eralization to test data. Experimental results based on real data demon- strate the advantage of the proposed algorithms and support our conclu- sions. Xuejun Liao, Lawrence Carin |
NIPS | 2 |
| 2005 | Sparse Multinomial Logistic Regression: Fast Algorithms and Generalization BoundsabstractRecently developed methods for learning sparse classifiers are among the state-of-the-art in supervised learning. These methods learn classifiers that incorporate weighted sums of basis functions with sparsity-promoting priors encouraging the weight estimates to be either significantly large or exactly zero. From a learning-theoretic perspective, these methods control the capacity of the learned classifier by minimizing the number of basis functions used, resulting in better generalization. This paper presents three contributions related to learning sparse classifiers. First, we introduce a true multiclass formulation based on multinomial logistic regression. Second, by combining a bound optimization approach with a component-wise update procedure, we derive fast exact algorithms for learning sparse multiclass classifiers that scale favorably in both the number of training samples and the feature dimensionality, making them applicable even to large data sets in high-dimensional feature spaces. To the best of our knowledge, these are the first algorithms to perform exact multinomial logistic regression with a sparsity-promoting prior. Third, we show how nontrivial generalization bounds can be derived for our classifier in the binary case. Experimental results on standard benchmark data sets attest to the accuracy, sparsity, and efficiency of the proposed methods. Balaji Krishnapuram, Lawrence Carin, Mário A. T. Figueiredo, Alexander J. Hartemink |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Kernel matching pursuits prioritization of wavelet coefficients for SPIHT image codingabstractThe Set Partitioning In Hierarchical Trees (SPIHT), an efficient wavelet-based progressive image-compression scheme, is oriented to minimize the mean-squared error (MSE) between the original and decoded imagery. In this paper, we use the kernel matching pursuits (KMP) method to estimate the importance of each wavelet sub-band for distinguishing between different textures segmented by an HMT mixture model. Before the SPIHT coding, we weight the wavelet coefficients, with the goal of achieving improved image-classification results at low bit rates. A modified SPIHT algorithm is proposed to improve the coding efficiency. The performance of the original SPIHT and the modified SPIHT algorithms is compared. Shaorong Chang, Lawrence Carin |
ICASSP (3) | 2 |
| 2004 | Time-reversal imaging and classification for distant targets in a shallow water channelabstractTime-reversal imaging (TRI) is analogous to matched-field processing, although TRI is typically very wideband and is capable of performing target classification (in addition to localization). We apply the time-reversal technique to locate man-made cylindrical targets moving in a shallow ocean channel at long range, as well as to classify them from natural false targets like a school of fish. We present imaging and classification on simulated scattering data, for both target classes. In addition to the imaging, we explore extraction of features from the time-reversal data, with these applied to subsequent target classification. Time-reversal implementation requires a fast forward model, which we implement by a normal-mode model. We present the underlying theory of TRI, feature extraction and target classification via a relevance vector machine (RVM). Nilanjan Dasgupta, Lawrence Carin |
ICASSP (2) | 2 |
| 2004 | Adaptive multi-aspect target classification and detection with hidden Markov modelsabstractWe consider target classification and detection based on backscattered observations measured from a sequence of target-sensor orientations. The multi-aspect scattered waves from a given target are modeled with a hidden Markov model (HMM). The targets are assumed concealed and the absolute target-sensor orientation is assumed unknown; therefore, it is possible to control only the angular displacements (change in orientation) between consecutive measurements. The performance of the HMM classifiers/detectors is influenced by the choice of the angular displacements, the optimization of which motivates the developed adaptive search strategies, based on entropy-driven optimality criteria. The search proceeds in a sequential fashion. Based on the previous observations and their associated angular displacements, one determines the optimal next displacement to perform an associated observation. The search strategies are detailed and example results presented on adaptive classification and detection of underwater targets. Shihao Ji 0001, Xuejun Liao, Lawrence Carin |
ICASSP (2) | 3 |
| 2004 | Airport detection in large aerial optical imageryabstractA method to detect airports in large aerial optical imagery is considered. Combining texture segmentation and shape detection, this method shows advantages in analyzing large aerial imagery. First, large aerial images are segmented and interpreted according to textural features using a fast kernel matching pursuits (KMP) algorithm. As a result, attention is then paid to small regions of interest, extracted from the large images. Second, for each region of interest, a corresponding binary image is generated via the Canny edge operator, yielding a modified Hough transform image with which we search for elongated rectangles with desired dimensions (characteristic of runways). Those detected rectangles are declared as runways and the corresponding region of interest as an airport. Application on a dozen aerial images from southern California, demonstrates the effectiveness of the algorithm. Dehong Liu, Lihan He, Lawrence Carin |
ICASSP (5) | 3 |
| 2004 | Active selection of labeled data for target detectionabstractAn information-theoretic approach is developed for target detection, with active training set selection, directly from the site-specific measured data. For the proposed kernel-based algorithm, a set of basis functions to characterize the signature distribution of the site are defined first; then we determine a parsimonious set of data, for which knowledge of the associated labels would be most informative to determine the weights for the basis functions. Both of them utilize the Fisher information criteria. The proposed framework is applied to subsurface target detection, with example results presented for an actual buried unexploded ordnance site. Yan Zhang 0025, Xuejun Liao, Esther Durá, Lawrence Carin |
ICASSP (5) | 4 |
| 2004 | On Semi-Supervised ClassificationabstractA graph-based prior is proposed for parametric semi-supervised classi- fication. The prior utilizes both labelled and unlabelled data; it also in- tegrates features from multiple views of a given sample (e.g., multiple sensors), thus implementing a Bayesian form of co-training. An EM algorithm for training the classifier automatically adjusts the tradeoff be- tween the contributions of: (a) the labelled data; (b) the unlabelled data; and (c) the co-training information. Active label query selection is per- formed using a mutual information based criterion that explicitly uses the unlabelled data and the co-training information. Encouraging results are presented on public benchmarks and on measured data from single and multiple sensors. 1 Introduction In many pattern classification problems, the acquisition of labelled training data is costly and/or time consuming, whereas unlabelled samples can be obtained easily. Semi- supervised algorithms that learn from both labelled and unlabelled samples have been the focus of much research in the last few years; a comprehensive review up to 2001 can be found in [13], while more recent references include [1, 2, 6, 7, 1618]. Most recent semi-supervised learning algorithms work by formulating the assumption that "nearby" points, and points in the same structure (e.g., cluster), should have similar labels [6, 7, 16]. This can be seen as a form of regularization, pushing the class boundaries toward regions of low data density. This regularization is often implemented by associating the vertices of a graph to all the (labelled and unlabelled) samples, and then formulating the problem on the vertices of the graph [6, 1618]. While current graph-based algorithms are inherently transductive -- i.e., they cannot be used directly to classify samples not present when training -- our classifier is paramet- ric and the learned classifier can be used directly on new samples. Furthermore, our al- gorithm is trained discriminatively by maximizing a concave objective function; thus we avoid thorny local maxima issues that plague many earlier methods. Unlike existing methods, our algorithm automatically learns the relative importance of the labelled and unlabelled data. When multiple views of the same sample are provided (e.g. features from different sensors), we develop a new Bayesian form of co-training [4]. In addition, we also show how to exploit the unlabelled data and the redundant views of the sample (from co-training) in order to improve active label query selection [15]. The paper is organized as follows. Sec. 2 briefly reviews multinomial logistic regression. Sec. 3 describes the priors for semi-supervised learning and co-training. The EM algorithm derived to learn the classifiers is presented in Sec. 4. Active label selection is discussed in Sec. 5. Experimental results are shown in Sec. 6, followed by conclusions in Sec. 7. 2 Multinomial Logistic Regression In an m-class supervised learning problem, one is given a labelled training set DL = {(x d 1, y ), . . . , (x )} 1 L, yL , where xi R is a feature vector and yi the corresponding (1) (m) class label. In "1-of-m" encoding, y = [y , . . . , y ] i is a binary vector, such that i i (c) (j) y = 1 and y = 0, for j = c, indicates that sample i belongs to class c. In multinomial i i logistic regression [5], the posterior class probabilities are modelled as log P (y(c) = 1|x) = xT w(c) - log m exp(xT w(k)), for c = 1, . . . , m, (1) k=1 where w(c) d R is the class-c weight vector. Notice that since m P (y(c)= 1|x) = 1, c=1 one of the weight vectors is redundant; we arbitrarily choose to set w(m) = 0, and consider the (d (m-1))-dimensional vector w = [(w(1))T , ..., (w(m-1))T ]T . Estimation of w may be achieved by maximizing the log-likelihood (with Y {y , ..., y } 1 L ) [5] (c) (w) log P (Y|w) = L m y xT w(c) - log m exp(xT w(j)) . (2) i=1 c=1 i i j=1 i In the presence of a prior p(w), we seek a maximum a posteriori (MAP) estimate, w = arg max { (w) + log p(w)} w . Actually, if the training data is separable, (w) is unbounded, and a prior is crucial. Although we focus on linear classifiers, we may see the d-dimensional feature vectors x as having resulted from some deterministic, maybe nonlinear, transformation of an input raw feature vector r; e.g., in a kernel classifier, xi = [1, K(ri, r1), ..., K(ri, rL)] (d = L + 1). 3 Graph-Based Data-Dependent Priors 3.1 Graph Laplacians and Regularization for Semi-Supervised Learning Consider a scalar function f = [f1, ..., f|V |]T , defined on the set V = {1, 2, ..., |V |} of vertices of an undirected graph (V, E). Each edge of the graph, joining vertices i and j, is given a weight kij = kji 0, and we collect all the weights in a |V | |V | matrix K. A natural way to measure how much f varies across the graph is by the quantity kij(fi - fj)2 = 2 f T f , (3) i j where = diag{ k k j 1j , ..., j |V |j } - K is the so-called graph Laplacian [2]. Notice that kij 0 (for all i, j) guarantees that is positive semi-definite and also that has (at least) one null eigenvalue (1T1 = 0, where 1 has all elements equal to one). In semi-supervised learning, in addition to DL, we are given U unlabelled samples DU = {xL+1, . . . , xL+U }. To use (3) for semi-supervised learning, the usual choice is to assign one vertex of the graph to each sample in X = [x1, . . . , xL+U ]T (thus |V | = L + U ), and to let kij represent some (non-negative) measure of "similarity" between xi and xj. A Gaussian random field (GRF) is defined on the vertices of V (with inverse variance ) p(f ) exp{- f T f /2}, in which configurations that vary more (according to (3)) are less probable. Most graph- based approaches estimate the values of f , given the labels, using p(f ) (or some modifica- tion thereof) as a prior. Accordingly, they work in a strictly transductive manner. 3.2 Non-Transductive Semi-Supervised Learning We first consider two-class problems (m = 2, thus w d R ). In contrast to previous uses of graph-based priors, we define f as the real function f (defined over the entire observation space) evaluated at the graph nodes. Specifically, f is defined as a linear function of x, and at the graph node i, fi f (xi) = wT xi. Then, f = [f1, ..., f|V |]T = Xw, and p(f ) induces a Gaussian prior on w, with precision matrix A = XT X, p(w) exp{-(/2) wT XT Xw} = exp{-(/2) wT Aw}. (4) Notice that since is singular, A may also be singular, and the corresponding prior may therefore be improper. This is no problem for MAP estimation of w because (as is well known) the normalization factor of the prior plays no role in this estimate. If we include extra regularization, by adding a non-negative diagonal matrix to A, the prior becomes p(w) exp -(1/2) wT (0A + ) w , (5) where we may choose = diag{1, ..., d}, = 1I, or even = 0. For m > 2, we define (m-1) identical independent priors, one for each w(c), c = 1, ..., m. The joint prior on w = [(w(1))T , ..., (w(m-1))T ]T is then m-1 1 (c) 1 p(w|) exp{- (w(c))T A + (c) w(c)} = exp{- wT ()w}, (6) 2 0 2 c=1 (c) (c) (c) where is a vector containing all the parameters, (c) = diag{ , ..., }, and i 1 d (1) (m-1) () = diag{ , ..., } A + block-diag{(1), ..., (m-1)}. (7) 0 0 Finally, since all the 's are inverses of variances, the conjugate priors are Gamma [3]: (c) (c) (c) (c) p( | | | | 0 0, 0) = Ga(0 0, 0), and p(i 1, 1) = Ga(i 1, 1), for c = 1, ..., m - 1 and i = 1, ..., d. Usually, 0, 0, 1, and 1 are given small values indicating diffuse priors. In the zero limit, we obtain scale-invariant (improper) Jeffreys hyper-priors. Summarizing, our model for semi-supervised learning includes the log-likelihood (2), a prior (6), and Gamma hyper-priors. In Section 4, we present a simple and computationally efficient expectation-maximization (EM) algorithm for obtaining the MAP estimate of w. 3.3 Exploiting Features from Multiple Sensors: The Co-Training Prior In some applications several sensors are available, each providing a different set of features. For simplicity, we assume two sensors s {1, 2}, but everything discussed here is easily (s) extended to any number of sensors. Denote the features from sensor s, for sample i, as x , i and Ss as the set of sample indices for which we have features from sensor s (S1 S2 = {1, ..., L + U }). Let O = S1 S2 be the indices for which both sensors are available, and OU = O {L + 1, ..., L + U } the unlabelled subset of O. By using the samples in S1 and S2 as two independent training sets, we may obtain two sep- arate classifiers (denoted w1 and w2). However, we can coordinate the information from both sensors by using an idea known as co-training [4]: on the OU samples, classifiers w1 and w2 should agree as much as possible. Notice that, in a logistic regression framework, the disagreement between the two classifiers on the OU samples can be measured by (1) (2) [(w1)T x - (w2)T x ]2 = T C , (8) iOU i i where = [(w1)T (w2)T ]T and C = [(x1)T (-x2)T ]T [(x1)T (-x2)T ]. This iOU i i i i suggests the "co-training prior" (where co is an inverse variance): p(w1, w2) = p() exp -(co/2) TC . (9) This Gaussian prior can be combined with two smoothness Gaussian priors on w1 and w2 (obtained as described in Section 3.2); this leads to a prior which is still Gaussian, p(w1, w2) = p() exp -(1/2) T coC + block-diag{1, 2} , (10) where 1 and 2 are the two graph-based precision matrices (see (7)) for w1 and w2. We can again adopt a Gamma hyper-prior for co. Under this prior, and with a logistic regression likelihood as above, estimates of w1 and w2 can easily be found using minor modifications to the EM algorithm described in Section 4. Computationally, this is only slightly more expensive than separately training the two classifiers. 4 Learning Via EM To find the MAP estimate w, we use the EM algorithm, with as missing data, which is equivalent to integrating out from the full posterior before maximization [8]. For simplicity, we will only describe the single sensor case (no co-training). E-step: We compute the expected value of the complete log-posterior, given Y and the current parameter estimate w: Q(w|w) E[log p(w, |Y)|w]. Since log p(w, |Y) = log p(Y|w) - (1/2)wT ()w + K, (11) (where K collects all terms independent of w) is linear w.r.t. all the parameters (see (6) and (7)), we just have to plug their conditional expectations into (11): Q(w|w) = log p(Y|w) - (1/2)wT E[()|w] w = (w) - (1/2)wT (w) w. (12) We consider several different choices for the structure of the matrix. The necessary expectations have well-known closed forms, due to the use of conjugate Gamma hyper- (c) priors [3]. For example, if the are m - 1 free non-negative parameters, we have 0 (c) (c) E[ |w] = (2 0 0 0 + d) [2 0 + (w(c))T Aw(c)]-1. (c) for c = 1, ..., m - 1. For = 0 0, we still have a simple closed-form expres- (c) sion for E[0|w], and the same is true for the parameters, for i > 0. Finally, i (w) E[()|w] results from replacing the 's in (7) by the corresponding conditional expectations. M-step: Given matrix (w), the M-step reduces to a logistic regression problem with a quadratic regularizer, i.e., maximizing (12). To this end, we adopt the bound optimization approach (see details in [5, 11]). Let B be a positive definite matrix such that -B bounds below (in the matrix sense) the Hessian of (w), which is negative definite, and g(w) is the gradient of (w). Then, we have the following lower bound on Q(w|w): Q(w|w) l(w) + (w - w)T g(w) - [(w - w)T B(w - w) + wT (w)w]/2. - The maximizer of this lower bound, wnew = (B + (w)) 1 (Bw + g(w)), is guaranteed to increase the Q-function, Q(wnew|w) Q(w|w), and we thus obtain a monotonic gen- eralized EM algorithm [5, 11]. This (maybe costly) matrix inversion can be avoided by a sequential approach where we only maximize w.r.t. one element of w at a time, preserving the monotonicity of the procedure. The sequential algorithm visits one particular element of w, say wu, and updates its estimate by maximizing the bound derived above, while keeping all other variables fixed at their previous values. This leads to - wnew = w ] [(B + (w)) 1 , u u + [gu(w) - ((w)w) (13) u uu] and wnew = w v v , for v = u. The total time required by a full sweep for all u = 1, ..., d is O(md(L + d)); this may be much better than the O((dm)3) of the matrix inversion. 5 Active Label Selection If we are allowed to obtain the label for one of the unlabelled samples, the following ques- tion arises: which sample, if labelled, would provide the most information? Consider the MAP estimate w provided by EM. Our approach uses a Laplace approxima- tion of the posterior p(w|Y) N (w|w, H-1), where H is the posterior precision matrix, i.e., the Hessian of minus the log-posterior H = 2(- log p(w|Y)). This approximation is known to be accurate for logistic regression under a Gaussian prior [14]. By treating (w) (the expectation of ()) as deterministic, we obtain an evidence-type approximation [14] H = 2[- log(p(Y|w)p(w|(w)))] = (w) + L (diag{p } - p pT ) x , i=1 i i i ixT i where pi is the (m - 1)-dimensional vector computed from (1), the c-th element of which indicates the probability that sample xi belongs to class c. Now let x DU be an unlabelled sample and y its label. Assume that the MAP esti- mate w remains unchanged after including y. In Sec. 7 we will discuss the merits and shortcomings of this assumption, which is only strictly valid when L . Accepting it implies that after labeling x, and regardless of y, the posterior precision changes to H = H + (diag{p} - ppT ) xxT . (14) Since the entropy of a Gaussian with precision H is (-1/2) log |H| (up to an additive constant), the mutual information (MI) between y and w (i.e., the expected decrease in entropy of w when y is observed) is I(w; y) = (1/2) log {|H |/|H|}. Our criterion is then: the best sample to label is the one that maximizes I(w; y). Further insight into I(w; y) can be obtained in the binary case (where p is a scalar); here, the matrix identity |H + p(1 - p)xxT | = |H|(1 + p(1 - p)xT H-1x) yields I(w; y) = (1/2) log(1 + p(1 - p)xT H-1x). (15) This MI is larger when p 0.5, i.e., for samples with uncertain classifications. On the other hand, with p fixed, I(w; y) grows with xT H-1x, i.e., it is large for samples with high variance of the corresponding class probability estimate. Summarizing, (15) favors samples with uncertain class labels and high uncertainty in the class probability estimate. 6 Experimental Results We begin by presenting two-dimensional synthetic examples to visually illustrate our semi- supervised classifier. Fig. 1 shows the utility of using unlabelled data to improve the deci- Figure 1: Synthetic two-dimensional examples. (a) Comparison of the supervised logistic linear classifier (boundary shown as dashed line) learned only from the labelled data (shown in color) with the proposed semi-supervised classifier (boundary shown as solid line) which also uses the unlabelled samples (shown as dots). (b) A RBF kernel classifier obtained by our algorithm, using two labelled samples (shaded circles) and many unlabelled samples. Figure 2: (a)-(c) Accuracy (on UCI datasets) of the proposed method, the supervised SVM, and the other semi-supervised classifiers mentioned in the text; a subset of samples is la- belled and the others are treated as unlabelled samples. In (d), a separate holdout set is used to evaluate the accuracy of our method versus the amount of labelled and unlabelled data. sion boundary in linear and non-linear (kernel) classifiers (see figure caption for details). Next we show results with linear classifiers on three UCI benchmark datasets. Results with nonlinear kernels are similar, and therefore omitted to save space. We compare our method against state-of-the-art semi-supervised classifiers: the GRF method of [18], the SGT method of [10], and the transductive SVM (TSVM) of [9]. For reference, we also present results for a standard SVM. To avoid unduly helping our method, we always use a k=5 nearest neighbors graph, though our algorithm is not very sensitive to k. To avoid disadvantaging other methods that do depend on such parameters, we use their best settings. Since these adjustments cannot be made in practice, the difference between our algorithm and the others is under-represented. Each point on the plots in Fig. 2(a)-(c) is an average of 20 trials: we randomly select 20 labelled sets which are used by every method. All remaining samples are used as unlabelled by the semi-supervised algorithms. Figs. 2(a)-(c) are transductive, in the sense that the unlabelled and test data are the same. Our logistic GRF is non-transductive: after being trained, it may be applied to classify new data without re-training. In Fig. 2(d) we present non-transductive results for the Ionosphere data. Training took place using labelled and unlabelled data, and testing was performed on 200 new unseen samples. The results suggest that semi-supervised classifiers are most relevant when the labelled set is small relative to the unlabelled set (as is often the case). Our final set of results address co-training (Sec. 3.3) and active learning (Sec. 5), applied to airborne sensing data for the detection of surface and subsurface land mines. Two sensors were used: (1) a 70-band hyper-spectral electro-optic (EOIR) sensor; (2) an X-band syn- thetic aperture radar (SAR). A simple (energy) "prescreener" detected potential targets; for each of these, two feature vectors were extracted, of sizes 420 and 9, for the EOIR and SAR sensors, respectively. 123 samples have features from the EOIR sensor alone, 398 from the Figure 3: (a) Land mine detection ROC curves of classifiers designed using only hyper- spectral (EOIR) features, only SAR features, and both. (b) Number of landmines detected during the active querying process (dotted lines), for active training and random selection (for the latter the bars reflect one standard deviation about the mean). ROC curves (solid) are for the learned classifier as applied to the remaining samples. SAR sensor alone, and 316 from both. This data will be made available upon request. We first consider supervised and semi-supervised classification. For the purely supervised case, a sparseness prior is used (as in [14]). In both cases a linear classifier is employed. For the data for which only one sensor is available, 20% of it is labelled (selected randomly). For the data for which both sensors are available, 80% is labelled (again selected randomly). The results presented in Fig. 3(a) show that, in general, the semi-supervised classifiers outperform the corresponding supervised ones, and the classifier learned from both sensors is markedly superior to classifiers learned from either sensor alone. In a second illustration, we use the active-learning algorithm (Sec. 5) to only acquire the 100 most informative labels. For comparison, we also show average results over 100 in- dependent realizations for random label query selection (error bars indicate one standard deviation). The results in Fig. 3(b) are plotted in two stages: first, mines and clutter are se- lected during the labeling process (dashed curves); then, the 100 labelled examples are used to build the final semi-supervised classifier, for which the ROC curve is obtained using the remaining unlabelled data (solid curves). Interestingly, the active-learning algorithm finds almost half of the mines while querying for labels. Due to physical limitations of the sen- sors, the rate at which mines are detected drops precipitously after approximately 90 mines are detected -- i.e., the remaining mines are poorly matched to the sensor physics. Balaji Krishnapuram, Ya Xue, Alexander J. Hartemink, Lawrence Carin, Mário A. T. Figueiredo |
NIPS | 5 |
| 2004 | A Bayesian Approach to Joint Feature Selection and Classifier DesignabstractThis paper adopts a Bayesian approach to simultaneously learn both an optimal nonlinear classifier and a subset of predictor variables (or features) that are most relevant to the classification task. The approach uses heavy-tailed priors to promote sparsity in the utilization of both basis functions and features; these priors act as regularizers for the likelihood function that rewards good classification on the training data. We derive an expectation-maximization (EM) algorithm to efficiently compute a maximum a posteriori (MAP) point estimate of the various parameters. The algorithm is an extension of recent state-of-the-art sparse Bayesian classifiers, which in turn can be seen as Bayesian counterparts of support vector machines. Experimental comparisons using kernel classifiers demonstrate both parsimonious feature selection and excellent classification accuracy on a range of synthetic and benchmark data sets. Balaji Krishnapuram, Alexander J. Hartemink, Lawrence Carin, Mário A. T. Figueiredo |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | Application of the Theory of Optimal Experiments to Adaptive Electromagnetic-Induction Sensing of Buried TargetsabstractA mobile electromagnetic-induction (EMI) sensor is considered for detection and characterization of buried conducting and/or ferrous targets. The sensor may be placed on a robot and, here, we consider design of an optimal adaptive-search strategy. A frequency-dependent magnetic-dipole model is used to characterize the target at EMI frequencies. The goal of the search is accurate characterization of the dipole-model parameters, denoted bythe vector theta; the target position and orientation are a subset of theta. The sensor position and operating frequency are denoted by the parameter vector p and a measurement is represented by the pair (p, O), where O denotes the observed data. The parametersp are fixed for a given measurement, but, in the context of a sequence of measurements p may be changed adaptively. In a locally optimal sequence of measurements, we desire the optimal sensor parameters, P(N+1) for estimation of theta, based on the previous measurements (p(n), On)n=1,N. The search strategy is based on the theory of optimal experiments, as discussed in detail and demonstrated via several numerical examples. Xuejun Liao, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Three-dimensional inverse scattering of a dielectric target embedded in a lossy half-spaceabstractA modified iterative Born method is applied for three-dimensional inversion of a lossless dielectric target embedded in a lossy half-space. The forward solver employs a modified form of the extended Born method, and the half-space Green's function is computed efficiently via the complex-image technique. Example results are shown, with all scattering data based on a computational model, utilizing a rigorous forward solver distinct from that employed in the inversion. In addition, distinct gridding schemes are used in the forward and inverse solvers. Simple Tikhonov regularization is found to yield adequate results for inversion of noisy data. Yijun Yu 0004, Tiejun Yu, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2004 | Detection of buried targets via active selection of labeled data: application to sensing subsurface UXOabstractWhen sensing subsurface targets, such as landmines and unexploded ordnance (UXO), the target signatures are typically a strong function of environmental and historical circumstances. Consequently, it is difficult to constitute a universal training set for design of detection or classification algorithms. In this paper, we develop an efficient procedure by which information-theoretic concepts are used to design the basis functions and training set, directly from the site-specific measured data. Specifically, assume that measured data (e.g., induction and/or magnetometer) are available from a given site, unlabeled in the sense that it is not known a priori whether a given signature is associated with a target or clutter. For N signatures, the data may be expressed as {x/sub i/,y/sub i/}/sub i=1,N/, where x/sub i/ is the measured data for buried object i, and y/sub i/ is the associated unknown binary label (target/nontarget). Let the N x/sub i/ define the set X. The algorithm works in four steps: 1) the Fisher information matrix is used to select a set of basis functions for the kernel-based algorithm, this step defining a set of n signatures B/sub n//spl sube/X that are most informative in characterizing the signature distribution of the site; 2) the Fisher information matrix is used again to define a small subset X/sub s//spl sube/X, composed of those x/sub i/ for which knowledge of the associated labels y/sub i/ would be most informative in defining the weights for the basis functions in B/sub n/; 3) the buried objects associated with the signatures in X/sub s/ are excavated, yielding the associated labels y/sub i/, represented by the set Y/sub s/; and 4) using B/sub n/,X/sub s/, and Y/sub s/, a kernel-based classifier is designed for use in classifying all remaining buried objects. This framework is discussed in detail, with example results presented for an actual buried-UXO site. Yan Zhang 0025, Xuejun Liao, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2004 | Application of the biorthogonal multiresolution time-domain method to the analysis of elastic-wave interactions with buried targetsabstractThe biorthogonal multiresolution time-domain (Bi-MRTD) method is introduced for the analysis of elastic-wave interaction with buried targets. We provide a detailed discussion on implementation of the perfectly matched layer and on treatment of the interface between two different materials. The algorithm has also been parallelized by the use of the message-passing interface. The numerical results show that numerical dispersion can be significantly improved by using biorthogonal wavelets as bases, as compared to the conventional pulse expansion employed in the finite-difference time-domain (FDTD) method. We demonstrate that with comparison to the second-order FDTD, the Bi-MRTD yields significant CPU time and memory savings for large problems, for a fixed level of accuracy. Xianyang Zhu, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2003 | Rate-Distortion Bound for Joint Compression and ClassificationabstractSummary form only given. Rate-distortion theory is applied to the problem of joint compression and classification. A Lagrangian distortion measure is used to consider both the Euclidean error in reconstructing the original data as well as the classification performance. The bound is calculated based on an alternating-minimization procedure, representing an extension of the Blahut-Arimoto algorithm. A hidden Markov model (HMM) source was considered as an example application and the objective is to quantize the source outputs and estimate the underlying HMM state sequence. Bounds on the minimum rate are required was presented to achieve desired average distortion on signal reconstruction and state-estimation accuracy. Yanting Dong, Lawrence Carin |
DCC | 2 |
| 2003 | Context-based graphical modeling for wavelet domain signal processingabstractWavelet-domain hidden Markov tree (HMT) modeling provides a powerful approach to capture the underlying statistics of the wavelet coefficients. We develop a mutual information-based information-theoretic approach to quantify the interactions between the wavelet coefficients within a wavelet tree. This graphical method enables the design of a context-specific hidden Markov tree (HMT) by adding or deleting links from the traditional tree structure. The performance of the model is demonstrated on segmenting two-dimensional synthetic textures having intricate substructures, although the method can be used for signals of arbitrary dimensions. Nilanjan Dasgupta, Lawrence Carin |
ICASSP (3) | 2 |
| 2003 | ICA with multiple quadratic constraintsabstractThe independent component analysis (ICA) with a single quadratic constraint on each source signal or column of the mixing matrix is extended to the case of multiple quadratic constraints. The criterion of Joint Approximate Diagonalization of Eigen-matrices (JADE) is used to measure the statistical independence. A new algorithm is derived to maximize the JADE criterion subject to the multiple quadratic constraints, using the augmented Lagrangian method. The extension offers the freedom to design various combinations of quadratic constraints. Examples include simultaneously constraining a source signal and the corresponding column of the mixing matrix, and two-sided constraints on the source signals or columns of the mixing matrix. Example results are provided to demonstrate the effectiveness of the algorithm. Xuejun Liao, Lawrence Carin |
ICASSP (5) | 2 |
| 2003 | Time-reversal imaging for wideband underwater target classificationabstractTime-reversal imaging is addressed for sensing an elastic target situated in an acoustic waveguide. It is demonstrated that the channel parameters associated with a given measurement may be determined via a genetic-algorithm (GA) search in parameter space. Target classification based on time-reversal imagery is considered, with this implemented via a relevance-vector machine. Nilanjan Dasgupta, Lawrence Carin |
ICASSP (5) | 3 |
| 2003 | Joint classifier and feature optimization for cancer diagnosis using gene expression dataabstractRecent research has demonstrated quite convincingly that accurate cancer diagnosis can be achieved by constructing classifiers that are designed to compare the gene expression profile of a tissue of unknown cancer status to a database of stored expression profiles from tissues of known cancer status. This paper introduces the JCFO, a novel algorithm that uses a sparse Bayesian approach to jointly identify both the optimal nonlinear classifier for diagnosis and the optimal set of genes on which to base that diagnosis. We show that the diagnostic classification accuracy of the proposed algorithm is superior to a number of current state-of-the-art methods in a full leave-one-out cross-validation study of two widely used benchmark datasets. In addition to its superior classification accuracy, the algorithm is designed to automatically identify a small subset of genes (typically around twenty in our experiments) that are capable of providing complete discriminatory information for diagnosis. Focusing attention on a small subset of genes is not only useful because it produces a classifier with good generalization capacity, but also because this set of genes may provide insights into the mechanisms responsible for the disease itself. A number of the genes identified by the JCFO in our experiments are already in use as clinical markers for cancer diagnosis; some of the remaining genes may be excellent candidates for further clinical investigation. If it is possible to identify a small set of genes that is indeed capable of providing complete discrimination, inexpensive diagnostic assays might be widely deployable in clinical settings. Balaji Krishnapuram, Lawrence Carin, Alexander J. Hartemink |
RECOMB | 2 |
| 2003 | Rate-Distortion Analysis of Discrete-HMM Pose Estimation via Multiaspect Scattering DataabstractWe consider the problem of estimating the pose of a target based on a sequence of scattered waveforms measured at multiple target-sensor orientations. Using a hidden Markov model (HMM) representation of the scattered-waveform sequence, pose estimation reduces to estimating the underlying HMM states from a sequence of observations. It is assumed that each scattered waveform must be quantized via an encoding procedure. A distortion D is defined as the error in estimating the underlying HMM states, and the rate R represents the size of the discrete-HMM codebook. Rate-distortion theory is applied to define the minimum rate required to achieve a desired distortion, denoted as R(D). After deriving the rate-distortion function R(D), we demonstrate that discrete-HMM performance based on Lloyd encoding is far from this bound. Performance is improved via block coding, based on Bayes VQ. Results are presented for a canonical HMM problem, and then for multiaspect acoustic scattering from underwater elastic targets. Although the examples presented here are for multiaspect scattering and pose estimation, the results are of general applicability to discrete-HMM state estimation. Yanting Dong, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Scalable multilevel fast multipole method for multiple targets in the vicinity of a half spaceabstractWe extend the multilevel fast multipole algorithm (MLFMA) to the case of electromagnetic scattering from an arbitrary number of dielectric and/or perfectly conducting targets in the presence of a half space. This multitarget MLFMA is implemented in an iterative fashion, in which the fields incident on and scattered from each target are updated sequentially by considering each target in isolation, with appropriate field updating to account for intertarget scattering. Each target is analyzed in parallel on a separate computer node, and intertarget interaction is addressed via message passaging between the processors. We also utilize the aforementioned iterative formulation employed for handling interactions between multiple targets to develop a new means of solving the MLFMA matrix equation for an isolated target. This new formulation generally results in significant acceleration in the analysis of scattering from single targets, thereby also accelerating the analysis of scattering from multiple targets (within the context of the iterative multitarget analysis developed). Xiaolong Dong, James A. Thompson, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2003 | Sensing of unexploded ordnance with magnetometer and induction data: theory and signal processingabstractWe consider the detection of subsurface unexploded ordnance via magnetometer and electromagnetic-induction (EMI) sensors. Detection performance is presented, using model-based signal processing algorithms. We first develop and validate the parametric models, using both numerical and measured data. These models are then applied in the context of feature extraction, and the features are processed via two signal-processing algorithms. The detection algorithms are discussed in detail, with comparisons made based on performance with measured magnetometer and EMI data. Yan Zhang 0025, Leslie M. Collins, Haitao Yu 0014, Carl E. Baum, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2002 | Infrared-image classification using support vector machinesabstractA target recognition classifier for forward-looking infrared (FUR) imagery is developed. A target class is defined as a set of contiguous target-sensor orientations (aspects) for which the associated FLIR imagery is stationary. We designed four sets of templates for each target class, to represent the overall image as well as three class-dependent subcomponents. The templates are designed by using expansion matching (EXM) filters and the Karhunen-Loeve transform (KLT). The feature vectors obtained with these eigen templates are used in the context of a support vector machine (SVM). The performance of the SVM classifier is presented and compared with other competitive classifiers. Shaorong Chang, Nasser M. Nasrabadi, Lawrence Carin |
ICASSP | 3 |
| 2002 | Support Vector Machines for improved multiaspect target recognition using the fisher kernel scores of Hidden Markov ModelsabstractIn conjunction with physics-based feature extraction, Hidden Markov Model. (HMM) classifiers have been used successfully to fuse scattering data from multiple target orientations where the target-sensor orientation is generally unknown or “hidden” [1]. The use of prior knowledge concerning sensor motion is employed in modeling the sequential data, improving classification performance. However, the assumptions of first order Markovian state transitions state-dependent statistics constrain the intrinsic class of pdf structures admitted by the HMM, for use in classification. In-this paper we overcome the above limitation by using the local variations in the HMMs induced by each sequence of observations as the feature vector for a support vector machine. (SYM) classifier. Improved discrimination results are presented for measured acoustic scattering data. Balaji Krishnapuram, Lawrence Carin |
ICASSP | 2 |
| 2002 | ICA and PLS modeling for functional analysis and drug sensitivity for DNA microarray signalsabstractThe DNA microarray technique offers an ability to analyze the expression profile of a genome. The complex correlation between the large number of genes present in the genome undermines straightforward understanding of their functionality. In this paper, we have proposed a pair of modeling schemes to recognize the functional identities of the known genes. In Independent Component Analysis (ICA), each of the microarray signals is modeled as a linear combination of some underlying independent components having specific biological interpretation. The second algorithm, Partial Least Squares (PLS) is proposed to identify the latent functional units contributing to drug sensitivity from the microarray data. Applications of this research include prediction of drug responses based on gene expressions, and also to identify the function(s) of a new gene. We consider the Rosetta compendium data set with yeast gene profiles, and the NCI-60 data set of human gene expressions as a function of drug type (cancer drugs are considered). Xuejun Liao, Nilanjan Dasgupta, Simon M. Lin, Lawrence Carin |
ICASSP | 4 |
| 2002 | Class-based target classification in shallow water channel based on Hidden Markov ModelabstractThis paper presents a class-based classification approach for targets in a shallow water channel, based on a waveguide propagation model and a Hidden Markov Model (HMM). We utilize the time-frequency properties of wave propagation in a shallow water channel to extract the channel-parameter-independent features, then build a class-based HMM to realize target classification based on multi-aspect scattering data. The accuracy of this approach is demonstrated by simulated results. Lawrence Carin |
ICASSP | 2 |
| 2002 | HMM-based multiresolution image segmentationabstractA texture segmentation algorithm is developed, utilizing a wavelet-based multi-resolution analysis of general imagery. The wavelet analysis yields a set of quadtrees, each composed of high-high (HH), high-low (HL) and low-high (LH) wavelet coefficients. Hidden Markov trees (HMTs) are designed for the quadtrees. For a given texture we define a set of “hidden” states, and a hidden Markov model (HMM) is developed to characterize the statistics of a given quadtree with respect to the statistics of surrounding quadtrees. Each HMM state is characterized by a unique set of HMTs. An HMM-HMT model is developed for each texture of interest, with which image segmentation is achieved. Several numerical examples are presented to demonstrate the model, with comparisons to alternative approaches. Jiuliu Lu, Lawrence Carin |
ICASSP | 2 |
| 2002 | Model-based statistical sensor fusion for unexploded ordnance detectionabstractDetection and remediation of unexploded ordnance (UXO) represents a major challenge on closed, closing, and transferred military ranges as well as on active installations. The detection problem is exacerbated by the fact that on sites contaminated with UXO, extensive surface and sub-surface clutter and shrapnel is also present. Traditional methods used for UXO remediation have difficulty distinguishing buried UXO from these anthropic clutter items as well as from naturally occurring magnetic geologic noise, and thus incur prohibitively high false alarm rates. The reduction of the false alarm rate has proven to be the greatest challenge for UXO remediation. In this paper, sensor fusion techniques are applied to field data from magnetometer and electromagnetic induction (EMI) sensors in order to determine to what degree such an approach results in false alarm mitigation. The adoption of a model consisting of multiple non-colocated dipoles is shown to improve our ability to predict measured signatures. The results indicate that performance can be improved by limiting the processing bandwidth to those frequencies that are the most robust to naturally occurring geological noise. Leslie M. Collins, Yizhe Zhang 0002, Lawrence Carin |
IGARSS | 3 |
| 2002 | Infrared-Image Classification Using Hidden Markov TreesabstractAn image of a three-dimensional target is generally characterized by the visible target subcomponents, with these dictated by the target-sensor orientation (target pose). An image often changes quickly with variable pose. We define a class as a set of contiguous target-sensor orientations over which the associated target image is relatively stationary with aspect. Each target is in general characterized by multiple classes. A distinct set of Wiener filters are employed for each class of images, to identify the presence of target subcomponents. A Karhunen-Loeve representation is used to minimize the number of filters (templates) associated with a given subcomponent. The statistical relationships between the different target subcomponents are modeled via a hidden Markov tree (HMT). The HMT classifier is discussed and example results are presented for forward-looking-infrared (FLIR) imagery of several vehicles. Priya Bharadwaj, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Sequential modeling for identifying CpG island locations in human genomeabstractWe consider several sequential processing algorithms for identifying genes in human DNA, based on detecting CpG ("C proceeds G") islands. The algorithms are designed to capture the underlying statistical structure in a DNA sequence. Sequential processing using a Markov model and a hidden Markov model are shown to identify most CpG islands in annotated (marked) DNA subsequences available from publicly available DNA datasets. We also consider a wavelet-based hidden Markov tree (HMT). In the context of the HMT, we address design of adaptive wavelets matched to CpG islands, this accomplished via lifting and genetic-algorithm optimization. Nilanjan Dasgupta, Simon M. Lin, Lawrence Carin |
IEEE Signal Process. Lett. | 3 |
| 2001 | Infrared-image classification using expansion matching filters and hidden Markov treesabstractForward-looking infrared (FLIR) images of targets are characterized by the different target components visible in the image, with are dependent on the target-sensor orientation and target history (i.e., which target components are hot). We define a target class as a set of contiguous target-sensor orientations over which the associated image is relatively invariant, or statistically stationary. Given an image from an unknown target, the objective is proper target-class association (target identity and pose). Our principal contribution is an image classifier in which a distinct set of templates is designed for each image class, with templates linked to the object sub-components, and the associated statistics are characterized via a hidden Markov model. In particular, we employ expansion matching (EXM) filters to identify the presence of the target components in the image, and use a hidden Markov tree (HMT) to characterize the statistics of the correlation of the image with the various templates. We achieve a successful classification rate of 92% on a data set of FLIR vehicle images, compared with 72% for a previously developed wavelet-feature-based HMT technique. Priya Bharadwaj, Paul Runkle, Lawrence Carin |
ICASSP | 3 |
| 2001 | Class-based identification of underwater targets using hidden Markov modelsabstractIt has been demonstrated that hidden Markov models (HMM) provide an effective architecture for classification of distinct targets from multiple target-sensor orientations. We present a methodology for designing class-based HMM that are well suited to the identification of targets with common physical attributes. This approach provides a means to form associations between existing target classes and data from targets never observed in training. After performing a wavefront-resonance matching-pursuits feature extraction, we present an information theoretic tree-based state-parsing algorithm to define the HMM state structure for each target class. In training, class association is determined by minimizing the statistical divergence between the target under consideration and each existing class, with a new class defined when the target is poorly matched to each existing class The class-based HMM are trained with data from the members of its corresponding class, and tested on previously unobserved data. Results are presented for simulated acoustic scattering data. Nilanjan Dasgupta, Paul Runkle, Lawrence Carin |
ICASSP | 3 |
| 2001 | Markov modeling of transient scattering and its application in multi-aspect target classificationabstractTransient scattered fields from a general target are composed of wavefronts, resonances and time delays, with these constituents linked to the target geometry. A classifier applied to transient scattering data requires a statistical model for such fundamental constituents. A Markov model is employed to characterize the transient scattered fields-for a set of target-sensor orientation over which the transient scattering is stationary-utilizing a wavefront, resonance, time-delay "alphabet". The Markov model is utilized in a classifier developed for multi-aspect transient scattering data, with a hidden Markov model (HMM) employed to address the generally non-stationary nature of the multi-aspect waveforms. Each state of the HMM is characteristic of a set of target-sensor orientations for which the scattering statistics are stationary, the statistics of which are characterized via the Markov model. The wavefront, resonance and time-delay features are extracted via a modified matching-pursuits algorithm. Yanting Dong, Paul Runkle, Lawrence Carin |
ICASSP | 3 |
| 2001 | Identification of ground targets from sequential HRR radar signaturesabstractAn approach to identifying ground targets from sequential high-range-resolution (HRR) radar signatures is presented. A hidden Markov model (HMM) is employed to model the sequential information contained in multi-aspect target signatures. Dominant range-amplitude features are extracted via RELAX for dimension reduction. A new distance measure is incorporated into the HMM to allow a direct matching operation in the feature domain without requiring interpolation. The approach is applied to the dataset of ten MSTAR targets and is shown to yield an average identification rate of 90.3% using sequential information from 6 degree angular spans. Xuejun Liao, Paul Runkle, Yan Jiao, Lawrence Carin |
ICASSP | 4 |
| 2001 | Genetic Algorithm Wavelet Design for Signal ClassificationabstractBiorthogonal wavelets are applied to parse multiaspect transient scattering data in the context of signal classification. A language-based genetic algorithm is used to design wavelet filters that enhance classification performance. The biorthogonal wavelets are implemented via the lifting procedure and the optimization is carried out using a classification-based cost function. Example results are presented for target classification using measured scattering data. Paul Runkle, Nilanjan Dasgupta, Luise Couchman, Lawrence Carin |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2001 | Dual hidden Markov model for characterizing wavelet coefficients from multi-aspect scattering data
Nilanjan Dasgupta, Paul Runkle, Luise Couchman, Lawrence Carin |
Signal Process. | 4 |
| 2001 | A comparison of the performance of statistical and fuzzy algorithms for unexploded ordnance detectionabstractWe focus on the development of signal processing algorithms that incorporate the underlying physics characteristic of the sensor and of the anticipated unexploded ordnance (UXO) target, in order to address the false alarm issue. In this paper, we describe several algorithms for discriminating targets from clutter that have been applied to data obtained with the multisensor towed array detection system (MTADS). This sensor suite includes both electromagnetic induction (EMI) and magnetometer sensors. We describe four signal processing techniques: a generalized likelihood ratio technique, a maximum likelihood estimation-based clustering algorithm, a probabilistic neural network, and a subtractive fuzzy clustering technique. These algorithms have been applied to the data measured by MTADS in a magnetically clean test pit and at a field demonstration. The results indicate that the application of advanced signal processing algorithms could provide up to a factor of two reduction in false alarm probability for the UXO detection problem. Leslie M. Collins, Yan Zhang 0025, Lawrence Carin, Sean J. Hart, Susan L. Rose-Pehrsson, Herbert H. Nelson, James R. McDonald |
IEEE Trans. Fuzzy Syst. | 5 |
| 2001 | Foreword
Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2001 | On the wideband EMI response of a rotationally symmetric permeable and conducting targetabstractA simple and accurate model is presented for computation of the electromagnetic induction (EMI) resonant frequencies of canonical conducting and ferrous targets, in particular, finite-length cylinders and rings. The imaginary resonant frequencies correspond to the well known exponential decay constants of interest for time-domain EMI interaction with conducting and ferrous targets. The results of the simple model are compared to data computed numerically, via method-of-moments (MoM) and finite-element models. Moreover, the simple model is used to fit measured wideband EMI data from ferrous cylindrical targets (in terms of a small number of parameters). It is also demonstrated that the general model for the magnetic-dipole magnetization, in terms of a frequency-dependent diagonal dyadic, is applicable to general rotationally symmetric targets (not just cylinders and rings). Lawrence Carin, Haitao Yu 0014, Yacine Dalichaouch, Alexander R. Perry, Peter V. Czipott, Carl E. Baum |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2001 | Time-domain sensing of targets buried under a Gaussian, exponential, or fractal rough interfaceabstractThe authors numerically examine subsurface sensing via an ultrawideband ground penetrating radar (GPR) system. The target is assumed to reside under a randomly rough air-ground interface and is illuminated by a pulsed plane wave. The underlying wave physics is addressed through application of the multiresolution time-domain (MRTD) algorithm. The scattered time-domain fields are parametrized as a random process and an optimal detection scheme is formulated, accounting for the clutter and target signature statistics. Detector performance is evaluated via receiver operating characteristics (ROCs) for variable sensor parameters (polarization and incident angle) and for several rough-surface statistical models. Traian Dogaru, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2001 | Multi-aspect detection of surface and shallow-buried unexploded ordnance via ultra-wideband synthetic aperture radarabstractAn ultra-wideband (UWB) synthetic aperture radar (SAR) system is investigated for the detection of former bombing ranges, littered by unexploded ordnance (UXO). The objective is detection of a high enough percentage of surface and shallow-buried UXO, with a low enough false-alarm rate, such that a former range can be detected. The physics of UWB SAR scattering is exploited in the context of a hidden Markov model (HMM), which explicitly accounts for the multiple aspects at which a SAR system views a given target. The HMM is trained on computed data, using SAR imagery synthesized via a validated physical-optics solution. The performance of the HMM is demonstrated by performing testing on measured UWB SAR data for many surface and shallow UXO buried in soil in the vicinity of naturally occurring clutter. Yanting Dong, Paul Runkle, Lawrence Carin, Raju Damarla, Anders Sullivan, Marc A. Ressler, Jeffrey Sichina |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2001 | Rigorous modeling of ultrawideband VHF scattering from tree trunks over flat and. sloped terrainabstractThree electromagnetic models are employed for the investigation of ultrawideband VHF scattering from tree trunks situated over flat and sloped terrain. Two of the models are numerical, each employing a frequency-domain integral-equation formulation solved via the method of moments (MoM). A body-of-revolution (BoR) Mote formulation is applied for a tree trunk on a flat terrain, implying that the BoR axis is perpendicular to the layers of an arbitrary layered-earth model. For the case of sloped terrain, the BoR model is inapplicable, and therefore the MoM solution is performed via general triangular-patch basis functions. Both MoM models are very accurate but are computationally expensive. Consequently, the authors also consider a third model, employing approximations based on the closed-form solution for scattering from an infinite dielectric cylinder in free space. The third model is highly efficient computationally and, despite the significant approximations, often yields accurate results relative to data computed via the reference MoM solutions. Data from the three models are considered, and several examples of application to remote sensing are addressed. Jiangqi He, Norbert Geng, Lam H. Nguyen, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2001 | Multi-aspect target detection for SAR imagery using hidden Markov modelsabstractRadar scattering from an illuminated object is often highly dependent on the target-sensor orientation. In typical synthetic aperture radar (SAR) imagery, the information in the multi-aspect target signatures is diffused in the image-formation process. In an effort to exploit the aspect dependence of the target signature, the authors employ a sequence of directional filters to the SAR imagery, thereby generating a sequence of subaperture images that recover the directional dependence of the target scattering. The scattering statistics are then used to design a hidden Markov model (HMM), wherein the orientation-dependent scattering statistics are exploited explicitly. This approach fuses information embodied in the orientation-dependent target signature under the assumption that. Both the target identity and orientation are unknown. Performance is assessed by considering the detection of tactical targets concealed in foliage, using measured foliage-penetrating (FOPEN) SAR data. Paul Runkle, Lam H. Nguyen, James H. McClellan, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2000 | Classification of landmine-like metal targets using wideband electromagnetic inductionabstractIn their previous work, the authors have shown that the detectability of landmines can be improved dramatically by the careful application of signal detection theory to time-domain electromagnetic induction (EMI) data using a purely statistical approach. In this paper, classification of various metallic land-mine-like targets via signal detection theory is investigated using a prototype wideband frequency-domain EMI sensor. An algorithm that incorporates both a theoretical model of the response of such a sensor and the uncertainties regarding the target/sensor orientation is developed. This allows the algorithms to be trained without an extensive data collection. The performance of this approach is evaluated using both simulated and experimental data. The results show that this approach affords substantial classification performance gains over a standard approach, which utilizes the signature obtained when the sensor is centered over the target and located at the mean expected target/sensor distance, and thus ignores the uncertainties inherent in the problem. On the average, a 60% improvement is obtained. Ping Gao 0007, Leslie M. Collins, Philip M. Garber, Norbert Geng, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2000 | Multilevel fast-multipole algorithm for scattering from conducting targets above or embedded in a lossy half spaceabstractAn extension of the multilevel fast multipole algorithm (MLFMA), originally developed for targets in free space, is presented for the electromagnetic scattering from arbitrarily shaped three-dimensional (3-D), electrically large, perfectly conducting targets above or embedded within a lossy half space. We have developed and implemented electric-field, magnetic-field, and combined-field integral equations for this purpose. The nearby terms in the MLFMA framework are evaluated by using the rigorous half-space dyadic Green's function, computed via the method of complex images. Non-nearby (far) MLFMA interactions, handled efficiently within the multilevel clustering construct, employ an approximate dyadic Green's function. This is expressed in terms of a direct-radiation term plus a single real image (representing the asymptotic far-field Green's function), with the image amplitude characterized by the polarization-dependent Fresnel reflection coefficient. Examples are presented to validate the code through comparison with a rigorous method-of-moments (MoM) solution. Finally, results are presented for scattering from a model unexploded ordnance (UXO) embedded in soil and for a realistic 3-D vehicle over soil. Norbert Geng, Anders Sullivan, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2000 | Analysis of the electromagnetic inductive response of a void in a conducting-soil backgroundabstractA lossless dielectric object situated in a lossy dielectric medium (soil) constitutes a void in a conducting background, which can be detected via an electromagnetic-induction (EMI) sensor operating at appropriate frequencies. The electromagnetic character of this void is dependent on the target and soil properties, as well as on the frequency of operation. We utilize the rigorous method of moments (MoM) and the approximate extended-Born technique to model this three-dimensional (3-D) problem. The modeling algorithms are discussed in detail, with a focus on efficient computation of the dyadic Green's function at the frequencies of interest. The MoM results are used to calibrate the accuracy of the approximate extended-Born solution, over a wide range of operating conditions. Furthermore, the computer simulations are used to perform a detailed phenomenological study. Tiejun Yu, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 1999 | Physics-based classification of targets in SAR imagery using subaperture sequencesabstractIt is well known that radar scattering from an illuminated object is often highly aspect dependent. We have developed a multiaspect target classification technique for SAR imagery that incorporates matching-pursuits feature extraction from each of a sequence of subaperture images, in conjunction with a hidden Markov model that explicitly incorporates the target-sensor motion represented by the image sequence. This approach exploits the aspect dependence of the signal features to facilitate maximum-likelihood identification. We consider SAR imagery containing targets concealed by foliage. Lawrence Carin, Gary Ybarra, Priya Bharadwaj, Paul Runkle |
ICASSP | 1 |
| 1999 | Classification of landmine-like metal targets using wideband electromagnetic inductionabstractOur previous work has indicated that the careful application of signal detection theory can dramatically improve detectability of landmines using time-domain electromagnetic induction (EMI) data. In this paper, classification of various metal targets via signal detection theory is investigated using a prototype wideband frequency-domain EMI sensor. An algorithm that incorporates both the uncertainties regarding the target-sensor orientation and a theoretical model of the response of such a sensor is developed. The performance of this approach is evaluated using both simulated and experimental data. The results show that this approach affords substantial classification performance gains over the traditional matched filter approach, on average by 60%. Ping Gao 0007, Leslie M. Collins, Norbert Geng, Lawrence Carin, Dean A. Keiswetter, I. J. Won |
ICASSP | 4 |
| 1999 | Multi-aspect target identification with wave-based matching pursuits and continuous hidden Markov modelsabstractA wave-based matching-pursuits algorithm is used to parse multi-aspect time-domain backscattering data into its underlying wavefront-resonance constituents, or features. Consequently, the N multi-aspect waveforms under test are mapped into N feature vectors, y/sub n/. Target identification is effected by fusing these N vectors in a maximum-likelihood sense, which we show, under reasonable assumptions, can be implemented via a hidden Markov model (HMM). The algorithm performance is assessed by considering measured acoustic scattering data from five similar submerged elastic targets. Paul Runkle, Lawrence Carin |
ICASSP | 2 |
| 1999 | Multiaspect Target Identification with Wave-Based Matched Pursuits and Continuous Hidden Markov ModelsabstractMultiaspect target identification is effected by fusing the features extracted from multiple scattered waveforms; these waveforms are characteristic of viewing the target from a sequence of distinct orientations. Classification is performed in the maximum-likelihood sense, which we show, under reasonable assumptions, can be implemented via a hidden Markov model (HMM). We utilize a continuous-HMM paradigm and compare its performance to its discrete counterpart. The feature parsing is performed via wave-based matched pursuits. Algorithm performance is assessed by considering measured acoustic scattering data from five similar submerged elastic targets. Paul Runkle, Lawrence Carin, Luise Couchman, Timothy J. Yoder, Joseph A. Bucaro |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1999 | An improved Bayesian decision theoretic approach for land mine detectionabstractA rigorous signal detection theoretic analysis is used to improve detectability of land mines. The development is performed for sensors that integrate time-domain information to provide a single data point (standard metal detector), those that provide a sampled portion of the time-domain waveform, and those that operate at several discrete frequencies. This approach is compared to standard thresholding techniques, and it is shown to provide substantial improvements when evaluated on measured data. Leslie M. Collins, Ping Gao 0007, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 1999 | On the low-frequency natural response of conducting and permeable targetsabstractThe low-frequency natural response of conducting, permeable targets is investigated. The authors demonstrate that the source-free response is characterized by a sum of nearly purely damped exponentials, with the damping constants strongly dependent on the target shape, conductivity, and permeability, thereby representing a potential tool for pulsed electromagnetic induction (EMI) identification (discrimination) of conducting and permeable targets. This general concept is then specialized to the particular case of a body of revolution (BOR), for which the method-of-moments (MoM)-computed natural damping constants from several targets are compared with measurements. Moreover, theoretical natural (equivalent) surface currents and damping coefficients are shown for other targets of interest. Finally, the authors investigate the practical use of such natural signatures in the context of identification, wherein Cramer-Rao bound (CRB) studies address signal-to-noise ratio (SNR) considerations. Norbert Geng, Carl E. Baum, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 1999 | Short-pulse electromagnetic scattering from arbitrarily oriented subsurface ordnanceabstractA rigorous method-of-moments (MoM) analysis is used to model wide-band scattering from general three-dimensional perfectly conducting objects buried in a lossy layered medium. The authors focus on ordnance buried in a half space (soil). The time-domain fields scattered from a tilted antitank mine are examined in detail as a function of polarization and observation position. Norbert Geng, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 1999 | Wide-band VHF scattering from a trihedral reflector situated above a lossy dispersive halfspaceabstractThe method of moments (MoM) is used to rigorously analyze wide-band VHF scattering from a perfectly conducting trihedral placed above a lossy, dispersive half space. The method of complex images is employed to evaluate the layered-medium Green's function efficiently and is applied here to the half-space problem. Particular attention is placed on the physics that underlie scattering from such targets (as a function of frequency, incidence angle, polarization, and soil type) for cases in which the target is small or of moderate size relative to wavelength (where high-frequency techniques fail). This problem is of interest when conventional trihedrals are employed to calibrate VHF synthetic aperture radar (SAR) systems. Norbert Geng, Marc A. Ressler, Lawrence Carin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 1998 | Polarimetric SAR imaging of buried landminesabstractIf the fields incident on a buried body of revolution are polarized vertically or horizontally (relative to the ground), the backscattered fields are exclusively copolarized (i.e., there are no cross-polarized backscattered fields). After substantiating this theoretically, measured ultrawideband (UWB) synthetic aperture radar (SAR) data are used for corroboration, considering real, buried landmines that approximate bodies of revolution. Lawrence Carin, Ravinder Kapoor, Carl E. Baum |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 1997 | Random Neural Network Recognition of Shaped Objects in Strong Clutter
Hakan Bakircioglu, Erol Gelenbe, Lawrence Carin |
ICANN | 3 |
| 1997 | Ultra-wideband, short-pulse ground-penetrating radar: simulation and measurementabstractUltra-wideband (UWB), short-pulse (SP) radar is investigated theoretically and experimentally for the detection and identification of targets buried in and placed atop soil. The calculations are performed using a rigorous, three-dimensional (3D) method of moments algorithm for perfectly conducting bodies of revolution. Particular targets investigated theoretically include anti-personnel mines, anti-tank mines, and a 55-gallon drum, for which the authors model the time-domain scattered fields and the buried-target late-time resonant frequencies. With regard to the latter, the computed resonant frequencies are utilized to assess the feasibility of resonance-based buried-target identification for this class of targets. The measurements are performed using a novel UWB, SP synthetic aperture radar (SAR) implemented on a mobile boom. Experimental and theoretical results are compared. Stanislav Vitebskiy, Lawrence Carin, Marc A. Ressler, Francis H. Le |
IEEE Trans. Geosci. Remote. Sens. | 2 |