VLDB 2026 Research / reviewers in the wild / expert
Jennifer G. Dy
dblp:24/6000
· DBLP profile ↗
123ranked-venue papers
5as first author
34since 2021 · last 2026
0000-0002-8430-134XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 89 · 5 first-author · 25 since 2021Databases, data management, data science and information retrieval · 30 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 since 2021Computer networks · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSplain: Sparse and Smooth Explainer for Retinopathy of Prematurity ClassificationabstractNeural networks are frequently used in medical diagnosis. However, due to their black-box nature, model explainers are used to help clinicians understand better and trust model outputs. This paper introduces an explainer method for classifying Retinopathy of Prematurity (ROP) from fundus images. Previous methods fail to generate explanations that preserve input image structures such as smoothness and sparsity. We introduce Sparse and Smooth Explainer (SSplain), a method that generates pixel-wise explanations while preserving image structures by enforcing smoothness and sparsity. This results in realistic explanations to enhance the understanding of the given black-box model. To achieve this goal, we define an optimization problem with combinatorial constraints and solve it using the Alternating Direction Method of Multipliers (ADMM). Experimental results show that SSplain outperforms commonly used explainers in terms of both post-hoc accuracy and smoothness analyses. Additionally, SSplain identifies features that are consistent with domain-understandable features that clinicians consider as discriminative factors for ROP. We also show SSplain’s generalization by applying it to additional publicly available datasets. Code is available at https://github.com/neu-spiral/SSplain. Elifnur Sunger, Tales Imbiriba, J. Peter Campbell, Deniz Erdogmus, Stratis Ioannidis, Jennifer G. Dy |
WACV | 6 |
| 2025 | Axiomatic Explainer Globalness via Optimal TransportabstractExplainability methods are often challenging to evaluate and compare. With a multitude of explainers available, practitioners must often compare and select explainers based on quantitative evaluation metrics. One particular differentiator between explainers is the diversity of explanations for a given dataset; i.e. whether all explanations are identical, unique and uniformly distributed, or somewhere between these two extremes. In this work, we define a complexity measure for explainers, globalness, which enables deeper understanding of the distribution of explanations produced by feature attribution and feature selection methods for a given dataset. We establish the axiomatic properties that any such measure should possess and prove that our proposed measure, Wasserstein Globalness, meets these criteria. We validate the utility of Wasserstein Globalness using image, tabular, and synthetic datasets, empirically showing that it both facilitates meaningful comparison between explainers and improves the selection process for explainability methods. Davin Hill, Joshua T. Bone, Aria Masoomi, Max Torop, Jennifer G. Dy |
AISTATS | 5 |
| 2025 | Linear-Time Demonstration Selection for In-Context Learning via Gradient EstimationabstractThis paper introduces an algorithm to select demonstration examples for in-context learning of a query set.Given a set of n examples, how can we quickly select k out of n to best serve as the conditioning for downstream inference?This problem has broad applications in prompt tuning and chain-of-thought reasoning.Since model weights remain fixed during in-context learning, previous work has sought to design methods based on the similarity of token embeddings.This work proposes a new approach based on gradients of the output taken in the input embedding space.Our approach estimates model outputs through a first-order approximation using the gradients.Then, we apply this estimation to multiple randomly sampled subsets.Finally, we aggregate the sampled subset outcomes to form an influence score for each demonstration, and select k most relevant examples.This procedure only requires precomputing model outputs and gradients once, resulting in a linear-time algorithm relative to model and training set sizes.Extensive experiments across various models and datasets validate the efficiency of our approach.We show that the gradient estimation procedure yields approximations of full inference with less than 1% error across six datasets.This allows us to scale up subset selection that would otherwise run full inference by up to 37.7× on models with up to 34 billion parameters, and outperform existing selection methods based on input embeddings by 11% on average. Ziniu Zhang, Zhenshuo Zhang, Lu Wang 0008, Jennifer G. Dy, Hongyang R. Zhang |
EMNLP | 5 |
| 2025 | STAR: Stability-Inducing Weight Perturbation for Continual LearningabstractHumans can naturally learn new and varying tasks in a sequential manner.
Continual learning is a class of learning algorithms that updates its learned model as it sees new data (on potentially new tasks) in a sequence.
A key challenge in continual learning is that as the model is updated to learn new tasks, it becomes susceptible to \textit{catastrophic forgetting}, where knowledge of previously learned tasks is lost. A popular approach to mitigate forgetting during continual learning is to maintain a small buffer of previously-seen samples, and to replay them during training. However, this approach is limited by the small buffer size and, while forgetting is reduced, it is still present. In this paper, we propose
a novel loss function STAR that exploits the worst-case parameter perturbation that reduces the KL-divergence of model predictions with that of its local parameter neighborhood to promote stability and alleviate forgetting. STAR can be combined with almost any existing rehearsal-based methods as a plug-and-play component. We empirically show that STAR consistently improves performance of existing methods by up to $\sim15\\%$ across varying baselines, and achieves superior or competitive accuracy to that of state-of-the-art methods aimed at improving rehearsal-based continual learning. Our implementation is available at https://github.com/Gnomy17/STAR_CL. Masih Eskandar, Tooba Imtiaz, Davin Hill, Zifeng Wang 0002, Jennifer G. Dy |
ICLR | 5 |
| 2025 | OrdShap: Feature Position Importance for Sequential Black-Box ModelsabstractSequential deep learning models excel in domains with temporal or sequential dependencies, but their complexity necessitates post-hoc feature attribution methods for understanding their predictions. While existing techniques quantify feature importance, they inherently assume fixed feature ordering — conflating the effects of (1) feature values and (2) their positions within input sequences. To address this gap, we introduce OrdShap, a novel attribution method that disentangles these effects by quantifying how a model's predictions change in response to permuting feature position. We establish a game-theoretic connection between OrdShap and Sanchez-Bergantiños values, providing a theoretically grounded approach to position-sensitive attribution. Empirical results from health, natural language, and synthetic datasets highlight OrdShap's effectiveness in capturing feature value and feature position attributions, and provide deeper insight into model behavior. Davin Hill, Brian L. Hill, Aria Masoomi, Vijay S. Nori, Robert E. Tillman, Jennifer G. Dy |
NeurIPS | 6 |
| 2025 | H-SPLID: HSIC-based Saliency Preserving Latent Information DecompositionabstractWe introduce H-SPLID, a novel algorithm for learning salient feature representations through the explicit decomposition of salient and non-salient features into separate spaces. We show that H-SPLID promotes learning low-dimensional, task-relevant features. We prove that the expected prediction deviation under input perturbations is upper-bounded by the dimension of the salient subspace and the Hilbert-Schmidt Independence Criterion (HSIC) between inputs and representations. This establishes a link between robustness and latent representation compression in terms of the dimensionality and information preserved. Empirical evaluations on image classification tasks show that models trained with H-SPLID primarily rely on salient input components, as indicated by reduced sensitivity to perturbations affecting non-salient features, such as image backgrounds. Lukas Miklautz, Chengzhi Shi, Andrii Shkabrii, Theodoros-Thirimachos Davarakis, Prudence Lam, Claudia Plant, Jennifer G. Dy, Stratis Ioannidis |
NeurIPS | 7 |
| 2025 | DISCO: Disentangled Communication Steering for Large Language ModelsabstractA variety of recent methods guide large language model outputs via the inference-time addition of *steering vectors* to residual-stream or attention-head representations. In contrast, we propose to inject steering vectors directly into the query and value representation spaces within attention heads. We provide evidence that a greater portion of these spaces exhibit high linear discriminability of concepts --a key property motivating the use of steering vectors-- than attention head outputs. We analytically characterize the effect of our method, which we term *DISentangled COmmunication (DISCO) Steering*, on attention head outputs. Our analysis reveals that DISCO disentangles a strong but underutilized baseline, steering attention head inputs, which implicitly modifies queries and values in a rigid manner. In contrast, DISCO's direct modulation of these components enables more granular control. We find that DISCO achieves superior performance over a number of steering vector baselines across multiple datasets on LLaMA 3.1 8B and Gemma 2 9B, with steering efficacy scoring up to $19.1$% higher than the runner-up. Our results support the conclusion that the query and value spaces are powerful building blocks for steering vector methods. Our code is publicly available at https://github.com/MaxTorop/DISCO. Max Torop, Aria Masoomi, Masih Eskandar, Jennifer G. Dy |
NeurIPS | 4 |
| 2025 | LVT: Large-Scale Scene Reconstruction via Local View TransformersabstractLarge transformer models are proving to be a powerful tool for 3D vision and novel view synthesis. However, the standard Transformer’s well-known quadratic complexity makes it difficult to scale these methods to large scenes. To address this challenge, we propose the Local View Transformer (LVT), a large-scale scene reconstruction and novel view synthesis architecture that circumvents the need for the quadratic attention operation. Motivated by the insight that spatially nearby views provide more useful signal about the local scene composition than distant views, our model processes all information in a local neighborhood around each view. To attend to tokens in nearby views, we leverage a novel positional encoding that conditions on the relative geometric transformation between the query and nearby views. We decode the output of our model into a 3D Gaussian Splat scene representation that includes both color and opacity view-dependence. Taken together, the Local View Transformer enables reconstruction of arbitrarily large, high-resolution scenes in a single forward pass. See our project page for results and interactive demos: https://toobaimt.github.io/lvt/. Tooba Imtiaz, Lucy Chai, Kathryn Heal, Jungyeon Park, Jennifer G. Dy, John Flynn |
SIGGRAPH Asia | 6 |
| 2024 | Boundary-Aware Uncertainty for Feature Attribution ExplainersabstractPost-hoc explanation methods have become a critical tool for understanding black-box classifiers in high-stakes applications. However, high-performing classifiers are often highly nonlinear and can exhibit complex behavior around the decision boundary, leading to brittle or misleading local explanations. Therefore there is an impending need to quantify the uncertainty of such explanation methods in order to understand when explanations are trustworthy. In this work we propose the Gaussian Process Explanation unCertainty (GPEC) framework, which generates a unified uncertainty estimate combining decision boundary-aware uncertainty with explanation function approximation uncertainty. We introduce a novel geodesic-based kernel, which captures the complexity of the target black-box decision boundary. We show theoretically that the proposed kernel similarity increases with decision boundary complexity. The proposed framework is highly flexible; it can be used with any black-box classifier and feature attribution method. Empirical results on multiple tabular and image datasets show that the GPEC uncertainty estimate improves understanding of explanations as compared to existing methods. Davin Hill, Aria Masoomi, Max Torop, Sandesh Ghimire, Jennifer G. Dy |
AISTATS | 5 |
| 2024 | Analyzing Explainer Robustness via Probabilistic Lipschitzness of Prediction Functions
Zulqarnain Khan, Davin Hill, Aria Masoomi, Joshua T. Bone, Jennifer G. Dy |
AISTATS | 5 |
| 2024 | Multiverse at the Edge: Interacting Real World and Digital Twins for Wireless BeamformingabstractCreating a digital world that closely mimics the real world with its many complex interactions and outcomes is possible today through advanced emulation software and ubiquitous computing power. Such a software-based emulation of an entity that exists in the real world is called a ‘digital twin’. In this paper, we consider a twin of a wireless millimeter-wave band radio that is mounted on a vehicle and show how it speeds up directional beam selection in mobile environments. To achieve this, we go beyond instantiating a single twin and propose the ‘$\MV$’ paradigm, with several possible digital twins attempting to capture the real world at different levels of fidelity. Towards this goal, this paper describes (i) a decision strategy at the vehicle that determines which twin must be used given the latency limitation, and (ii) a self-learning scheme that uses the$\MV$-guided beam outcomes to enhance DL-based decision-making in the real world over time. Our work is distinguished from prior works as follows: First, we use a publicly available RF dataset collected from an autonomous car for creating different twins. Second, we present a framework with continuous interaction between the real world and$\MV$of twins at the edge, as opposed to a one-time emulation that is completed prior to actual deployment. Results reveal that$\MV$offers up to$79.43\%$and$85.22\%$top-$10$beam selection accuracy for LOS and NLOS scenarios, respectively. Moreover, we observe$67.70-90.79\%$improvement in beam selection time compared to 802.11ad standard and 5G-NR standards. Batool Salehi, Utku Demir, Debashri Roy, Suyash Pradhan, Jennifer G. Dy, Stratis Ioannidis, Kaushik R. Chowdhury |
IEEE/ACM Trans. Netw. | 5 |
| 2023 | DualHSIC: HSIC-Bottleneck and Alignment for Continual LearningabstractRehearsal-based approaches are a mainstay of continual learning (CL). They mitigate the catastrophic forgetting problem by maintaining a small fixed-size buffer with a subset of data from past tasks. While most rehearsal-based approaches exploit the knowledge from buffered past data, little attention is paid to inter-task relationships and to critical task-specific and task-invariant knowledge. By appropriately leveraging inter-task relationships, we propose a novel CL method, named DualHSIC, to boost the performance of existing rehearsal-based methods in a simple yet effective way. DualHSIC consists of two complementary components that stem from the so-called Hilbert Schmidt independence criterion (HSIC): HSIC-Bottleneck for Rehearsal (HBR) lessens the inter-task interference and HSIC Alignment (HA) promotes task-invariant knowledge sharing. Extensive experiments show that DualHSIC can be seamlessly plugged into existing rehearsal-based methods for consistent performance improvements, outperforming recent state-of-the-art regularization-enhanced rehearsal methods. Zifeng Wang 0002, Zheng Zhan 0001, Yifan Gong 0004, Yucai Shao, Stratis Ioannidis, Yanzhi Wang 0001, Jennifer G. Dy |
ICML | 7 |
| 2023 | SmoothHess: ReLU Network Feature Interactions via Stein's LemmaabstractSeveral recent methods for interpretability model feature interactions by looking at the Hessian of a neural network. This poses a challenge for ReLU networks, which are piecewise-linear and thus have a zero Hessian almost everywhere. We propose SmoothHess, a method of estimating second-order interactions through Stein's Lemma. In particular, we estimate the Hessian of the network convolved with a Gaussian through an efficient sampling algorithm, requiring only network gradient calls. SmoothHess is applied post-hoc, requires no modifications to the ReLU network architecture, and the extent of smoothing can be controlled explicitly. We provide a non-asymptotic bound on the sample complexity of our estimation procedure. We validate the superior ability of SmoothHess to capture interactions on benchmark datasets and a real-world medical spirometry dataset. Max Torop, Aria Masoomi, Davin Hill, Kivanç Köse, Stratis Ioannidis, Jennifer G. Dy |
NeurIPS | 6 |
| 2023 | Graph transfer learning
Andrey Gritsenko, Kimia Shayestehfard, Armin Moharrer, Jennifer G. Dy, Stratis Ioannidis |
Knowl. Inf. Syst. | 5 |
| 2023 | Transforming Complex Problems Into K-Means SolutionsabstractK-means is a fundamental clustering algorithm widely used in both academic and industrial applications. Its popularity can be attributed to its simplicity and efficiency. Studies show the equivalence of K-means to principal component analysis, non-negative matrix factorization, and spectral clustering. However, these studies focus on standard K-means with squared euclidean distance. In this review paper, we unify the available approaches in generalizing K-means to solve challenging and complex problems. We show that these generalizations can be seen from four aspects: data representation, distance measure, label assignment, and centroid updating. As concrete applications of transforming problems into modified K-means formulation, we review the following applications: iterative subspace projection and clustering, consensus clustering, constrained clustering, domain adaptation, and outlier detection. Hongfu Liu 0001, Junxiang Chen, Jennifer G. Dy, Yun Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Deep Layer-wise Networks Have Closed-Form WeightsabstractThere is currently a debate within the neuroscience community over the likelihood of the brain performing backpropagation (BP). To better mimic the brain, training a network one layer at a time with only a "single forward pass" has been proposed as an alternative to bypass BP; we refer to these networks as "layer-wise" networks. We continue the work on layer-wise networks by answering two outstanding questions. First, do they have a closed-form solution? Second, how do we know when to stop adding more layers? This work proves that the "Kernel Mean Embedding" is the closed-form solution that achieves the network global optimum while driving these networks to converge towards a highly desirable kernel for classification; we call it the Neural Indicator Kernel. Chieh Wu, Aria Masoomi, Arthur Gretton, Jennifer G. Dy |
AISTATS | 4 |
| 2022 | Learning to Prompt for Continual LearningabstractThe mainstream paradigm behind continual learning has been to adapt the model parameters to non-stationary data distributions, where catastrophic forgetting is the central challenge. Typical methods rely on a rehearsal buffer or known task identity at test time to retrieve learned knowl-edge and address forgetting, while this work presents a new paradigm for continual learning that aims to train a more succinct memory system without accessing task identity at test time. Our method learns to dynamically prompt (L2P) a pre-trained model to learn tasks sequen-tially under different task transitions. In our proposed framework, prompts are small learnable parameters, which are maintained in a memory space. The objective is to optimize prompts to instruct the model prediction and ex-plicitly manage task-invariant and task-specific knowledge while maintaining model plasticity. We conduct comprehen-sive experiments under popular image classification bench-marks with different challenging continual learning set-tings, where L2P consistently outperforms prior state-of-the-art methods. Surprisingly, L2P achieves competitive results against rehearsal-based methods even without a re-hearsal buffer and is directly applicable to challenging task-agnostic continual learning. Source code is available at https://github.com/google-research/12p. Zifeng Wang 0002, Chen-Yu Lee, Han Zhang 0010, Ruoxi Sun 0002, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer G. Dy, Tomas Pfister |
CVPR | 9 |
| 2022 | DualPrompt: Complementary Prompting for Rehearsal-Free Continual Learning
Zifeng Wang 0002, Sayna Ebrahimi, Ruoxi Sun 0002, Han Zhang 0010, Chen-Yu Lee, Xiaoqi Ren, Guolong Su, Vincent Perot, Jennifer G. Dy, Tomas Pfister |
ECCV (26) | 10 |
| 2022 | Pruning Adversarially Robust Neural Networks without Adversarial ExamplesabstractAdversarial pruning compresses models while preserving robustness. Current methods require access to adversarial examples during pruning. This significantly hampers training efficiency. Moreover, as new adversarial attacks and training methods develop at a rapid rate, adversarial pruning methods need to be modified accordingly to keep up. In this work, we propose a novel framework to prune a previously trained robust neural network while maintaining adversarial robustness, without further generating adversarial examples. We leverage concurrent self-distillation and pruning to preserve knowledge in the original model as well as regularizing the pruned model via the Hilbert-Schmidt Information Bottleneck. We comprehensively evaluate our proposed framework and show its superior performance in terms of both adversarial robustness and efficiency when pruning architectures trained on the MNIST, CIFAR-10, and CIFAR-100 datasets against five state-of-the-art attacks.. Tong Jian, Zifeng Wang 0002, Yanzhi Wang 0001, Jennifer G. Dy, Stratis Ioannidis |
ICDM | 4 |
| 2022 | Explanations of Black-Box Models based on Directional Feature Interactions
Aria Masoomi, Davin Hill, Zhonghui Xu, Craig P. Hersh, Edwin K. Silverman, Peter J. Castaldi, Stratis Ioannidis, Jennifer G. Dy |
ICLR | 8 |
| 2022 | SparCL: Sparse Continual Learning on the EdgeabstractExisting work in continual learning (CL) focuses on mitigating catastrophic forgetting, i.e., model performance deterioration on past tasks when learning a new task. However, the training efficiency of a CL system is under-investigated, which limits the real-world application of CL systems under resource-limited scenarios. In this work, we propose a novel framework called Sparse Continual Learning (SparCL), which is the first study that leverages sparsity to enable cost-effective continual learning on edge devices. SparCL achieves both training acceleration and accuracy preservation through the synergy of three aspects: weight sparsity, data efficiency, and gradient sparsity. Specifically, we propose task-aware dynamic masking (TDM) to learn a sparse network throughout the entire CL process, dynamic data removal (DDR) to remove less informative training data, and dynamic gradient masking (DGM) to sparsify the gradient updates. Each of them not only improves efficiency, but also further mitigates catastrophic forgetting. SparCL consistently improves the training efficiency of existing state-of-the-art (SOTA) CL methods by at most 23X less training FLOPs, and, surprisingly, further improves the SOTA accuracy by at most 1.7%. SparCL also outperforms competitive baselines obtained from adapting SOTA sparse training methods to the CL setting in both efficiency and accuracy. We also evaluate the effectiveness of SparCL on a real mobile phone, further indicating the practical potential of our method. Zifeng Wang 0002, Zheng Zhan 0001, Yifan Gong 0004, Geng Yuan, Wei Niu 0002, Tong Jian, Bin Ren 0002, Stratis Ioannidis, Yanzhi Wang 0001, Jennifer G. Dy |
NeurIPS | 10 |
| 2022 | Deep Bayesian Unsupervised Lifelong Learning
Zifeng Wang 0002, Aria Masoomi, Jennifer G. Dy |
Neural Networks | 4 |
| 2022 | Sample complexity of rank regression using pairwise comparisons
Berkan Kadioglu, Jennifer G. Dy, Deniz Erdogmus, Stratis Ioannidis |
Pattern Recognit. | 3 |
| 2022 | Spectral Ranking RegressionabstractWe study the problem of ranking regression, in which a dataset of rankings is used to learn Plackett–Luce scores as functions of sample features. We propose a novel spectral algorithm to accelerate learning in ranking regression. Our main technical contribution is to show that the Plackett–Luce negative log-likelihood augmented with a proximal penalty has stationary points that satisfy the balance equations of a Markov Chain. This allows us to tackle the ranking regression problem via an efficient spectral algorithm by using the Alternating Directions Method of Multipliers (ADMM). ADMM separates the learning of scores and model parameters, and in turn, enables us to devise fast spectral algorithms for ranking regression via both shallow and deep neural network (DNN) models. For shallow models, our algorithms are up to 579 times faster than the Newton’s method. For DNN models, we extend the standard ADMM via a Kullback–Leibler proximal penalty and show that this is still amenable to fast inference via a spectral approach. Compared to a state-of-the-art siamese network, our resulting algorithms are up to 175 times faster and attain better predictions by up to 26% Top-1 Accuracy and 6% Kendall-Tau correlation over five real-life ranking datasets. Ilkay Yildiz, Jennifer G. Dy, Deniz Erdogmus, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
ACM Trans. Knowl. Discov. Data | 2 |
| 2022 | Radio Frequency Fingerprinting on the EdgeabstractDeep learning methods have been very successful at radio frequency fingerprinting tasks, predicting the identity of transmitting devices with high accuracy. We study radio frequency fingerprinting deployments at resource-constrained edge devices. We use structured pruning to jointly train and sparsify neural networks tailored to edge hardware implementations. We compress convolutional layers by a$27.2\times$factor while incurring a negligible prediction accuracy decrease (less than 1 percent). We demonstrate the efficacy of our approach over multiple edge hardware platforms, including a Samsung Gallaxy S10 phone and a Xilinx-ZCU104 FPGA. Our method yields significant inference speedups,$11.5\times$on the FPGA and$3\times$on the smartphone, as well as high efficiency: the FPGA processing time is$17\times$smaller than in a V100 GPU. To the best of our knowledge, we are the first to explore the possibility of compressing networks for radio frequency fingerprinting; as such, our experiments can be seen as a means of characterizing the informational capacity associated with this specific learning task. Tong Jian, Yifan Gong 0004, Zheng Zhan 0001, Runbin Shi, Nasim Soltani, Zifeng Wang 0002, Jennifer G. Dy, Kaushik R. Chowdhury, Yanzhi Wang 0001, Stratis Ioannidis |
IEEE Trans. Mob. Comput. | 7 |
| 2021 | Faster & More Reliable Tuning of Neural Networks: Bayesian Optimization with Importance SamplingabstractMany contemporary machine learning models require extensive tuning of hyperparameters to perform well. A variety of methods, such as Bayesian optimization, have been developed to automate and expedite this process. However, tuning remains extremely costly as it typically requires repeatedly fully training models. To address this issue, Bayesian optimization methods have been extended to use cheap, partially trained models to extrapolate to expensive complete models. While this approach enlarges the set of explored hyperparameters, including many low-fidelity observations adds to the intrinsic randomness of the procedure and makes extrapolation challenging. We propose to accelerate hyperparameter tuning for neural networks in a robust way by taking into account the relative amount of information contributed by each training example. To do so, we integrate importance sampling with Bayesian optimization, which significantly increases the quality of the black-box function evaluations and their runtime. To overcome the additional overhead cost of using importance sampling, we cast hyperparameter search as a multi-task Bayesian optimization problem over both hyperparameters and importance sampling design, which achieves the best of both worlds. Through learning a trade-off between training complexity and quality, our method improves upon validation error, in the average and worst-case. We show that this results in more reliable performance of our method in less wall-clock time across a variety of and datasets complex neural architectures. Setareh Ariafar, Zelda Mariet, Dana H. Brooks, Jennifer G. Dy, Jasper Snoek |
AISTATS | 4 |
| 2021 | Rate-Regularization and Generalization in Variational AutoencodersabstractVariational autoencoders (VAEs) optimize an objective that comprises a reconstruction loss (the distortion) and a KL term (the rate). The rate is an upper bound on the mutual information, which is often interpreted as a regularizer that controls the degree of compression. We here examine whether inclusion of the rate term also improves generalization. We perform rate-distortion analyses in which we control the strength of the rate term, the network capacity, and the difficulty of the generalization problem. Lowering the strength of the rate term paradoxically improves generalization in most settings, and reducing the mutual information typically leads to underfitting. Moreover, we show that generalization performance continues to improve even after the mutual information saturates, indicating that the gap on the bound (i.e. the KL divergence relative to the inference marginal) affects generalization. This suggests that the standard spherical Gaussian prior is not an inductive bias that typically improves generalization, prompting further work to understand what choices of priors improve generalization in VAEs. Alican Bozkurt, Babak Esmaeili 0001, Jean-Baptiste Tristan, Dana H. Brooks, Jennifer G. Dy, Jan-Willem van de Meent |
AISTATS | 5 |
| 2021 | Deep Spectral RankingabstractLearning from ranking observations arises in many domains, and siamese deep neural networks have shown excellent inference performance in this setting. However, SGD does not scale well, as an epoch grows exponentially with the ranking observation size. We show that a spectral algorithm can be combined with deep learning methods to significantly accelerate training. We combine a spectral estimate of Plackett-Luce ranking scores with a deep model via the Alternating Directions Method of Multipliers with a Kullback-Leibler proximal penalty. Compared to a state-of-the-art siamese network, our algorithms are up to 175 times faster and attain better predictions by up to 26% Top-1 Accuracy and 6% Kendall-Tau correlation over five real-life ranking datasets. Ilkay Yildiz, Jennifer G. Dy, Deniz Erdogmus, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
AISTATS | 2 |
| 2021 | Graph Transfer LearningabstractGraph embeddings have been tremendously successful at producing node representations that are discriminative for downstream tasks. In this paper, we study the problem of graph transfer learning: given two graphs and labels in the nodes of the first graph, we wish to predict the labels on the second graph. We propose a tractable, non-combinatorial method for solving the graph transfer learning problem by combining classification and embedding losses with a continuous, convex penalty motivated by tractable graph distances. We demonstrate that our method successfully predicts labels across graphs with almost perfect accuracy; in the same scenarios, training embeddings through standard methods leads to predictions that are no better than random. Andrey Gritsenko, Kimia Shayestehfard, Armin Moharrer, Jennifer G. Dy, Stratis Ioannidis |
ICDM | 5 |
| 2021 | Deep Learning on Visual and Location Data for V2I mmWave BeamformingabstractAccurate beam alignment in the millimeter-wave (mmWave) band introduces considerable overheads involving brute-force exploration of multiple beam-pair combinations and beam retraining due to mobility. This cost becomes often intractable under high mobility scenarios, where fast beamforming algorithms that can quickly adapt the beam configurations are still under development for 5G and beyond. Besides, blockage prediction is a key capability in order to establish mmWave reliable links. In this paper, we propose a data fusion approach that takes inputs from visual edge devices and localization sensors to (i) reduce the beam selection overhead by narrowing down the search to a small set containing the best possible beam-pairs and (ii) detect blockage conditions between transmitters and receivers. We evaluate our approach through joint simulation of multi-modal data from vision and localization sensors and RF data. Additionally, we show how deep learning based fusion of images and Global Positioning System (GPS) data can play a key role in configuring vehicle-to-infrastructure (V2I) mmWave links. We show a 90% top-10 beam selection accuracy and a 92.86% blockage prediction accuracy. Furthermore, the proposed approach achieves a 99.7% reduction on the beam selection time while keeping a 94.86% of the maximum achievable throughput. Guillem Reus Muns, Batool Salehi, Debashri Roy, Tong Jian, Zifeng Wang 0002, Jennifer G. Dy, Stratis Ioannidis, Kaushik R. Chowdhury |
MSN | 6 |
| 2021 | Reliable Estimation of KL Divergence using a Discriminator in Reproducing Kernel Hilbert SpaceabstractEstimating Kullback–Leibler (KL) divergence from samples of two distributions is essential in many machine learning problems. Variational methods using neural network discriminator have been proposed to achieve this task in a scalable manner. However, we noticed that most of these methods using neural network discriminators suffer from high fluctuations (variance) in estimates and instability in training. In this paper, we look at this issue from statistical learning theory and function space complexity perspective to understand why this happens and how to solve it. We argue that the cause of these pathologies is lack of control over the complexity of the neural network discriminator function and could be mitigated by controlling it. To achieve this objective, we 1) present a novel construction of the discriminator in the Reproducing Kernel Hilbert Space (RKHS), 2) theoretically relate the error probability bound of the KL estimates to the complexity of the discriminator in the RKHS space, 3) present a scalable way to control the complexity (RKHS norm) of the discriminator for a reliable estimation of KL divergence, and 4) prove the consistency of the proposed estimator. In three different applications of KL divergence -- estimation of KL, estimation of mutual information and Variational Bayes -- we show that by controlling the complexity as developed in the theory, we are able to reduce the variance of KL estimates and stabilize the training. Sandesh Ghimire, Aria Masoomi, Jennifer G. Dy |
NeurIPS | 3 |
| 2021 | Revisiting Hilbert-Schmidt Information Bottleneck for Adversarial RobustnessabstractWe investigate the HSIC (Hilbert-Schmidt independence criterion) bottleneck as a regularizer for learning an adversarially robust deep neural network classifier. In addition to the usual cross-entropy loss, we add regularization terms for every intermediate layer to ensure that the latent representations retain useful information for output prediction while reducing redundant information. We show that the HSIC bottleneck enhances robustness to adversarial attacks both theoretically and experimentally. In particular, we prove that the HSIC bottleneck regularizer reduces the sensitivity of the classifier to adversarial examples. Our experiments on multiple benchmark datasets and architectures demonstrate that incorporating an HSIC bottleneck regularizer attains competitive natural accuracy and improves adversarial robustness, both with and without adversarial examples during training. Our code and adversarially robust models are publicly available. Zifeng Wang 0002, Tong Jian, Aria Masoomi, Stratis Ioannidis, Jennifer G. Dy |
NeurIPS | 5 |
| 2021 | Segmentation of cellular patterns in confocal images of melanocytic lesions in vivo via a multiscale encoder-decoder network (MED-Net)
Kivanç Köse, Alican Bozkurt, Christi Alessi-Fox, Melissa Gill, Caterina Longo, Giovanni Pellacani, Jennifer G. Dy, Dana H. Brooks, Milind Rajadhyaksha |
Medical Image Anal. | 7 |
| 2021 | Improved prediction of smoking status via isoform-aware RNA-seq deep learning modelsabstractMost predictive models based on gene expression data do not leverage information related to gene splicing, despite the fact that splicing is a fundamental feature of eukaryotic gene expression. Cigarette smoking is an important environmental risk factor for many diseases, and it has profound effects on gene expression. Using smoking status as a prediction target, we developed deep neural network predictive models using gene, exon, and isoform level quantifications from RNA sequencing data in 2,557 subjects in the COPDGene Study. We observed that models using exon and isoform quantifications clearly outperformed gene-level models when using data from 5 genes from a previously published prediction model. Whereas the test set performance of the previously published model was 0.82 in the original publication, our exon-based models including an exon-to-isoform mapping layer achieved a test set AUC (area under the receiver operating characteristic) of 0.88, which improved to an AUC of 0.94 using exon quantifications from a larger set of genes. Isoform variability is an important source of latent information in RNA-seq data that can be used to improve clinical prediction models. Zifeng Wang 0002, Aria Masoomi, Zhonghui Xu, Adel Boueiz, Sool Lee, Russell Bowler, Michael H. Cho, Edwin K. Silverman, Craig P. Hersh, Jennifer G. Dy, Peter J. Castaldi |
PLoS Comput. Biol. | 11 |
| 2020 | Fast and Accurate Ranking RegressionabstractWe consider a ranking regression problem in which we use a dataset of ranked choices to learn Plackett-Luce scores as functions of sample features. We solve the maximum likelihood estimation problem by using the Alternating Directions Method of Multipliers (ADMM), effectively separating the learning of scores and model parameters. This separation allows us to express scores as the stationary distribution of a continuous-time Markov Chain. Using this equivalence, we propose two spectral algorithms for ranking regression that learn model parameters up to 579 times faster than the Newton’s method. Ilkay Yildiz, Jennifer G. Dy, Deniz Erdogmus, Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
AISTATS | 2 |
| 2020 | A Quantitative Machine Learning Approach to Master Students Admission for Professional Institutions
Bryan Lackaye, Jennifer G. Dy, Carla E. Brodley |
EDM | 3 |
| 2020 | Learn-Prune-Share for Lifelong LearningabstractIn lifelong learning, we wish to maintain and update a model (e.g., a neural network classifier) in the presence of new classification tasks that arrive sequentially. In this paper, we propose a learn-prune-share (LPS) algorithm which addresses the challenges of catastrophic forgetting, parsimony, and knowledge reuse simultaneously. LPS splits the network into task-specific partitions via an ADMM-based pruning strategy. This leads to no forgetting, while maintaining parsimony. Moreover, LPS integrates a novel selective knowledge sharing scheme into this ADMM optimization framework. This enables adaptive knowledge sharing in an end-to-end fashion. Comprehensive experimental results on two lifelong learning benchmark datasets and a challenging real world radio frequency fingerprinting dataset are provided to demonstrate the effectiveness of our approach. Our experiments show that LPS consistently outperforms multiple state-of-the-art competitors. Zifeng Wang 0002, Tong Jian, Kaushik R. Chowdhury, Yanzhi Wang 0001, Jennifer G. Dy, Stratis Ioannidis |
ICDM | 5 |
| 2020 | Open-World Class Discovery with Kernel NetworksabstractWe study an Open-World Class Discovery problem in which, given labeled training samples from old classes, we need to discover new classes from unlabeled test samples. There are two critical challenges to addressing this paradigm: (a) transferring knowledge from old to new classes, and (b) incorporating knowledge learned from new classes back to the original model. We propose Class Discovery Kernel Network with Expansion (CD-KNet-Exp), a deep learning framework, which utilizes the Hilbert Schmidt Independence Criterion to bridge supervised and unsupervised information together in a systematic way, such that the learned knowledge from old classes is distilled appropriately for discovering new classes. Compared to competing methods, CD-KNet-Exp shows superior performance on three publicly available benchmark datasets and a challenging real-world radio frequency fingerprinting dataset. Zifeng Wang 0002, Batool Salehi, Andrey Gritsenko, Kaushik R. Chowdhury, Stratis Ioannidis, Jennifer G. Dy |
ICDM | 6 |
| 2020 | Using Undersampling with Ensemble Learning to Identify Factors Contributing to Preterm BirthabstractIn this paper, we propose Ensemble Learning models to identify factors contributing to preterm birth. Our work leverages a rich dataset collected by a NIEHS P42 Center that is trying to identify the dominant factors responsible for the high rate of premature births in northern Puerto Rico. We investigate analytical models addressing two major challenges present in the dataset: 1) the significant amount of incomplete data in the dataset, and 2) class imbalance in the dataset. First, we leverage and compare two types of missing data imputation methods: 1) mean-based and 2) similarity-based, increasing the completeness of this dataset. Second, we propose a feature selection and evaluation model based on using undersampling with Ensemble Learning to address class imbalance present in the dataset. We leverage and compare multiple Ensemble Feature selection methods, including Complete Linear Aggregation (CLA), Weighted Mean Aggregation (WMA), Feature Occurrence Frequency (OFA) and Classification Accuracy Based Aggregation (CAA). To further address missing data present in each feature, we propose two novel methods: 1) Missing Data Rate and Accuracy Based Aggregation (MAA), and 2) Entropy and Accuracy Based Aggregation (EAA). Both proposed models balance the degree of data variance introduced by the missing data handling during the feature selection process, while maintaining model performance. Our results show a 42% improvement in sensitivity versus fallout over previous state-of-the-art methods. Shi Dong 0002, Zlatan Feric, Chieh Wu, April Z. Gu, Jennifer G. Dy, John Meeker, Ingrid Y. Padilla, José Cordero, Carmen Velez Vega, Zaira Rosario, Akram Alshawabkeh, David R. Kaeli |
ICMLA | 6 |
| 2020 | Exposing the Fingerprint: Dissecting the Impact of the Wireless Channel on Radio FingerprintingabstractRadio fingerprinting uniquely identifies wireless devices by leveraging tiny hardware-level imperfections inevitably present in off-the-shelf radio circuitry. This way, devices can be directly identified at the physical layer by analyzing the unprocessed received waveform - thus avoiding energy-expensive upper-layer cryptography that resource-challenged embedded devices may not be able to afford. Recent advances have proven that convolutional neural networks (CNNs) - thanks to their multidimensional mappings - can achieve fingerprinting accuracy levels impossible to achieve by traditional low-dimensional algorithms. The same research, however, has also suggested that the wireless channel may negatively impact the accuracy of CNN-based radio fingerprinting algorithms by making device-unique hardware imperfections much harder to recognize.In spite of the growing interest in radio fingerprinting research by academia and DARPA, the wireless research community still lacks (i) a large-scale open dataset for radio fingerprinting collected in diverse environments and rich, diverse, channel conditions; and (ii) a full-fledged, systematic, quantitative investigation of the impact of the wireless channel on the accuracy of CNN-based radio fingerprinting algorithms. The key contribution of this paper is to bridge this gap by (i) collecting and sharing with the community more than 7TB of wireless data obtained from 20 wireless devices with identical RF circuitry (and thus, worst-case scenario for fingerprinting) over the course of several days in (a) an anechoic chamber, (b) in-the-wild testbed, and (c) with cable connections; and (ii) providing a first-of-its-kind evaluation of the impact of the wireless channel on CNN-based fingerprinting algorithms through (a) the 7TB experimental dataset and (b) a 400GB dataset provided by DARPA containing hundreds of thousands of transmissions from thousands of WiFi and ADS-B devices with different SNR conditions. Experimental results conclude that (i) the wireless channel impacts the classification accuracy significantly, i.e., from 85% to 9% and from 30% to 17% in the experimental and DARPA dataset, respectively; and that (ii) equalizing I/Q data can increase the accuracy to a significant extent (i.e., by up to 23%) when the number of devices increases significantly. Amani Al-Shawabka, Francesco Restuccia 0001, Salvatore D'Oro, Tong Jian, Bruno Costa Rendon, Nasim Soltani, Jennifer G. Dy, Stratis Ioannidis, Kaushik R. Chowdhury, Tommaso Melodia |
INFOCOM | 7 |
| 2020 | Climate Downscaling Using YNet: A Deep Convolutional Network with Skip Connections and FusionabstractClimate change is one of the major challenges to human beings in our time. It brings many unexpected disasters which cause drastic losses including lives and properties. To better understand climate change, scientists developed various Global Climate Models (GCMs) to simulate the global climate and make projections for future climate values. These global climate models have coarse grids (i.e., low resolutions both in space and time) due to limitations of computing power and simulation time. Although they are helpful in predicting large scale long term trend in climate, they are too coarse for impact analysis in smaller scales such as in regional or local scale. However, climate conditions in regional or local scale are very important in making decisions related to climate conditions such as infrastructure, transportation and evacuation, as they highly depend on small scale climate conditions. In this paper, we proposed YNet, a novel deep convolutional neural network (CNN) with skip connections and fusion capabilities to perform downscaling for climate variables, on multiple GCMs directly rather than on reanalysis data. We analyzed and compared our proposed method with four other methods on datasets of three climate variables: mean precipitation, and extreme values (maximum temperature and minimum temperature). The results show the effectiveness of the proposed method. Auroop R. Ganguly, Jennifer G. Dy |
KDD | 3 |
| 2020 | Machine Learning on Camera Images for Fast mmWave BeamformingabstractPerfect alignment in chosen beam sectors at both transmit- and receive-nodes is required for beamforming in mmWave bands. Current 802.11ad WiFi and emerging 5G cellular standards spend up to several milliseconds exploring different sector combinations to identify the beam pair with the highest SNR. In this paper, we propose a machine learning (ML) approach with two sequential convolutional neural networks (CNN) that uses out-of-band information, in the form of camera images, to (i) rapidly identify the locations of the transmitter and receiver nodes, and then (ii) return the optimal beam pair. We experimentally validate this intriguing concept for indoor settings using the NI 60GHz mmwave transceiver. Our results reveal that our ML approach reduces beamforming related exploration time by 93% under different ambient lighting conditions, with an error of less than 1% compared to the time-intensive deterministic method defined by the current standards. Batool Salehi, Mauro Belgiovine, Sara Garcia Sanchez, Jennifer G. Dy, Stratis Ioannidis, Kaushik R. Chowdhury |
MASS | 4 |
| 2020 | Instance-wise Feature GroupingabstractIn many learning problems, the domain scientist is often interested in discovering the groups of features that are redundant and are important for classification. Moreover, the features that belong to each group, and the important feature groups may vary per sample. But what do we mean by feature redundancy? In this paper, we formally define two types of redundancies using information theory: \textit{Representation} and \textit{Relevant redundancies}. We leverage these redundancies to design a formulation for instance-wise feature group discovery and reveal a theoretical guideline to help discover the appropriate number of groups. We approximate mutual information via a variational lower bound and learn the feature group and selector indicators with Gumbel-Softmax in optimizing our formulation. Experiments on synthetic data validate our theoretical claims. Experiments on MNIST, Fashion MNIST, and gene expression datasets show that our method discovers feature groups with high classification accuracies. Aria Masoomi, Chieh Wu, Zifeng Wang 0002, Peter J. Castaldi, Jennifer G. Dy |
NeurIPS | 6 |
| 2020 | Neural Topographic Factor Analysis for fMRI DataabstractNeuroimaging studies produce gigabytes of spatio-temporal data for a small number of participants and stimuli. Recent work increasingly suggests that the common practice of averaging across participants and stimuli leaves out systematic and meaningful information. We propose Neural Topographic Factor Analysis (NTFA), a probabilistic factor analysis model that infers embeddings for participants and stimuli. These embeddings allow us to reason about differences between participants and stimuli as signal rather than noise. We evaluate NTFA on data from an in-house pilot experiment, as well as two publicly available datasets. We demonstrate that inferring representations for participants and stimuli improves predictive generalization to unseen data when compared to previous topographic methods. We also demonstrate that the inferred latent factor representations are useful for downstream tasks such as multivoxel pattern analysis and functional connectivity. Eli Sennesh, Zulqarnain Khan, J. Benjamin Hutchinson, Ajay B. Satpute, Jennifer G. Dy, Jan-Willem van de Meent |
NeurIPS | 6 |
| 2020 | Deep Kernel Learning for ClusteringabstractWe propose a deep learning approach for discovering kernels tailored to identifying clusters over sample data. Our neural network produces sample embeddings that are motivated by and are at least as expressive as spectral clustering. Our training objective, based on the Hilbert Schmidt Independence Criterion, can be optimized via gradient adaptations on the Stiefel manifold, leading to significant acceleration over spectral methods relying on eigen-decompositions. Finally, our trained embedding can be directly applied to out-of-sample data. We show experimentally that our approach outperforms several state-of-the-art deep clustering methods, as well as traditional approaches such as k-means and spectral clustering over a broad array of real and synthetic datasets. Chieh Wu, Zulqarnain Khan, Stratis Ioannidis, Jennifer G. Dy |
SDM | 4 |
| 2019 | Variational Inference from Ranked Samples with FeaturesabstractIn many supervised learning settings, elicited labels comprise pairwise comparisons or rankings of samples. We propose a Bayesian inference model for ranking datasets, allowing us to take a probabilistic approach to ranking inference. Our probabilistic assumptions are motivated by, and consistent with, the so-called Plackett-Luce model. We propose a variational inference method to extract a closed-form Gaussian posterior distribution. We show experimentally that the resulting posterior yields more reliable ranking predictions compared to predictions via point estimates. Jennifer G. Dy, Deniz Erdogmus, Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
ACML | 2 |
| 2019 | Structured Disentangled RepresentationsabstractDeep latent-variable models learn representations of high-dimensional data in an unsupervised manner. A number of recent efforts have focused on learning representations that disentangle statistically independent axes of variation by introducing modifications to the standard objective function. These approaches generally assume a simple diagonal Gaussian prior and as a result are not able to reliably disentangle discrete factors of variation. We propose a two-level hierarchical objective to control relative degree of statistical independence between blocks of variables and individual variables within blocks. We derive this objective as a generalization of the evidence lower bound, which allows us to explicitly represent the trade-offs between mutual information between data and representation, KL divergence between representation and prior, and coverage of the support of the empirical data distribution. Experiments on a variety of datasets demonstrate that our objective can not only disentangle discrete variables, but that doing so also improves disentanglement of other variables and, importantly, generalization even to unseen combinations of factors. Babak Esmaeili 0001, Hao Wu 0020, Alican Bozkurt, N. Siddharth 0001, Brooks Paige, Dana H. Brooks, Jennifer G. Dy, Jan-Willem van de Meent |
AISTATS | 8 |
| 2019 | Nonparametric Mixture of Sparse Regressions on Spatio-Temporal Data - An Application to Climate PredictionabstractClimate prediction is a very challenging problem. Many institutes around the world try to predict climate variables by building climate models called General Circulation Models (GCMs), which are based on mathematical equations that describe the physical processes. The prediction abilities of different GCMs may vary dramatically across different regions and time. Motivated by the need of identifying which GCMs are more useful for a particular region and time, we introduce a clustering model combining Dirichlet Process (DP) mixture of sparse linear regression with Markov Random Fields (MRFs). This model incorporates DP to automatically determine the number of clusters, imposes MRF constraints to guarantee spatio-temporal smoothness, and selects a subset of GCMs that are useful for prediction within each spatio-temporal cluster with a spike-and-slab prior. We derive an effective Gibbs sampling method for this model. Experimental results are provided for both synthetic and real-world climate data. Junxiang Chen, Auroop R. Ganguly, Jennifer G. Dy |
KDD | 4 |
| 2019 | A Severity Score for Retinopathy of PrematurityabstractRetinopathy of Prematurity (ROP) is a leading cause for childhood blindness worldwide. An automated ROP detection system could significantly improve the chance of a child receiving proper diagnosis and treatment. We propose a means of producing a continuous severity score in an automated fashion, regressed from both (a) diagnostic class labels as well as (b) comparison outcomes. Our generative model combines the two sources, and successfully addresses inherent variability in diagnostic outcomes. In particular, our method exhibits an excellent predictive performance of both diagnostic and comparison outcomes over a broad array of metrics, including AUC, precision, and recall. Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Jennifer G. Dy, Deniz Erdogmus, Stratis Ioannidis |
KDD | 7 |
| 2019 | Solving Interpretable Kernel Dimensionality ReductionabstractKernel dimensionality reduction (KDR) algorithms find a low dimensional representation of the original data by optimizing kernel dependency measures that are capable of capturing nonlinear relationships. The standard strategy is to first map the data into a high dimensional feature space using kernels prior to a projection onto a low dimensional space. While KDR methods can be easily solved by keeping the most dominant eigenvectors of the kernel matrix, its features are no longer easy to interpret. Alternatively, Interpretable KDR (IKDR) is different in that it projects onto a subspace \textit{before} the kernel feature mapping, therefore, the projection matrix can indicate how the original features linearly combine to form the new features. Unfortunately, the IKDR objective requires a non-convex manifold optimization that is difficult to solve and can no longer be solved by eigendecomposition. Recently, an efficient iterative spectral (eigendecomposition) method (ISM) has been proposed for this objective in the context of alternative clustering. However, ISM only provides theoretical guarantees for the Gaussian kernel. This greatly constrains ISM's usage since any kernel method using ISM is now limited to a single kernel. This work extends the theoretical guarantees of ISM to an entire family of kernels, thereby empowering ISM to solve any kernel method of the same objective. In identifying this family, we prove that each kernel within the family has a surrogate $\Phi$ matrix and the optimal projection is formed by its most dominant eigenvectors. With this extension, we establish how a wide range of IKDR applications across different learning paradigms can be solved by ISM. To support reproducible results, the source code is made publicly available on \url{https://github.com/ANONYMIZED}. Chieh Wu, Jared Miller, Yale Chang, Mario Sznaier, Jennifer G. Dy |
NeurIPS | 5 |
| 2019 | Accelerated Experimental Design for Pairwise ComparisonsabstractPairwise comparison labels are more informative and less variable than class labels, but generating them poses a challenge: their number grows quadratically in the dataset size. We study a natural experimental design objective, namely, D-optimality, that can be used to identify which K pairwise comparisons to generate. This objective is known to perform well in practice, and is submodular, making the selection approximable via the greedy algorithm. A naïve greedy implementation has O(N2 d2 K) complexity, where N is the dataset size, d is the feature space dimension, and K is the number of generated comparisons. We show that, by exploiting the inherent geometry of the dataset–namely, that it consists of pairwise comparisons–the greedy algorithm's complexity can be reduced to O(N2 (K + d) + N(dK + d2) + d2 K). We apply the same acceleration also to the so-called lazy greedy algorithm. When combined, the above improvements lead to an execution time of less than 1 hour for a dataset with 108 comparisons; the naïve greedy algorithm on the same dataset would require more than 10 days to terminate. Jennifer G. Dy, Deniz Erdogmus, Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
SDM | 2 |
| 2019 | ADMMBO: Bayesian Optimization with Unknown Constraints using ADMMabstractThere exist many problems in science and engineering that involve optimization of an unknown or partially unknown objective function. Recently, Bayesian Optimization (BO) has emerged as a powerful tool for solving optimization problems whose objective functions are only available as a black box and are expensive to evaluate. Many practical problems, however, involve optimization of an unknown objective function subject to unknown constraints. This is an important yet challenging problem for which, unlike optimizing an unknown function, existing methods face several limitations. In this paper, we present a novel constrained Bayesian optimization framework to optimize an unknown objective function subject to unknown constraints. We introduce an equivalent optimization by augmenting the objective function with constraints, introducing auxiliary variables for each constraint, and forcing the new variables to be equal to the main variable. Building on the Alternating Direction Method of Multipliers (ADMM) algorithm, we propose ADMM-Bayesian Optimization (ADMMBO) to solve the problem in an iterative fashion. Our framework leads to multiple unconstrained subproblems with unknown objective functions, which we then solve via BO. Our method resolves several challenges of state-of-the-art techniques: it can start from infeasible points, is insensitive to initialization, can efficiently handle `decoupled problems' and has a concrete stopping criterion. Extensive experiments on a number of challenging BO benchmark problems show that our proposed approach outperforms the state-of-the-art methods in terms of the speed of obtaining a feasible solution and convergence to the global optimum as well as minimizing the number of total evaluations of unknown objective and constraints functions. Setareh Ariafar, Jaume Coll-Font, Dana H. Brooks, Jennifer G. Dy |
J. Mach. Learn. Res. | 4 |
| 2019 | Classification and comparison via neural networks
Ilkay Yildiz, Jennifer G. Dy, Deniz Erdogmus, James M. Brown 0001, Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
Neural Networks | 3 |
| 2019 | Intelligent Labeling Based on Fisher Information for Medical Image Segmentation Using Deep LearningabstractDeep convolutional neural networks (CNN) have recently achieved superior performance at the task of medical image segmentation compared to classic models. However, training a generalizable CNN requires a large amount of training data, which is difficult, expensive, and time-consuming to obtain in medical settings. Active Learning (AL) algorithms can facilitate training CNN models by proposing a small number of the most informative data samples to be annotated to achieve a rapid increase in performance. We proposed a new active learning method based on Fisher information (FI) for CNNs for the first time. Using efficient backpropagation methods for computing gradients together with a novel low-dimensional approximation of FI enabled us to compute FI for CNNs with a large number of parameters. We evaluated the proposed method for brain extraction with a patch-wise segmentation CNN model in two different learning scenarios: universal active learning and active semi-automatic segmentation. In both scenarios, an initial model was obtained using labeled training subjects of a source data set and the goal was to annotate a small subset of new samples to build a model that performs well on the target subject(s). The target data sets included images that differed from the source data by either age group (e.g. newborns with different image contrast) or underlying pathology that was not available in the source data. In comparison to several recently proposed AL methods and brain extraction baselines, the results showed that FI-based AL outperformed the competing methods in improving the performance of the model after labeling a very small portion of target data set (<0.25%). Jamshid Sourati, Ali Gholipour, Jennifer G. Dy, Xavier Tomas-Fernandez, Sila Kurugol, Simon K. Warfield |
IEEE Trans. Medical Imaging | 3 |
| 2018 | Crowdclustering with Partition LabelsabstractCrowdclustering is a practical way to incorporate domain knowledge into clustering, by combining opinions from multiple domain experts. Existing crowdclustering methods analyze binary pairwise similarity labels. However, in some applications, experts might provide partition labels. If we convert partition labels into pairwise similarity, then it would be difficult to understand the relationships between clustering solutions from different experts. In this paper, we propose a crowdclustering model that directly analyzes partition labels. The proposed model adopts a novel approach based on a modified multinomial logistic regression model, which simultaneously learns the number of clusters and determines hyper-planes that partition samples into clusters. The proposed model also learns a mapping between the latent clusters and expert labels, revealing the agreements and disagreements between experts. Experiments on benchmark data demonstrate that the proposed model simultaneously learns the number of clusters and discovers the clustering structure. An experiment on disease subtyping problem illustrates that the proposed model helps us understand the agreement and disagreement between experts. Junxiang Chen, Yale Chang, Peter J. Castaldi, Michael H. Cho, Brian D. Hobbs, Jennifer G. Dy |
AISTATS | 6 |
| 2018 | Iterative Spectral Method for Alternative ClusteringabstractGiven a dataset and an existing clustering as input, alternative clustering aims to find an alternative partition. One of the state-of-the-art approaches is Kernel Dimension Alternative Clustering (KDAC). We propose a novel Iterative Spectral Method (ISM) that greatly improves the scalability of KDAC. Our algorithm is intuitive, relies on easily implementable spectral decompositions, and comes with theoretical guarantees. Its computation time improves upon existing implementations of KDAC by as much as 5 orders of magnitude. Chieh Wu, Stratis Ioannidis, Mario Sznaier, Xiangyu Li 0006, David R. Kaeli, Jennifer G. Dy |
AISTATS | 6 |
| 2018 | Interactive Kernel Dimension Alternative Clustering on GPUsabstractMachine learning has seen tremendous growth in recent years thanks to two key advances in technology: massive data generation and highly-parallel accelerator architectures. The rate that data is being generated is exploding across multiple domains, including medical research, environmental science, web-search, and e-commerce. Many of these advances have benefited from emergent web-based applications, and improvements in data storage and sensing technologies. Innovations in parallel accelerator hardware, such as GPUs, has made it possible to process massive amounts of data in a timely fashion. Given these advanced data acquisition technology and hardware, machine learning researchers are equipped to generate and sift through much larger and complex datasets quickly. In this work, we focus on accelerating Kernel Dimension Alternative Clustering algorithms using GPUs. We conduct a thorough performance analysis by using both synthetic and real-world datasets, while also modifying both the structure of the data, and the size of the datasets. Our GPU implementation reduces execution time from minutes to seconds, which enables us to develop a web-based application for users to, interactively, view alternative clustering solutions. Xiangyu Li 0006, Chieh Wu, Shi Dong 0002, Jennifer G. Dy, David R. Kaeli |
ASONAM | 4 |
| 2018 | A Hybrid Approach to Identifying Key Factors in Environmental Health StudiesabstractIn recent years, the availability of data-driven analytics has become a key tool in discovery in public health and environmental science research. As a result, these communities have looked to leverage recent advances in machine learning algorithms. This class of algorithms are able to find hidden patterns and develop new knowledge in complex data, accelerating the rate of discovery in multiple research domains. In this paper, we present our methodology of applying machine learning algorithms to health outcomes, chemical exposures, and social behavior data from expectant mothers, as part of the NIEHS-supported PROTECT Center. The ultimate goal is to determine the dominant factors/features potentially responsible for the high rate of premature births in Puerto Rico.Many commonly-used machine learning algorithms can be used for feature selection. However, given the imbalance in our birth outcome data, with many more term (i.e., 37 weeks or longer) versus preterm pregnancies (i.e., less than 37 weeks), analysis of the PROTECT dataset presents many unique challenges. In addition to outcome imbalance, our database contains both quantitative and categorical data variables, adding some complexity to the analytical methods used. Applying straightforward correlation or regression analysis would be insufficient. Our datasets also contain a significant amount of missing data (incomplete records), providing noisy input to our algorithms. A further challenge is that we are working with a relatively limited set of complex data (only 2000 participants to date), so our models must be able to be built with a relatively small number of data samples.To overcome these challenges, we have implemented a cus-tomized end-to-end analytical toolchain which forms a pre-processing pipeline. Our framework performs general data filtering and handles missing data fields using a similarity-based approach. Next, we apply one of a number of different machine learning algorithms, including Linear Correlation, Normalized Mutual Information, Logistic Regression, and Decision Trees. We use these during both feature selection and model performance evaluation. Finally, we present top-ranked features produced by our model as potential key contributors of high preterm birth rates in Puerto Rico, and discuss results across these algorithms. Shi Dong 0002, Zlatan Feric, Xiangyu Li 0006, Sheikh Mokhlesur Rahman, Chieh Wu, April Z. Gu, Jennifer G. Dy, David R. Kaeli, John Meeker, Ingrid Y. Padilla, José Cordero, Carmen Velez Vega, Zaira Rosario, Akram Alshawabkeh |
IEEE BigData | 8 |
| 2018 | Experimental Design under the Bradley-Terry ModelabstractLabels generated by human experts via comparisons exhibit smaller variance compared to traditional sample labels. Collecting comparison labels is challenging over large datasets, as the number of comparisons grows quadratically with the dataset size. We study the following experimental design problem: given a budget of expert comparisons, and a set of existing sample labels, we determine the comparison labels to collect that lead to the highest classification improvement. We study several experimental design objectives motivated by the Bradley-Terry model. The resulting optimization problems amount to maximizing submodular functions. We experimentally evaluate the performance of these methods over synthetic and real-life datasets. Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Deniz Erdogmus, Jennifer G. Dy, Stratis Ioannidis |
IJCAI | 8 |
| 2018 | Quantifying Uncertainty in Discrete-Continuous and Skewed Data with Bayesian Deep LearningabstractDeep Learning (DL) methods have been transforming computer vision with innovative adaptations to other domains including climate change. For DL to pervade Science and Engineering (S&EE) applications where risk management is a core component, well-characterized uncertainty estimates must accompany predictions. However, S&E observations and model-simulations often follow heavily skewed distributions and are not well modeled with DL approaches, since they usually optimize a Gaussian, or Euclidean, likelihood loss. Recent developments in Bayesian Deep Learning (BDL), which attempts to capture uncertainties from noisy observations, aleatoric, and from unknown model parameters, epistemic, provide us a foundation. Here we present a discrete-continuous BDL model with Gaussian and lognormal likelihoods for uncertainty quantification (UQ). We demonstrate the approach by developing UQ estimates on "DeepSD'', a super-resolution based DL model for Statistical Downscaling (SD) in climate applied to precipitation, which follows an extremely skewed distribution. We find that the discrete-continuous models outperform a basic Gaussian distribution in terms of predictive accuracy and uncertainty calibration. Furthermore, we find that the lognormal distribution, which can handle skewed distributions, produces quality uncertainty estimates at the extremes. Such results may be important across S&E, as well as other domains such as finance and economics, where extremes are often of significant interest. Furthermore, to our knowledge, this is the first UQ model in SD where both aleatoric and epistemic uncertainties are characterized. Thomas Vandal, Evan Kodra, Jennifer G. Dy, Sangram Ganguly, Ramakrishna R. Nemani, Auroop R. Ganguly |
KDD | 3 |
| 2018 | A Multiresolution Convolutional Neural Network with Partial Label Training for Annotating Reflectance Confocal Microscopy Images of Skin
Alican Bozkurt, Kivanç Köse, Christi Alessi-Fox, Melissa Gill, Jennifer G. Dy, Dana H. Brooks, Milind Rajadhyaksha |
MICCAI (2) | 5 |
| 2018 | A Probabilistic Active Learning Algorithm Based on Fisher Information RatioabstractThe task of labeling samples is demanding and expensive. Active learning aims to generate the smallest possible training data set that results in a classifier with high performance in the test phase. It usually consists of two steps of selecting a set of queries and requesting their labels. Among the suggested objectives to score the query sets, information theoretic measures have become very popular. Yet among them, those based on Fisher information (FI) have the advantage of considering the diversity among the queries and tractable computations. In this work, we provide a practical algorithm based on Fisher information ratio to obtain query distribution for a general framework where, in contrast to the previous FI-based querying methods, we make no assumptions over the test distribution. The empirical results on synthetic and real-world data sets indicate that this algorithm gives competitive results. Jamshid Sourati, Murat Akçakaya, Deniz Erdogmus, Todd K. Leen, Jennifer G. Dy |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | Informative Subspace Learning for Counterfactual InferenceabstractInferring causal relations from observational data is widely used for knowledge discovery in healthcare and economics. To investigate whether a treatment can affect an outcome of interest, we focus on answering counterfactual questions of this type: what would a patient’s blood pressure be had he/she received a different treatment? Nearest neighbor matching (NNM) sets the counterfactual outcome of any treatment (control) sample to be equal to the factual outcome of its nearest neighbor in the control (treatment) group. Although being simple, flexible and interpretable, most NNM approaches could be easily misled by variables that do not affect the outcome. In this paper, we address this challenge by learning subspaces that are predictive of the outcome variable for both the treatment group and control group. Applying NNM in the learned subspaces leads to more accurate estimation of the counterfactual outcomes and therefore treatment effects. We introduce an informative subspace learning algorithm by maximizing the nonlinear dependence between the candidate subspace and the outcome variable measured by the Hilbert-Schmidt Independence Criterion (HSIC). We propose a scalable estimator of HSIC, called HSIC-RFF that reduces the quadratic computational and storage complexities (with respect to the sample size) of the naive HSIC implementation to linear through constructing random Fourier features. We also prove an upper bound on the approximation error of the HSIC-RFF estimator. Experimental results on simulated datasets and real-world datasets demonstrate our proposed approach outperforms existing NNM approaches and other commonly used regression-based methods for counterfactual inference. Yale Chang, Jennifer G. Dy |
AAAI | 2 |
| 2017 | Rate Optimal Estimation for High Dimensional Spatial Covariance MatricesabstractSpatial covariance matrix estimation is of great significance in many applications in climatology, econometrics and many other fields with complex data structures involving spatial dependencies. High dimensionality brings new challenges to this problem, and no theoretical optimal estimator has been proved for the spatial high-dimensional covariance matrix. Over the past decade, the method of regularization has been introduced to high-dimensional covariance estimation for various structured matrices, to achieve rate optimal estimators. In this paper, we aim to bridge the gap in these two research areas. We use a structure of block bandable covariance matrices to incorporate spatial dependence information, and study rate optimal estimation of this type of structured high dimensional covariance matrices. A double tapering estimator is proposed, and is shown to achieve the asymptotic minimax error bound. Numerical studies on both synthetic and real data are conducted showing the improvement of the double tapering estimator over the sample covariance matrix estimator. A. Adam Ding, Jennifer G. Dy |
ACML | 3 |
| 2017 | Clustering from Multiple Uncertain ExpertsabstractUtilizing expert input often improves clustering performance. However in a knowledge discovery problem, ground truth is unknown even to an expert. Thus, instead of one expert, we solicit the opinion from multiple experts. The key question motivating this work is: which experts should be assigned higher weights when there is disagreement on whether to put a pair of samples in the same group? To model the uncertainty in constraints from different experts, we build a probabilistic model for pairwise constraints through jointly modeling each expert’s accuracy and the mapping from features to latent cluster assignments. After learning our probabilistic discriminative clustering model and accuracies of different experts, 1) samples that were not annotated by any expert can be clustered using the discriminative clustering model; and 2) experts with higher accuracies are automatically assigned higher weights in determining the latent cluster assignments. Experimental results on UCI benchmark datasets and a real-world disease subtyping dataset demonstrate that our proposed approach outperforms competing alternatives, including semi-crowdsourced clustering, semi-supervised clustering with constraints from majority voting, and consensus clustering. Yale Chang, Junxiang Chen, Michael H. Cho, Peter J. Castaldi, Edwin K. Silverman, Jennifer G. Dy |
AISTATS | 6 |
| 2017 | Multiple Clustering Views from Multiple Uncertain ExpertsabstractExpert input can improve clustering performance. In today’s collaborative environment, the availability of crowdsourced multiple expert input is becoming common. Given multiple experts’ inputs, most existing approaches can only discover one clustering structure. However, data is multi-faced by nature and can be clustered in different ways (also known as views). In an exploratory analysis problem where ground truth is not known, different experts may have diverse views on how to cluster data. In this paper, we address the problem on how to automatically discover multiple ways to cluster data given potentially diverse inputs from multiple uncertain experts. We propose a novel Bayesian probabilistic model that automatically learns the multiple expert views and the clustering structure associated with each view. The benefits of learning the experts’ views include 1) enabling the discovery of multiple diverse clustering structures, and 2) improving the quality of clustering solution in each view by assigning higher weights to experts with higher confidence. In our approach, the expert views, multiple clustering structures and expert confidences are jointly learned via variational inference. Experimental results on synthetic datasets, benchmark datasets and a real-world disease subtyping problem show that our proposed approach outperforms competing baselines, including meta clustering, semi-supervised clustering, semi-crowdsourced clustering and consensus clustering. Yale Chang, Junxiang Chen, Michael H. Cho, Peter J. Castaldi, Edwin K. Silverman, Jennifer G. Dy |
ICML | 6 |
| 2017 | Clustering with Domain-Specific Usefulness ScoresabstractClustering is a challenging problem because given the same data set, it can be grouped in multiple different ways. Which of these clustering solutions is interesting depends on its domain application. Thus, incorporating domain expert input often improves clustering performance. However, most existing semi-supervised clustering techniques can only incorporate instance-level constraints (a few labels or must-link/cannot-link constraints), which domain experts may not be comfortable providing in knowledge discovery problems because categories are not known. Fortunately, domain experts often have an idea regarding properties that clustering solutions should have in order to be useful in domain application based on domain relevant scores. In this paper, we provide a framework for jointly optimizing the usefulness and quality of a clustering solution. Experiments on a synthetic data, a benchmark data, and a real-world disease subtyping problem demonstrate the usefulness of our proposed approach. Yale Chang, Junxiang Chen, Michael H. Cho, Peter J. Castaldi, Edwin K. Silverman, Jennifer G. Dy |
SDM | 6 |
| 2017 | A Robust-Equitable Measure for Feature Ranking and SelectionabstractIn many applications, not all the features used to represent data samples are important. Often only a few features are relevant for the prediction task. The choice of dependence measures often affect the final result of many feature selection methods. To select features that have complex nonlinear relationships with the response variable, the dependence measure should be equitable, a concept proposed by Reshef et al. (2011); that is, the dependence measure treats linear and nonlinear relationships equally. Recently, Kinney and Atwal (2014) gave a mathematical definition of self- equitability. In this paper, we introduce a new concept of robust-equitability and identify a robust- equitable copula dependence measure, the robust copula dependence (RCD) measure. RCD is based on the $L_1$-distance of the copula density from uniform and we show that it is equitable under both equitability definitions. We also prove theoretically that RCD is much easier to estimate than mutual information. Because of these theoretical properties, the RCD measure has the following advantages compared to existing dependence measures: it is robust to different relationship forms and robust to unequal sample sizes of different features. Experiments on both synthetic and real-world data sets confirm the theoretical analysis, and illustrate the advantage of using the dependence measure RCD for feature selection. A. Adam Ding, Jennifer G. Dy, Yale Chang |
J. Mach. Learn. Res. | 2 |
| 2017 | Asymptotic Analysis of Objectives Based on Fisher Information in Active LearningabstractObtaining labels can be costly and time-consuming. Active learning allows a learning algorithm to intelligently query samples to be labeled for a more efficient learning. Fisher information ratio (FIR) has been used as an objective for selecting queries. However, little is known about the theory behind the use of FIR for active learning. There is a gap between the underlying theory and the motivation of its usage in practice. In this paper, we attempt to fill this gap and provide a rigorous framework for analyzing existing FIR-based active learning methods. In particular, we show that FIR can be asymptotically viewed as an upper bound of the expected variance of the log-likelihood ratio. Additionally, our analysis suggests a unifying framework that not only enables us to make theoretical comparisons among the existing querying methods based on FIR, but also allows us to give insight into the development of new active learning approaches based on this objective. Jamshid Sourati, Murat Akçakaya, Todd K. Leen, Deniz Erdogmus, Jennifer G. Dy |
J. Mach. Learn. Res. | 5 |
| 2017 | Subject-specific abnormal region detection in traumatic brain injury using sparse model selection on high dimensional diffusion data
Matineh Shaker, Deniz Erdogmus, Jennifer G. Dy, Sylvain Bouix |
Medical Image Anal. | 3 |
| 2017 | Automated Target Detection for Geophysical ApplicationsabstractIn many geophysical surveys, there is a predefined goal-to detect and locate very specific anomalies, those that correspond to buried objects (targets). The types of targets range from various types of pipes (metallic or not), to rebars or wires in walls to land mines. This paper presents a novel unsupervised method for automatically detecting targets, and extracting information about them and the medium in which they reside. Most existing detection methods are supervised, which means that one has to provide a training set (which can be labor expensive) in order to train a classifier. By contrast, the method presented here is unsupervised and is model based, which alleviates the need to manually annotate a training set. Another drawback of many existing methods is the underlying assumption of a homogeneous medium. This assumption is greatly relaxed for this method, since it assumes no a priori knowledge of the medium. Instead, it learns the medium's properties from the targets themselves. Furthermore, our method is designed to be computationally efficient and applicable in real-time applications. It was implemented on the StructureScan Mini XT system (Geophysical Survey Systems, Inc.), and the runtime on that system was measured to be 20 μs per scan of 512 samples. Experiments on 50 ground penetrating radar images with 278 targets show that our method is able to detect the targets with high positioning accuracy, with a 95.3% detection rate and near-zero false alarm rate. Uri Pe'er, Jennifer G. Dy |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | A Marked Poisson Process Driven Latent Shape Model for 3D Segmentation of Reflectance Confocal Microscopy Image Stacks of Human SkinabstractSegmenting objects of interest from 3D data sets is a common problem encountered in biological data. Small field of view and intrinsic biological variability combined with optically subtle changes of intensity, resolution, and low contrast in images make the task of segmentation difficult, especially for microscopy of unstained living or freshly excised thick tissues. Incorporating shape information in addition to the appearance of the object of interest can often help improve segmentation performance. However, the shapes of objects in tissue can be highly variable and design of a flexible shape model that encompasses these variations is challenging. To address such complex segmentation problems, we propose a unified probabilistic framework that can incorporate the uncertainty associated with complex shapes, variable appearance, and unknown locations. The driving application that inspired the development of this framework is a biologically important segmentation problem: the task of automatically detecting and segmenting the dermal-epidermal junction (DEJ) in 3D reflectance confocal microscopy (RCM) images of human skin. RCM imaging allows noninvasive observation of cellular, nuclear, and morphological detail. The DEJ is an important morphological feature as it is where disorder, disease, and cancer usually start. Detecting the DEJ is challenging, because it is a 2D surface in a 3D volume which has strong but highly variable number of irregularly spaced and variably shaped "peaks and valleys." In addition, RCM imaging resolution, contrast, and intensity vary with depth. Thus, a prior model needs to incorporate the intrinsic structure while allowing variability in essentially all its parameters. We propose a model which can incorporate objects of interest with complex shapes and variable appearance in an unsupervised setting by utilizing domain knowledge to build appropriate priors of the model. Our novel strategy to model this structure combines a spatial Poisson process with shape priors and performs inference using Gibbs sampling. Experimental results show that the proposed unsupervised model is able to automatically detect the DEJ with physiologically relevant accuracy in the range 10- 20 μm . Sindhu Ghanta, Michael I. Jordan, Kivanç Köse, Dana H. Brooks, Milind Rajadhyaksha, Jennifer G. Dy |
IEEE Trans. Image Process. | 6 |
| 2017 | A Bayesian Nonparametric Model for Disease Subtyping: Application to Emphysema PhenotypesabstractWe introduce a novel Bayesian nonparametric model that uses the concept of disease trajectories for disease subtype identification. Although our model is general, we demonstrate that by treating fractions of tissue patterns derived from medical images as compositional data, our model can be applied to study distinct progression trends between population subgroups. Specifically, we apply our algorithm to quantitative emphysema measurements obtained from chest CT scans in the COPDGene Study and show several distinct progression patterns. As emphysema is one of the major components of chronic obstructive pulmonary disease (COPD), the third leading cause of death in the United States [1], an improved definition of emphysema and COPD subtypes is of great interest. We investigate several models with our algorithm, and show that one with age , pack years (a measure of cigarette exposure), and smoking status as predictors gives the best compromise between estimated predictive performance and model complexity. This model identified nine subtypes which showed significant associations to seven single nucleotide polymorphisms (SNPs) known to associate with COPD. Additionally, this model gives better predictive accuracy than multiple, multivariate ordinary least squares regression as demonstrated in a five-fold cross validation analysis. We view our subtyping algorithm as a contribution that can be applied to bridge the gap between CT-level assessment of tissue composition to population-level analysis of compositional trends that vary between disease subtypes. James C. Ross, Peter J. Castaldi, Michael H. Cho, Junxiang Chen, Yale Chang, Jennifer G. Dy, Edwin K. Silverman, George R. Washko, Raúl San José Estépar |
IEEE Trans. Medical Imaging | 6 |
| 2016 | A Robust-Equitable Copula Dependence Measure for Feature SelectionabstractFeature selection aims to select relevant features to improve the performance of predictors. Many feature selection methods depend on the choice of dependence measures. To select features that have complex nonlinear relationships with the response variable, the dependence measure should be equitable: treating linear and nonlinear relationships equally. In this paper we introduce the concept of robust-equitability and a robust-equitable dependence measure copula correlation (Ccor). This measure has the following advantages compared to existing dependence measures: it is robust to different relationship forms and robust to unequal sample sizes of different features. In contrast, existing dependence measures cannot take these factors into account simultaneously. Experiments on synthetic and real-world datasets confirm our theoretical analysis, and illustrates its advantage in feature selection. Yale Chang, A. Adam Ding, Jennifer G. Dy |
AISTATS | 4 |
| 2016 | Interpretable Clustering via Discriminative Rectangle Mixture ModelabstractClustering is a technique that is usually applied as a tool for exploratory data analysis. Because of the exploratory nature of this task, it would be beneficial if a clustering method generates interpretable results, and allows incorporating domain knowledge. This motivates us to develop a probabilistic discriminative model that learns a rectangular decision rule for each cluster, we call Discriminative Rectangle Mixture (DReaM) model. DReaM gives interpretable clustering results, because the rectangular decision rules discovered explicitly illustrate how one cluster is defined and differs from other clusters. It also facilitates us to take advantage of existing rules because we can choose informative prior distributions for the rectangular rules. Moreover, DReaM allows that the features for generating rules do not have to be the same as the features for discovering cluster structure. We approximate the distribution for the rules discovered via variational inference. Experimental results demonstrate that DReaM gives more interpretable clustering results, and yet its performance is comparable to existing clustering methods when solving traditional clustering. Furthermore, in real applications, DReaM is able to effectively take advantage of domain knowledge, and to generate reasonable clustering results. Junxiang Chen, Yale Chang, Brian D. Hobbs, Peter J. Castaldi, Michael H. Cho, Edwin K. Silverman, Jennifer G. Dy |
ICDM | 7 |
| 2016 | A Non-parametric Approach to Detect Epileptogenic Lesions using Restricted Boltzmann MachinesabstractVisual detection of lesional areas on a cortical surface is critical in rendering a successful surgical operation for Treatment Resistant Epilepsy (TRE) patients. Unfortunately, 45% of Focal Cortical Dysplasia (FCD, the most common kind of TRE) patients have no visual abnormalities in their brains' 3D-MRI images. We collaborate with doctors from NYU Langone's Comprehensive Epilepsy Center and apply machine learning methodologies to identify the resective zones for these {MRI-negative} FCD patients. Our task is particularly challenging because MRI images can only provide a limited number of features. Furthermore, data from different patients often exhibit inter-patient variabilities due to age, gender, left/right handedness, etc. In this paper, we introduce a new approach which combines the restricted Boltzmann machines and a Bayesian non-parametric mixture model to address these issues. We demonstrate the efficacy of our model by applying it to a retrospective dataset of MRI-negative FCD patients who are seizure free after surgery. Thomas Thesen, Karen E. Blackmon, Jennifer G. Dy, Carla E. Brodley, Ruben Kuzniecky, Orrin Devinsky |
KDD | 5 |
| 2016 | A Generative Block-Diagonal Model for Clustering
Junxiang Chen, Jennifer G. Dy |
UAI | 2 |
| 2015 | A Sparse Combined Regression-Classification Formulation for Learning a Physiological Alternative to Clinical Post-Traumatic Stress Disorder ScoresabstractCurrent diagnostic methods for mental pathologies, including Post-Traumatic Stress Disorder (PTSD), involve a clinician-coded interview, which can be subjective. Heart rate and skin conductance, as well as other peripheral physiology measures, have previously shown utility in predicting binary diagnostic decisions. The binary decision problem is easier, but misses important information on the severity of the patient’s condition. This work utilizes a novel experimental set-up that exploits virtual reality videos and peripheral physiology for PTSD diagnosis. In pursuit of an automated physiology-based objective diagnostic method, we propose a learning formulation that integrates the description of the experimental data and expert knowledge on desirable properties of a physiological diagnostic score. From a list of desired criteria, we derive a new cost function that combines regression and classification while learning the salient features for predicting physiological score. The physiological score produced by Sparse Combined Regression-Classification (SCRC) is assessed with respect to three sets of criteria chosen to reflect design goals for an objective, physiological PTSD score: parsimony and context of selected features, diagnostic score validity, and learning generalizability. For these criteria, we demonstrate that Sparse Combined Regression-Classification performs better than more generic learning approaches. Sarah Marie Brown, Andrea Webb, Rami Mangoubi, Jennifer G. Dy |
AAAI | 4 |
| 2015 | Domain Induced Dirichlet Mixture of Gaussian Processes: An Application to Predicting Disease Progression in Multiple Sclerosis PatientsabstractPredicting disease course is critical in chronic progressive diseases such as multiple sclerosis (MS) for determining treatment. Forming an accurate predictive model based on clinical data is particularly challenging when data is gathered from multiple clinics/physicians as the labels vary with physicians' subjective judgment about clinical tests and further we have no a priori knowledge of the various types of physician subjectivity. At the same time, we often have some (limited) domain knowledge on how to group patients into disease progression subgroups. In this paper, we first present our rationale for choosing a Dirichlet mixture of Gaussian processes (DPMGP) model to address the subjectivity in our data. We then introduce a new approach to incorporating domain knowledge into the non-parametric mixture model. We demonstrate the efficacy of our model by applying it to two medical datasets to predict disease progression in MS patients and disability levels in early Parkinson's patients. Tanuja Chitnis, Brian C. Healy, Jennifer G. Dy, Carla E. Brodley |
ICDM | 4 |
| 2015 | Clustering and Ranking in Heterogeneous Information Networks via Gamma-Poisson ModelabstractClustering and ranking have been successfully applied independently to homogeneous information networks, containing only one type of objects. However, real-world information networks are oftentimes heterogeneous, containing multiple types of objects and links. Recent research has shown that clustering and ranking can actually mutually enhance each other, and several techniques have been developed to integrate clustering and ranking together on a heterogeneous information network. To the best our knowledge, however, all of such techniques assume the network follows a certain schema. In this paper, we propose a probabilistic generative model that simultaneously achieves clustering and ranking on a heterogeneous network that can follow arbitrary schema, where the edges from different types are sampled from a Poisson distribution with the parameters determined by the ranking scores of the nodes in each cluster. A variational Bayesian inference method is proposed to learn these parameters, which can be used to output ranking and clusters simultaneously. Our method is evaluated on both synthetic and real-world networks extracted from the DBLP and YELP data. Experimental results show that our method outperforms the state-of-the-art baselines. Junxiang Chen, Yizhou Sun, Jennifer G. Dy |
SDM | 4 |
| 2015 | MultiClust special issue on discovering, summarizing and using multiple clusterings
Emmanuel Müller, Ira Assent, Stephan Günnemann, Thomas Seidl 0001, Jennifer G. Dy |
Mach. Learn. | 5 |
| 2014 | Dual beta process priors for latent cluster discovery in chronic obstructive pulmonary diseaseabstractChronic obstructive pulmonary disease (COPD) is a lung disease characterized by airflow limitation usually associated with an inflammatory response to noxious particles, such as cigarette smoke. COPD is currently the third leading cause of death in the United States and is the only leading cause of death that is increasing in prevalence. It also represents an enormous financial burden to society, costing tens of billions of dollars annually in the U.S. It is widely accepted by the medical community that COPD is a heterogeneous disease, with substantial evidence indicating that genetic variation contributes to varying levels of disease susceptibility. This heterogeneity makes it difficult to predict health decline and develop targeted treatments for better patient care. Although researchers have made several attempts to discover disease subtypes, results have been inconclusive, in part because standard clustering methods have not properly dealt with disease manifestations that may worsen with increased exposure. In this paper we introduce a transformative way of looking at the COPD subtyping task. Specifically, we model the relationship between risk factors (such as age and smoke exposure) and manifestations of disease severity using Gaussian Processes, which allow us to represent so-called "disease trajectories". We also posit that individuals can be associated with multiple disease types (latent clusters), which we assume are influenced by genetics. Furthermore, we predict that only subsets of the numerous disease-related quantitative features are useful for describing each latent subtype. We model these associations using two separate beta process priors, and we describe a variational inference approach to discover the most probable latent cluster assignments. Results are validated with associations to genetic markers. James C. Ross, Peter J. Castaldi, Michael H. Cho, Jennifer G. Dy |
KDD | 4 |
| 2014 | Harnessing the Power of GPUs to Speed Up Feature Selection for Outlier Detection
Fatemeh Azmandian, Ayse Yilmazer, Jennifer G. Dy, Javed A. Aslam, David R. Kaeli |
J. Comput. Sci. Technol. | 3 |
| 2014 | Learning from multiple annotators with varying expertise
Yan Yan 0024, Rómer Rosales, Glenn Fung, Subramanian Ramanathan, Jennifer G. Dy |
Mach. Learn. | 5 |
| 2014 | Iterative Discovery of Multiple AlternativeClustering ViewsabstractComplex data can be grouped and interpreted in many different ways. Most existing clustering algorithms, however, only find one clustering solution, and provide little guidance to data analysts who may not be satisfied with that single clustering and may wish to explore alternatives. We introduce a novel approach that provides several clustering solutions to the user for the purposes of exploratory data analysis. Our approach additionally captures the notion that alternative clusterings may reside in different subspaces (or views). We present an algorithm that simultaneously finds these subspaces and the corresponding clusterings. The algorithm is based on an optimization procedure that incorporates terms for cluster quality and novelty relative to previously discovered clustering solutions. We present a range of experiments that compare our approach to alternatives and explore the connections between simultaneous and iterative modes of discovery of multiple clusterings. Donglin Niu, Jennifer G. Dy, Michael I. Jordan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Accelerated Learning-Based Interactive Image Segmentation Using Pairwise ConstraintsabstractAlgorithms for fully automatic segmentation of images are often not sufficiently generic with suitable accuracy, and fully manual segmentation is not practical in many settings. There is a need for semiautomatic algorithms, which are capable of interacting with the user and taking into account the collected feedback. Typically, such methods have simply incorporated user feedback directly. Here, we employ active learning of optimal queries to guide user interaction. Our work in this paper is based on constrained spectral clustering that iteratively incorporates user feedback by propagating it through the calculated affinities. The original framework does not scale well to large data sets, and hence is not straightforward to apply to interactive image segmentation. In order to address this issue, we adopt advanced numerical methods for eigen-decomposition implemented over a subsampling scheme. Our key innovation, however, is an active learning strategy that chooses pairwise queries to present to the user in order to increase the rate of learning from the feedback. Performance evaluation is carried out on the Berkeley segmentation and Graz-02 image data sets, confirming that convergence to high accuracy levels is realizable in relatively few iterations. Jamshid Sourati, Deniz Erdogmus, Jennifer G. Dy, Dana H. Brooks |
IEEE Trans. Image Process. | 3 |
| 2013 | Nonparametric Mixture of Gaussian Processes with ConstraintsabstractMotivated by the need to identify new and clinically relevant categories of lung disease, we propose a novel clustering with constraints method using a Dirichlet process mixture of Gaussian processes in a variational Bayesian nonparametric framework. We claim that individuals should be grouped according to biological and/or genetic similarity regardless of their level of disease severity; therefore, we introduce a new way of looking at subtyping/clustering by recasting it in terms of discovering associations of individuals to disease trajectories (i.e., grouping individuals based on their similarity in response to environmental and/or disease causing variables). The nonparametric nature of our algorithm allows for learning the unknown number of meaningful trajectories. Additionally, we acknowledge the usefulness of expert guidance by providing for their input using must-link and cannot- link constraints. These constraints are encoded with Markov random fields. We also provide an efficient variational approach for performing inference on our model. James C. Ross, Jennifer G. Dy |
ICML (3) | 2 |
| 2012 | Unsupervised wrinkle detection in reflectance confocal microscopy images of the human skinabstractReflectance confocal microscopy (RCM) is a non-invasive and in-vivo imaging modality, which can take images from different depths of the human skin. A challenging problem is to detect a clinically important subsurface section of the skin, the Dermis/Epidermis junction, in RCM images. This is a tough problem because of the huge variation of texture and intensity features across both intersubject and intrasubject tissues. On the other hand, there's almost no wrinkle-free part of the skin. This well-known phenomenon can be used as a histological clue for guessing the probability of being Dermis or Epidermis in the neighboring regions. In this paper, we develop a two-step wrinkle detector for RCM images. By analyzing the results on different RCM images, we conclude it has high sensitivity and specificity, but a relatively lower Jaccard index. Jamshid Sourati, Dana H. Brooks, Jennifer G. Dy, Esra Ataer Cansizoglu, Deniz Erdogmus, Milind Rajadhyaksha |
ICASSP | 3 |
| 2012 | Feature Weighting and Selection Using Hypothesis Margin of BoostingabstractUtilizing the concept of hypothesis margins to measure the quality of a set of features has been a growing line of research in the last decade. However, most previous algorithms have been developed under the large hypothesis margin principles of the 1-NN algorithm, such as Simba. Little attention has been paid so far to exploiting the hypothesis margins of boosting to evaluate features. Boosting is well known to maximize the training examples' hypothesis margins, in particular, the average margins which are known to be the first statistics that considers the whole margin distribution. In this paper, we describe how to utilize the training examples' mean margins of boosting to select features. A weight criterion, termed Margin Fraction (MF), is assigned to each feature that contributes to the average margin distribution combined in the final output produced by boosting. Applying the idea of MF to a sequential backward selection method, a new embedded selection algorithm is proposed, called SBS-MF. Experimentation is carried out using different data sets, which compares the proposed SBS-MF with two boosting based feature selection approaches, as well as to Simba. The results show that SBS-MF is effective in most of the cases. Malak Alshawabkeh, Javed A. Aslam, Jennifer G. Dy, David R. Kaeli |
ICDM | 3 |
| 2012 | GPU-Accelerated Feature Selection for Outlier Detection Using the Local Kernel Density RatioabstractEffective outlier detection requires the data to be described by a set of features that captures the behavior of normal data while emphasizing those characteristics of outliers which make them different than normal data. In this work, we present a novel non-parametric evaluation criterion for filter-based feature selection which caters to outlier detection problems. The proposed method seeks the subset of features that represents the inherent characteristics of the normal dataset while forcing outliers to stand out, making them more easily distinguished by outlier detection algorithms. Experimental results on real datasets show the advantage of our feature selection algorithm compared to popular and state-of-the-art methods. We also show that the proposed algorithm is able to overcome the small sample space problem and perform well on highly imbalanced datasets. Furthermore, due to the highly parallelizable nature of the feature selection, we implement the algorithm on a graphics processing unit (GPU) to gain significant speedup over the serial version. The benefits of the GPU implementation are two-fold, as its performance scales very well in terms of the number of features, as well as the number of data points. Fatemeh Azmandian, Ayse Yilmazer, Jennifer G. Dy, Javed A. Aslam, David R. Kaeli |
ICDM | 3 |
| 2012 | A Computational model for compressed sensing RNAi cellular screeningabstractBACKGROUND: RNA interference (RNAi) becomes an increasingly important and effective genetic tool to study the function of target genes by suppressing specific genes of interest. This system approach helps identify signaling pathways and cellular phase types by tracking intensity and/or morphological changes of cells. The traditional RNAi screening scheme, in which one siRNA is designed to knockdown one specific mRNA target, needs a large library of siRNAs and turns out to be time-consuming and expensive. RESULTS: In this paper, we propose a conceptual model, called compressed sensing RNAi (csRNAi), which employs a unique combination of group of small interfering RNAs (siRNAs) to knockdown a much larger size of genes. This strategy is based on the fact that one gene can be partially bound with several small interfering RNAs (siRNAs) and conversely, one siRNA can bind to a few genes with distinct binding affinity. This model constructs a multi-to-multi correspondence between siRNAs and their targets, with siRNAs much fewer than mRNA targets, compared with the conventional scheme. Mathematically this problem involves an underdetermined system of equations (linear or nonlinear), which is ill-posed in general. However, the recently developed compressed sensing (CS) theory can solve this problem. We present a mathematical model to describe the csRNAi system based on both CS theory and biological concerns. To build this model, we first search nucleotide motifs in a target gene set. Then we propose a machine learning based method to find the effective siRNAs with novel features, such as image features and speech features to describe an siRNA sequence. Numerical simulations show that we can reduce the siRNA library to one third of that in the conventional scheme. In addition, the features to describe siRNAs outperform the existing ones substantially. CONCLUSIONS: This csRNAi system is very promising in saving both time and cost for large-scale RNAi screening experiments which may benefit the biological research with respect to cellular processes and pathways. Hua Tan, Jiguang Bao, Jennifer G. Dy, Xiaobo Zhou 0001 |
BMC Bioinform. | 4 |
| 2011 | A Unified Probabilistic Model for Global and Local Unsupervised Feature Selection
Yue Guan 0001, Jennifer G. Dy, Michael I. Jordan |
ICML | 2 |
| 2011 | Active Learning from Crowds
Yan Yan 0024, Rómer Rosales, Glenn Fung, Jennifer G. Dy |
ICML | 4 |
| 2011 | A Novel Feature Selection for Intrusion Detection in Virtual Machine EnvironmentsabstractIntrusion detection systems (IDSs) are continuously evolving, with the goal of improving the security of computer infrastructures. However, one of the most significant challenges in this area is the poor detection rate, due to the presence of excessive features in a data set whose class distributions are imbalanced. Despite the relatively long existence and the promising nature of feature selection methods, most of them fail to account for imbalance class distributions, particularly, for intrusion data, leading to poor predictions for minority class samples. In this paper, we propose a new feature selection algorithm to enhance the accuracy of IDS of virtual server environments. Our algorithm assigns weights to subsets of features according to the maximized area under the ROC curve (AUC) margin it induces during the boosting process over the minority and the majority examples. The best subset of features is then selected by a greedy search strategy. The empirical experiments are carried out on multiple intrusion data sets using different commercial virtual appliances and real malwares. Malak Alshawabkeh, Javed A. Aslam, David R. Kaeli, Jennifer G. Dy |
ICTAI | 4 |
| 2011 | Workload Characterization at the Virtualization LayerabstractVirtualization technology has many attractive qualities including improved security, reliability, scalability, and resource sharing/management. As a result, virtualization has been deployed on an array of platforms, from mobile devices to high end enterprise servers. In this paper, we present a novel approach to working at a virtualization interface, performing workload characterization equipped with the information available at the virtual machine monitor (VMM) interface. Due to the semantic gap between the raw VMM-level data available and the true application behavior, we employ the power of regression techniques to extract meaningful information about a workload's behavior. We also demonstrate that the information available at the VMM level still retains rich workload characteristics that can be used to identify application behavior. We show that we are able to capture enough information about a workload to characterize and decompose it into a combination of CPU, memory, disk I/O, and network I/O-intensive components. Dissecting the behavior of a workload in terms of these components, we can develop significant insight into the behavior of any application. Workload characterization can be used for online performance monitoring, workload scheduling, workload trending, virtual machine (VM)health monitoring, and security analysis. We can also consider how VMM-based workload profiles can be used to detect anomalous behavior in virtualized environments by comparing a model of potentially malicious execution to that of normal execution. Fatemeh Azmandian, Micha Moffie, Jennifer G. Dy, Javed A. Aslam, David R. Kaeli |
MASCOTS | 3 |
| 2010 | From Transformation-Based Dimensionality Reduction to Feature Selection
Mahdokht Masaeli, Glenn Fung, Jennifer G. Dy |
ICML | 3 |
| 2010 | Multiple Non-Redundant Spectral Clustering Views
Donglin Niu, Jennifer G. Dy, Michael I. Jordan |
ICML | 2 |
| 2010 | Effective Virtual Machine Monitor Intrusion Detection Using Feature Selection on Highly Imbalanced DataabstractVirtualization is becoming an increasingly popular service hosting platform. Recently, intrusion detection systems (IDSs) which utilize virtualization have been introduced. One particular challenge present in current virtualization-based IDS systems is considered in this paper. IDS systems are commonly faced with high-dimensionality imbalanced data. Improved feature selection methods are needed to achieve more accurate detection when presented with imbalanced data. These methods must select the right set of features which will lead to a lower number of false alarms and higher correct detection rates. In this paper we propose a new Boosting-based feature selection that evaluates the relative importance of individual features using the fractional absolute confidence that Boosting produces. Our approach accounts for the sample distributions by optimizing for the area under the Receive Operating Characteristic (ROC) curve (i.e., Area Under the Curve(AUC)). Empirical results on different commercial virtual appliances and malwares indicate that proper input feature selection is key if we want an effective virtualization-based IDS that is lightweight, efficient and effective. Malak Alshawabkeh, Micha Moffie, Fatemeh Azmandian, Javed A. Aslam, Jennifer G. Dy, David R. Kaeli |
ICMLA | 5 |
| 2010 | Locally Deformable Shape Model to Improve 3D Level Set Based Esophagus SegmentationabstractIn this paper we propose a supervised 3D segmentation algorithm to locate the esophagus in thoracic CT scans using a variational framework. To address challenges due to low contrast, several priors are learned from a training set of segmented images. Our algorithm first estimates the centerline based on a spatial model learned at a few manually marked anatomical reference points. Then an implicit shape model is learned by subtracting the centerline and applying PCA to these shapes. To allow local variations in the shapes, we propose to use nonlinear smooth local deformations. Finally, the esophageal wall is located within a 3D level set framework by optimizing a cost function including terms for appearance, the shape model, smoothness constraints and an air/contrast model. Sila Kurugol, Necmiye Ozay, Jennifer G. Dy, Gregory C. Sharp, Dana H. Brooks |
ICPR | 3 |
| 2010 | Medical coding classification by leveraging inter-code relationshipsabstractMedical coding or classification is the process of transforming information contained in patient medical records into standard predefined medical codes. There are several worldwide accepted medical coding conventions associated with diagnoses and medical procedures; however, in the United States the Ninth Revision of ICD(ICD-9) provides the standard for coding clinical records. Accurate medical coding is important since it is used by hospitals for insurance billing purposes. Since after discharge a patient can be assigned or classified to several ICD-9 codes, the coding problem can be seen as a multi-label classification problem. In this paper, we introduce a multi-label large-margin classifier that automatically learns the underlying inter-code structure and allows the controlled incorporation of prior knowledge about medical code relationships. In addition to refining and learning the code relationships, our classifier can also utilize this shared information to improve its performance. Experiments on a publicly available dataset containing clinical free text and their associated medical codes showed that our proposed multi-label classifier outperforms related multi-label models in this problem. Yan Yan 0024, Glenn Fung, Jennifer G. Dy, Rómer Rosales |
KDD | 3 |
| 2010 | Convex Principal Feature SelectionabstractA popular approach for dimensionality reduction and data analysis is principal component analysis (PCA). A limiting factor with PCA is that it does not inform us on which of the original features are important. There is a recent interest in sparse PCA (SPCA). By applying an L1 regularizer to PCA, a sparse transformation is achieved. However, true feature selection may not be achieved as non-sparse coefficients may be distributed over several features. Feature selection is an NP-hard combinatorial optimization problem. This paper relaxes and re-formulates the feature selection problem as a convex continuous optimization problem that minimizes a mean-squared-reconstruction error (a criterion optimized by PCA) and considers feature redundancy into account (an important property in PCA and feature selection). We call this new method Convex Principal Feature Selection (CPFS). Experiments show that CPFS performed better than SPCA in selecting features that maximize variance or minimize the mean-squared-reconstruction error. Mahdokht Masaeli, Yan Yan 0024, Glenn Fung, Jennifer G. Dy |
SDM | 5 |
| 2010 | Modeling Multiple Annotator Expertise in the Semi-Supervised Learning Scenario
Yan Yan 0024, Rómer Rosales, Glenn Fung, Jennifer G. Dy |
UAI | 4 |
| 2010 | A Novel Approach to Monitor Rehabilitation Outcomes in Stroke Survivors Using Wearable TechnologyabstractQuantitative assessment of motor abilities in stroke survivors can provide valuable feedback to guide clinical interventions. Numerous clinical scales were developed in the past to assess levels of impairment and functional limitation in individuals after stroke. The Functional Ability Scale is one of these clinical scales. It is a 75-point scale used to evaluate the functional ability of subjects by grading movement quality during performance of 15 motor tasks. Performance of these motor tasks requires subjects to reach for objects (e.g., a pencil on a table) and manipulate them (e.g., lift the pencil). In this paper, we show that accelerometer data recorded during performance of a subset of the motor tasks pertaining to the Functional Ability Scale can be relied upon to derive accurate estimates of the scores provided by a clinician using this scale. Accelerometer-based estimates of clinical scores were obtained by segmenting the recordings into movement components (reaching, manipulation, release/return), extracting data features, selecting features that maximized the separation among classes associated with different clinical scores, feeding these features to Random Forests to estimate scores for individual motor tasks, and using a linear equation to estimate the total Functional Ability Scale score based on the sum of the clinical scores for individual motor tasks derived from the accelerometer data. Results showed that it is possible to achieve estimates of the total Functional Ability Scale score marked by a bias of only 0.04 points of the scale and a standard deviation of only 2.43 points when using as few as three sensors to collect data during performance of only six motor tasks. Shyamal Patel, Richard Hughes, Todd Hester, Joel Stein, Metin Akay, Jennifer G. Dy, Paolo Bonato |
Proc. IEEE | 6 |
| 2010 | Learning multiple nonredundant clusteringsabstractReal-world applications often involve complex data that can be interpreted in many different ways. When clustering such data, there may exist multiple groupings that are reasonable and interesting from different perspectives. This is especially true for high-dimensional data, where different feature subspaces may reveal different structures of the data. However, traditional clustering is restricted to finding only one single clustering of the data. In this article, we propose a new clustering paradigm for exploratory data analysis: find all non-redundant clustering solutions of the data, where data points in the same cluster in one solution can belong to different clusters in other partitioning solutions. We present a framework to solve this problem and suggest two approaches within this framework: (1) orthogonal clustering, and (2) clustering in orthogonal subspaces. In essence, both approaches find alternative ways to partition the data by projecting it to a space that is orthogonal to the current solution. The first approach seeks orthogonality in the cluster space, while the second approach seeks orthogonality in the feature space. We study the relationship between the two approaches. We also combine our framework with techniques for automatically finding the number of clusters in the different solutions, and study stopping criteria for determining when all meaningful solutions are discovered. We test our framework on both synthetic and high-dimensional benchmark data sets, and the results show that indeed our approaches were able to discover varied clustering solutions that are interesting and meaningful. Xiaoli Z. Fern, Jennifer G. Dy |
ACM Trans. Knowl. Discov. Data | 3 |
| 2009 | Multi-Class Classifiers and their Underlying Shared Structure
Volkan Vural, Glenn Fung, Rómer Rosales, Jennifer G. Dy |
IJCAI | 4 |
| 2009 | A fuzzy logics clustering approach to computing human attention allocation using eyegaze movement cue
Yingzi Lin, Wenjun Zhang 0005, Chieh Wu, Jennifer G. Dy |
Int. J. Hum. Comput. Stud. | 5 |
| 2009 | Using Local Dependencies within Batches to Improve Large Margin Classifiers
Volkan Vural, Glenn Fung, Balaji Krishnapuram, Jennifer G. Dy, R. Bharat Rao |
J. Mach. Learn. Res. | 4 |
| 2009 | Monitoring Motor Fluctuations in Patients With Parkinson's Disease Using Wearable SensorsabstractThis paper presents the results of a pilot study to assess the feasibility of using accelerometer data to estimate the severity of symptoms and motor complications in patients with Parkinson's disease. A support vector machine (SVM) classifier was implemented to estimate the severity of tremor, bradykinesia and dyskinesia from accelerometer data features. SVM-based estimates were compared with clinical scores derived via visual inspection of video recordings taken while patients performed a series of standardized motor tasks. The analysis of the video recordings was performed by clinicians trained in the use of scales for the assessment of the severity of Parkinsonian symptoms and motor complications. Results derived from the accelerometer time series were analyzed to assess the effect on the estimation of clinical scores of the duration of the window utilized to derive segments (to eventually compute data features) from the accelerometer data, the use of different SVM kernels and misclassification cost values, and the use of data features derived from different motor tasks. Results were also analyzed to assess which combinations of data features carried enough information to reliably assess the severity of symptoms and motor complications. Combinations of data features were compared taking into consideration the computational cost associated with estimating each data feature on the nodes of a body sensor network and the effect of using such data features on the reliability of SVM-based estimates of the severity of Parkinsonian symptoms and motor complications. Shyamal Patel, Konrad Lorincz, Richard Hughes, Nancy Huggins, John Growdon, David G. Standaert, Metin Akay, Jennifer G. Dy, Matt Welsh, Paolo Bonato |
IEEE Trans. Inf. Technol. Biomed. | 8 |
| 2008 | Learning methods for lung tumor markerless gating in image-guided radiotherapyabstractIn an idealized gated radiotherapy treatment, radiation is delivered only when the tumor is at the right position. For gated lung cancer radiotherapy, it is difficult to generate accurate gating signals due to the large uncertainties when using external surrogates and the risk of pneumothorax when using implanted fiducial markers. In this paper, we investigate machine learning algorithms for markerless gated radiotherapy with fluoroscopic images. Previous approach utilizes template matching to localize the tumor position. Here, we investigate two ways to improve the precision of tumor target localization by applying: (1) an ensemble of templates where the representative templates are selected by Gaussian mixture clustering, and (2) a support vector machine (SVM) classifier with radial basis kernels. Template matching only considers images inside the gating window, but images outside the gating window might provide additional information. We take advantage of both states and re-cast the gating problem into a classification problem. Thus, we are able to use the SVM classifier for gated radiotherapy. To verify the effectiveness of the two proposed techniques, we apply them on five sequences of fluoroscopic images from five lung cancer patients against the gating signal of manually contoured tumors as ground truth. Our five-patient case study shows that both ensemble template matching and SVM are reasonable tools for image-guided markerless gated radiotherapy with an average of approximately 95% precision in terms of delivered target dose at approximately 35% duty cycle. Jennifer G. Dy, Gregory C. Sharp, Brian M. Alexander, Steve B. Jiang |
KDD | 2 |
| 2008 | Impact of imputation of missing values on classification error for discrete data
Alireza Farhangfar, Lukasz A. Kurgan, Jennifer G. Dy |
Pattern Recognit. | 3 |
| 2007 | Non-redundant Multi-view Clustering via OrthogonalizationabstractTypical clustering algorithms output a single clustering of the data. However, in real world applications, data can often be interpreted in many different ways; data can have different groupings that are reasonable and interesting from different perspectives. This is especially true for high-dimensional data, where different feature subspaces may reveal different structures of the data. Why commit to one clustering solution while all these alternative clustering views might be interesting to the user. In this paper, we propose a new clustering paradigm for explorative data analysis: find all non-redundant clustering views of the data, where data points of one cluster can belong to different clusters in other views. We present a framework to solve this problem and suggest two approaches within this framework: (1) orthogonal clustering, and (2) clustering in orthogonal subspaces. In essence, both approaches find alternative ways to partition the data by projecting it to a space that is orthogonal to our current solution. The first approach seeks orthogonality in the cluster space, while the second approach seeks orthogonality in the feature space. We test our framework on both synthetic and high-dimensional benchmark data sets, and the results show that indeed our approaches were able to discover varied solutions that are interesting and meaningful. keywords: multi-view clustering, non-redundant clustering, orthogonalization Xiaoli Z. Fern, Jennifer G. Dy |
ICDM | 3 |
| 2007 | In search of deterministic methods for initializing K-means and Gaussian mixture clustering
Ting Su 0002, Jennifer G. Dy |
Intell. Data Anal. | 2 |
| 2006 | Batch Classification with Applications in Computer Aided Diagnosis
Volkan Vural, Glenn Fung, Balaji Krishnapuram, Jennifer G. Dy, R. Bharat Rao |
ECML | 4 |
| 2005 | Enabling a RealTime Solution for Neuron Detection with Reconfigurable Hardware (abstract only)abstractFPGAs provide a speed advantage in processing for embedded systems, especially when processing is moved close to the sensors. Perhaps the ultimate embedded system is a neural prosthetic, where probes are inserted into the brain and recorded electrical activity is analyzed to determine which neurons have fired. In turn, this information can be used to manipulate an external device such as a robot arm or a computer mouse. To make the detection of these signals possible, some baseline data must be processed to correlate impulses to particular neurons. One method for processing this data uses a statistical clustering algorithm called Expectation Maximization, or EM. In this paper, we examine the EM clustering algorithm, determine the most computationally intensive portion, map it onto a reconfigurable device, and show several areas of performance gain. Ben Cordes, Jennifer G. Dy, Miriam Leeser, James Goebel |
FPGA | 2 |
| 2005 | A multinomial clustering model for fast simulation of computer architecture designsabstractComputer architects utilize simulation tools to evaluate the merits of a new design feature. The time needed to adequately evaluate the tradeoffs associated with adding any new feature has become a critical issue. Recent work has found that by identifying execution phases present in common workloads used in simulation studies, we can apply clustering algorithms to significantly reduce the amount of time needed to complete the simulation. Our goal in this paper is to demonstrate the value of this approach when applied to the set of industry-standard benchmarks most commonly used in computer architecture studies. We also look to improve upon prior work by applying more appropriate clustering algorithms to identify phases, and to further reduce simulation time.We find that the phase clustering in computer architecture simulation has many similarities to text clustering. In prior work on clustering techniques to reduce simulation time, K-means clustering was used to identify representative program phases. In this paper we apply a mixture of multinomials to the clustering problem and show its advantages over using K-means on simulation data. We have implemented these two clustering algorithms and evaluate how well they can characterize program behavior. By adopting a mixture of multinomials model, we find that we can maintain simulation result fidelity, while greatly reducing overall simulation time. We report results for a range of applications taken from the SPEC2000 benchmark suite. Kaushal Sanghai, Ting Su 0002, Jennifer G. Dy, David R. Kaeli |
KDD | 3 |
| 2004 | Automated hierarchical mixtures of probabilistic principal component analyzersabstractMany clustering algorithms fail when dealing with high dimensional data. Principal component analysis (PCA) is a popular dimensionality reduction algorithm. However, it assumes a single multivariate Gaussian model, which provides a global linear projection of the data. Mixture of probabilistic principal component analyzers (PPCA) provides a better model to the clustering paradigm. It provides a local linear PCA projection for each multivariate Gaussian cluster component. We extend this model to build hierarchical mixtures of PPCA. Hierarchical clustering provides a flexible representation showing relationships among clusters in various perceptual levels. We introduce an automated hierarchical mixture of PPCA algorithm, which utilizes the integrated classification likelihood as a criterion for splitting and stopping the addition of hierarchical levels. An automated approach requires automated methods for initialization, determining the number of principal component dimensions, and determining when to split clusters. We address each of these in the paper. This automated approach results in a coarse to fine local component model with varying projections and with different number of dimensions for each cluster. Ting Su 0002, Jennifer G. Dy |
ICML | 2 |
| 2004 | A hierarchical method for multi-class support vector machinesabstractWe introduce a framework, which we call Divide-by-2 (DB2), for extending support vector machines (SVM) to multi-class prob-lems. DB2 offers an alternative to the stan-dard one-against-one and one-against-rest al-gorithms. For an N class problem, DB2 pro-duces an N − 1 node binary decision tree where nodes represent decision boundaries formed by N−1 SVM binary classifiers. This tree structure allows us to present a gener-alization and a time complexity analysis of DB2. Our analysis and related experiments show that, DB2 is faster than one-against-one and one-against-rest algorithms in terms of testing time, significantly faster than one-against-rest in terms of training time, and that the cross-validation accuracy of DB2 is comparable to these two methods. 1. Volkan Vural, Jennifer G. Dy |
ICML | 2 |
| 2004 | A Deterministic Method for Initializing K-Means ClusteringabstractThe performance of K-means clustering depends on the initial guess of partition. We motivate theoretically and experimentally the use of a deterministic divisive hierarchical method, which we refer to as PCA-Part (principal components analysis partitioning) for initialization. The criterion that K-means clustering minimizes is the SSE (sum-squared-error) criterion. The first principal direction (the eigenvector corresponding to the largest eigenvalue of the covariance matrix) is the direction which contributes the largest SSE. Hence, a good candidate direction to project a cluster for splitting is, then, the first principal direction. This is the basis for PCA-Part initialization method. Our experiments reveal that generally PCA-Part leads K-means to generate clusters with SSE values close to the minimum SSE values obtained by one hundred random start runs. In addition, this deterministic initialization method often leads K-means to faster convergence (less iterations) compared to random methods. Furthermore, we also theoretically show and confirm experimentally on synthetic data when PCA-Part may fail. Ting Su 0002, Jennifer G. Dy |
ICTAI | 2 |
| 2004 | Feature Selection for Unsupervised Learning
Jennifer G. Dy, Carla E. Brodley |
J. Mach. Learn. Res. | 1 |
| 2003 | Unsupervised Feature Selection Applied to Content-Based Retrieval of Lung ImagesabstractThis paper describes a new hierarchical approach to content-based image retrieval called the "customized-queries" approach (CQA). Contrary to the single feature vector approach which tries to classify the query and retrieve similar images in one step, CQA uses multiple feature sets and a two-step approach to retrieval. The first step classifies the query according to the class labels of the images using the features that best discriminate the classes. The second step then retrieves the most similar images within the predicted class using the features customized to distinguish "subclasses" within that class. Needing to find the customized feature subset for each class led us to investigate feature selection for unsupervised learning. As a result, we developed a new algorithm called FSSEM (feature subset selection using expectation-maximization clustering). We applied our approach to a database of high resolution computed tomography lung images and show that CQA radically improves the retrieval precision over the single feature vector approach. To determine whether our CBIR system is helpful to physicians, we conducted an evaluation trial with eight radiologists. The results show that our system using CQA retrieval doubled the doctors' diagnostic accuracy. Jennifer G. Dy, Carla E. Brodley, Avinash C. Kak, Lynn S. Broderick, Alex M. Aisen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Feature Subset Selection and Order Identification for Unsupervised Learning
Jennifer G. Dy, Carla E. Brodley |
ICML | 1 |
| 2000 | Visualization and interactive feature selection for unsupervised dataabstractFor many feature selection problems, a human denes the features that are potentially useful, and then a subset is chosen from the original pool of features using an automated feature selection algorithm. In contrast to supervised learning, class information is not available to guide the feature search for unsupervised learning tasks. In this paper, we introduce Visual-FSSEM (Visual Feature Subset Selection using Expectation-Maximization Clustering), which incorporates visualization techniques, clustering, and user interaction to guide the feature subset search and to enable a deeper understanding of the data. Visual-FSSEM, serves both as an exploratory and multivariate-data visualization tool. We illustrate Visual-FSSEM on a high-resolution computed tomography lung image data set. 1. INTRODUCTION Most research in unsupervised clustering assumes that when creating the target data set, the data analyst in conjunction with the domain expert was able to identify a small relevant set of ... Jennifer G. Dy, Carla E. Brodley |
KDD | 1 |
| 1999 | The Customized-Queries Approach to CBIR Using EMabstractThis paper makes two contributions. The first contribution is an approach called the "customized-queries" approach (CQA) to content-based image retrieval. The second is an algorithm called FSSEM that performs feature selection and clustering simultaneously. The customized queries approach first classifies a query using the features that best differentiate the major classes and then customizes the query to that class by using the features that best distinguish the images within the chosen major class. This approach is motivated by the observation that the features that are most effective in discriminating among images from different classes may not be the most effective for retrieval of visually similar images within a class. This occurs for domains in which not all pairs of images within one class have equivalent visual similarity, i.e., subclasses exists. Because we are not given subclass labels, we must simultaneously find the features that best discriminate the subclasses and at the same time find these subclasses. We use FSSEM to find these features. We apply this approach to content-based retrieval of high-resolution tomographic images of patients with lung disease and show that this approach radically improves the retrieval precision over the traditional approach that performs retrieval using a single feature vector. Jennifer G. Dy, Carla E. Brodley, Avinash C. Kak, Chi-Ren Shyu, Lynn S. Broderick |
CVPR | 1 |