VLDB 2026 Research / reviewers in the wild / expert
Artur Dubrawski
dblp:76/48 · also Artur W. Dubrawski
· DBLP profile ↗
73ranked-venue papers
7as first author
25since 2021 · last 2025
0000-0002-2372-0831ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 51 · 6 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 1 first-author · 5 since 2021Systems, architecture and hardware · 7 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring Representations and Interventions in Time Series Foundation ModelsabstractTime series foundation models (TSFMs) promise to be powerful tools for a wide range of applications. However, their internal representations and learned concepts are still not well understood. In this study, we investigate the structure and redundancy of representations across various TSFMs, examining the self-similarity of model layers within and across different model sizes. This analysis reveals block-like redundancy in the representations, which can be utilized for informed pruning to improve inference speed and efficiency. We also explore the concepts learned by these models, such as periodicity and trends. We demonstrate how conceptual priors can be derived from TSFM representations and leveraged to steer its outputs toward concept-informed predictions. Our work bridges representational analysis from language and vision models to TSFMs, offering new methods for building more computationally efficient and transparent TSFMs. Michal Wilinski, Mononito Goswami, Willa Potosnak, Nina Zukowska, Artur Dubrawski |
ICML | 5 |
| 2024 | JoLT: Jointly Learned Representations of Language and Time-Series for Clinical Time-Series Interpretation (Student Abstract)abstractTime-series and text data are prevalent in healthcare and frequently co-exist, yet they are typically modeled in isolation. Even studies that jointly model time-series and text, do so by converting time-series to images or graphs. We hypothesize that explicitly modeling time-series jointly with text can improve tasks such as summarization and question answering for time-series data, which have received little attention so far. To address this gap, we introduce JoLT to jointly learn desired representations from pre-trained time-series and text models. JoLT utilizes a Querying Transformer (Q-Former) to align the time-series and text representations. Our experiments on a large real-world electrocardiography dataset for medical time-series summarization show that JoLT outperforms state-of-the-art image captioning approaches. Yifu Cai, Arvind Srinivasan 0002, Mononito Goswami, Arjun Choudhry, Artur Dubrawski |
AAAI | 5 |
| 2024 | Data-Driven Discovery of Design Specifications (Student Abstract)abstractEnsuring a machine learning model’s trustworthiness is crucial to prevent potential harm. One way to foster trust is through the formal verification of the model’s adherence to essential design requirements. However, this approach relies on well-defined, application-domain-centric criteria with which to test the model, and such specifications may be cumbersome to collect in practice. We propose a data-driven approach for creating specifications to evaluate a trained model effectively. Implementing this framework allows us to prove that the model will exhibit safe behavior while minimizing the false-positive prediction rate. This strategy enhances predictive accuracy and safety, providing deeper insight into the model’s strengths and weaknesses, and promotes trust through a systematic approach. Angela Chen, Nicholas Gisolfi, Artur Dubrawski |
AAAI | 3 |
| 2024 | PICSR: Prototype-Informed Cross-Silo Router for Federated Learning (Student Abstract)abstractFederated Learning is an effective approach for learning from data distributed across multiple institutions. While most existing studies are aimed at improving predictive accuracy of models, little work has been done to explain knowledge differences between institutions and the benefits of collaboration. Understanding these differences is critical in cross-silo federated learning domains, e.g., in healthcare or banking, where each institution or silo has a different underlying distribution and stakeholders want to understand how their institution compares to their partners. We introduce Prototype-Informed Cross-Silo Router (PICSR) which utilizes a mixture of experts approach to combine local models derived from multiple silos. Furthermore, by computing data similarity to prototypical samples from each silo, we are able to ground the router’s predictions in the underlying dataset distributions. Experiments on a real-world heart disease prediction dataset show that PICSR retains high performance while enabling further explanations on the differences among institutions compared to a single black-box model. Eric Enouen, Sebastian Caldas, Mononito Goswami, Artur Dubrawski |
AAAI | 4 |
| 2024 | Adapting Animal Models to Assess Sufficiency of Fluid Resuscitation in Humans (Student Abstract)abstractFluid resuscitation is an initial treatment frequently employed to treat shock, restore lost blood, protect tissues from injury, and prevent organ dysfunction in critically ill patients. However, it is not without risk (e.g., overly aggressive resuscitation may cause organ damage and even death). We leverage machine learning models trained to assess sufficiency of resuscitation in laboratory animals subjected to induced hemorrhage and transfer them to use with human trauma patients. Our key takeaway is that animal experiments and models can inform human healthcare, especially when human data is limited or when collecting relevant human data via potentially harmful protocols is unfeasible. Ryan Schuerkamp, Brian Kunzer, Leonard S. Weiss, Hernando Gómez, Francis Guyette, Michael R. Pinsky, Artur Dubrawski |
AAAI | 8 |
| 2024 | A Rate-Distortion View of Uncertainty QuantificationabstractIn supervised learning, understanding an input’s proximity to the training data can help a model decide whether it has sufficient evidence for reaching a reliable prediction. While powerful probabilistic models such as Gaussian Processes naturally have this property, deep neural networks often lack it. In this paper, we introduce Distance Aware Bottleneck (DAB), i.e., a new method for enriching deep neural networks with this property. Building on prior information bottleneck approaches, our method learns a codebook that stores a compressed representation of all inputs seen during training. The distance of a new example from this codebook can serve as an uncertainty estimate for the example. The resulting model is simple to train and provides deterministic uncertainty estimates by a single forward pass. Finally, our method achieves better out-of-distribution (OOD) detection and misclassification prediction than prior methods, including expensive ensemble methods, deep kernel Gaussian Processes, and approaches based on the standard information bottleneck. Ifigeneia Apostolopoulou, Benjamin Eysenbach, Frank Nielsen, Artur Dubrawski |
ICML | 4 |
| 2024 | MOMENT: A Family of Open Time-series Foundation ModelsabstractWe introduce MOMENT, a family of open-source foundation models for general-purpose time series analysis. Pre-training large models on time series data is challenging due to (1) the absence of a large and cohesive public time series repository, and (2) diverse time series characteristics which make multi-dataset training onerous. Additionally, (3) experimental benchmarks to evaluate these models, especially in scenarios with limited resources, time, and supervision, are still in their nascent stages. To address these challenges, we compile a large and diverse collection of public time series, called the Time series Pile, and systematically tackle time series-specific challenges to unlock large-scale multi-dataset pre-training. Finally, we build on recent work to design a benchmark to evaluate time series foundation models on diverse tasks and datasets in limited supervision settings. Experiments on this benchmark demonstrate the effectiveness of our pre-trained models with minimal data and task-specific fine-tuning. Finally, we present several interesting empirical observations about large pre-trained time series models. Pre-trained models (AutonLab/MOMENT-1-large) and Time Series Pile (AutonLab/Timeseries-PILE) are available on Huggingface. Mononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai, Artur Dubrawski |
ICML | 6 |
| 2024 | Bifurcation Identification for Ultrasound-driven Robotic CannulationabstractIn trauma and critical care settings, rapid and precise intravascular access is key to patients’ survival. Our research aims at ensuring this access, even when skilled medical personnel are not readily available. Vessel bifurcations are anatomical landmarks that can guide the safe placement of catheters or needles during medical procedures. Although ultrasound is advantageous in navigating anatomical landmarks in emergency scenarios due to its portability and safety, to our knowledge no existing algorithm can autonomously extract vessel bifurcations using ultrasound images. This is primarily due to the limited availability of ground truth data, in particular, data from live subjects, needed for training and validating reliable models. We introduce BIFURC (Bifurcation Identification For Ultrasound-driven Robot Cannulation), a novel algorithm that identifies vessel bifurcations and provides optimal needle insertion sites for an autonomous robotic cannulation system. BIFURC integrates expert knowledge with deep learning techniques to efficiently detect vessel bifurcations within the femoral region and can be trained on a limited amount of in-vivo data. We evaluated our algorithm using a medical phantom as well as real-world experiments involving live pigs. In all cases, BIFURC consistently identified bifurcation points and needle insertion locations in alignment with those identified by expert clinicians. Cecilia G. Morales, Dhruv Srikanth, Jack H. Good, Keith Dufendach, Artur Dubrawski |
IROS | 5 |
| 2023 | NHITS: Neural Hierarchical Interpolation for Time Series ForecastingabstractRecent progress in neural forecasting accelerated improvements in the performance of large-scale forecasting systems. Yet, long-horizon forecasting remains a very difficult task. Two common challenges afflicting the task are the volatility of the predictions and their computational complexity. We introduce NHITS, a model which addresses both challenges by incorporating novel hierarchical interpolation and multi-rate data sampling techniques. These techniques enable the proposed method to assemble its predictions sequentially, emphasizing components with different frequencies and scales while decomposing the input signal and synthesizing the forecast. We prove that the hierarchical interpolation technique can efficiently approximate arbitrarily long horizons in the presence of smoothness. Additionally, we conduct extensive large-scale dataset experiments from the long-horizon forecasting literature, demonstrating the advantages of our method over the state-of-the-art methods, where NHITS provides an average accuracy improvement of almost 20% over the latest Transformer architectures while reducing the computation time by an order of magnitude (50 times). Our code is available at https://github.com/Nixtla/neuralforecast. Cristian Challu, Kin G. Olivares, Boris N. Oreshkin, Federico Garza Ramírez, Max Mergenthaler Canseco, Artur Dubrawski |
AAAI | 6 |
| 2023 | Ordinal Programmatic Weak Supervision and Crowdsourcing for Estimating Cognitive States (Student Abstract)abstractCrowdsourcing and weak supervision offer methods to efficiently label large datasets. Our work builds on existing weak supervision models to accommodate ordinal target classes, in an effort to recover ground truth from weak, external labels. We define a parameterized factor function and show that our approach improves over other baselines. Prakruthi Pradeep, Benedikt Boecking, Nicholas Gisolfi, Jacob R. Kintz, Torin K. Clark, Artur Dubrawski |
AAAI | 6 |
| 2023 | Generative Modeling Helps Weak Supervision (and Vice Versa)
Benedikt Boecking, Nicholas Carl Roberts, Willie Neiswanger, Stefano Ermon, Frederic Sala, Artur Dubrawski |
ICLR | 6 |
| 2023 | Reslicing Ultrasound Images for Data Augmentation and Vessel ReconstructionabstractRobot-guided vascular access has the potential to deliver urgent medical care in situations where medical personnel are unavailable. However, this technique requires accurate and reliable segmentation of anatomical landmarks in the body. For the ultrasound imaging modality, obtaining large amounts of training data for a segmentation model is time-consuming and expensive. This paper introduces RESUS (RESlicing of UltraSound Images), a weak supervision data augmentation technique for ultrasound images based on slicing reconstructed 3D volumes from tracked 2D images. This technique allows us to generate views which cannot be easily obtained in vivo due to physical constraints of ultrasound imaging, and use these augmented ultrasound images to train a semantic segmentation model. We demonstrate that RESUS achieves statistically significant improvement over training with non-augmented images and highlight qualitative improvements through vessel reconstruction. Cecilia G. Morales, Jason Yao, Tejas Rane, Robert Edman, Howie Choset, Artur Dubrawski |
ICRA | 6 |
| 2023 | Feature Learning for Interpretable, Performant Decision TreesabstractDecision trees are regarded for high interpretability arising from their hierarchical partitioning structure built on simple decision rules. However, in practice, this is not realized because axis-aligned partitioning of realistic data results in deep trees, and because ensemble methods are used to mitigate overfitting. Even then, model complexity and performance remain sensitive to transformation of the input, and extensive expert crafting of features from the raw data is common. We propose the first system to alternate sparse feature learning with differentiable decision tree construction to produce small, interpretable trees with good performance. We benchmark against conventional tree-based models and demonstrate several notions of interpretation of a model and its predictions. Jack H. Good, Torin Kovach, James Kyle Miller, Artur Dubrawski |
NeurIPS | 4 |
| 2023 | AQuA: A Benchmarking Tool for Label Quality AssessmentabstractMachine learning (ML) models are only as good as the data they are trained on. But recent studies have found datasets widely used to train and evaluate ML models, e.g. ImageNet, to have pervasive labeling errors. Erroneous labels on the train set hurt ML models' ability to generalize, and they impact evaluation and model selection using the test set. Consequently, learning in the presence of labeling errors is an active area of research, yet this field lacks a comprehensive benchmark to evaluate these methods. Most of these methods are evaluated on a few computer vision datasets with significant variance in the experimental protocols. With such a large pool of methods and inconsistent evaluation, it is also unclear how ML practitioners can choose the right models to assess label quality in their data. To this end, we propose a benchmarking environment AQuA to rigorously evaluate methods that enable machine learning in the presence of label noise. We also introduce a design space to delineate concrete design choices of label error detection models. We hope that our proposed design space and benchmark enable practitioners to choose the right tools to improve their label quality and that our benchmark enables objective and rigorous evaluation of machine learning tools facing mislabeled data. Mononito Goswami, Vedant Sanil, Arjun Choudhry, Arvind Srinivasan 0002, Chalisa Udompanyawit, Artur Dubrawski |
NeurIPS | 6 |
| 2023 | Verification of Fuzzy Decision TreesabstractIn recent years, there have been major strides in the safety verification of machine learning models such as neural networks and tree ensembles. However, fuzzy decision trees (FDT), also called soft or differentiable decision trees, are yet unstudied in the context of verification. They present unique verification challenges resulting from multiplications of input values; in the simplest case with a piecewise-linear splitting function, an FDT is piecewise-polynomial with degree up to the depth of the tree. We propose an abstraction-refinement algorithm for verification of properties of FDTs. We show that the problem is NP-Complete, like many other machine learning verification problems, and that our algorithm is complete in a finite precision setting. We benchmark on a selection of public data sets against an off-the-shelf SMT solver and a baseline variation of our algorithm that uses a refinement strategy from similar methods for neural network verification, finding the proposed method to be the fastest. Code for our algorithm along with our experiments and demos are available on GitHub athttps://github.com/autonlab/fdt_verification. Jack H. Good, Nicholas Gisolfi, James Kyle Miller, Artur Dubrawski |
IEEE Trans. Software Eng. | 4 |
| 2022 | Actionable Model-Centric Explanations (Student Abstract)abstractWe recommend using a model-centric, Boolean Satisfiability (SAT) formalism to obtain useful explanations of trained model behavior, different and complementary to what can be gleaned from LIME and SHAP, popular data-centric explanation tools in Artificial Intelligence (AI).We compare and contrast these methods, and show that data-centric methods may yield brittle explanations of limited practical utility.The model-centric framework, however, can offer actionable insights into risks of using AI models in practice. For critical applications of AI, split-second decision making is best informed by robust explanations that are invariant to properties of data, the capability offered by model-centric frameworks. Cecilia G. Morales, Nicholas Gisolfi, Robert Edman, James Kyle Miller, Artur Dubrawski |
AAAI | 5 |
| 2022 | Weakly Supervised Classification of Vital Sign Alerts as Real or Artifact
Mononito Goswami, Joo Heung Yoon, Gilles Clermont, Michael R. Pinsky, Marilyn Hravnak, Artur Dubrawski |
AMIA | 7 |
| 2022 | Deep Attentive Variational Inference
Ifigeneia Apostolopoulou, Ian Char, Elan Rosenfeld, Artur Dubrawski |
ICLR | 4 |
| 2022 | Counterfactual Phenotyping with Censored Time-to-EventsabstractEstimation of treatment efficacy of real-world clinical interventions involves working with continuous time-to-event outcomes such as time-to-death, re-hospitalization, or a composite event that may be subject to censoring. Counterfactual reasoning in such scenarios requires decoupling the effects of confounding physiological characteristics that affect baseline survival rates from the effects of the interventions being assessed. In this paper, we present a latent variable approach to model heterogeneous treatment effects by proposing that an individual can belong to one of latent clusters with distinct response characteristics. We show that this latent structure can mediate the base survival rates and help determine the effects of an intervention. We demonstrate the ability of our approach to discover actionable phenotypes of individuals based on their treatment response on multiple large randomized clinical trials originally conducted to assess appropriate treatment strategies to reduce cardiovascular risk. Chirag Nagpal, Mononito Goswami, Keith Dufendach, Artur Dubrawski |
KDD | 4 |
| 2021 | Using Machine Learning to Support Transfer of Best Practices in Healthcare
Sebastian Caldas, Jieshi Chen, Artur Dubrawski |
AMIA | 3 |
| 2021 | Weak Supervision for Affordable Modeling of Electrocardiogram Data
Mononito Goswami, Benedikt Boecking, Artur Dubrawski |
AMIA | 3 |
| 2021 | Affect, Support and Personal Factors: Multimodal Causal Models of One-on-one Coaching
Lujie Karen Chen, Joseph D. Ramsey, Artur Dubrawski |
EDM | 3 |
| 2021 | Interactive Weak Supervision: Learning Useful Heuristics for Data Labeling
Benedikt Boecking, Willie Neiswanger, Eric P. Xing, Artur Dubrawski |
ICLR | 4 |
| 2021 | End-to-End Weak SupervisionabstractAggregating multiple sources of weak supervision (WS) can ease the data-labeling bottleneck prevalent in many machine learning applications, by replacing the tedious manual collection of ground truth labels. Current state of the art approaches that do not use any labeled training data, however, require two separate modeling steps: Learning a probabilistic latent variable model based on the WS sources -- making assumptions that rarely hold in practice -- followed by downstream model training. Importantly, the first step of modeling does not consider the performance of the downstream model.To address these caveats we propose an end-to-end approach for directly learning the downstream model by maximizing its agreement with probabilistic labels generated by reparameterizing previous probabilistic posteriors with a neural network. Our results show improved performance over prior work in terms of end model performance on downstream test sets, as well as in terms of improved robustness to dependencies among weak supervision sources. Salva Rühling Cachay, Benedikt Boecking, Artur Dubrawski |
NeurIPS | 3 |
| 2021 | Deep Survival Machines: Fully Parametric Survival Regression and Representation Learning for Censored Data With Competing RisksabstractWe describe a new approach to estimating relative risks in time-to-event prediction problems with censored data in a fully parametric manner. Our approach does not require making strong assumptions of constant proportional hazards of the underlying survival distribution, as required by the Cox-proportional hazard model. By jointly learning deep nonlinear representations of the input covariates, we demonstrate the benefits of our approach when used to estimate survival risks through extensive experimentation on multiple real world datasets with different levels of censoring. We further demonstrate advantages of our model in the competing risks scenario. To the best of our knowledge, this is the first work involving fully parametric estimation of survival times with competing risks in the presence of censoring. Chirag Nagpal, Artur Dubrawski |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Discriminating Cognitive Disequilibrium and Flow in Problem Solving: A Semi-Supervised Approach Using Involuntary Dynamic Behavioral SignalsabstractProblem solving is one of the most important 21st century skills. However, effectively coaching young students in problem solving is challenging because teachers must continuously monitor their cognitive and affective states, and make real-time pedagogical interventions to maximize their learning outcomes. It is an even more challenging task in social environments with limited human coaching resources. To lessen the cognitive load on a teacher and enable affect-sensitive intelligent tutoring, many researchers have investigated automated cognitive and affective detection methods. However, most of the studies use culturally-sensitive indices of affect that are prone to social editing such as facial expressions, and only few studies have explored involuntary dynamic behavioral signals such as gross body movements. In addition, most current methods rely on expensive labelled data from trained annotators for supervised learning. In this paper, we explore a semi-supervised learning framework that can learn low-dimensional representations of involuntary dynamic behavioral signals (mainly gross-body movements) from a modest number of short time series segments. Experiments on a real-world dataset reveal a significant advantage of these representations in discriminating cognitive disequilibrium and flow, as compared to traditional complexity measures from dynamical systems literature, and demonstrate their potential in transferring learned models to previously unseen subjects. Mononito Goswami, Lujie Chen, Artur Dubrawski |
AAAI | 3 |
| 2020 | Modeling Involuntary Dynamic Behaviors to Support Intelligent Tutoring (Student Abstract)abstractProblem solving is one of the most important 21st century skills. However, effectively coaching young students in problem solving is challenging because teachers must continuously monitor their cognitive and affective states and make real-time pedagogical interventions to maximize students' learning outcomes. It is an even more challenging task in social environments with limited human coaching resources. To lessen the cognitive load on a teacher and enable affect-sensitive intelligent tutoring, many researchers have investigated automated cognitive and affective detection methods. However, most of the studies use culturally-sensitive indices of affect that are prone to social editing such as facial expressions, and only few studies have explored involuntary dynamic behavioral signals such as gross body movements. In addition, most current methods rely on expensive labelled data from trained annotators for supervised learning. In this paper, we explore a semi-supervised learning framework that can learn low-dimensional representations of involuntary dynamic behavioral signals (mainly gross-body movements) from a modest number of short time series segments. Experiments on a real-world dataset reveal a significant utility of these representations in discriminating cognitive disequilibrium and flow and demonstrate their potential in transferring learned models to previously unseen subjects. Mononito Goswami, Lujie Chen, Chufan Gao, Artur Dubrawski |
AAAI | 4 |
| 2020 | Robust Multi-View Representation Learning (Student Abstract)abstractMulti-view data has become ubiquitous, especially with multi-sensor systems like self-driving cars or medical patient-side monitors. We propose two methods to approach robust multi-view representation learning with the aim of leveraging local relationships between views.The first is an extension of Canonical Correlation Analysis (CCA) where we consider multiple one-vs-rest CCA problems, one for each view. We use a group-sparsity penalty to encourage finding local relationships. The second method is a straightforward extension of a multi-view AutoEncoder with view-level drop-out.We demonstrate the effectiveness of these methods in simple synthetic experiments. We also describe heuristics and extensions to improve and/or expand on these methods. Sibi Venkatesan, James Kyle Miller, Artur Dubrawski |
AAAI | 3 |
| 2020 | Thresholding Bandit Problem with Both Duels and PullsabstractThe Thresholding Bandit Problem (TBP) aims to find the set of arms with mean rewards greater than a given threshold. We consider a new setting of TBP, where in addition to pulling arms, one can also duel two arms and get the arm with a greater mean. In our motivating application from crowdsourcing, dueling two arms can be more cost-effective and time-efficient than direct pulls. We refer to this problem as TBP with Dueling Choices (TBP-DC). This paper provides an algorithm called Rank-Search (RS) for solving TBP-DC by alternating between ranking and binary search. We prove theoretical guarantees for RS, and also give lower bounds to show the optimality of it. Experiments show that RS outperforms previous baseline algorithms that only use pulls or duels. Yichong Xu, Aarti Singh, Artur Dubrawski |
AISTATS | 4 |
| 2020 | High Resolution Diffuse Optical Tomography using Short Range Indirect Subsurface ImagingabstractDiffuse optical tomography (DOT) is an approach to recover subsurface structures beneath the skin by measuring light propagation beneath the surface. The method is based on optimizing the difference between the images collected and a forward model that accurately represents diffuse photon propagation within a heterogeneous scattering medium. However, to date, most works have used a few source-detector pairs and recover the medium at only a very low resolution. And increasing the resolution requires prohibitive computations/storage. In this work, we present a fast imaging and algorithm for high resolution diffuse optical tomography with a line imaging and illumination system. Key to our approach is a convolution approximation of the forward heterogeneous scattering model that can be inverted to produce deeper than ever before structured beneath the surface. We show that our proposed method can detect reasonably accurate boundaries and relative depth of heterogeneous structures up to a depth of 8 mm below highly scattering medium such as milk. This work can extend the potential of DOT to recover more intricate structures (vessels, tissue, tumors, etc.) beneath the skin for diagnosing many dermatological and cardio-vascular conditions. Chao Liu 0064, Akash K. Maity, Artur Dubrawski, Ashutosh Sabharwal, Srinivasa G. Narasimhan |
ICCP | 3 |
| 2020 | Preference-based Reinforcement Learning with Finite-Time GuaranteesabstractPreference-based Reinforcement Learning (PbRL) replaces reward values in traditional reinforcement learning by preferences to better elicit human opinion on the target objective, especially when numerical reward values are hard to design or interpret. Despite promising results in applications, the theoretical understanding of PbRL is still in its infancy. In this paper, we present the first finite-time analysis for general PbRL problems. We first show that a unique optimal policy may not exist if preferences over trajectories are deterministic for PbRL. If preferences are stochastic, and the preference probability relates to the hidden reward values, we present algorithms for PbRL, both with and without a simulator, that are able to identify the best policy up to accuracy $\varepsilon$ with high probability. Our method explores the state space by navigating to under-explored states, and solves PbRL using a combination of dueling bandits and policy search. Experiments show the efficacy of our method when it is applied to real-world problems. Yichong Xu, Ruosong Wang, Lin Yang 0011, Aarti Singh, Artur Dubrawski |
NeurIPS | 5 |
| 2020 | Zeroth Order Non-convex optimization with Dueling-Choice BanditsabstractWe consider a novel setting of zeroth order non-convex optimization, where in addition to querying the function value at a given point, we can also duel two points and get the point with the larger function value. We refer to this setting as optimization with dueling-choice bandits, since both direct queries and duels are available for optimization. We give the COMP-GP-UCB algorithm based on GP-UCB (Srinivas et al., 2009),, where instead of directly querying the point with the maximum Upper Confidence Bound (UCB), we perform constrained optimization and use comparisons to filter out suboptimal points. COMP-GP-UCB comes with theoretical guarantee of $O(\frac{\Phi}{\sqrt{T}})$ on simple regret where $T$ is the number of direct queries and $\Phi$ is an improved information gain stemming from a comparison-based constraint set that restricts the space for optimum search. In contrast, in the plain direct query setting, $\Phi$ depends on the entire domain. We discuss theoretical aspects and show experimental results to demonstrate efficacy of our algorithm. Yichong Xu, Aparna Joshi, Aarti Singh, Artur Dubrawski |
UAI | 4 |
| 2020 | Regression with Comparisons: Escaping the Curse of Dimensionality with Ordinal InformationabstractIn supervised learning, we typically leverage a fully labeled dataset to design methods for function estimation or prediction. In many practical situations, we are able to obtain alternative feedback, possibly at a low cost. A broad goal is to understand the usefulness of, and to design algorithms to exploit, this alternative feedback. In this paper, we consider a semi-supervised regression setting, where we obtain additional ordinal (or comparison) information for the unlabeled samples. We consider ordinal feedback of varying qualities where we have either a perfect ordering of the samples, a noisy ordering of the samples or noisy pairwise comparisons between the samples. We provide a precise quantification of the usefulness of these types of ordinal feedback in both nonparametric and linear regression, showing that in many cases it is possible to accurately estimate an underlying function with a very small labeled set, effectively escaping the curse of dimensionality. We also present lower bounds, that establish fundamental limits for the task and show that our algorithms are optimal in a variety of settings. Finally, we present extensive experiments on new datasets that demonstrate the efficacy and practicality of our algorithms and investigate their robustness to various sources of noise and model misspecification. Yichong Xu, Sivaraman Balakrishnan, Aarti Singh, Artur Dubrawski |
J. Mach. Learn. Res. | 4 |
| 2019 | On the Interaction Effects Between Prediction and ClusteringabstractMachine learning systems increasingly depend on pipelines of multiple algorithms to provide high quality and well structured predictions. This paper argues interaction effects between clustering and prediction (e.g. classification, regression) algorithms can cause subtle adverse behaviors during cross-validation that may not be initially apparent. In particular, we focus on the problem of estimating the out-of-cluster (OOC) prediction loss given an approximate clustering with probabilistic error rate p_0. Traditional cross-validation techniques exhibit significant empirical bias in this setting, and the few attempts to estimate and correct for these effects are intractable on larger datasets. Further, no previous work has been able to characterize the conditions under which these empirical effects occur, and if they do, what properties they have. We precisely answer these questions by providing theoretical properties which hold in various settings, and prove that expected out-of-cluster loss behavior rapidly decays with even minor clustering errors. Fortunately, we are able to leverage these same properties to construct hypothesis tests and scalable estimators necessary for correcting the problem. Empirical results on benchmark datasets validate our theoretical results and demonstrate how scaling techniques provide solutions to new classes of problems. Matt Barnes 0001, Artur Dubrawski |
AISTATS | 2 |
| 2019 | Parent as a Companion for Solving Challenging Math Problems: Insights from Multi-modal Observational Data
Lujie Chen, Eva Gjekmarkaj, Artur Dubrawski |
EDM | 3 |
| 2019 | Mutually Regressive Point ProcessesabstractMany real-world data represent sequences of interdependent events unfolding over time. They can be modeled naturally as realizations of a point process. Despite many potential applications, existing point process models are limited in their ability to capture complex patterns of interaction. Hawkes processes admit many efficient inference algorithms, but are limited to mutually excitatory effects. Non- linear Hawkes processes allow for more complex influence patterns, but for their estimation it is typically necessary to resort to discrete-time approximations that may yield poor generative models. In this paper, we introduce the first general class of Bayesian point process models extended with a nonlinear component that allows both excitatory and inhibitory relationships in continuous time. We derive a fully Bayesian inference algorithm for these processes using Polya-Gamma augmentation and Poisson thinning. We evaluate the proposed model on single and multi-neuronal spike train recordings. Results demonstrate that the proposed model, unlike existing point process models, can generate biologically-plausible spike trains, while still achieving competitive predictive likelihoods. Ifigeneia Apostolopoulou, Scott W. Linderman, James Kyle Miller, Artur Dubrawski |
NeurIPS | 4 |
| 2019 | Statistical outbreak detection by joining medical records and pathogen similarity
James Kyle Miller, Jieshi Chen, Alexander Sundermann, Jane W. Marsh, Melissa I. Saul, Kathleen A. Shutt, Marissa Pacey, Mustapha M. Mustapha, Lee H. Harrison, Artur Dubrawski |
J. Biomed. Informatics | 10 |
| 2018 | Near-light photometric stereo using circularly placed point light sourcesabstractMost photometric stereo approaches assume distant or directional lighting and orthographic imaging. However, when the source is divergent and is near the object and the camera is projective, the image intensity of a Lambertian object is a non-linear function of both the unknown surface normals and the unknown distances of the source to the surface points. The resulting non-linear optimization is non-convex and highly sensitive to the initial guess. In this paper, we propose a two-stage near-light photometric stereo method using circularly placed point light sources (commonly seen in recent consumer imaging devices like NESTcam, Amazon Cloudcam, etc). We represent the scene using a 3D mesh and directly optimize the vertices of the mesh. This reduces the complexity of the relationship between surface normals and depths in the image formation model. In the first stage, we optimize the vertex positions using the differential images induced by small changes in light source position. This procedure yields a strong initial guess for the second stage that refines the estimations using the raw captured images. We propose an accurate calibration approach to estimate the positions of the sources. Our approach performs better on simulations and on real Lambertian scenes with complex shapes than the state-of-the-art method with near-field lighting. Chao Liu 0064, Srinivasa G. Narasimhan, Artur Dubrawski |
ICCP | 3 |
| 2018 | Nonparametric Regression with Comparisons: Escaping the Curse of Dimensionality with Ordinal InformationabstractIn supervised learning, we leverage a labeled dataset to design methods for function estimation. In many practical situations, we are able to obtain alternative feedback, possibly at a low cost. A broad goal is to understand the usefulness of, and to design algorithms to exploit, this alternative feedback. We focus on a semi-supervised setting where we obtain additional ordinal (or comparison) information for potentially unlabeled samples. We consider ordinal feedback of varying qualities where we have either a perfect ordering of the samples, a noisy ordering of the samples or noisy pairwise comparisons between the samples. We provide a precise quantification of the usefulness of these types of ordinal feedback in non-parametric regression, showing that in many cases it is possible to accurately estimate an underlying function with a very small labeled set, effectively escaping the curse of dimensionality. We develop an algorithm called Ranking-Regression (RR) and analyze its accuracy as a function of size of the labeled and unlabeled datasets and various noise parameters. We also present lower bounds, that establish fundamental limits for the task and show that RR is optimal in a variety of settings. Finally, we present experiments that show the efficacy of RR and investigate its robustness to various sources of noise and model-misspecification. Yichong Xu, Hariank Muthakana, Sivaraman Balakrishnan, Aarti Singh, Artur Dubrawski |
ICML | 5 |
| 2018 | Active Search of Connections for Case Building and Combating Human TraffickingabstractHow can we help an investigator to efficiently connect the dots and uncover the network of individuals involved in a criminal activity based on the evidence of their connections, such as visiting the same address, or transacting with the same bank account? We formulate this problem as Active Search of Connections, which finds target entities that share evidence of different types with a given lead, where their relevance to the case is queried interactively from the investigator. We present RedThread, an efficient solution for inferring related and relevant nodes while incorporating the user's feedback to guide the inference. Our experiments focus on case building for combating human trafficking, where the investigator follows leads to expose organized activities, i.e. different escort advertisements that are connected and possibly orchestrated. RedThread is a local algorithm and enables online case building when mining millions of ads posted in one of the largest classified advertising websites. The results of RedThread are interpretable, as they explain how the results are connected to the initial lead. We experimentally show that RedThread learns the importance of the different types and different pieces of evidence, while the former could be transferred between cases. Reihaneh Rabbany, David Bayani, Artur Dubrawski |
KDD | 3 |
| 2018 | Accelerated apprenticeship: teaching data science problem solving skills at scaleabstractIt often takes years of hands-on practice to build operational problem solving skills for a data scientist to be sufficiently competent to tackle real world problems. In this research, we explore a new scalable technology-enhanced learning (TEL) platform that enables accelerated apprenticeship process via a repository of caselets - small but focused case studies with scaffolding questions and feedback. In this paper, we report rationales of the design, caselet authoring process, and the planned experiment with cohorts of students who will use caselets while taking graduate level data science courses. Lujie Chen, Artur Dubrawski |
L@S | 2 |
| 2018 | Social-Affiliation Networks: Patterns and the SOAR Model
Dhivya Eswaran, Reihaneh Rabbany, Artur Dubrawski, Christos Faloutsos |
ECML/PKDD (2) | 3 |
| 2017 | Utility of Anti-hypertension Prescription Orders in Predicting Future Hypertensive Instability Events
Lujie Chen, Marilyn Hravnak, Gilles Clermont, Michael R. Pinsky, Artur Dubrawski |
AMIA | 6 |
| 2017 | Matting and Depth Recovery of Thin Structures Using a Focal StackabstractThin structures such as fence, grass and vessels are common in photography and scientific imaging. They exhibit complex 3D structures with sharp depth variations/discontinuities and mutual occlusions. In this paper, we develop a method to estimate the occlusion matte and depths of thin structures from a focal image stack, which is obtained either by varying the focus/aperture of the lens or computed from a one-shot light field image. We propose an image formation model that explicitly describes the spatially varying optical blur and mutual occlusions for structures located at different depths. Based on the model, we derive an efficient MCMC inference algorithm that enables direct and analytical computations of the iterative update for the model/images without re-rendering images in the sampling process. Then, the depths of the thin structures are recovered using gradient descent with the differential terms computed using the image formation model. We apply the proposed method to scenes at both macro and micro scales. For macro-scale, we evaluate our method on scenes with complex 3D thin structures such as tree branches and grass. For micro-scale, we apply our method to in-vivo microscopic images of micro-vessels with diameters less than 50 μm. To our knowledge, the proposed method is the first approach to reconstruct the 3D structures of micro-vessels from non-invasive in-vivo image measurements. Chao Liu 0064, Srinivasa G. Narasimhan, Artur Dubrawski |
CVPR | 3 |
| 2017 | Scaling Active Search using Linear Similarity FunctionsabstractActive Search has become an increasingly useful tool in information retrieval problems where the goal is to discover as many target elements as possible using only limited label queries. With the advent of big data, there is a growing emphasis on the scalability of such techniques to handle very large and very complex datasets. In this paper, we consider the problem of Active Search where we are given a similarity function between data points. We look at an algorithm introduced by Wang et al. [Wang et al., 2013] known as Active Search on Graphs and propose crucial modifications which allow it to scale significantly. Their approach selects points by minimizing an energy function over the graph induced by the similarity function on the data. Our modifications require the similarity function to be a dot-product between feature vectors of data points, equivalent to having a linear kernel for the adjacency matrix. With this, we are able to scale tremendously: for n data points, the original algorithm runs in O(n^2) time per iteration while ours runs in only O(nr + r^2) given r-dimensional features. We also describe a simple alternate approach using a weighted-neighbor predictor which also scales well. In our experiments, we show that our method is competitive with existing semi-supervised approaches. We also briefly discuss conditions under which our algorithm performs well. Sibi Venkatesan, James Kyle Miller, Jeff G. Schneider, Artur Dubrawski |
IJCAI | 4 |
| 2017 | Learning from learning curves: discovering interpretable learning trajectoriesabstractWe propose a data driven method for decomposing population level learning curve models into mutually exclusive distinctive groups each consisting of similar learning trajectories. We validate this method on six knowledge components from the log data from an online tutoring system ASSIST-ment. Preliminary analysis reveals interpretable patterns of "skill growth" that correlate with students' performance in the subsequently administered state standardized tests. Lujie Chen, Artur Dubrawski |
LAK | 2 |
| 2017 | Noise-Tolerant Interactive Learning Using Pairwise ComparisonsabstractWe study the problem of interactively learning a binary classifier using noisy labeling and pairwise comparison oracles, where the comparison oracle answers which one in the given two instances is more likely to be positive. Learning from such oracles has multiple applications where obtaining direct labels is harder but pairwise comparisons are easier, and the algorithm can leverage both types of oracles. In this paper, we attempt to characterize how the access to an easier comparison oracle helps in improving the label and total query complexity. We show that the comparison oracle reduces the learning problem to that of learning a threshold function. We then present an algorithm that interactively queries the label and comparison oracles and we characterize its query complexity under Tsybakov and adversarial noise conditions for the comparison and labeling oracles. Our lower bounds show that our label and total query complexity is almost optimal. Yichong Xu, Hongyang Zhang 0001, Aarti Singh, Artur Dubrawski, James Kyle Miller |
NIPS | 4 |
| 2017 | Beyond Assortativity: Proclivity Index for Attributed Networks (ProNe)
Reihaneh Rabbany, Dhivya Eswaran, Artur Dubrawski, Christos Faloutsos |
PAKDD (1) | 3 |
| 2017 | The Binomial Block Bootstrap Estimator for Evaluating Loss on Dependent Clusters
Matt Barnes 0001, Artur Dubrawski |
UAI | 2 |
| 2017 | Learning temporal rules to forecast instability in continuously monitored patientsabstractInductive machine learning, and in particular extraction of association rules from data, has been successfully used in multiple application domains, such as market basket analysis, disease prognosis, fraud detection, and protein sequencing. The appeal of rule extraction techniques stems from their ability to handle intricate problems yet produce models based on rules that can be comprehended by humans, and are therefore more transparent. Human comprehension is a factor that may improve adoption and use of data-driven decision support systems clinically via face validity. In this work, we explore whether we can reliably and informatively forecast cardiorespiratory instability (CRI) in step-down unit (SDU) patients utilizing data from continuous monitoring of physiologic vital sign (VS) measurements. We use a temporal association rule extraction technique in conjunction with a rule fusion protocol to learn how to forecast CRI in continuously monitored patients. We detail our approach and present and discuss encouraging empirical results obtained using continuous multivariate VS data from the bedside monitors of 297 SDU patients spanning 29 346 hours (3.35 patient-years) of observation. We present example rules that have been learned from data to illustrate potential benefits of comprehensibility of the extracted models, and we analyze the empirical utility of each VS as a potential leading indicator of an impending CRI event. Mathieu Guillame-Bert, Artur Dubrawski, Donghan Wang, Marilyn Hravnak, Gilles Clermont, Michael R. Pinsky |
J. Am. Medical Informatics Assoc. | 2 |
| 2017 | Classification of Time Sequences using Graphs of Temporal ConstraintsabstractWe introduce two algorithms that learn to classify Symbolic and Scalar Time Sequences (SSTS); an extension of multivariate time series. An SSTS is a set of \emph{events} and a set of scalars. An event is defined by a symbol and a time-stamp. A scalar is defined by a symbol and a function mapping a number for each possible time stamp of the data. The proposed algorithms rely on temporal patterns called Graph of Temporal Constraints (GTC). A GTC is a directed graph in which vertices express occurrences of specific events, and edges express temporal constraints between occurrences of pairs of events. Additionally, each vertex of a GTC can be augmented with numeric constraints on scalar values. We allow GTCs to be cyclic and/or disconnected. The first of the introduced algorithms extracts sets of co-dependent GTCs to be used in a voting mechanism. The second algorithm builds decision forest like representations where each node is a GTC. In both algorithms, extraction of GTCs and model building are interleaved. Both algorithms are closely related to each other and they exhibit complementary properties including complexity, performance, and interpretability. The main novelties of this work reside in direct building of the model and efficient learning of GTC structures. We explain the proposed algorithms and evaluate their performance against a diverse collection of 59 benchmark data sets. In these experiments, our algorithms come across as highly competitive and in most cases closely match or outperform state-of-the-art alternatives in terms of the computational speed while dominating in terms of the accuracy of classification of time sequences. Mathieu Guillame-Bert, Artur Dubrawski |
J. Mach. Learn. Res. | 2 |
| 2016 | A Framework for Visual Tracking of Risk and its Drivers in Monitoring Patients Susceptible for Cardiorespiratory Instability
Lujie Chen, Gilles Clermont, Marilyn Hravnak, Michael R. Pinsky, Artur Dubrawski |
AMIA | 5 |
| 2016 | Riding an emotional roller-coaster: A multimodal study of young child's math problem solving activities
Lujie Chen, Zhuyun Xia, Zhanmei Song, Louis-Philippe Morency, Artur Dubrawski |
EDM | 6 |
| 2016 | VIPR: An Interactive Tool for Meaningful Visualization of High-Dimensional Data
Donghan Wang, Madalina Fiterau, Artur Dubrawski |
IJCAI | 3 |
| 2016 | Detection of radioactive sources in urban scenes using Bayesian Aggregation of data from mobile spectrometers
Prateek Tandon 0002, Peter Huggins, Robert A. MacLachlan, Artur Dubrawski, Karl Nelson, Simon Labov |
Inf. Syst. | 4 |
| 2015 | Leveraging Common Structure to Improve Prediction across Related Datasets
Matt Barnes 0001, Nicholas Gisolfi, Madalina Fiterau, Artur Dubrawski |
AAAI | 4 |
| 2015 | Active Learning for Informative Projection RetrievalabstractWe introduce an active learning framework designed to train classification models which use informative projections. Our approach works with the obtained low-dimensional models in finding unlabeled data for annotation by experts. The advantage of our approach is that the labeling effort is expended mainly on samples which benefit models from the considered hypothesis class. This results in an improved learning rate over standard selection criteria for data from the clinical domain. Madalina Fiterau, Artur Dubrawski |
AAAI | 2 |
| 2015 | Finding Meaningful Gaps to Guide Data Acquisition for a Radiation Adjudication SystemabstractWe consider the problem of identifying discrepancies between training and test data which are responsible for the reduced performance of a classification system. Intended for use when data acquisition is an iterative process controlled by domain experts, our method exposes insufficiencies of training data and presents them in a user-friendly manner. The system is capable of working with any classification system which admits diagnostics on test data. We illustrate the usefulness of our approach in recovering compact representations of the revealed gaps in training data and show that predictive accuracy of the resulting models is improved once the gaps are filled through collection of additional training samples. Nicholas Gisolfi, Madalina Fiterau, Artur Dubrawski |
AAAI | 3 |
| 2015 | Modelling Risk of Cardio-Respiratory Instability as a Heterogeneous Process
Lujie Chen, Artur Dubrawski, Marilyn Hravnak, Gilles Clermont, Michael R. Pinsky |
AMIA | 2 |
| 2015 | Real-time visual analysis of microvascular blood flow for critical careabstractMicrocirculatory monitoring plays an important role in diagnosis and treatment of critical care patients. Sidestream Dark Field (SDF) imaging devices have been used to visualize and support interpretation of the micro-vascular blood flow. However, due to subsurface scattering within the tissue that embeds the capillaries, transparency of plasma, imaging noise and lack of features, it is difficult to obtain reliable physiological data from SDF videos. Therefore, thus far microcirculatory videos have been analyzed manually with significant input from expert clinicians. In this paper, we present a framework that automates the analysis process. It includes stages of video stabilization, enhancement, and micro-vessel extraction, in order to automatically estimate statistics of the micro blood flows from SDF videos. Our method has been validated in critical care experiments conducted carefully to record the microcirculatory blood flow in test animal subjects before, during and after induced bleeding episodes, as well as to study the effect of fluid resuscitation. Our method is able to extract microcirculatory measurements that are consistent with clinical intuition and it has a potential to become a useful tool in critical care medicine. Chao Liu 0064, Hernando Gómez, Srinivasa G. Narasimhan, Artur Dubrawski, Michael R. Pinsky, Brian Zuckerbraun |
CVPR | 4 |
| 2013 | Informative Projection Recovery for Classification, Clustering and RegressionabstractData driven decision support systems often benefit from human participation to validate outcomes produced by automated procedures. Perceived utility hinges on the system's ability to learn transparent, comprehensible models from data. We introduce and formalize Informative Projection Recovery: the problem of extracting a set of low-dimensional projections of data which jointly form an accurate solution to a given learning task. We approach this problem with RIPR: a regression-based algorithm that identifies informative projections by optimizing over a matrix of point-wise loss estimators. It generalizes from our previous algorithm, offering solutions to classification, clustering, and regression tasks. Experiments show that RIPR can discover and leverage structures of informative projections in data, if they exist, while yielding accurate and compact models. It is particularly useful in applications involving multivariate numeric data in which expert assessment of the results is of the essence. Madalina Fiterau, Artur Dubrawski |
ICMLA (1) | 2 |
| 2012 | Projection Retrieval for ClassificationabstractIn many applications classification systems often require in the loop human intervention. In such cases the decision process must be transparent and comprehensible simultaneously requiring minimal assumptions on the underlying data distribution. To tackle this problem, we formulate it as an axis-alligned subspacefinding task under the assumption that query specific information dictates the complementary use of the subspaces. We develop a regression-based approach called RECIP that efficiently solves this problem by finding projections that minimize a nonparametric conditional entropy estimator. Experiments show that the method is accurate in identifying the informative projections of the dataset, picking the correct ones to classify query points and facilitates visual evaluation by users. Madalina Fiterau, Artur Dubrawski |
NIPS | 2 |
| 2010 | Automatic state discovery for unstructured audio scene classificationabstractIn this paper we present a novel scheme for unstructured audio scene classification that possesses three highly desirable and powerful features: autonomy, scalability, and robustness. Our scheme is based on our recently introduced machine learning algorithm called Simultaneous Temporal And Contextual Splitting (STACS) that discovers the appropriate number of states and efficiently learns accurate Hidden Markov Model (HMM) parameters for the given data. STACS-based algorithms train HMMs up to five times faster than Baum-Welch, avoid the overfitting problem commonly encountered in learning large state-space HMMs using Expectation Maximization (EM) methods such as Baum-Welch, and achieve superior classification results on a very diverse dataset with minimal pre-processing. Furthermore, our scheme has proven to be highly effective for building real-world applications and has been integrated into a commercial surveillance system as an event detection component. Julian Ramos 0001, Sajid M. Siddiqi, Artur Dubrawski, Geoffrey J. Gordon |
ICASSP | 3 |
| 2010 | Learning Compressible ModelsabstractIn this paper, we study the combination of compression and ℓ1-norm regularization in a machine learning context: learning compressible models. By including a compression operation into the ℓ1 regularization, the assumption on model sparsity is relaxed to compressibility: model coefficients are compressed before being penalized, and sparsity is achieved in a compressed domain rather than the original space. We focus on the design of different compression operations, by which we can encode various compressibility assumptions and inductive biases, e.g., piecewise local smoothness, compacted energy in the frequency domain, and semantic correlation. We show that use of a compression operation provides an opportunity to leverage auxiliary information from various sources, e.g., domain knowledge, coding theories, unlabeled data. We conduct extensive experiments on brain-computer interfacing, handwritten character recognition and text classification. Empirical results show clear improvements in prediction performance by including compression in ℓ1 regularization. We also analyze the learned model coefficients under appropriate compressibility assumptions, which further demonstrate the advantages of learning compressible models instead of sparse models. Yi Zhang 0010, Jeff G. Schneider, Artur Dubrawski |
SDM | 3 |
| 2009 | Trade-offs between Agility and Reliability of Predictions in Dynamic Social Networks Used to Model Risk of Microbial Contamination of FoodabstractThis paper evaluates trade-offs between agility and reliability of predictions arising due to sparseness of data modeled with dynamic social networks. We use real field data from food safety domain to illustrate the discussion. We model food production facilities as one type of entities in a social network evolving in time. Another type of entities denotes various specific strains of Salmonella. Two entities are linked in the graph if a microbial test of food sample conducted at the specific food facility over specific period of time turns out positive for the particular pathogen. We use a computationally efficient latent space model to predict future occurrences of pathogens in individual facilities. Empirical results indicate predictive utility of the proposed representation. However, sparseness of data limits the attainable agility of predictions. We identify exploiting recency of data and using the known patterns in it, such as seasonality, as plausible means of battling the challenge of sparseness. Artur Dubrawski, Purnamrita Sarkar, Lujie Chen |
ASONAM | 1 |
| 2009 | T-Cube Web Interface in support of real-time bio-surveillance programabstractT-Cube Web Interface is a generic tool to visualize and manipulate large scale multivariate time series datasets. The interface allows the user to execute complex queries quickly and to run various types of statistical tests on the loaded data. We show its utility in an important application scenario: real-time bio-surveillance system designed to support rapid detection and mitigation of bio-medical threats in developing countries. Artur Dubrawski, Maheshkumar Sabhnani, Michael Knight, Michael Baysek, Daniel B. Neill, Saswati Ray, Anna Michalska, Nuwan Waidyanatha |
ICTD | 1 |
| 2008 | Learning Detectors of Events in Multivariate Time Series
Josep Roure Alcobé, Artur Dubrawski, Jeff G. Schneider |
AMIA | 2 |
| 2008 | Learning the Semantic Correlation: An Alternative Way to Gain from Unlabeled TextabstractIn this paper, we address the question of what kind of knowledge is generally transferable from unlabeled text. We suggest and analyze the semantic correlation of words as a generally transferable structure of the language and propose a new method to learn this structure using an appropriately chosen latent variable model. This semantic correlation contains structural information of the language space and can be used to control the joint shrinkage of model parameters for any specific task in the same space through regularization. In an empirical study, we construct 190 different text classification tasks from a real-world benchmark, and the unlabeled documents are a mixture from all these tasks. We test the ability of various algorithms to use the mixed unlabeled text to enhance all classification tasks. Empirical results show that the proposed approach is a reliable and scalable method for semi-supervised learning, regardless of the source of unlabeled data, the specific task to be enhanced, and the prediction model used. Yi Zhang 0010, Jeff G. Schneider, Artur Dubrawski |
NIPS | 3 |
| 2006 | Monitoring Food Safety by Detecting Patterns in Consumer Complaints
Artur Dubrawski, Kimberly Elenberg, Andrew W. Moore 0001, Maheshkumar Sabhnani |
AAAI | 1 |
| 1998 | A Method for Tracking Pose of a Mobile Robot Equipped with a Scanning Laser Range FinderabstractOne of the essential problems in navigation of a mobile robot is to accurately determine its location using data obtained by range sensors. In this paper we present a novel approach to track changes of orientation and position of a robot equipped with a scanning laser range finder and designed to work in partially structured environments. It is shown experimentally on a real robot that the proposed approach is more robust against sensory noise than the method of angle histograms, which is often used to track orientation and position of indoor vehicles equipped with scanning range sensors. Artur Dubrawski, Barbara Siemiatkowska |
ICRA | 1 |
| 1997 | Tuning neural networks with stochastic optimizationabstractThis paper describes a method for automated tuning of hyper-parameters of supervised learning systems. It emerges from stochastic aproximation, uses memory-based learning principles, follows certain ideas of experimental design and employs a particular approach to resampling called stochastic validation. Potential usefulness of the proposed approach is illustrated with the fuzzy-ARTMAP neural network application to learning a qualitative positioning of an indoor mobile robot equipped with ultrasonic range sensors. Automatically selected setpoints allow the system to reach a similar or better performance in comparison to that achieved by human experts in all studied cases. The presented method may serve as a design tool in practical applications of supervised learning algorithms. Artur Dubrawski |
IROS | 1 |
| 1994 | Self-Supervised Neural System for Reactive NavigationabstractThis paper deals with an artificial neural system for a mobile robot reactive navigation in an unknown, cluttered environment. A task of a presented system is to provide a steering angle signal letting a robot reach a goal while avoiding collisions with obstacles. Basic reactive navigation methods are briefly characterized, a special attention is paid to neural approaches. Then a qualitative description of a presented system is given. The main parts of the system are: the Fuzzy-ART classifier performing a perceptual space partitioning, and the neural associative memory, storing system's experience and superposing influences of different behaviours. Preliminary tests show that the learning by trial-and-error is efficient, as well in a case of beginning from scratch, as after some disturbances of either system's or environmental characteristics.> Artur Dubrawski, James L. Crowley |
ICRA | 1 |
| 1994 | Learning to categorize perceptual space of a mobile robot using fuzzy-ART neural networkabstractThis paper deals with an application of fuzzy-ART self-organizing neural classifier to adaptive categorization of the perceptual space of a mobile robot. The aim of the research is to develop a learning system for reactive locomotion control in an unknown, cluttered environment. A qualitative description of the proposed categorization technique for a trial-and-error learning paradigm is given. Experimental results show that the method of control is efficient, when learning starts from scratch, as well as after some major disturbances of an already experienced system.> Artur Dubrawski, Patrick Reignier |
IROS | 1 |