VLDB 2026 Research / reviewers in the wild / expert
Weng-Keen Wong
dblp:19/1015
· DBLP profile ↗
58ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-6673-343XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 10 · 1 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8Computer networks · 4 · 4 since 2021Software engineering, systems software and programming languages · 3Systems, architecture and hardware · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Localized Near Surface Temperature Inversion Forecasting Using Long Short-Term MemoryabstractNear surface temperature inversions are periods in which a low layer of warm air is trapped between cooler air higher up in the atmosphere and dense cooler air below it near the surface level. By causing cooler air to pool near the surface level, inversions can have detrimental effects for crop growers, including frost, increased moisture, and pesticide drift. As a result, predicting the occurrence and magnitude of these inversions yields substantial benefits for growers. We introduce a Long Short-Term Memory (LSTM) model for temperature inversion forecasting that is able to effectively predict localized, near surface temperature inversions in advance such that growers can take actions to mitigate the detrimental effects. We show a substantial performance gain over a deployed temperature inversion forecasting system, and include a series of ablations that show the benefit of using publicly available terrain-specific feature information when modeling inversions at this scale. Taylor Dinkins, Weng-Keen Wong, Basavaraj R. Amogi, Paola Pesantez-Cabrera, Jaitun Patel, Lav R. Khot, Alan Fern |
AAAI | 2 |
| 2025 | Data Augmentation Approaches for Satellite ImageryabstractDeep learning models commonly benefit from data augmentation techniques to diversify the set of training images. When working with satellite imagery, it is common for practitioners to apply a limited set of transformations developed for natural images (e.g., flip and rotate) to expand the training set without overly modifying the satellite images. There are many techniques for natural image data augmentation, but given the differences between the two domains, it is not clear whether data augmentation methods developed for natural images are well suited for satellite imagery. This paper presents an extensive experimental study on three classification and three regression tasks over four satellite image datasets. We compare common computer vision data augmentation techniques and propose three novel satellite-specific data augmentation strategies. Across tasks and datasets, we find that geometric transformations are beneficial for satellite imagery while color transformations generally are not. Additionally, our novel Sat-SlideMix, Sat-CutMix, and Sat-Trivial methods all exhibit strong performance across all tasks and datasets. Laurel M. Hopkins, Weng-Keen Wong, Hannah Kerner, Fuxin Li, Rebecca A. Hutchinson |
AAAI | 2 |
| 2025 | Incorporating Interactive User Feedback Through Precision Matrix Adjustments for High-Dimensional Anomaly DetectionabstractAnomaly detection involves finding unusual data instances of interest. This task has often been compared to looking for a needle in a haystack, especially in high dimensions. One strategy for improving anomaly detection is to leverage expert feedback in a human-in-the-loop process to label anomalies and nominal points. For high-dimensional data, the majority of existing anomaly detection approaches that incorporate expert feedback do not scale and often result in algorithms that are too slow to operate in an interactive setting with a human expert. We introduce a new anomaly detection algorithm intended for high dimensional data that can efficiently incorporate expert feedback on each round of querying. Our approach uses ideas from metric learning to perform an efficient incremental update when the user labels a data instance. We demonstrate through extensive experiments that our work provides the best tradeoff between performance and running time among existing approaches. Taylor Dinkins, Weng-Keen Wong, Yanchi Liu, Kai Ishikawa |
ICDM | 2 |
| 2024 | Unsupervised Contrastive Learning for Robust RF Device Fingerprinting Under Time-Domain ShiftabstractRadio Frequency (RF) device fingerprinting has been recognized as a potential technology for enabling automated wireless device identification and classification. However, it faces a key challenge due to the domain shift that could arise from variations in the channel conditions and environmental settings, potentially degrading the accuracy of RF-based device classification when testing and training data is collected in different domains. This paper introduces a novel solution that leverages contrastive learning to mitigate this domain shift problem. Contrastive learning, a state-of-the-art self-supervised learning approach from deep learning, learns a distance metric such that positive pairs are closer (i.e. more similar) in the learned metric space than negative pairs. When applied to RF fingerprinting, our model treats RF signals from the same transmission as positive pairs and those from different transmissions as negative pairs. Through experiments on wireless and wired RF datasets collected over several days, we demonstrate that our contrastive learning approach captures domain-invariant features, diminishing the effects of domain-specific variations. Our results show large and consistent improvements in accuracy (10.8% to 27.8%) over baseline models, thus underscoring the effectiveness of contrastive learning in improving device classification under domain shift. Weng-Keen Wong, Bechir Hamdaoui |
ICC | 2 |
| 2024 | Biharmonic Distance of Graphs and its Higher-Order Variants: Theoretical Properties with Applications to Centrality and ClusteringabstractEffective resistance is a distance between vertices of a graph that is both theoretically interesting and useful in applications. We study a variant of effective resistance called the biharmonic distance. While the effective resistance measures how well-connected two vertices are, we prove several theoretical results supporting the idea that the biharmonic distance measures how important an edge is to the global topology of the graph. Our theoretical results connect the biharmonic distance to well-known measures of connectivity of a graph like its total resistance and sparsity. Based on these results, we introduce two clustering algorithms using the biharmonic distance. Finally, we introduce a further generalization of the biharmonic distance that we call the $k$-harmonic distance. We empirically study the utility of biharmonic and $k$-harmonic distance for edge centrality and graph clustering. Mitchell Black 0002, Lucy Lin, Weng-Keen Wong, Amir Nayyeri |
ICML | 3 |
| 2023 | Cross-GAN Auditing: Unsupervised Identification of Attribute Level Similarities and Differences Between Pretrained Generative ModelsabstractGenerative Adversarial Networks (GANs) are notoriously difficult to train especially for complex distributions and with limited data. This has driven the need for tools to audit trained networks in human intelligible format, for example, to identify biases or ensure fairness. Existing GAN audit tools are restricted to coarse-grained, modeldata comparisons based on summary statistics such as FID or recall. In this paper, we propose an alternative approach that compares a newly developed GAN against a prior baseline. To this end, we introduce Cross-GAN Auditing (xGA) that, given an established “reference” GAN and a newly proposed “client” GAN, jointly identifies intelligible attributes that are either common across both GANs, novel to the client GAN, or missing from the client GAN. This provides both users and model developers an intuitive assessment of similarity and differences between GANs. We introduce novel metrics to evaluate attribute-based GAN auditing approaches and use these metrics to demonstrate quantitatively that xGA outperforms baseline approaches. We also include qualitative results that illustrate the common, novel and missing attributes identified by xGA from GANs trained on a variety of image datasets1 Matthew L. Olson, Shusen Liu 0001, Rushil Anirudh, Jayaraman J. Thiagarajan, Peer-Timo Bremer, Weng-Keen Wong |
CVPR | 6 |
| 2023 | ADL-ID: Adversarial Disentanglement Learning for Wireless Device Fingerprinting Temporal Domain AdaptationabstractAs the journey of 5G standardization is coming to an end, academia and industry have already begun to consider the sixth-generation (6G) wireless networks, with an aim to meet the service demands for the next decade. Deep learning-based RF fingerprinting (DL-RFFP) has recently been recognized as a potential solution for enabling key wireless network applications and services, such as spectrum policy enforcement and network access control. The state-of-the-art DL-RFFP frameworks suffer from a significant performance drop when tested with data drawn from a domain that is different from that used for training data. In this paper, we propose ADL-ID, an unsupervised domain adaption framework that is based on adversarial disentanglement representation to address the temporal domain adaptation for the RFFP task. Our framework has been evaluated on real LoRa and WiFi datasets and showed about 24% improvement in accuracy when compared to the baseline CNN network on short-term temporal adaptation. It also improves the classification accuracy by up to 9% on long-term temporal adaptation. Furthermore, we release a 5-day, 2.1TB, large-scale WiFi 802.11b dataset collected from 50 Pycom devices to support the research community efforts in developing and validating robust RFFP methods. Abdurrahman Elmaghbub, Bechir Hamdaoui, Weng-Keen Wong |
ICC | 3 |
| 2023 | HiNoVa: A Novel Open-Set Detection Method for Automating RF Device AuthenticationabstractNew capabilities in wireless network security have been enabled by deep learning, which leverages patterns in radio frequency (RF) data to identify and authenticate devices. Open-set detection is an area of deep learning that identifies samples captured from new devices during deployment that were not part of the training set. Past work in open-set detection has mostly been applied to independent and identically distributed data such as images. In contrast, RF signal data present a unique set of challenges as the data forms a time series with non-linear time dependencies among the samples. We introduce a novel open-set detection approach based on the patterns of the hidden state values within a Convolutional Neural Network Long Short-Term Memory model. Our approach greatly improves the Area Under the Precision-Recall Curve on LoRa, Wireless-WiFi, and Wired-WiFi datasets, and hence, can be used successfully to monitor and control unauthorized network access of wireless devices. Luke Puppo, Weng-Keen Wong, Bechir Hamdaoui, Abdurrahman Elmaghbub |
ISCC | 2 |
| 2022 | An Analysis of Complex-Valued CNNs for RF Data-Driven Wireless Device ClassificationabstractRecent deep neural network-based device classification studies show that complex-valued neural networks (CVNNs) yield higher classification accuracy than real-valued neural networks (RVNNs). Although this improvement is (intuitively) attributed to the complex nature of the input RF data (i.e., IQ symbols), no prior work has taken a closer look into analyzing such a trend in the context of wireless device identification. Our study provides a deeper understanding of this trend using real LoRa and WiFi RF datasets. We perform a deep dive into understanding the impact of (i) the input representation/type and (ii) the architectural layer of the neural network. For the input representation, we considered the IQ as well as the polar coordinates both partially and fully. For the architectural layer, we considered a series of ablation experiments that eliminate parts of the CVNN components. Our results show that CVNNs consistently outperform RVNNs counterpart in the various scenarios mentioned above, indicating that CVNNs are able to make better use of the joint information provided via the in-phase (I) and quadrature (Q) components of the signal. Weng-Keen Wong, Bechir Hamdaoui, Abdurrahman Elmaghbub, Kathiravetpillai Sivanesan, Richard Dorrance, Lily L. Yang |
ICC | 2 |
| 2021 | Counterfactual state explanations for reinforcement learning agents via generative deep learning
Matthew L. Olson, Roli Khanna, Lawrence Neal, Fuxin Li, Weng-Keen Wong |
Artif. Intell. | 5 |
| 2020 | The Quantile Snapshot Scan: Comparing Quantiles of Spatial Data from Two Snapshots in TimeabstractWe introduce the Quantile Snapshot Scan (Qsnap), a spatial scan algorithm which identifies spatial regions that differ the most between two snapshots in time. Qsnap is designed for spatial data with a numeric response and a vector of associated covariates for each spatial data point. Qsnap focuses on differences involving a specific quantile of the data distribution. A naive implementation of Qsnap is too computationally expensive for large datasets but our novel incremental update provides an order of magnitude speedup. We demonstrate Qsnap’s effectiveness over an extensive set of experiments on simulated data. In addition, we apply Qsnap to two real-world problems: discovering bird migration paths and identifying regions with dramatic changes in drought conditions. Travis Moore, Weng-Keen Wong |
AISTATS | 2 |
| 2020 | IoT Device Type Identification Using Hybrid Deep Learning Approach for Increased IoT SecurityabstractIoT networks can be viewed as collections of Internet-enabled physical devices and objects, embedded with sensor, actuator, computation, storage and communication components, that are capable of connecting and exchanging data to one another. In recent years, organizations have allowed more and more IoT devices to be connected to their networks, thereby increasing their risks of and exposure to security vulnerabilities and threats. Therefore, it is important for such organizations to be able to identify which devices are connected to their network and which ones are legitimate and pose no risk. Leveraging network traffic to identify devices through supervised learning has recently been gaining popularity, where feature information is first extracted by intercepting device traffic and then exploited to provide device classification. The main limitation of prior works is that they can only identify previously seen types of devices, and any newly added device types are treated as abnormal types. In the real world, hundreds of millions of new IoT devices are produced each year, and the lack of a large amount of training data makes a system based solely on supervised learning unrealistic. In this paper, we propose a hybrid supervised and unsupervised learning method that enables secondary classification of unseen device types. Our technique combines deep neural networks with clustering to enable both seen and unseen device classification, and employs autoencoder technique to reduce dimensionality of datasets, thereby providing a good balance between classification accuracy and overhead. Jiaqi Bao, Bechir Hamdaoui, Weng-Keen Wong |
IWCMC | 3 |
| 2020 | Discovering Anomalies by Incorporating Feedback from an ExpertabstractUnsupervised anomaly detection algorithms search for outliers and then predict that these outliers are the anomalies. When deployed, however, these algorithms are often criticized for high false-positive and high false-negative rates. One main cause of poor performance is that not all outliers are anomalies and not all anomalies are outliers. In this article, we describe the Active Anomaly Discovery (AAD) algorithm, which incorporates feedback from an expert user that labels a queried data instance as an anomaly or nominal point. This feedback is intended to adjust the anomaly detector so that the outliers it discovers are more in tune with the expert user’s semantic understanding of the anomalies. The AAD algorithm is based on a weighted ensemble of anomaly detectors. When it receives a label from the user, it adjusts the weights on each individual ensemble member such that the anomalies rank higher in terms of their anomaly score than the outliers. The AAD approach is designed to operate in an interactive data exploration loop. In each iteration of this loop, our algorithm first selects a data instance to present to the expert as a potential anomaly and then the expert labels the instance as an anomaly or as a nominal data point. When it receives the instance label, the algorithm updates its internal model and the loop continues until a budget of B queries is spent. The goal of our approach is to maximize the total number of true anomalies in the B instances presented to the expert. We show that the AAD method performs well and in some cases doubles the number of true anomalies found compared to previous methods. In addition we present approximations that make the AAD algorithm much more computationally efficient while maintaining a desirable level of performance. Shubhomoy Das, Weng-Keen Wong, Thomas G. Dietterich, Alan Fern, Andrew Emmott |
ACM Trans. Knowl. Discov. Data | 2 |
| 2019 | Sequential Feature Explanations for Anomaly DetectionabstractIn many applications, an anomaly detection system presents the most anomalous data instance to a human analyst, who then must determine whether the instance is truly of interest (e.g., a threat in a security setting). Unfortunately, most anomaly detectors provide no explanation about why an instance was considered anomalous, leaving the analyst with no guidance about where to begin the investigation. To address this issue, we study the problems of computing and evaluating sequential feature explanations (SFEs) for anomaly detectors. An SFE of an anomaly is a sequence of features, which are presented to the analyst one at a time (in order) until the information contained in the highlighted features is enough for the analyst to make a confident judgement about the anomaly. Since analyst effort is related to the amount of information that they consider in an investigation, an explanation’s quality is related to the number of features that must be revealed to attain confidence. In this article, we first formulate the problem of optimizing SFEs for a particular density-based anomaly detector. We then present both greedy algorithms and an optimal algorithm, based on branch-and-bound search, for optimizing SFEs. Finally, we provide a large scale quantitative evaluation of these algorithms using a novel framework for evaluating explanations. The results show that our algorithms are quite effective and that our best greedy algorithm is competitive with optimal solutions. Md Amran Siddiqui, Alan Fern, Thomas G. Dietterich, Weng-Keen Wong |
ACM Trans. Knowl. Discov. Data | 4 |
| 2018 | Open Set Learning with Counterfactual Images
Lawrence Neal, Matthew L. Olson, Xiaoli Z. Fern, Weng-Keen Wong, Fuxin Li |
ECCV (6) | 4 |
| 2018 | Discriminative Probabilistic Framework for Generalized Multi-Instance LearningabstractMultiple-instance learning is a framework for learning from data consisting of bags of instances labeled at the bag level. A common assumption in multi-instance learning is that a bag label is positive if and only if at least one instance in the bag is positive. In practice, this assumption may be violated. For example, experts may provide a noisy label to a bag consisting of many instances, to reduce labeling time. Here, we consider generalized multi-instance learning, which assumes that the bag label is non-deterministically determined based on the number of positive instances in the bag. The challenge in this setting is to simultaneous learn an instance classifier and the unknown bag-labeling probabilistic rule. This paper addresses the generalized multi-instance learning using a discriminative probabilistic graphical model with exact and efficient inference. Experiments on both synthetic and real data illustrate the effectiveness of the proposed method relative to other methods including those that follow the traditional multiple-instance learning assumption. Anh T. Pham 0002, Raviv Raich, Xiaoli Z. Fern, Weng-Keen Wong, Xinze Guan |
ICASSP | 4 |
| 2018 | An Efficient Quantile Spatial Scan Statistic for Finding Unusual Regions in Continuous Spatial Data with Covariates
Travis Moore, Weng-Keen Wong |
UAI | 2 |
| 2016 | Incorporating Expert Feedback into Active Anomaly DiscoveryabstractUnsupervised anomaly detection algorithms search for outliers and then predict that these outliers are the anomalies. When deployed, however, these algorithms are often criticized for high false positive and high false negative rates. One cause of poor performance is that not all outliers are anomalies and not all anomalies are outliers. In this paper, we describe an Active Anomaly Discovery (AAD) method for incorporating expert feedback to adjust the anomaly detector so that the outliers it discovers are more in tune with the expert user's semantic understanding of the anomalies. The AAD approach is designed to operate in an interactive data exploration loop. In each iteration of this loop, our algorithm first selects a data instance to present to the expert as a potential anomaly and then the expert labels the instance as an anomaly or as a nominal data point. Our algorithm updates its internal model with the instance label and the loop continues until a budget of B queries is spent. The goal of our approach is to maximize the total number of true anomalies in the B instances presented to the expert. We show that when compared to other state-of-the-art algorithms, AAD is consistently one of the best performers. Shubhomoy Das, Weng-Keen Wong, Thomas G. Dietterich, Alan Fern, Andrew Emmott |
ICDM | 2 |
| 2016 | Efficient Multi-Instance Learning for Activity Recognition from Time Series Data Using an Auto-Regressive Hidden Markov ModelabstractActivity recognition from sensor data has spurred a great deal of interest due to its impact on health care. Prior work on activity recognition from multivariate time series data has mainly applied supervised learning techniques which require a high degree of annotation effort to produce training data with the start and end times of each activity. In order to reduce the annotation effort, we present a weakly supervised approach based on multi-instance learning. We introduce a generative graphical model for multi-instance learning on time series data based on an auto-regressive hidden Markov model. Our model has a number of advantages, including the ability to produce both bag and instance-level predictions as well as an efficient exact inference algorithm based on dynamic programming. Xinze Guan, Raviv Raich, Weng-Keen Wong |
ICML | 3 |
| 2015 | Principles of Explanatory Debugging to Personalize Interactive Machine LearningabstractHow can end users efficiently influence the predictions that machine learning systems make on their behalf? This paper presents Explanatory Debugging, an approach in which the system explains to users how it made each of its predictions, and the user then explains any necessary corrections back to the learning system. We present the principles underlying this approach and a prototype instantiating it. An empirical evaluation shows that Explanatory Debugging increased participants' understanding of the learning system by 52% and allowed participants to correct its mistakes up to twice as efficiently as participants using a traditional learning system. Todd Kulesza, Margaret M. Burnett, Weng-Keen Wong, Simone Stumpf |
IUI | 3 |
| 2015 | TIPR: transcription initiation pattern recognition on a genome scaleabstractMOTIVATION: The computational identification of gene transcription start sites (TSSs) can provide insights into the regulation and function of genes without performing expensive experiments, particularly in organisms with incomplete annotations. High-resolution general-purpose TSS prediction remains a challenging problem, with little recent progress on the identification and differentiation of TSSs which are arranged in different spatial patterns along the chromosome. RESULTS: In this work, we present the Transcription Initiation Pattern Recognizer (TIPR), a sequence-based machine learning model that identifies TSSs with high accuracy and resolution for multiple spatial distribution patterns along the genome, including broadly distributed TSS patterns that have previously been difficult to characterize. TIPR predicts not only the locations of TSSs but also the expected spatial initiation pattern each TSS will form along the chromosome-a novel capability for TSS prediction algorithms. As spatial initiation patterns are associated with spatiotemporal expression patterns and gene function, this capability has the potential to improve gene annotations and our understanding of the regulation of transcription initiation. The high nucleotide resolution of this model locates TSSs within 10 nucleotides or less on average. AVAILABILITY AND IMPLEMENTATION: Model source code is made available online at http://megraw.cgrb.oregonstate.edu/software/TIPR/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Taj Morton, Weng-Keen Wong, Molly Megraw |
Bioinform. | 2 |
| 2014 | A Latent Variable Model for Discovering Bird Species Commonly Misidentified by Citizen ScientistsabstractData quality is a common source of concern for large-scale citizen science projects like eBird. In the case of eBird, a major cause of poor quality data is the misidentification of bird species by inexperienced contributors. A proactive approach for improving data quality is to discover commonly misidentified bird species and to teach inexperienced birders the differences between these species. To accomplish this goal, we develop a latent variable graphical model that can identify groups of bird species that are often confused for each other by eBird participants. Our model is a multi-species extension of the classic occupancy-detection model in the ecology literature. This multi-species extension requires a structure learning step as well as a computationally expensive parameter learning stage which we make efficient through a variational approximation. We show that our model can not only discover groups of misidentified species, but by including these misidentifications in the model, it can also achieve more accurate predictions of both species occupancy and detection. Rebecca A. Hutchinson, Weng-Keen Wong |
AAAI | 3 |
| 2014 | Clustering Species Accumulation Curves to Identify Skill Levels of Citizen Scientists Participating in the eBird ProjectabstractAlthough citizen science projects such as eBird can compile large volumes of data over broad spatial and temporal extents, the quality of this data is a concern due to differences in the skills of volunteers at identifying bird species. Species accumulation curves, which plot the number of unique species observed over time, are an effective way to quantify the skill level of an eBird participant. Intuitively, more skilled observers can identify a greater number of species per unit time than inexperienced birders, resulting in a steeper curve. We propose a mixture model for clustering species accumulation curves. These clusters enable the identification of distinct skill levels of eBird participants, which can then be used to build more accurate species distribution models and to develop automated data quality filters. Weng-Keen Wong, Steve Kelling |
AAAI | 2 |
| 2014 | Using Crowdsourcing to Generate Surrogate Training Data for Robotic Grasp PredictionabstractAs an alternative to the laborious process of collecting training data from physical robotic platforms for learning robotic grasp quality prediction, we explore the use of surrogate training data from crowd-sourced evaluations of images of robotic grasps. We show that in certain regions of the grasp feature space, grasp predictors trained with this surrogate data were almost as accurate as predictors built using data from physical testing with robots. Matt Unrath, Alex K. Goins, Ryan Carpenter, Weng-Keen Wong, Ravi Balasubramanian |
HCOMP | 5 |
| 2014 | Evaluating the efficacy of grasp metrics for utilization in a Gaussian Process-based grasp predictorabstractWith the goal of advancing the state of automatic robotic grasping, we present a novel approach that combines machine learning techniques and rigorous validation on a physical robotic platform in order to develop an algorithm that predicts the quality of a robotic grasp before execution. After collecting a large grasp sample set (522 grasps), we first conduct a thorough statistical analysis of the ability of grasp metrics that are commonly used in the robotics literature to discriminate between good and bad grasps. We then apply Principal Component Analysis and Gaussian Process algorithms on the discriminative grasp metrics to build a classifier that predicts grasp quality. The key findings are as follows: (i) several of the grasp metrics in the literature are weak predictors of grasp quality when implemented on a physical robotic platform; (ii) the Gaussian Process-based classifier significantly improves grasp prediction techniques by providing an absolute grasp quality prediction score from combining multiple grasp metrics. Specifically, the GP classifier showed a 66% percent improvement in the True Positive classification rate at a low False Positive rate of 5% when compared with classification based on thresholding of individual grasp metrics. Alex K. Goins, Ryan Carpenter, Weng-Keen Wong, Ravi Balasubramanian |
IROS | 3 |
| 2014 | Latent dirichlet allocation based diversified retrieval for e-commerce searchabstractDiversified retrieval is a very important problem on many e-commerce sites, e.g. eBay and Amazon. Using IR approaches without optimizing for diversity results in a clutter of redundant items that belong to the same products. Most existing product taxonomies are often too noisy, with overlapping structures and non-uniform granularity, to be used directly in diversified retrieval. To address this problem, we propose a Latent Dirichlet Allocation (LDA) based diversified retrieval approach that selects diverse items based on the hidden user intents. Our approach first discovers the hidden user intents of a query using the LDA model, and then ranks the user intents by making trade-offs between their relevance and information novelty. Finally, it chooses the most representative item for each user intent to display. To evaluate the diversity in the search results on e-commerce sites, we propose a new metric, average satisfaction, measuring user satisfaction with the search results. Through our empirical study on eBay, we show that the LDA model discovers meaningful user intents and the LDA-based approach provides significantly higher user satisfaction than the eBay production ranker and three other diversified retrieval approaches. Sunil Mohan, Duangmanee Putthividhya, Weng-Keen Wong |
WSDM | 4 |
| 2014 | You Are the Only Possible Oracle: Effective Test Selection for End Users of Interactive Machine Learning SystemsabstractHow do you test a program when only a single user, with no expertise in software testing, is able to determine if the program is performing correctly? Such programs are common today in the form of machine-learned classifiers. We consider the problem of testing this common kind of machine-generated program when the only oracle is an end user: e.g., only you can determine if your email is properly filed. We present test selection methods that provide very good failure rates even for small test suites, and show that these methods work in both large-scale random experiments using a “gold standard” and in studies with real users. Our methods are inexpensive and largely algorithm-independent. Key to our methods is an exploitation of properties of classifiers that is not possible in traditional software testing. Our results suggest that it is plausible for time-pressured end users to interactively detect failures-even very hard-to-find failures-without wading through a large number of successful (and thus less useful) tests. We additionally show that some methods are able to find the arguably most difficult-to-detect faults of classifiers: cases where machine learning algorithms have high confidence in an incorrect result. Alex Groce, Todd Kulesza, Chaoqiang Zhang, Shalini Shamasunder, Margaret M. Burnett, Weng-Keen Wong, Simone Stumpf, Shubhomoy Das, Amber Shinsel, Forrest Bice, Kevin McIntosh |
IEEE Trans. Software Eng. | 6 |
| 2013 | Physical Activity Recognition from Accelerometer Data Using a Multi-Scale Ensemble MethodabstractAccurate and detailed measurement of an individual’s physical activity is a key requirement for helping researchers understand the relationship between physical activity and health. Accelerometers have become the method of choice for measuring physical activity due to their small size, low cost, convenience and their ability to provide objective information about physical activity. However, interpreting accelerometer data once it has been collected can be challenging. In this work, we applied machine learning algorithms to the task of physical activity recognition from triaxial accelerometer data. We employed a simple but effective approach of dividing the accelerometer data into short non-overlapping windows, converting each window into a feature vector, and treating each feature vector as an i.i.d training instance for a supervised learning algorithm. In addition, we improved on this simple approach with a multi-scale ensemble method that did not need to commit to a single window size and was able to leverage the fact that physical activities produced time series with repetitive patterns and discriminative features for physical activity occurred at different temporal scales. Yonglei Zheng, Weng-Keen Wong, Xinze Guan, Stewart Trost |
IAAI | 2 |
| 2013 | Detecting insider threats in a real corporate database of computer usage activityabstractThis paper reports on methods and results of an applied research project by a team consisting of SAIC and four universities to develop, integrate, and evaluate new approaches to detect the weak signals characteristic of insider threats on organizations' information systems. Our system combines structural and semantic information from a real corporate database of monitored activity on their users' computers to detect independently developed red team inserts of malicious insider activities. We have developed and applied multiple algorithms for anomaly detection based on suspected scenarios of malicious insider behavior, indicators of unusual activities, high-dimensional statistical patterns, temporal sequences, and normal graph evolution. Algorithms and representations for dynamic graph processing provide the ability to scale as needed for enterprise-level deployments on real-time data streams. We have also developed a visual language for specifying combinations of features, baselines, peer groups, time periods, and algorithms to detect anomalies suggestive of instances of insider threat behavior. We defined over 100 data features in seven categories based on approximately 5.5 million actions per day from approximately 5,500 users. We have achieved area under the ROC curve values of up to 0.979 and lift values of 65 on the top 50 user-days identified on two months of real data. Ted E. Senator, Henry G. Goldberg, Alex Memory, William T. Young, Bradley Rees, Robert Pierce, Daniel Huang 0003, Matthew Reardon, David A. Bader, Edmond Chow, Irfan A. Essa, Joshua Jones, Vinay Bettadapura, Polo Chau, Oded Green, Oguz Kaya, Anita Zakrzewska, Erica Briscoe, Rudolph Louis Mappus IV, Robert McColl, Lora Weiss, Thomas G. Dietterich, Alan Fern, Weng-Keen Wong, Shubhomoy Das, Andrew Emmott, Jed Irvine, Jay-Yoon Lee, Danai Koutra, Christos Faloutsos, Daniel D. Corkill, Lisa Friedland, Amanda Gentzel, David D. Jensen |
KDD | 24 |
| 2013 | Taming compiler fuzzersabstractAggressive random testing tools ("fuzzers") are impressively effective at finding compiler bugs. For example, a single test-case generator has resulted in more than 1,700 bugs reported for a single JavaScript engine. However, fuzzers can be frustrating to use: they indiscriminately and repeatedly find bugs that may not be severe enough to fix right away. Currently, users filter out undesirable test cases using ad hoc methods such as disallowing problematic features in tests and grepping test results. This paper formulates and addresses the fuzzer taming problem: given a potentially large number of random test cases that trigger failures, order them such that diverse, interesting test cases are highly ranked. Our evaluation shows our ability to solve the fuzzer taming problem for 3,799 test cases triggering 46 bugs in a C compiler and 2,603 test cases triggering 28 bugs in a JavaScript engine. Yang Chen 0024, Alex Groce, Chaoqiang Zhang, Weng-Keen Wong, Xiaoli Z. Fern, Eric Eide, John Regehr |
PLDI | 4 |
| 2013 | Too much, too little, or just right? Ways explanations impact end users' mental modelsabstractResearch is emerging on how end users can correct mistakes their intelligent agents make, but before users can correctly “debug” an intelligent agent, they need some degree of understanding of how it works. In this paper we consider ways intelligent agents should explain themselves to end users, especially focusing on how the soundness and completeness of the explanations impacts the fidelity of end users' mental models. Our findings suggest that completeness is more important than soundness: increasing completeness via certain information types helped participants' mental models and, surprisingly, their perception of the cost/benefit tradeoff of attending to the explanations. We also found that oversimplification, as per many commercial agents, can be a problem: when soundness was very low, participants experienced more mental demand and lost trust in the explanations, thereby reducing the likelihood that users will pay attention to such explanations at all. Todd Kulesza, Simone Stumpf, Margaret M. Burnett, Sherry Yang 0002, Irwin Kwan, Weng-Keen Wong |
VL/HCC | 6 |
| 2013 | End-user feature labeling: Supervised and semi-supervised approaches based on locally-weighted logistic regression
Shubhomoy Das, Travis Moore, Weng-Keen Wong, Simone Stumpf, Ian Oberst, Kevin McIntosh, Margaret M. Burnett |
Artif. Intell. | 3 |
| 2012 | Automated data verification in a large-scale citizen science project: A case studyabstractAlthough citizen science projects can engage a very large number of volunteers to collect volumes of data, they are susceptible to issues with data quality. Our experience with eBird, which is a broad-scale citizen science project to collect bird observations, has shown that a massive effort by volunteer experts is needed to screen data, identify outliers and flag them in the database. The increasing volume of data being collected by eBird places a huge burden on these volunteer experts and other automated approaches to improve data quality are needed. In this work, we describe a case study in which we evaluate an automated data quality filter that improves data quality by identifying outliers and categorizing these outliers as either unusual valid observations or mis-identified (invalid) observations. This automated data filter involves a two-step process: first, a data-driven method detects outliers (ie. observations that are unusual for a given region and date). Next, we use a data quality model based on an observer's predicted expertise to decide if an outlier should be flagged for review. We applied this automated data filter retrospectively to eBird data from Tompkins Co., NY and found that that this automated process significantly reduced the workload of reviewers by as much as 43% and identifies 52% more potentially invalid observations. Steve Kelling, Jeff Gerbracht, Weng-Keen Wong |
eScience | 4 |
| 2012 | eBird: A Human/Computer Learning Network for Biodiversity Conservation and ResearchabstractIn this paper we describe eBird, a citizen science project that takes advantage of human observational capacity and machine learning methods to explore the synergies between human computation and mechanical computation. We call this model a Human/Computer Learning Network, whose core is an active learning feedback loop between humans and machines that dramatically improves the quality of both, and thereby continually improves the effectiveness of the network as a whole. Human/Computer Learning Networks leverage the contributions of a broad recruitment of human observers and processes their contributed data with Artificial Intelligence algorithms leading to a computational power that far exceeds the sum of the individual parts. Steve Kelling, Jeff Gerbracht, Daniel Fink 0002, Carl Lagoze, Weng-Keen Wong, Theodoros Damoulas, Carla P. Gomes |
IAAI | 5 |
| 2012 | Towards recognizing "cool": can end users help computer vision recognize subjective attributes of objects in images?abstractRecent computer vision approaches are aimed at richer image interpretations that extend the standard recognition of objects in images (e.g., cars) to also recognize object attributes (e.g., cylindrical, has-stripes, wet). However, the more idiosyncratic and abstract the notion of an object attribute (e.g., cool car), the more challenging the task of attribute recognition. This paper considers whether end users can help vision algorithms recognize highly idiosyncratic attributes, referred to here as subjective attributes. We empirically investigated how end users recognized three subjective attributes of carscool, cute, and classic. Our results suggest the feasibility of vision algorithms recognizing subjective attributes of objects, but an interactive approach beyond standard supervised learning from labeled training examples is needed. William Curran, Travis Moore, Todd Kulesza, Weng-Keen Wong, Sinisa Todorovic, Simone Stumpf, Rachel White, Margaret M. Burnett |
IUI | 4 |
| 2012 | An Ensemble Architecture for Learning Complex Problem-Solving Techniques from DemonstrationabstractWe present a novel ensemble architecture for learning problem-solving techniques from a very small number of expert solutions and demonstrate its effectiveness in a complex real-world domain. The key feature of our “Generalized Integrated Learning Architecture” (GILA) is a set of heterogeneous independent learning and reasoning (ILR) components, coordinated by a central meta-reasoning executive (MRE). The ILRs are weakly coupled in the sense that all coordination during learning and performance happens through the MRE. Each ILR learns independently from a small number of expert demonstrations of a complex task. During performance, each ILR proposes partial solutions to subproblems posed by the MRE, which are then selected from and pieced together by the MRE to produce a complete solution. The heterogeneity of the learner-reasoners allows both learning and problem solving to be more effective because their abilities and biases are complementary and synergistic. We describe the application of this novel learning and problem solving architecture to the domain of airspace management, where multiple requests for the use of airspaces need to be deconflicted, reconciled, and managed automatically. Formal evaluations show that our system performs as well as or better than humans after learning from the same training data. Furthermore, GILA outperforms any individual ILR run in isolation, thus demonstrating the power of the ensemble architecture for learning and problem solving. Xiaoqin Zhang 0001, Bhavesh Shrestha, Subbarao Kambhampati, Phillip DiBona, Jinhong K. Guo, Daniel McFarlane, Martin O. Hofmann, Kenneth R. Whitebread, Darren Scott Appling, Elizabeth T. Whitaker, Ethan Trewhitt, Li Ding 0001, James Michaelis, Deborah L. McGuinness, James A. Hendler, Janardhan Rao Doppa, Thomas G. Dietterich, Prasad Tadepalli, Weng-Keen Wong, Derek T. Green, Antons Rebguns, Diana F. Spears, Ugur Kuter, Geoffrey Levine, Gerald DeJong, Reid MacTavish, Santiago Ontañón, Jainarayan Radhakrishnan, Ashwin Ram 0001, Hala Mostafa, Huzaifa Zafar, Chongjie Zhang, Daniel D. Corkill, Victor R. Lesser, Zhexuan Song |
ACM Trans. Intell. Syst. Technol. | 21 |
| 2011 | End-User Feature Labeling via Locally Weighted Logistic RegressionabstractApplications that adapt to a particular end user often make inaccurate predictions during the early stages when training data is limited. Although an end user can improve the learning algorithm by labeling more training data, this process is time consuming and too ad hoc to target a particular area of inaccuracy. To solve this problem, we propose a new learning algorithm based on Locally Weighted Logistic Regression for feature labeling by end users, enabling them to point out which features are important for a class, rather than provide new training instances. In our user study, the first allowing ordinary end users to freely choose features to label directly from text documents, our algorithm was more effective than others at leveraging end users’ feature labels to improve the learning algorithm. Our results strongly suggest that allowing users to freely choose features to label is a promising method for allowing end users to improve learning algorithms effectively. Weng-Keen Wong, Ian Oberst, Shubhomoy Das, Travis Moore, Simone Stumpf, Kevin McIntosh, Margaret M. Burnett |
AAAI | 1 |
| 2011 | End-user feature labeling: a locally-weighted regression approachabstractWhen intelligent interfaces, such as intelligent desktop assistants, email classifiers, and recommender systems, customize themselves to a particular end user, such customizations can decrease productivity and increase frustration due to inaccurate predictions - especially in early stages, when training data is limited. The end user can improve the learning algorithm by tediously labeling a substantial amount of additional training data, but this takes time and is too ad hoc to target a particular area of inaccuracy. To solve this problem, we propose a new learning algorithm based on locally weighted regression for feature labeling by end users, enabling them to point out which features are important for a class, rather than provide new training instances. In our user study, the first allowing ordinary end users to freely choose features to label directly from text documents, our algorithm was both more effective than others at leveraging end users' feature labels to improve the learning algorithm, and more robust to real users' noisy feature labels. These results strongly suggest that allowing users to freely choose features to label is a promising method for allowing end users to improve learning algorithms effectively. Weng-Keen Wong, Ian Oberst, Shubhomoy Das, Travis Moore, Simone Stumpf, Kevin McIntosh, Margaret M. Burnett |
IUI | 1 |
| 2011 | Mini-crowdsourcing end-user assessment of intelligent assistants: A cost-benefit studyabstractIntelligent assistants sometimes handle tasks too important to be trusted implicitly. End users can establish trust via systematic assessment, but such assessment is costly. This paper investigates whether, when, and how bringing a small crowd of end users to bear on the assessment of an intelligent assistant is useful from a cost/benefit perspective. Our results show that a mini-crowd of testers supplied many more benefits than the obvious decrease in workload, but these benefits did not scale linearly as mini-crowd size increased - there was a point of diminishing returns where the cost-benefit ratio became less attractive. Amber Shinsel, Todd Kulesza, Margaret M. Burnett, William Curran, Alex Groce, Simone Stumpf, Weng-Keen Wong |
VL/HCC | 7 |
| 2011 | Why-oriented end-user debugging of naive Bayes text classificationabstractMachine learning techniques are increasingly used in intelligent assistants , that is, software targeted at and continuously adapting to assist end users with email, shopping, and other tasks. Examples include desktop SPAM filters, recommender systems, and handwriting recognition. Fixing such intelligent assistants when they learn incorrect behavior, however, has received only limited attention. To directly support end-user “debugging” of assistant behaviors learned via statistical machine learning, we present a Why-oriented approach which allows users to ask questions about how the assistant made its predictions, provides answers to these “why” questions, and allows users to interactively change these answers to debug the assistant's current and future predictions. To understand the strengths and weaknesses of this approach, we then conducted an exploratory study to investigate barriers that participants could encounter when debugging an intelligent assistant using our approach, and the information those participants requested to overcome these barriers. To help ensure the inclusiveness of our approach, we also explored how gender differences played a role in understanding barriers and information needs. We then used these results to consider opportunities for Why-oriented approaches to address user barriers and information needs. Todd Kulesza, Simone Stumpf, Weng-Keen Wong, Margaret M. Burnett, Stephen Perona, Amy J. Ko, Ian Oberst |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2010 | Modeling Experts and Novices in Citizen Science Data for Species Distribution ModelingabstractCitizen scientists, who are volunteers from the community that participate as field assistants in scientific studies, enable research to be performed at much larger spatial and temporal scales than trained scientists can cover. Species distribution modeling, which involves understanding species-habitat relationships, is a research area that can benefit greatly from citizen science. The eBird project is one of the largest citizen science programs in existence. By allowing birders to upload observations of bird species to an online database, eBird can provide useful data for species distribution modeling. However, since birders vary in their levels of expertise, the quality of data submitted to eBird is often questioned. In this paper, we develop a probabilistic model called the Occupancy-Detection-Expertise (ODE) model that incorporates the expertise of birders submitting data to eBird. We show that modeling the expertise of birders can improve the accuracy of predicting observations of a bird species at a site. In addition, we can use the ODE model for two other tasks: predicting birder expertise given their history of eBird checklists and identifying bird species that are difficult for novices to detect. Weng-Keen Wong, Rebecca A. Hutchinson |
ICDM | 2 |
| 2010 | Explanatory Debugging: Supporting End-User Debugging of Machine-Learned ProgramsabstractMany machine-learning algorithms learn rules of behavior from individual end users, such as task-oriented desktop organizers and handwriting recognizers. These rules form a “program” that tells the computer what to do when future inputs arrive. Little research has explored how an end user can debug these programs when they make mistakes. We present our progress toward enabling end users to debug these learned programs via a Natural Programming methodology. We began with a formative study exploring how users reason about and correct a text-classification program. From the results, we derived and prototyped a concept based on “explanatory debugging”, then empirically evaluated it. Our results contribute methods for exposing a learned program's logic to end users and for eliciting user corrections to improve the program's predictions. Todd Kulesza, Simone Stumpf, Margaret M. Burnett, Weng-Keen Wong, Yann Riche, Travis Moore, Ian Oberst, Amber Shinsel, Kevin McIntosh |
VL/HCC | 4 |
| 2010 | Supersplat - spliced RNA-seq alignmentabstractMOTIVATION: High-throughput sequencing technologies have recently made deep interrogation of expressed transcript sequences practical, both economically and temporally. Identification of intron/exon boundaries is an essential part of genome annotation, yet remains a challenge. Here, we present supersplat, a method for unbiased splice-junction discovery through empirical RNA-seq data. RESULTS: Using a genomic reference and RNA-seq high-throughput sequencing datasets, supersplat empirically identifies potential splice junctions at a rate of approximately 11.4 million reads per hour. We further benchmark the performance of the algorithm by mapping Illumina RNA-seq reads to identify introns in the genome of the reference dicot plant Arabidopsis thaliana and we demonstrate the utility of supersplat for de novo empirical annotation of splice junctions using the reference monocot plant Brachypodium distachyon. AVAILABILITY: Implemented in C++, supersplat source code and binaries are freely available on the web at http://mocklerlab-tools.cgrb.oregonstate.edu/. Douglas W. Bryant Jr., Rongkun Shen, Henry D. Priest, Weng-Keen Wong, Todd C. Mockler |
Bioinform. | 4 |
| 2010 | Machine learning algorithms for event detection
Dragos D. Margineantu, Weng-Keen Wong, Denver Dash |
Mach. Learn. | 2 |
| 2009 | An Ensemble Learning and Problem Solving Architecture for Airspace Management
Xiaoqin Zhang 0001, Phillip DiBona, Darren Scott Appling, Li Ding 0001, Janardhan Rao Doppa, Derek T. Green, Jinhong K. Guo, Ugur Kuter, Geoffrey Levine, Reid MacTavish, Daniel McFarlane, James Michaelis, Hala Mostafa, Santiago Ontañón, Jainarayan Radhakrishnan, Antons Rebguns, Bhavesh Shrestha, Zhexuan Song, Ethan Trewhitt, Huzaifa Zafar, Chongjie Zhang, Daniel D. Corkill, Gerald DeJong, Thomas G. Dietterich, Subbarao Kambhampati, Victor R. Lesser, Deborah L. McGuinness, Ashwin Ram 0001, Diana F. Spears, Prasad Tadepalli, Elizabeth T. Whitaker, Weng-Keen Wong, James A. Hendler, Martin O. Hofmann, Kenneth R. Whitebread |
IAAI | 34 |
| 2009 | Fixing the program my computer learned: barriers for end users, challenges for the machineabstractThe results of a machine learning from user behavior can be thought of as a program, and like all programs, it may need to be debugged. Providing ways for the user to debug it matters, because without the ability to fix errors users may find that the learned program's errors are too damaging for them to be able to trust such programs. We present a new approach to enable end users to debug a learned program. We then use an early prototype of our new approach to conduct a formative study to determine where and when debugging issues arise, both in general and also separately for males and females. The results suggest opportunities to make machine-learned programs more effective tools. Todd Kulesza, Weng-Keen Wong, Simone Stumpf, Stephen Perona, Rachel White, Margaret M. Burnett, Ian Oberst, Amy J. Ko |
IUI | 2 |
| 2009 | Category detection using hierarchical mean shiftabstractMany applications in surveillance, monitoring, scientific discovery, and data cleaning require the identification of anomalies. Although many methods have been developed to identify statistically significant anomalies, a more difficult task is to identify anomalies that are both interesting and statistically significant. Category detection is an emerging area of machine learning that can help address this issue using a ”human-in-the-loop”approach. In this interactive setting, the algorithm asks the user to label a query data point under Pavan Vatturi, Weng-Keen Wong |
KDD | 2 |
| 2009 | QSRA - a quality-value guided de novo short read assemblerabstractBACKGROUND: New rapid high-throughput sequencing technologies have sparked the creation of a new class of assembler. Since all high-throughput sequencing platforms incorporate errors in their output, short-read assemblers must be designed to account for this error while utilizing all available data. RESULTS: We have designed and implemented an assembler, Quality-value guided Short Read Assembler, created to take advantage of quality-value scores as a further method of dealing with error. Compared to previous published algorithms, our assembler shows significant improvements not only in speed but also in output quality. CONCLUSION: QSRA generally produced the highest genomic coverage, while being faster than VCAKE. QSRA is extremely competitive in its longest contig and N50/N80 contig lengths, producing results of similar quality to those of EDENA and VELVET. QSRA provides a step closer to the goal of de novo assembly of complex genomes, improving upon the original VCAKE algorithm by not only drastically reducing runtimes but also increasing the viability of the assembly algorithm through further error handling capabilities. Douglas W. Bryant Jr., Weng-Keen Wong, Todd C. Mockler |
BMC Bioinform. | 2 |
| 2009 | Interacting meaningfully with machine learning systems: Three experiments
Simone Stumpf, Vidya Rajaram, Lida Li, Weng-Keen Wong, Margaret M. Burnett, Thomas G. Dietterich, Erin Sullivan, Jon Herlocker |
Int. J. Hum. Comput. Stud. | 4 |
| 2008 | Markov Blanket Feature Selection for Support Vector Machines
Jianqiang Shen, Lida Li, Weng-Keen Wong |
AAAI | 3 |
| 2008 | Logical Hierarchical Hidden Markov Models for Modeling User Activities
Sriraam Natarajan, Hung Hai Bui, Prasad Tadepalli, Kristian Kersting, Weng-Keen Wong |
ILP | 5 |
| 2008 | Integrating rich user feedback into intelligent user interfacesabstractThe potential for machine learning systems to improve via a mutually beneficial exchange of information with users has yet to be explored in much detail. Previously, we found that users were willing to provide a generous amount of rich feedback to machine learning systems, and that the types of some of this rich feedback seem promising for assimilation by machine learning algorithms. Following up on those findings, we ran an experiment to assess the viability of incorporating real-time keyword-based feedback in initial training phases when data is limited. We found that rich feedback improved accuracy but an initial unstable period often caused large fluctuations in classifier behavior. Participants were able to give feedback by relying heavily on system communication in order to respond to changes. The results show that in order to benefit from the user's knowledge, machine learning systems must be able to absorb keyword-based rich feedback in a graceful manner and provide clear explanations of their predictions. Simone Stumpf, Erin Sullivan, Erin Fitzhenry, Ian Oberst, Weng-Keen Wong, Margaret M. Burnett |
IUI | 5 |
| 2005 | What's Strange About Recent Events (WSARE): An Algorithm for the Early Detection of Disease OutbreaksabstractTraditional biosurveillance algorithms detect disease outbreaks by looking for peaks in a univariate time series of health-care data. Current health-care surveillance data, however, are no longer simply univariate data streams. Instead, a wealth of spatial, temporal, demographic and symptomatic information is available. We present an early disease outbreak detection algorithm called What's Strange About Recent Events (WSARE), which uses a multivariate approach to improve its timeliness of detection. WSARE employs a rule-based technique that compares recent health-care data against data from a baseline distribution and finds subgroups of the recent data whose proportions have changed the most from the baseline data. In addition, health-care data also pose difficulties for surveillance algorithms because of inherent temporal trends such as seasonal effects and day of week variations. WSARE approaches this problem using a Bayesian network to produce a baseline distribution that accounts for these temporal trends. The algorithm itself incorporates a wide range of ideas, including association rules, Bayesian networks, hypothesis testing and permutation tests to produce a detection algorithm that is careful to evaluate the significance of the alarms that it raises. Weng-Keen Wong, Andrew W. Moore 0001, Gregory F. Cooper, Michael M. Wagner 0001 |
J. Mach. Learn. Res. | 1 |
| 2004 | Bayesian Biosurveillance of Disease Outbreaks
Gregory F. Cooper, Denver Dash, John D. Levander, Weng-Keen Wong, William R. Hogan, Michael M. Wagner 0001 |
UAI | 4 |
| 2003 | Optimal Reinsertion: A New Search Operator for Accelerated and More Accurate Bayesian Network Structure Learning
Andrew W. Moore 0001, Weng-Keen Wong |
ICML | 2 |
| 2003 | Bayesian Network Anomaly Pattern Detection for Disease Outbreaks
Weng-Keen Wong, Andrew W. Moore 0001, Gregory F. Cooper, Michael M. Wagner 0001 |
ICML | 1 |
| 2002 | Data, network, and application: technical description of the Utah RODS Winter Olympic Biosurveillance System
Fu-Chiang Tsui, Jeremy U. Espino, Michael M. Wagner 0001, Per H. Gesteland, Oleg Ivanov, Robert T. Olszewski, Xiaoming Zeng, Wendy W. Chapman, Weng-Keen Wong, Andrew W. Moore 0001 |
AMIA | 10 |
| 1999 | Distributed Value Functions
Jeff G. Schneider, Weng-Keen Wong, Andrew W. Moore 0001, Martin A. Riedmiller |
ICML | 2 |