EDBT 2026 Demo / reviewers in the wild / expert
Patrick Butler
dblp:55/6669
· DBLP profile ↗
19ranked-venue papers
2as first author
1since 2021 · last 2024
0000-0003-0468-6794ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 13 · 1 since 2021Artificial intelligence and machine learning · 9 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Security and privacy · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
6 papers |
Data mining · 92% Information retrieval · 4% Web and social media mining · 4% | |
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Computational social science and digital humanities · 48% Medical and health informatics · 34% Computational finance and economics · 18% | |
| Network and information security
2 papers |
Network security · 88% Digital forensics and information hiding · 12% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 70% Distributed systems · 30% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 100% | |
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% |
Topics — the 22 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Medical and health informatics › public health › public health informatics
disease surveillance |
0.2 | 1 | 2014 | Flu Gone Viral: Syndromic Surveillance of Flu on Twitter Using Temporal Topic Models · ICDM 2014 |
Data mining › text mining › topic modeling
dynamic topic model |
0.2 | 1 | 2014 | Flu Gone Viral: Syndromic Surveillance of Flu on Twitter Using Temporal Topic Models · ICDM 2014 |
Data mining › predictive modeling
event prediction |
0.2 | 1 | 2014 | 'Beating the news' with EMBERS: forecasting civil unrest using open source indicators · KDD 2014 |
Data mining
predictive modeling |
0.2 | 1 | 2014 | 'Beating the news' with EMBERS: forecasting civil unrest using open source indicators · KDD 2014 |
Data mining › text mining
topic model |
0.2 | 1 | 2014 | Flu Gone Viral: Syndromic Surveillance of Flu on Twitter Using Temporal Topic Models · ICDM 2014 |
Natural language and speech › Information extraction and text analysis
sentiment analysis |
0.2 | 1 | 2013 | Forex-foreteller: currency trend modeling using news articles · KDD 2013 |
Computational finance and economics › financial market prediction › stock prediction › stock movement prediction
news-based prediction |
0.2 | 1 | 2013 | Forex-foreteller: currency trend modeling using news articles · KDD 2013 |
Data mining
clustering |
0.2 | 1 | 2013 | Clustering with Complex Constraints - Algorithms and Applications · AAAI 2013 |
Data mining › clustering
constrained clustering |
0.2 | 1 | 2013 | Clustering with Complex Constraints - Algorithms and Applications · AAAI 2013 |
Network security › intrusion detection and prevention › intrusion detection › malicious traffic detection
botnet detection |
0.2 | 1 | 2013 | DNS for Massive-Scale Command and Control · IEEE Trans. Dependable Secur. Comput. 2013 |
Network security
traffic analysis |
0.2 | 1 | 2013 | DNS for Massive-Scale Command and Control · IEEE Trans. Dependable Secur. Comput. 2013 |
Data mining › structured data mining
graph mining |
0.1 | 1 | 2012 | Storytelling in entity networks to support intelligence analysts · KDD 2012 |
Medical and health informatics › electronic health records
electronic health record analysis |
0.1 | 1 | 2011 | Experiences with mining temporal event sequences from electronic medical records: initial successes and some challenges · KDD 2011 |
Data mining › predictive modeling
forecasting |
0.1 | 1 | 2016 | EMBERS at 4 years: Experiences operating an Open Source Indicators Forecasting System · KDD 2016 |
Storage systems
distributed storage |
0.1 | 1 | 2007 | PeerStripe: a p2p-based large-file storage for desktop grids · HPDC 2007 |
Storage systems › distributed storage
peer-to-peer storage |
0.1 | 1 | 2007 | PeerStripe: a p2p-based large-file storage for desktop grids · HPDC 2007 |
Distributed systems
peer-to-peer systems |
0.1 | 1 | 2007 | PeerStripe: a p2p-based large-file storage for desktop grids · HPDC 2007 |
Web and social media mining
social media analysis |
0.1 | 1 | 2014 | 'Beating the news' with EMBERS: forecasting civil unrest using open source indicators · KDD 2014 |
Information retrieval
text analysis |
0.1 | 1 | 2014 | 'Beating the news' with EMBERS: forecasting civil unrest using open source indicators · KDD 2014 |
Visualization and visual analytics › information visualization › quantitative data visualization
financial visualization |
0.0 | 1 | 2013 | Forex-foreteller: currency trend modeling using news articles · KDD 2013 |
Internet architecture and protocols
domain name system |
0.0 | 1 | 2013 | DNS for Massive-Scale Command and Control · IEEE Trans. Dependable Secur. Comput. 2013 |
Data mining › pattern mining
sequential pattern mining |
0.0 | 1 | 2011 | Experiences with mining temporal event sequences from electronic medical records: initial successes and some challenges · KDD 2011 |
Methods — techniques the papers use, named apart from their topics
topic clustering · 0.5sentiment analysis · 0.5linear regression · 0.5language model · 0.5optimization · 0.4concept lattice mining · 0.4topic modeling · 0.4suppression engine · 0.4state aggregation · 0.4data fusion · 0.4statistical analysis · 0.3spectral clustering · 0.3quadratic programming · 0.3network trace analysis · 0.3conjunctive normal form encoding · 0.3partial order mining · 0.2structured p2p routing · 0.1striping · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Forecasting Migration Patterns and Land Border EncountersabstractThis paper leverages open source “big data” intelligence to develop predictive models that can provide timely, relevant and accurate indications, warning, and tracking of migration flows / movements of large groups (> 100 persons) through South and Central America to the southwest border of the United States. We describe experiments with a live forecasting setup, development and refinement of predictive models, and how machine learning models can yield insight into the factors underlying mass migration. Raquib Bin Yousuf, Shengzhe Xu, Patrick Butler, Brian Mayer, Nathan Self, David Mares, Naren Ramakrishnan |
IEEE Big Data | 3 |
| 2020 | Detecting Media Self-Censorship without Explicit Training DataabstractThe motives and means of explicit state censorship have been well studied, both quantitatively and qualitatively. Self-censorship by media outlets, however, has not received nearly as much attention, mostly because it is difficult to systematically detect. We develop a novel approach to identify news media self-censorship by using social media as a sensor. We develop a hypothesis testing framework to identify and evaluate censored clusters of keywords and a near-linear-time algorithm (called GraphDPD) to identify the highest scoring clusters as indicators of censorship. We evaluate the accuracy of our framework, versus other state-of-the-art algorithms, using both semi-synthetic and real-world data from Mexico and Venezuela during Year 2014. These tests demonstrate the capacity of our framework to identify self-censorship, and provide an indicator of broader media freedom. The results of this study lay the foundation for detection, study, and policy-response to self-censorship. Rongrong Tao, Baojian Zhou, Feng Chen 0001, David Mares, Patrick Butler, Naren Ramakrishnan, Ryan Kennedy |
SDM | 5 |
| 2016 | EMBERS at 4 years: Experiences operating an Open Source Indicators Forecasting SystemabstractEMBERS is an anticipatory intelligence system forecasting population-level events in multiple countries of Latin America. A deployed system from 2012, EMBERS has been generating alerts 24x7 by ingesting a broad range of data sources including news, blogs, tweets, machine coded events,currency rates, and food prices. In this paper, we describe our experiences operating EMBERS continuously for nearly 4 years, with specific attention to the discoveries it has enabled, correct as well as missed forecasts, lessons learnt from participating in a forecasting tournament, and our perspectives on the limits of forecasting including ethical considerations. Sathappan Muthiah, Patrick Butler, Rupinder Paul Khandpur, Parang Saraf, Nathan Self, Alla Rozovskaya, Liang Zhao 0002, Jose Cadena, Chang-Tien Lu, Anil Vullikanti, Achla Marathe, Kristen Maria Summers, Graham Katz, Andy Doyle, Jaime Arredondo, Dipak Gupta, David Mares, Naren Ramakrishnan |
KDD | 2 |
| 2016 | Syndromic surveillance of Flu on Twitter using weakly supervised temporal topic models
Liangzhe Chen, K. S. M. Tozammel Hossain, Patrick Butler, Naren Ramakrishnan, B. Aditya Prakash |
Data Min. Knowl. Discov. | 3 |
| 2016 | A framework for intelligence analysis using spatio-temporal storytelling
Raimundo F. Dos Santos, Sumit Shah, Arnold P. Boedihardjo, Feng Chen 0001, Chang-Tien Lu, Patrick Butler, Naren Ramakrishnan |
GeoInformatica | 6 |
| 2014 | The EMBERS architecture for streaming predictive analyticsabstractDeveloped under the IARPA Open Source Initiative program, EMBERS (Early Model Based Event Recognition using Surrogates) is a large-scale Big-Data analytics system for forecasting significant societal events, such as civil unrest incidents and disease outbreaks on the basis of continuous, automated analysis of large volumes of publicly available data. It has been operational since November of 2012, delivering approximately 50 predictions each day. EMBERS is built on a streaming, scalable, share-nothing architecture and is deployed on Amazon Web Services (AWS). Andy Doyle, Graham Katz, Kristen Maria Summers, Chris Ackermann, Ilya Zavorin, Zunsik Lim, Sathappan Muthiah, Liang Zhao 0002, Chang-Tien Lu, Patrick Butler, Rupinder Paul Khandpur, Youssef Fayed, Naren Ramakrishnan |
IEEE BigData | 10 |
| 2014 | Flu Gone Viral: Syndromic Surveillance of Flu on Twitter Using Temporal Topic ModelsabstractSurveillance of epidemic outbreaks and spread from social media is an important tool for governments and public health authorities. Machine learning techniques for now casting the flu have made significant inroads into correlating social media trends to case counts and prevalence of epidemics in a population. There is a disconnect between data-driven methods for forecasting flu incidence and epidemiological models that adopt a state based understanding of transitions, that can lead to sub-optimal predictions. Furthermore, models for epidemiological activity and social activity like on Twitter predict different shapes and have important differences. We propose a temporal topic model to capture hidden states of a user from his tweets and aggregate states in a geographical region for better estimation of trends. We show that our approach helps fill the gap between phenomenological methods for disease surveillance and epidemiological models. We validate this approach by modeling the flu using Twitter in multiple countries of South America. We demonstrate that our model can consistently outperform plain vocabulary assessment in flu case-count predictions, and at the same time get better flu-peak predictions than competitors. We also show that our fine-grained modeling can reconcile some contrasting behaviors between epidemiological and social models. Liangzhe Chen, K. S. M. Tozammel Hossain, Patrick Butler, Naren Ramakrishnan, B. Aditya Prakash |
ICDM | 3 |
| 2014 | 'Beating the news' with EMBERS: forecasting civil unrest using open source indicatorsabstractWe describe the design, implementation, and evaluation of EMBERS, an automated, 24x7 continuous system for forecasting civil unrest across 10 countries of Latin America using open source indicators such as tweets, news sources, blogs, economic indicators, and other data sources. Unlike retrospective studies, EMBERS has been making forecasts into the future since Nov 2012 which have been (and continue to be) evaluated by an independent T&E team (MITRE). Of note, EMBERS has successfully forecast the June 2013 protests in Brazil and Feb 2014 violent protests in Venezuela. We outline the system architecture of EMBERS, individual models that leverage specific data sources, and a fusion and suppression engine that supports trading off specific evaluation criteria. EMBERS also provides an audit trail interface that enables the investigation of why specific predictions were made along with the data utilized for forecasting. Through numerous evaluations, we demonstrate the superiority of EMBERS over baserate methods and its capability to forecast significant societal happenings. Naren Ramakrishnan, Patrick Butler, Sathappan Muthiah, Nathan Self, Rupinder Paul Khandpur, Parang Saraf, Wei Wang 0064, Jose Cadena, Anil Vullikanti, Gizem Korkmaz, Chris J. Kuhlman, Achla Marathe, Liang Zhao 0002, Ting Hua, Feng Chen 0001, Chang-Tien Lu, Bert Huang, Aravind Srinivasan, Khoa Trinh, Lise Getoor, Graham Katz, Andy Doyle, Chris Ackermann, Ilya Zavorin, Jim Ford, Kristen Maria Summers, Youssef Fayed, Jaime Arredondo, Dipak Gupta, David Mares |
KDD | 2 |
| 2014 | Forecasting a Moving Target: Ensemble Models for ILI Case Count PredictionsabstractModern epidemiological forecasts of common illnesses, such as the flu, rely on both traditional surveillance sources as well as digital surveillance data. However, most published studies have been retrospective. Concurrently, the reports about flu activity generally lags by several weeks and even when published are revised for several weeks more. We posit that effectively handling this uncertainty is one of the key challenges for a real-time prediction system in this sphere. In this paper, we present a detailed prospective analysis on the generation of robust quantitative predictions about temporal trends of flu activity, using several surrogate data sources for 15 Latin American countries. We present our findings about the limitations and possible advantages of correcting the uncertainty associated with official flu estimates. We also compare the prediction accuracy between model-level fusion of different surrogate data sources against data-level fusion. Finally, we present a novel matrix factorization approach using neighborhood embedding to predict flu case counts. Comparing our proposed ensemble method against several baseline methods helps us demarcate the importance of different data sources for the countries under consideration. Prithwish Chakraborty, Pejman Khadivi, Bryan L. Lewis, Aravindan Mahendiran, Jiangzhuo Chen, Patrick Butler, Elaine O. Nsoesie, Sumiko R. Mekaru, John S. Brownstein, Madhav V. Marathe, Naren Ramakrishnan |
SDM | 6 |
| 2014 | Charging and Storage Infrastructure Design for Electric VehiclesabstractUshered by recent developments in various areas of science and technology, modern energy systems are going to be an inevitable part of our societies. Smart grids are one of these modern systems that have attracted many research activities in recent years. Before utilizing the next generation of smart grids, we should have a comprehensive understanding of the interdependent energy networks and processes. Next-generation energy systems networks cannot be effectively designed, analyzed, and controlled in isolation from the social, economic, sensing, and control contexts in which they operate. In this article, we present a novel framework to support charging and storage infrastructure design for electric vehicles. We develop coordinated clustering techniques to work with network models of urban environments to aid in placement of charging stations for an electrical vehicle deployment scenario. Furthermore, we evaluate the network before and after the deployment of charging stations, to recommend the installation of appropriate storage units to overcome the extra load imposed on the network by the charging stations. We demonstrate the multiple factors that can be simultaneously leveraged in our framework to achieve practical urban deployment. Our ultimate goal is to help realize sustainable energy system management in urban electrical infrastructure by modeling and analyzing networks of interactions between electric systems and urban populations. Marjan Momtazpour, Patrick Butler, Naren Ramakrishnan, Mahmud Shahriar Hossain, Mohammad Chehreghani Bozchalui, Ratnesh K. Sharma |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2013 | Clustering with Complex Constraints - Algorithms and ApplicationsabstractClustering with constraints is an important and developing area. However, most work is confined to conjunctions of simple together and apart constraints which limit their usability. In this paper, we propose a new formulation of constrained clustering that is able to incorporate not only existing types of constraints but also more complex logical combinations beyond conjunctions. We first show how any statement in conjunctive normal form (CNF) can be represented as a linear inequality. Since existing clustering formulations such as spectral clustering cannot easily incorporate these linear inequalities, we propose a quadratic programming (QP) clustering formulation to accommodate them. This new formulation allows us to have much more complex guidance in clustering. We demonstrate the effectiveness of our approach in two applications on text and personal information management. We also compare our algorithm against existing constrained spectral clustering algorithm to show its efficiency in computational time. Weifeng Zhi, Xiang Wang 0001, Buyue Qian, Patrick Butler, Naren Ramakrishnan, Ian Davidson |
AAAI | 4 |
| 2013 | Forex-foreteller: currency trend modeling using news articlesabstractFinancial markets are quite sensitive to unanticipated news and events. Identifying the effect of news on the market is a challenging task. In this demo, we present Forex-foreteller (FF) which mines news articles and makes forecasts about the movement of foreign currency markets. The system uses a combination of language models, topic clustering, and sentiment analysis to identify relevant news articles. These articles along with the historical stock index and currency exchange values are used in a linear regression model to make forecasts. The system has an interactive visualizer designed specifically for touch-sensitive devices which depicts forecasts along with the chronological news events and financial data used for making the forecasts. Fang Jin, Nathan Self, Parang Saraf, Patrick Butler, Wei Wang 0064, Naren Ramakrishnan |
KDD | 4 |
| 2013 | DNS for Massive-Scale Command and ControlabstractAttackers, in particular botnet controllers, use stealthy messaging systems to set up large-scale command and control. To systematically understand the potential capability of attackers, we investigate the feasibility of using domain name service (DNS) as a stealthy botnet command-and-control channel. We describe and quantitatively analyze several techniques that can be used to effectively hide malicious DNS activities at the network level. Our experimental evaluation makes use of a two-month-long 4.6-GB campus network data set and 1 million domain names obtained from alexa.com. We conclude that the DNS-based stealthy command-and-control channel (in particular, the codeword mode) can be very powerful for attackers, showing the need for further research by defenders in this direction. The statistical analysis of DNS payload as a countermeasure has practical limitations inhibiting its large-scale deployment. Kui Xu 0002, Patrick Butler, Sudip Saha, Danfeng Yao |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2012 | MANTICORE: Masking All Network Traffic via IP Concealment with OpenVPN Relaying to EC2abstractMalware and computer forensic researchers often communicate with malicious servers, either directly or indirectly, through the web browser or other ports utilized by malicious software. Communication with this form of adversary can sometimes necessitate the use of a proxy server in order to conceal the true origin of the researcher's traffic. Open source projects such as OpenVPN currently offer a structured method for establishing software based virtual private networks (VPNs) between arbitrary clients and servers. Likewise, paradigms exist which allow a user to proxy traffic from one end of a VPN to another, effectively masking the origin of traffic being sent to and from the client system. In this paper, we present MANTICORE - a system that combines ideas from VPN with the instancing functionality of a cloud computing system in order to dynamically mask and reassign the apparent IP address of a researcher's system. We also present experimental evaluation of our system on Amazon's Elastic Compute Cloud (EC2). Patrick Butler, Adam Rhodes, Ragib Hasan |
IEEE CLOUD | 1 |
| 2012 | Storytelling in entity networks to support intelligence analystsabstractIntelligence analysts grapple with many challenges, chief among them is the need for software support in storytelling, i.e., automatically 'connecting the dots' between disparate entities (e.g., people, organizations) in an effort to form hypotheses and suggest non-obvious relationships. We present a system to automatically construct stories in entity networks that can help form directed chains of relationships, with support for co-referencing, evidence marshaling, and imposing syntactic constraints on the story generation process. A novel optimization technique based on concept lattice mining enables us to rapidly construct stories on massive datasets. Using several public domain datasets, we illustrate how our approach overcomes many limitations of current systems and enables the analyst to efficiently narrow down to hypotheses of interest and reason about alternative explanations. Mahmud Shahriar Hossain, Patrick Butler, Arnold P. Boedihardjo, Naren Ramakrishnan |
KDD | 2 |
| 2011 | Quantitatively Analyzing Stealthy Communication Channels
Patrick Butler, Kui Xu 0002, Danfeng Yao |
ACNS | 1 |
| 2011 | Experiences with mining temporal event sequences from electronic medical records: initial successes and some challengesabstractThe standardization and wider use of electronic medical records (EMR) creates opportunities for better understanding patterns of illness and care within and across medical systems. Our interest is in the temporal history of event codes embedded in patients' records, specifically investigating frequently occurring sequences of event codes across patients. In studying data from more than 1.6 million patient histories at the University of Michigan Health system we quickly realized that frequent sequences, while providing one level of data reduction, still constitute a serious analytical challenge as many involve alternate serializations of the same sets of codes. To further analyze these sequences, we designed an approach where a partial order is mined from frequent sequences of codes. We demonstrate an EMR mining system called EMRView that enables exploration of the precedence relationships to quickly identify and visualize partial order information encoded in key classes of patients. We demonstrate some important nuggets learned through our approach and also outline key challenges for future research based on our experiences. Debprakash Patnaik, Patrick Butler, Naren Ramakrishnan, Laxmi Parida, Benjamin J. Keller, David A. Hanauer |
KDD | 2 |
| 2008 | On utilization of contributory storage in desktop gridsabstractModern desktop grid environments and shared computing platforms have popularized the use of contributory resources, such as desktop computers, as computing substrates for a variety of applications. However, addressing the exponentially growing storage demands of applications, especially in a contributory environment, remains a challenging research problem. In this paper, we propose a transparent distributed storage system that harnesses the storage contributed by desktop grid participants arranged in a peer-to-peer network to yield a scalable, robust, and self- organizing system. The novelty of our work lies in (i) design simplicity to facilitate actual use; (ii) support for easy integration with grid platforms; (Hi) innovative use of striping and error coding techniques to support very large data files; and (iv) the use of multicast techniques for data replication. Experimental results through large-scale simulations, verification on PlanetLab, and an actual implementation show that our system can provide reliable and efficient storage with support for large files for desktop grid applications. Chreston A. Miller, Ali Raza Butt, Patrick Butler |
IPDPS | 3 |
| 2007 | PeerStripe: a p2p-based large-file storage for desktop gridsabstractIn desktop grids the use of off-the-shelf shared components makes the use of dedicated resources economically nonviable and increases the complexity of design of efficient storage systems that are required to address the exponentially growing storage demands of modern applications that run on these platforms. To address this challenge, we present PeerStripe, a storage system that transparently distributes files to storage space contributed by participants that have joined a peer-to-peer (p2p) network. PeerStripe uses structured p2p routing to yield a scalable, robust, reliable, and self-organizing storage system. The novelty of PeerStripe lies in its ingenious use of striping and error coding techniques in a heterogeneous distributed environment to stor every large data files. Our evaluation of PeerStripe shows that it can achieve acceptable performance for applications in desktop grids. Chreston A. Miller, Patrick Butler, Ankur Shah, Ali Raza Butt |
HPDC | 2 |