VLDB 2026 Research / reviewers in the wild / expert
Tyson Condie
dblp:88/2774
· DBLP profile ↗
35ranked-venue papers
9as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 15 · 4 first-authorSystems, architecture and hardware · 6Software engineering, systems software and programming languages · 6 · 2 first-authorComputer networks · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 3Security and privacy · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
14 papers |
Distributed systems · 41% Cloud and datacenter computing · 35% High-performance computing · 14% | |
| Software engineering, system software, and programming languages
7 papers |
Debugging and program repair · 80% Operating systems · 11% Programming languages and type systems · 6% | |
| Databases, data mining, and information retrieval
10 papers |
Query processing and optimization · 24% Distributed and cloud data management · 16% Machine learning and data management · 13% | |
| Human-computer interaction and pervasive computing
1 paper |
Health and well-being technologies · 100% |
Topics — the 30 heaviest of 51, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Debugging and program repair › software debugging
interactive debugging |
0.8 | 3 | 2017 | Debugging Big Data Analytics in Spark with BigDebug · SIGMOD Conference 2017 BigDebug: interactive debugger for big data analytics in Apache Spark · SIGSOFT FSE 2016 BigDebug: debugging primitives for interactive big data processing in spark · ICSE 2016 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.5 | 3 | 2017 | Apache REEF: Retainable Evaluator Execution Framework · ACM Trans. Comput. Syst. 2017 REEF: Retainable Evaluator Execution Framework · Proc. VLDB Endow. 2013 Titian: Data Provenance Support in Spark · Proc. VLDB Endow. 2015 |
Debugging and program repair
fault localization |
0.5 | 2 | 2016 | BigDebug: debugging primitives for interactive big data processing in spark · ICSE 2016 Titian: Data Provenance Support in Spark · Proc. VLDB Endow. 2015 |
Distributed systems
data provenance |
0.3 | 1 | 2018 | Adding data provenance support to Apache Spark · VLDB J. 2018 |
Parallel and multicore computing › data-parallel programming
mapreduce |
0.3 | 3 | 2011 | Online aggregation and continuous query support in MapReduce · SIGMOD Conference 2010 MapReduce Online · NSDI 2010 Online Aggregation for Large MapReduce Jobs · Proc. VLDB Endow. 2011 |
Data models and query languages
datalog |
0.2 | 1 | 2016 | Big Data Analytics with Datalog Queries on Spark · SIGMOD Conference 2016 |
Query processing and optimization › recursive query
recursive query evaluation |
0.2 | 1 | 2016 | Big Data Analytics with Datalog Queries on Spark · SIGMOD Conference 2016 |
Debugging and program repair › concurrent program debugging
distributed debugging |
0.2 | 1 | 2016 | BigDebug: debugging primitives for interactive big data processing in spark · ICSE 2016 |
Operating systems
provenance |
0.2 | 1 | 2016 | BigDebug: debugging primitives for interactive big data processing in spark · ICSE 2016 |
High-performance computing › data-intensive computing
large-scale data analytics |
0.2 | 1 | 2016 | Big Data Analytics with Datalog Queries on Spark · SIGMOD Conference 2016 |
Data integration and cleaning
data provenance |
0.2 | 1 | 2015 | Titian: Data Provenance Support in Spark · Proc. VLDB Endow. 2015 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.2 | 1 | 2015 | REEF: Retainable Evaluator Execution Framework · SIGMOD Conference 2015 |
Graph data management
distributed graph processing |
0.2 | 1 | 2014 | Pregelix: Big(ger) Graph Analytics on a Dataflow Engine · Proc. VLDB Endow. 2014 |
Distributed systems
fault tolerance |
0.2 | 3 | 2017 | Apache REEF: Retainable Evaluator Execution Framework · ACM Trans. Comput. Syst. 2017 REEF: Retainable Evaluator Execution Framework · SIGMOD Conference 2015 Online aggregation and continuous query support in MapReduce · SIGMOD Conference 2010 |
Machine learning › Efficient and distributed learning
large-scale learning |
0.2 | 1 | 2013 | Machine learning for big data · SIGMOD Conference 2013 |
Distributed and cloud data management
declarative networking |
0.1 | 2 | 2008 | Evita raced: metacompilation for declarative networks · Proc. VLDB Endow. 2008 Declarative networking: language, execution and optimization · SIGMOD Conference 2006 |
Query processing and optimization › aggregation
online aggregation |
0.1 | 1 | 2011 | Online Aggregation for Large MapReduce Jobs · Proc. VLDB Endow. 2011 |
Programming languages and type systems › programming paradigms
declarative programming |
0.1 | 1 | 2010 | Boom analytics: exploring data-centric, declarative programming for the cloud · EuroSys 2010 |
Distributed systems › stream processing
continuous query |
0.1 | 1 | 2010 | Online aggregation and continuous query support in MapReduce · SIGMOD Conference 2010 |
Distributed systems
distributed programming |
0.1 | 1 | 2010 | Boom analytics: exploring data-centric, declarative programming for the cloud · EuroSys 2010 |
Cloud and datacenter computing
cluster computing framework |
0.1 | 1 | 2018 | Adding data provenance support to Apache Spark · VLDB J. 2018 |
Health and well-being technologies › mobile health
mobile health sensing |
0.1 | 1 | 2017 | mCerebrum: A Mobile Sensing Software Platform for Development and Validation of Digital Biomarkers and Interventions · SenSys 2017 |
High-performance computing
data-intensive computing |
0.1 | 1 | 2017 | Debugging Big Data Analytics in Spark with BigDebug · SIGMOD Conference 2017 |
Compilers and program optimization › compiler construction
extensible compiler |
0.1 | 1 | 2008 | Evita raced: metacompilation for declarative networks · Proc. VLDB Endow. 2008 |
Data mining
big data analytics |
0.1 | 1 | 2016 | BigDebug: interactive debugger for big data analytics in Apache Spark · SIGSOFT FSE 2016 |
Query processing and optimization
recursive query |
0.1 | 1 | 2006 | Declarative networking: language, execution and optimization · SIGMOD Conference 2006 |
Internet architecture and protocols › naming and addressing
locator/identifier separation |
0.1 | 1 | 2006 | ROFL: routing on flat labels · SIGCOMM 2006 |
Routing and switching
routing |
0.1 | 1 | 2006 | ROFL: routing on flat labels · SIGCOMM 2006 |
Indexing and storage engines › multidimensional indexing
high-dimensional indexing |
0.1 | 1 | 2005 | LSH forest: self-tuning indexes for similarity search · WWW 2005 |
Information retrieval › hashing › hashing for nearest neighbor search
locality-sensitive hashing |
0.1 | 1 | 2005 | LSH forest: self-tuning indexes for similarity search · WWW 2005 |
Methods — techniques the papers use, named apart from their topics
log analysis · 0.6interactive debugging · 0.6watchpoints · 0.5simulated breakpoints · 0.5recursion optimization · 0.5provenance tracking · 0.5on-demand watchpoints · 0.5compilation · 0.5breakpoints · 0.5statistical machine learning · 0.5provenance capture · 0.3lineage tracking · 0.3reconfigurable scheduling · 0.3micro-batching · 0.3collaborative filtering · 0.3message passing · 0.2iterative dataflow · 0.2tera-scale learning · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | Data Deduplication with Edit ErrorsabstractIn this paper we tackle the problem of file deduplication for efficient data storage. We consider the case where the deduplication is performed on files that are modified by edit errors relative to the original version. We propose a novel block-level deduplication algorithm with variable-lengths in the case of non-binary alphabets. Compared to hash-based deduplication algorithms where file deduplication depends on the content of the hash keys or to brute force methods that compare files symbol-by- symbol, our algorithm significantly reduces the number of symbol comparisons and achieves high deduplication ratios. We present a theoretical analysis on the cost of the algorithm compared to naive methods and experimental results to evaluate the efficiency of our deduplication algorithm. Laura Conde-Canencia, Tyson Condie, Lara Dolecek |
GLOBECOM | 2 |
| 2018 | Scaling-up reasoning and advanced analytics on BigDataabstractAbstract BigDatalog is an extension of Datalog that achieves performance and scalability on both Apache Spark and multicore systems to the point that its graph analytics outperform those written in GraphX. Looking back, we see how this realizes the ambitious goal pursued by deductive database researchers beginning 40 years ago: this is the goal of combining the rigor and power of logic in expressing queries and reasoning with the performance and scalability by which relational databases managed BigData. This goal led to Datalog which is based on Horn Clauses like Prolog but employs implementation techniques, such as semi-naïve fixpoint and magic sets, that extend the bottom-up computation model of relational systems, and thus obtain the performance and scalability that relational systems had achieved, as far back as the 80s, using data-parallelization on shared-nothing architectures. But this goal proved difficult to achieve because of major issues at (i) the language level and (ii) at the system level. The paper describes how (i) was addressed by simple rules under which the fixpoint semantics extends to programs using count, sum and extrema in recursion, and (ii) was tamed by parallel compilation techniques that achieve scalability on multicore systems and Apache Spark. This paper is under consideration for acceptance in Theory and Practice of Logic Programming. Tyson Condie, Ariyam Das, Matteo Interlandi, Alexander Shkapsky, Mohan Yang, Carlo Zaniolo |
Theory Pract. Log. Program. | 1 |
| 2018 | Adding data provenance support to Apache Spark
Matteo Interlandi, Ari Ekmekji, Kshitij Shah, Muhammad Ali Gulzar, Sai Deep Tetali, Miryung Kim, Todd D. Millstein, Tyson Condie |
VLDB J. | 8 |
| 2017 | mCerebrum and Cerebral Cortex: A Real-time Collection, Analytic, and Intervention Platform for High-frequency Mobile Sensor Data
Timothy Hnat, Syed Monowar Hossain, Nasir Ali, Simona Carini, Tyson Condie, Ida Sim, Mani Srivastava 0001, Santosh Kumar 0001 |
AMIA | 5 |
| 2017 | Automated debugging in data-intensive scalable computingabstractDeveloping Big Data Analytics workloads often involves trial and error debugging, due to the unclean nature of datasets or wrong assumptions made about data. When errors (e.g., program crash, outlier results, etc.) arise, developers are often interested in identifying a subset of the input data that is able to reproduce the problem. BigSift is a new faulty data localization approach that combines insights from automated fault isolation in software engineering and data provenance in database systems to find a minimum set of failure-inducing inputs. BigSift redefines data provenance for the purpose of debugging using a test oracle function and implements several unique optimizations, specifically geared towards the iterative nature of automated debugging workloads. BigSift improves the accuracy of fault localizability by several orders-of-magnitude (∼103 to 107×) compared to Titian data provenance, and improves performance by up to 66× compared to Delta Debugging, an automated fault-isolation technique. For each faulty output, BigSift is able to localize fault-inducing data within 62% of the original job running time. Muhammad Ali Gulzar, Matteo Interlandi, Xueyuan Han, Tyson Condie, Miryung Kim |
SoCC | 5 |
| 2017 | mCerebrum: A Mobile Sensing Software Platform for Development and Validation of Digital Biomarkers and InterventionsabstractThe development and validation studies of new multisensory biomarkers and sensor-triggered interventions requires collecting raw sensor data with associated labels in the natural field environment. Unlike platforms for traditional mHealth apps, a software platform for such studies needs to not only support high-rate data ingestion, but also share raw high-rate sensor data with researchers, while supporting high-rate sense-analyze-act functionality in real-time. We present mCerebrum, a realization of such a platform, which supports high-rate data collections from multiple sensors with realtime assessment of data quality. A scalable storage architecture (with near optimal performance) ensures quick response despite rapidly growing data volume. Micro-batching and efficient sharing of data among multiple source and sink apps allows reuse of computations to enable real-time computation of multiple biomarkers without saturating the CPU or memory. Finally, it has a reconfigurable scheduler which manages all prompts to participants that is burden- and context-aware. With a modular design currently spanning 23+ apps, mCerebrum provides a comprehensive ecosystem of system services and utility apps. The design of mCerebrum has evolved during its concurrent use in scientific field studies at ten sites spanning 106,806 person days. Evaluations show that compared with other platforms, mCerebrum's architecture and design choices support 1.5 times higher data rates and 4.3 times higher storage throughput, while causing 8.4 times lower CPU usage. Syed Monowar Hossain, Timothy Hnat, Nazir Saleheen, Nusrat Jahan Nasrin, Joseph Noor, Bo-Jhang Ho, Tyson Condie, Mani Srivastava 0001, Santosh Kumar 0001 |
SenSys | 7 |
| 2017 | Debugging Big Data Analytics in Spark with BigDebugabstractTo process massive quantities of data, developers leverage Data-Intensive Scalable Computing (DISC) systems such as Apache Spark. In terms of debugging, DISC systems support only post-mortem log analysis and do not provide any debugging functionality. This demonstration paper showcases BigDebug: a tool enhancing Apache Spark with a set of interactive debugging features that can help users in debug their Big Data Applications. Muhammad Ali Gulzar, Matteo Interlandi, Tyson Condie, Miryung Kim |
SIGMOD Conference | 3 |
| 2017 | Apache REEF: Retainable Evaluator Execution FrameworkabstractResource Managers like YARN and Mesos have emerged as a critical layer in the cloud computing system stack, but the developer abstractions for leasing cluster resources and instantiating application logic are very low level. This flexibility comes at a high cost in terms of developer effort, as each application must repeatedly tackle the same challenges (e.g., fault tolerance, task scheduling and coordination) and reimplement common mechanisms (e.g., caching, bulk-data transfers). This article presents REEF, a development framework that provides a control plane for scheduling and coordinating task-level (data-plane) work on cluster resources obtained from a Resource Manager. REEF provides mechanisms that facilitate resource reuse for data caching and state management abstractions that greatly ease the development of elastic data processing pipelines on cloud platforms that support a Resource Manager service. We illustrate the power of REEF by showing applications built atop: a distributed shell application, a machine-learning framework, a distributed in-memory caching system, and a port of the CORFU system. REEF is currently an Apache top-level project that has attracted contributors from several institutions and it is being used to develop several commercial offerings such as the Azure Stream Analytics service. Byung-Gon Chun, Tyson Condie, Yingda Chen, Carlo Curino, Chris Douglas, Matteo Interlandi, Beomyeol Jeon, Joo Seong Jeong, Gyewon Lee, Yunseong Lee, Tony Majestro, Dahlia Malkhi, Sergiy Matusevych, Brandon Myers, Mariia Mykhailova, Shravan M. Narayanamurthy, Joseph Noor, Raghu Ramakrishnan 0001, Sriram Rao, Russell Sears, Beysim Sezgin, Taegeon Um, Julia Wang, Markus Weimer, Youngseok Yang |
ACM Trans. Comput. Syst. | 2 |
| 2017 | Fixpoint semantics and optimization of recursive Datalog programs with aggregatesabstractAbstract A very desirable Datalog extension investigated by many researchers in the last 30 years consists in allowing the use of the basic SQL aggregates min, max, count and sum in recursive rules. In this paper, we propose a simple comprehensive solution that extends the declarative least-fixpoint semantics of Horn Clauses, along with the optimization techniques used in the bottom-up implementation approach adopted by many Datalog systems. We start by identifying a large class of programs of great practical interest in which the use of min or max in recursive rules does not compromise the declarative fixpoint semantics of the programs using those rules. Then, we revisit the monotonic versions of count and sum aggregates proposed by Mazuran et al. (2013b, The VLDB Journal 22, 4, 471–493) and named, respectively, mcount and msum. Since mcount, and also msum on positive numbers, are monotonic in the lattice of set-containment, they preserve the fixpoint semantics of Horn Clauses. However, in many applications of practical interest, their use can lead to inefficiencies, that can be eliminated by combining them with max, whereby mcount and msum become the standard count and sum. Therefore, the semantics and optimization techniques of Datalog are extended to recursive programs with min, max, count and sum, making possible the advanced applications of superior performance and scalability demonstrated by BigDatalog (Shkapsky et al. 2016. In SIGMOD. ACM, 1135–1149) and Datalog-MC (Yang et al. 2017. The VLDB Journal 26, 2, 229–248). Carlo Zaniolo, Mohan Yang, Ariyam Das, Alexander Shkapsky, Tyson Condie, Matteo Interlandi |
Theory Pract. Log. Program. | 5 |
| 2016 | Programming and Runtime Support to Blaze FPGA Accelerator Deployment at Datacenter ScaleabstractWith the end of CPU core scaling due to dark silicon limitations, customized accelerators on FPGAs have gained increased attention in modern datacenters due to their lower power, high performance and energy efficiency. Evidenced by Microsoft's FPGA deployment in its Bing search engine and Intel's 16.7 billion acquisition of Altera, integrating FPGAs into datacenters is considered one of the most promising approaches to sustain future datacenter growth. However, it is quite challenging for existing big data computing systems-like Apache Spark and Hadoop-to access the performance and energy benefits of FPGA accelerators. In this paper we design and implement Blaze to provide programming and runtime support for enabling easy and efficient deployments of FPGA accelerators in datacenters. In particular, Blaze abstracts FPGA accelerators as a service (FaaS) and provides a set of clean programming APIs for big data processing applications to easily utilize those accelerators. Our Blaze runtime implements an FaaS framework to efficiently share FPGA accelerators among multiple heterogeneous threads on a single node, and extends Hadoop YARN with accelerator-centric scheduling to efficiently share them among multiple computing tasks in the cluster. Experimental results using four representative big data applications demonstrate that Blaze greatly reduces the programming efforts to access FPGA accelerators in systems like Apache Spark and YARN, and improves the system throughput by 1.7 × to 3× (and energy efficiency by 1.5× to 2.7×) compared to a conventional CPU-only cluster. Muhuan Huang, Di Wu 0010, Cody Hao Yu, Zhenman Fang, Matteo Interlandi, Tyson Condie, Jason Cong |
SoCC | 6 |
| 2016 | Optimizing Interactive Development of Data-Intensive ApplicationsabstractModern Data-Intensive Scalable Computing (DISC) systems are designed to process data through batch jobs that execute programs (e.g., queries) compiled from a high-level language. These programs are often developed interactively by posing ad-hoc queries over the base data until a desired result is generated. We observe that there can be significant overlap in the structure of these queries used to derive the final program. Yet, each successive execution of a slightly modified query is performed anew, which can significantly increase the development cycle. Vega is an Apache Spark framework that we have implemented for optimizing a series of similar Spark programs, likely originating from a development or exploratory data analysis session. Spark developers (e.g., data scientists) can leverage Vega to significantly reduce the amount of time it takes to re-execute a modified Spark program, reducing the overall time to market for their Big Data applications. Matteo Interlandi, Sai Deep Tetali, Muhammad Ali Gulzar, Joseph Noor, Tyson Condie, Miryung Kim, Todd D. Millstein |
SoCC | 5 |
| 2016 | BigDebug: debugging primitives for interactive big data processing in sparkabstractDevelopers use cloud computing platforms to process a large quantity of data in parallel when developing big data analytics. Debugging the massive parallel computations that run in today's data-centers is time consuming and error-prone. To address this challenge, we design a set of interactive, real-time debugging primitives for big data processing in Apache Spark, the next generation data-intensive scalable cloud computing platform. This requires re-thinking the notion of step-through debugging in a traditional debugger such as gdb, because pausing the entire computation across distributed worker nodes causes significant delay and naively inspecting millions of records using a watchpoint is too time consuming for an end user. First, BIGDEBUG's simulated breakpoints and on-demand watchpoints allow users to selectively examine distributed, intermediate data on the cloud with little overhead. Second, a user can also pinpoint a crash-inducing record and selectively resume relevant sub-computations after a quick fix. Third, a user can determine the root causes of errors (or delays) at the level of individual records through a fine-grained data provenance capability. Our evaluation shows that BIGDEBUG scales to terabytes and its record-level tracing incurs less than 25% overhead on average. It determines crash culprits orders of magnitude more accurately and provides up to 100% time saving compared to the baseline replay debugger. The results show that BIGDEBUG supports debugging at interactive speeds with minimal performance impact. Muhammad Ali Gulzar, Matteo Interlandi, Seunghyun Yoo, Sai Deep Tetali, Tyson Condie, Todd D. Millstein, Miryung Kim |
ICSE | 5 |
| 2016 | Big Data Analytics with Datalog Queries on SparkabstractThere is great interest in exploiting the opportunity provided by cloud computing platforms for large-scale analytics. Among these platforms, Apache Spark is growing in popularity for machine learning and graph analytics. Developing efficient complex analytics in Spark requires deep understanding of both the algorithm at hand and the Spark API or subsystem APIs (e.g., Spark SQL, GraphX). Our BigDatalog system addresses the problem by providing concise declarative specification of complex queries amenable to efficient evaluation. Towards this goal, we propose compilation and optimization techniques that tackle the important problem of efficiently supporting recursion in Spark. We perform an experimental comparison with other state-of-the-art large-scale Datalog systems and verify the efficacy of our techniques and effectiveness of Spark in supporting Datalog-based analytics. Alexander Shkapsky, Mohan Yang, Matteo Interlandi, Hsuan Chiu, Tyson Condie, Carlo Zaniolo |
SIGMOD Conference | 5 |
| 2016 | BigDebug: interactive debugger for big data analytics in Apache SparkabstractTo process massive quantities of data, developers leverage data-intensive scalable computing (DISC) systems in the cloud, such as Google's MapReduce, Apache Hadoop, and Apache Spark. In terms of debugging, DISC systems support post-mortem log analysis but do not provide interactive debugging features in realtime. This tool demonstration paper showcases a set of concrete usecases on how BigDebug can help debug Big Data Applications by providing interactive, realtime debug primitives. To emulate interactive step-wise debugging without reducing throughput, BigDebug provides simulated breakpoints to enable a user to inspect a program without actually pausing the entire computation. To minimize unnecessary communication and data transfer, BigDebug provides on-demand watchpoints that enable a user to retrieve intermediate data using a guard and transfer the selected data on demand. To support systematic and efficient trial-and-error debugging, BigDebug also enables users to change program logic in response to an error at runtime and replay the execution from that step. BigDebug is available for download at http://web.cs.ucla.edu/~miryung/software.html Muhammad Ali Gulzar, Matteo Interlandi, Tyson Condie, Miryung Kim |
SIGSOFT FSE | 3 |
| 2015 | REEF: Retainable Evaluator Execution FrameworkabstractResource Managers like Apache YARN have emerged as a critical layer in the cloud computing system stack, but the developer abstractions for leasing cluster resources and instantiating application logic are very low-level. This flexibility comes at a high cost in terms of developer effort, as each application must repeatedly tackle the same challenges (e.g., fault-tolerance, task scheduling and coordination) and re-implement common mechanisms (e.g., caching, bulk-data transfers). This paper presents REEF, a development framework that provides a control-plane for scheduling and coordinating task-level (data-plane) work on cluster resources obtained from a Resource Manager. REEF provides mechanisms that facilitate resource re-use for data caching, and state management abstractions that greatly ease the development of elastic data processing work-flows on cloud platforms that support a Resource Manager service. REEF is being used to develop several commercial offerings such as the Azure Stream Analytics service. Furthermore, we demonstrate REEF development of a distributed shell application, a machine learning algorithm, and a port of the CORFU [4] system. REEF is also currently an Apache Incubator project that has attracted contributors from several instititutions. Markus Weimer, Yingda Chen, Byung-Gon Chun, Tyson Condie, Carlo Curino, Chris Douglas, Yunseong Lee, Tony Majestro, Dahlia Malkhi, Sergiy Matusevych, Brandon Myers, Shravan M. Narayanamurthy, Raghu Ramakrishnan 0001, Sriram Rao, Russell Sears, Beysim Sezgin, Julia Wang |
SIGMOD Conference | 4 |
| 2015 | Center of excellence for mobile sensor data-to-knowledge (MD2K)abstractMobile sensor data-to-knowledge (MD2K) was chosen as one of 11 Big Data Centers of Excellence by the National Institutes of Health, as part of its Big Data-to-Knowledge initiative. MD2K is developing innovative tools to streamline the collection, integration, management, visualization, analysis, and interpretation of health data generated by mobile and wearable sensors. The goal of the big data solutions being developed by MD2K is to reliably quantify physical, biological, behavioral, social, and environmental factors that contribute to health and disease risk. The research conducted by MD2K is targeted at improving health through early detection of adverse health events and by facilitating prevention. MD2K will make its tools, software, and training materials widely available and will also organize workshops and seminars to encourage their use by researchers and clinicians. Santosh Kumar 0001, Gregory D. Abowd, William T. Abraham, Mustafa al'Absi, J. Gayle Beck, Polo Chau, Tyson Condie, David E. Conroy, Emre Ertin, Deborah Estrin, Deepak Ganesan, Cho Lam, Benjamin M. Marlin, Clay B. Marsh, Susan A. Murphy, Inbal Nahum-Shani, Kevin Patrick 0001, James M. Rehg, Moushumi Sharmin, Vivek Shetty, Ida Sim, Bonnie Spring, Mani Srivastava 0001, David W. Wetter |
J. Am. Medical Informatics Assoc. | 7 |
| 2015 | Titian: Data Provenance Support in SparkabstractDebugging data processing logic in Data-Intensive Scalable Computing (DISC) systems is a difficult and time consuming effort. Today's DISC systems offer very little tooling for debugging programs, and as a result programmers spend countless hours collecting evidence ( e.g. , from log files) and performing trial and error debugging. To aid this effort, we built Titian , a library that enables data provenance ---tracking data through transformations---in Apache Spark. Data scientists using the Titian Spark extension will be able to quickly identify the input data at the root cause of a potential bug or outlier result. Titian is built directly into the Spark platform and offers data provenance support at interactive speeds---orders-of-magnitude faster than alternative solutions---while minimally impacting Spark job performance; observed overheads for capturing data lineage rarely exceed 30% above the baseline job execution time. Matteo Interlandi, Kshitij Shah, Sai Deep Tetali, Muhammad Ali Gulzar, Seunghyun Yoo, Miryung Kim, Todd D. Millstein, Tyson Condie |
Proc. VLDB Endow. | 8 |
| 2014 | Pregelix: Big(ger) Graph Analytics on a Dataflow EngineabstractThere is a growing need for distributed graph processing systems that are capable of gracefully scaling to very large graph datasets. Unfortunately, this challenge has not been easily met due to the intense memory pressure imposed by process-centric, message passing designs that many graph processing systems follow. Pregelix is a new open source distributed graph processing system that is based on an iterative dataflow design that is better tuned to handle both in-memory and out-of-core workloads. As such, Pregelix offers improved performance characteristics and scaling properties over current open source systems (e.g., we have seen up to 15X speedup compared to Apache Giraph and up to 35X speedup compared to distributed GraphLab), and more effective use of available machine resources to support Big(ger) Graph Analytics. Yingyi Bu, Vinayak R. Borkar, Jianfeng Jia, Michael J. Carey 0001, Tyson Condie |
Proc. VLDB Endow. | 5 |
| 2013 | Machine learning on Big DataabstractStatistical Machine Learning has undergone a phase transition from a pure academic endeavor to being one of the main drivers of modern commerce and science. Even more so, recent results such as those on tera-scale learning [1] and on very large neural networks [2] suggest that scale is an important ingredient in quality modeling. This tutorial introduces current applications, techniques and systems with the aim of cross-fertilizing research between the database and machine learning communities. The tutorial covers current large scale applications of Machine Learning, their computational model and the workflow behind building those. Based on this foundation, we present the current state-of-the-art in systems support in the bulk of the tutorial. We also identify critical gaps in the state-of-the-art. This leads to the closing of the seminar, where we introduce two sets of open research questions: Better systems support for the already established use cases of Machine Learning and support for recent advances in Machine Learning research. Tyson Condie, Paul Mineiro, Neoklis Polyzotis, Markus Weimer |
ICDE | 1 |
| 2013 | Machine learning for big dataabstractStatistical Machine Learning has undergone a phase transition from a pure academic endeavor to being one of the main drivers of modern commerce and science. Even more so, recent results such as those on tera-scale learning [1] and on very large neural networks [2] suggest that scale is an important ingredient in quality modeling. This tutorial introduces current applications, techniques and systems with the aim of cross-fertilizing research between the database and machine learning communities. Tyson Condie, Paul Mineiro, Neoklis Polyzotis, Markus Weimer |
SIGMOD Conference | 1 |
| 2013 | REEF: Retainable Evaluator Execution FrameworkabstractIn this demo proposal, we describe REEF, a framework that makes it easy to implement scalable, fault-tolerant runtime environments for a range of computational models. We will demonstrate diverse workloads, including extract-transform-load MapReduce jobs, iterative machine learning algorithms, and ad-hoc declarative query processing. At its core, REEF builds atop YARN (Apache Hadoop 2's resource manager) to provide retainable hardware resources with lifetimes that are decoupled from those of computational tasks. This allows us to build persistent (cross-job) caches and cluster-wide services, but, more importantly, supports high-performance iterative graph processing and machine learning algorithms. Unlike existing systems, REEF aims for composability of jobs across computational models, providing significant performance and usability gains, even with legacy code. REEF includes a library of interoperable data management primitives optimized for communication and data movement (which are distinct from storage locality). The library also allows REEF applications to access external services, such as user-facing relational databases. We were careful to decouple lower levels of REEF from the data models and semantics of systems built atop it. The result was two new standalone systems: Tang, a configuration manager and dependency injector, and Wake, a state-of-the-art event-driven programming and data movement framework. Both are language independent, allowing REEF to bridge the JVM and .NET. Byung-Gon Chun, Tyson Condie, Carlo Curino, Raghu Ramakrishnan 0001, Russell Sears, Markus Weimer |
Proc. VLDB Endow. | 2 |
| 2011 | Online Aggregation for Large MapReduce Jobs
Niketan Pansare, Vinayak R. Borkar, Chris Jermaine, Tyson Condie |
Proc. VLDB Endow. | 4 |
| 2010 | Boom analytics: exploring data-centric, declarative programming for the cloudabstractBuilding and debugging distributed software remains extremely difficult. We conjecture that by adopting a data-centric approach to system design and by employing declarative programming languages, a broad range of distributed software can be recast naturally in a data-parallel programming model. Our hope is that this model can significantly raise the level of abstraction for programmers, improving code simplicity, speed of development, ease of software evolution, and program correctness. Peter Alvaro, Tyson Condie, Neil Conway, Khaled Elmeleegy, Joseph M. Hellerstein, Russell Sears |
EuroSys | 2 |
| 2010 | MapReduce Online
Tyson Condie, Neil Conway, Peter Alvaro, Joseph M. Hellerstein, Khaled Elmeleegy, Russell Sears |
NSDI | 1 |
| 2010 | Online aggregation and continuous query support in MapReduceabstractMapReduce is a popular framework for data-intensive distributed computing of batch jobs. To simplify fault tolerance, the output of each MapReduce task and job is materialized to disk before it is consumed. In this demonstration, we describe a modified MapReduce architecture that allows data to be pipelined between operators. This extends the MapReduce programming model beyond batch processing, and can reduce completion times and improve system utilization for batch jobs as well. We demonstrate a modified version of the Hadoop MapReduce framework that supports online aggregation, which allows users to see "early returns" from a job as it is being computed. Our Hadoop Online Prototype (HOP) also supports continuous queries, which enable MapReduce programs to be written for applications such as event monitoring and stream processing. HOP retains the fault tolerance properties of Hadoop, and can run unmodified user-defined MapReduce programs. Tyson Condie, Neil Conway, Peter Alvaro, Joseph M. Hellerstein, John Gerth, Justin Talbot, Khaled Elmeleegy, Russell Sears |
SIGMOD Conference | 1 |
| 2008 | Evita raced: metacompilation for declarative networksabstractDeclarative languages have recently been proposed for many new applications outside of traditional data management. Since these are relatively early research efforts, it is important that the architectures of these declarative systems be extensible, in order to accommodate unforeseen needs in these new domains. In this paper, we apply the lessons of declarative systems to the internals of a declarative engine. Specifically, we describe our design and implementation of Evita Raced , an extensible compiler for the OverLog language used in our declarative networking system, P2. Evita Raced is a metacompiler : an OverLog compiler written in OverLog. We describe the minimalist architecture of Evita Raced, including its extensibility interfaces and its reuse of P2's data model and runtime engine. We demonstrate that a declarative language like OverLog is well-suited to expressing traditional and novel query optimizations as well as other query manipulations, in a compact and natural fashion. Finally, we present initial results of Evita Raced extended with various optimization programs, running on both Internet overlay networks and wireless sensor networks. Tyson Condie, David Chu, Joseph M. Hellerstein, Petros Maniatis |
Proc. VLDB Endow. | 1 |
| 2007 | Public Health for the Internet (PHI)
Joseph M. Hellerstein, Tyson Condie, Minos N. Garofalakis, Boon Thau Loo, Petros Maniatis, Timothy Roscoe, Nina Taft |
CIDR | 2 |
| 2006 | Induced Churn as Shelter from Routing-Table Poisoning
Tyson Condie, Varun Kacholia, Sriram Sank, Joseph M. Hellerstein, Petros Maniatis |
NDSS | 1 |
| 2006 | ROFL: routing on flat labelsabstractIt is accepted wisdom that the current Internet architecture conflates network locations and host identities, but there is no agreement on how a future architecture should distinguish the two. One could sidestep this quandary by routing directly on host identities themselves, and eliminating the need for network-layer protocols to include any mention of network location. The key to achieving this is the ability to route on flat labels. In this paper we take an initial stab at this challenge, proposing and analyzing our ROFL routing algorithm. While its scaling and efficiency properties are far from ideal, our results suggest that the idea of routing on flat labels cannot be immediately dismissed. Matthew Caesar 0001, Tyson Condie, Jayanthkumar Kannan, Karthik Lakshminarayanan, Ion Stoica |
SIGCOMM | 2 |
| 2006 | Declarative networking: language, execution and optimizationabstractThe networking and distributed systems communities have recently explored a variety of new network architectures, both for application-level overlay networks, and as prototypes for a next-generation Internet architecture. In this context, we have investigated declarative networking: the use of a distributed recursive query engine as a powerful vehicle for accelerating innovation in network architectures [23, 24, 33]. Declarative networking represents a significant new application area for database research on recursive query processing. In this paper, we address fundamental database issues in this domain. First, we motivate and formally define the Network Datalog (NDlog) language for declarative network specifications. Second, we introduce and prove correct relaxed versions of the traditional semi-naïve query evaluation technique, to overcome fundamental problems of the traditional technique in an asynchronous distributed setting. Third, we consider the dynamics of network state, and formalize the iheventual consistencyl. of our programs even when bursts of updates can arrive in the midst of query execution. Fourth, we present a number of query optimization opportunities that arise in the declarative networking context, including applications of traditional techniques as well as new optimizations. Last, we present evaluation results of the above ideas implemented in our P2 declarative networking system, running on 100 machines over the Emulab network testbed. Boon Thau Loo, Tyson Condie, Minos N. Garofalakis, David E. Gay, Joseph M. Hellerstein, Petros Maniatis, Raghu Ramakrishnan 0001, Timothy Roscoe, Ion Stoica |
SIGMOD Conference | 2 |
| 2005 | Non-Cooperation in Competitive P2P NetworksabstractLarge-scale competitive P2P networks are threatened by the non-cooperation problem, where peers do not forward queries to potential competitors. Non-cooperation will be a growing problem in such applications as pay-per-transaction file-sharing, P2P auctions, and P2P service discovery networks, where peers are in competition with each other to provide services. Here, the authors showed how non-cooperation causes unacceptable degradation in quality of results, and present an economic protocol to address this problem. This protocol, called the RTR protocol, is based on the buying and selling of the right-to-respond (RTR) to each query in the network. Through simulations it is shown how the RTR protocol not only overcomes non-cooperation by providing proper incentives to peers, but also results in a network that is even more effective and efficient through intelligent, incentive-compatible routing of messages Beverly Yang, Tyson Condie, Sepandar D. Kamvar, Hector Garcia-Molina |
ICDCS | 2 |
| 2005 | A need for componentized transport protocolsabstractThere has been a steady stream of research over the years into componentized network protocols: protocol implementations assembled from a variety of building blocks. A promise of such frameworks has generally been flexibility: a protocol stack tailored for a particular application can be easily assembled, usually without writing any new code, by binding protocol objects together. Tyson Condie, Joseph M. Hellerstein, Petros Maniatis, Sean C. Rhea, Timothy Roscoe |
SOSP | 1 |
| 2005 | Implementing declarative overlaysabstractOverlay networks are used today in a variety of distributed systems ranging from file-sharing and storage systems to communication infrastructures. However, designing, building and adapting these overlays to the intended application and the target environment is a difficult and time consuming process.To ease the development and the deployment of such overlay networks we have implemented P2, a system that uses a declarative logic language to express overlay networks in a highly compact and reusable form. P2 can express a Narada-style mesh network in 16 rules, and the Chord structured overlay in only 47 rules. P2 directly parses and executes such specifications using a dataflow architecture to construct and maintain overlay networks. We describe the P2 approach, how our implementation works, and show by experiment its promising trade-off point between specification complexity and performance. Boon Thau Loo, Tyson Condie, Joseph M. Hellerstein, Petros Maniatis, Timothy Roscoe, Ion Stoica |
SOSP | 2 |
| 2005 | LSH forest: self-tuning indexes for similarity searchabstractWe consider the problem of indexing high-dimensional data for answering (approximate) similarity-search queries. Similarity indexes prove to be important in a wide variety of settings: Web search engines desire fast, parallel, main-memory-based indexes for similarity search on text data; database systems desire disk-based similarity indexes for high-dimensional data, including text and images; peer-to-peer systems desire distributed similarity indexes with low communication cost. We propose an indexing scheme called LSH Forest which is applicable in all the above contexts. Our index uses the well-known technique of locality-sensitive hashing (LSH), but improves upon previous designs by (a) eliminating the different data-dependent parameters for which LSH must be constantly hand-tuned, and (b) improving on LSH's performance guarantees for skewed data distributions while retaining the same storage and query overhead. We show how to construct this index in main memory, on disk, in parallel systems, and in peer-to-peer systems. We evaluate the design with experiments on multiple text corpora and demonstrate both the self-tuning nature and the superior performance of LSH Forest. Mayank Bawa, Tyson Condie, Prasanna Ganesan |
WWW | 2 |
| 2004 | Adaptive Peer-to-Peer TopologiesabstractWe present a peer-level protocol for forming adaptive, self-organizing topologies for data-sharing P2P networks. This protocol is based on the idea that a peer should directly connect to those peers from which it is most likely to download satisfactory content. We show that the resulting topologies are more efficient than standard Gnutella topologies. Furthermore, we show that these adaptive topologies have the added benefits of increased resistance to certain types of attacks, intrinsic rewards for active peers and punishments for malicious peers and free riders. Tyson Condie, Sepandar D. Kamvar, Hector Garcia-Molina |
Peer-to-Peer Computing | 1 |