Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Daniel Mills

dblp:86/5332 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 2 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Data stream processing · 55% Query processing and optimization · 45%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 100%
Human-computer interaction and pervasive computing
2 papers
Human-robot interaction · 83% Wearable and physiological sensing · 17%

Topics — the 8 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
incremental computation
0.712023
What's the Difference? Incremental Processing with Change Queries in Snowflake · Proc. ACM Manag. Data 2023
Query processing and optimization › view maintenance
incremental view maintenance
0.712023
What's the Difference? Incremental Processing with Change Queries in Snowflake · Proc. ACM Manag. Data 2023
Data stream processing
stream processing systems
0.512021
Watermarks in Stream Processing Systems: Semantics and Comparative Analysis of Apache Flink and Google Cloud Dataflow · Proc. VLDB Endow. 2021
Human-robot interaction
animal-computer interaction
0.212014
UbiComp for animal welfare: envisioning smart environments for kenneled dogs · UbiComp 2014
Data stream processing
fault tolerance
0.212013
MillWheel: Fault-Tolerant Stream Processing at Internet Scale · Proc. VLDB Endow. 2013
Data stream processing › stream processing systems
low-latency stream processing
0.212013
MillWheel: Fault-Tolerant Stream Processing at Internet Scale · Proc. VLDB Endow. 2013
Distributed systems › fault tolerance
exactly-once processing
0.212013
MillWheel: Fault-Tolerant Stream Processing at Internet Scale · Proc. VLDB Endow. 2013
Distributed systems
fault tolerance
0.212013
MillWheel: Fault-Tolerant Stream Processing at Internet Scale · Proc. VLDB Endow. 2013

Methods — techniques the papers use, named apart from their topics

stream objects · 0.7DML · 0.7cubic spline models · 0.5comparative analysis · 0.5persistent state management · 0.3directed computation graph · 0.3windowing semantics · 0.2dataflow model · 0.2ethnographic study · 0.2user study · 0.1
YearPublicationVenuePosition
2023 What's the Difference? Incremental Processing with Change Queries in Snowflake
abstract
Incremental algorithms are the heart and soul of stream processing. Low latency results depend on the ability to react to the subset of changes in a dataset over time rather than reprocessing the entirety of a dataset as it evolves. But while the SQL language is well suited for representing streams of changes (via tables) and their application to tables over time (via DML), it entirely lacks a method to query the changes to a table or view in the first place. In this paper, we present CHANGES queries and STREAM objects, Snowflake's primitives for querying and consuming incremental changes to table objects over time. CHANGES queries and STREAMs have been in use within Snowflake for three years, and see broad adoption across our customers. We describe the semantics of these primitives, discuss the implementation challenges, present an analysis of their usage at Snowflake, and contrast with other offerings.
Tyler Akidau, Paul Barbier, Istvan Cseri, Fabian Hueske, Tyler Jones, Sasha Lionheart, Daniel Mills, Dzmitry Pauliukevich, Lukas Probst, Niklas Semmler, Dan Sotolongo, Boyuan Zhang 0004
Proc. ACM Manag. Data7
2021 Watermarks in Stream Processing Systems: Semantics and Comparative Analysis of Apache Flink and Google Cloud Dataflow
abstract
Streaming data processing is an exercise in taming disorder: from oftentimes huge torrents of information, we hope to extract powerful and timely analyses. But when dealing with streaming data, the unbounded and temporally disordered nature of real-world streams introduces a critical challenge: how does one reason about the completeness of a stream that never ends? In this paper, we present a comprehensive definition and analysis of watermarks , a key tool for reasoning about temporal completeness in infinite streams. First, we describe what watermarks are and why they are important, highlighting how they address a suite of stream processing needs that are poorly served by eventually-consistent approaches: • Computing a single correct answer, as in notifications. • Reasoning about a lack of data, as in dip detection. • Performing non-incremental processing over temporal subsets of an infinite stream, as in statistical anomaly detection with cubic spline models. • Safely and punctually garbage collecting obsolete inputs and intermediate state. • Surfacing a reliable signal of overall pipeline health . Second, we describe, evaluate, and compare the semantically equivalent, but starkly different, watermark implementations in two modern stream processing engines: Apache Flink and Google Cloud Dataflow.
Edmon Begoli, Tyler Akidau, Slava Chernyak, Fabian Hueske, Kathryn Knight, Kenneth L. Knowles, Daniel Mills, Dan Sotolongo
Proc. VLDB Endow.7
2015 The Dataflow Model: A Practical Approach to Balancing Correctness, Latency, and Cost in Massive-Scale, Unbounded, Out-of-Order Data Processing
abstract
Unbounded, unordered, global-scale datasets are increasingly common in day-to-day business (e.g. Web logs, mobile usage statistics, and sensor networks). At the same time, consumers of these datasets have evolved sophisticated requirements, such as event-time ordering and windowing by features of the data themselves, in addition to an insatiable hunger for faster answers. Meanwhile, practicality dictates that one can never fully optimize along all dimensions of correctness, latency, and cost for these types of input. As a result, data processing practitioners are left with the quandary of how to reconcile the tensions between these seemingly competing propositions, often resulting in disparate implementations and systems. We propose that a fundamental shift of approach is necessary to deal with these evolved requirements in modern data processing. We as a field must stop trying to groom unbounded datasets into finite pools of information that eventually become complete, and instead live and breathe under the assumption that we will never know if or when we have seen all of our data, only that new data will arrive, old data may be retracted, and the only way to make this problem tractable is via principled abstractions that allow the practitioner the choice of appropriate tradeoffs along the axes of interest: correctness, latency, and cost. In this paper, we present one such approach, the Dataflow Model, along with a detailed examination of the semantics it enables, an overview of the core principles that guided its design, and a validation of the model itself via the real-world experiences that led to its development.
Tyler Akidau, Robert Bradshaw, Craig Chambers, Slava Chernyak, Rafael Fernández-Moctezuma, Reuven Lax, Sam McVeety, Daniel Mills, Frances Perry, Eric Schmidt 0001, Sam Whittle
Proc. VLDB Endow.8
2014 UbiComp for animal welfare: envisioning smart environments for kenneled dogs
abstract
Whilst the ubicomp community has successfully embraced a number of societal challenges for human benefit, including healthcare and sustainability, the well-being of other animals is hitherto underrepresented. We argue that ubicomp technologies, including sensing and monitoring devices as well as tangible and embodied interfaces, could make a valuable contribution to animal welfare. This paper particularly focuses on dogs in kenneled accommodation, as we investigate the opportunities and challenges for a smart kennel aiming to foster canine welfare. We conducted an in-depth ethnographic study of a dog rehoming center over four months; based on our findings, we propose a welfare-centered framework for designing smart environments, integrating monitoring and interaction with information management. We discuss the methodological issues we encountered during the research and propose a smart ethnographic approach for similar projects.
Clara Mancini, Janet van der Linden, Gerd Kortuem, Guy Dewsbury, Daniel Mills, Paula Boyden
UbiComp5
2013 MillWheel: Fault-Tolerant Stream Processing at Internet Scale
abstract
MillWheel is a framework for building low-latency data-processing applications that is widely used at Google. Users specify a directed computation graph and application code for individual nodes, and the system manages persistent state and the continuous flow of records, all within the envelope of the framework's fault-tolerance guarantees. This paper describes MillWheel's programming model as well as its implementation. The case study of a continuous anomaly detector in use at Google serves to motivate how many of MillWheel's features are used. MillWheel's programming model provides a notion of logical time, making it simple to write time-based aggregations. MillWheel was designed from the outset with fault tolerance and scalability in mind. In practice, we find that MillWheel's unique combination of scalability, fault tolerance, and a versatile programming model lends itself to a wide variety of problems at Google.
Tyler Akidau, Alex Balikov, Kaya Bekiroglu, Slava Chernyak, Josh Haberman, Reuven Lax, Sam McVeety, Daniel Mills, Paul Nordstrom, Sam Whittle
Proc. VLDB Endow.8
2008 Interaction with a zoomorphic robot that exhibits canid mechanisms of behaviour
abstract
Despite parallels between the cooperative use of domestic dogs in human society today, the predicted similar deployment of robots in the future, and the plethora of superficially dog-like robotic entertainment devices, very little effort has been directed at exploiting any understanding of social cognition between dogs and humans when designing interactive robotic systems. This paper describes an experiment in which we gave interactive robots zoomorphic appearances and dog-like behavioural properties. We analysed human reactions to robots exhibiting differing levels of zoomorphism and dog-like behaviour during an interaction task; we were particularly interested to determine whether behaviour and/or appearance that mimicked that of dogs facilitated increased satisfaction in robot performance and a willingness to persevere with a robot that made mistakes. Our findings show that neither the appearance or behaviour of a robot had an impact on the participants' rating of robot performance whilst there was also no significant difference in the self-reported categories of frustration, excitement and desire to persist with an interaction. However, our findings suggest that differences in individual preferences are revealed when people are asked to interact with robots that exhibit dog-like behaviours and other zoomorphic characteristics and that further research is required in order to better understand these differences.
Trevor D. Jones, Shaun W. Lawson, Daniel Mills
ICRA3