VLDB 2026 Research / reviewers in the wild / expert
Markus Mock
dblp:03/5001
· DBLP profile ↗
14ranked-venue papers
4as first author
5since 2021 · last 2024
0000-0003-1553-4948ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Distributed Dataflow Across the Edge-Cloud ContinuumabstractInternet of Things (IoT) applications span the edge-cloud continuum to form multiscale distributed systems. The heterogeneity that defines this architecture, coupled with the asynchronous, event-triggered and failure-prone nature of these deployments create significant programming and maintenance challenges for developers of IoT applications. To address this impediment to innovation, we present Lam-1nar’a dataflow programming model for IoT applications implemented using a novel log-based and concurrent runtime system that spans all resource scales. We describe the properties that underpin Laminar'sdesign and compare it to a lower-level event-based approach. We show that Laminar'sdataflow model hides many of the complexities of “lock-free” event-driven programming. Through an empirical evaluation of Laminar,we find its design and implementation are both more straightforward for developers and more performant. Tyler Ekaireb, Lukas Brand, Nagarjun Avaraddy, Markus Mock, Chandra Krintz, Richard Wolski |
CLOUD | 4 |
| 2023 | Machine Learning for Predicting Photovoltaic Power Generation: an Application on a University CampusabstractThis paper describes the application of machine learning in an energy management system set up to monitor the use and generation of energy at a university campus. We explain how we built precise forecasting models for producing photovoltaic energy produced by solar panels at the University of Applied Sciences Landshut, Germany. We describe the practical challenges of dealing with data dropout and erroneous data when working with data from an actual in-production setting outside a simple lab scenario. Applying several data cleaning methods based on statistical methods, we obtained models that allow us to predict electrical energy produced by the solar panels with a precision of 0.97 as measured by the R2metric. More importantly, the process described to arrive at this result, more than the result per se, serves as a real-world example of applying data science and machine learning methodology in an in-production setting. Fabio López, Markus Mock, Abraham Dávila |
CLEI | 2 |
| 2023 | Depot: Dependency-Eager Platform of TransformationsabstractThis paper presents a new model for a data management system specifically designed to enable community-curated data repositories and collaboration. Depot (a Dependence-Eager Platform of Transformations) is based on a data-lake approach that eases the technological burdens associated with data contribution while providing an interactive programming environment for developing transformations that result in structured tables supporting SQL database operations. Crucially, Depot implements lazy evaluation of these transformations so that only the structured data that is demanded by a data consumer is generated. Until the structured data is "materialized," Depot tracks and maintains the dependencies that are required to perform the eventual materialization. This lazy approach to creating structured data allows Depot to maintain a smaller resource footprint compared to a typical data warehouse approach while maintaining the flexibility of the data lake model. Furthermore, Depot is designed as a community-sustainable platform. The initial prototype is implemented for cloud deployment and it distributes the storage and ETL workload cost among the data consumers. Performance results of the early prototype are encouraging, making Depot a new infrastructure for creating data lakes that foster contributed-consumer collaboration. Kerem Çelik, Samridhi Maheshwari, Shereen ElSayed, Markus Mock, Chandra Krintz, Richard Wolski |
CloudCom | 4 |
| 2023 | GreenCoin: A Renewable Energy-Aware CryptocurrencyabstractIn this paper, we propose GreenCoin – an energy-efficient cryptocurrency system with mining protocols designed to favor locations with relatively higher availability of renewable energy. Traditionally, crypto coin mining involves solving complex mathematical problems by high-end computing devices consuming an enormous amount of electricity, thus adversely affecting net carbon emissions. To reduce cost and emissions, GreenCoin uses a modified proof of stake (PoS) consensus algorithm, which itself is more energy efficient compared to other state-of-the-art methods. Our modified PoS algorithm, called Green PoS (GPoS), allows GreenCoin to favor nodes (with reward and privilege) located in regions with higher availability of renewable energy. We present a detailed system architecture of GreenCoin and explain the operating method of GPoS. We also provide results from empirical studies demonstrating the renewable energy-aware approach of GreenCoin. Shivaansh Kapoor, Chandra Krintz, Richard Wolski, Markus Mock |
IC2E | 5 |
| 2021 | An Evaluation of WebAssembly in Non-Web EnvironmentsabstractIn 2017, WebAssembly, a portable low-level byte-code, was released by the four major web browser makers to address the challenges presented by the maturation of the web and the rise of sophisticated and interactive applications such as 3D visualization, audio, and video streaming, and online games. JavaScript heretofore was the only built-in language of the web, and unfortunately, it is not well outfitted for the rich applications that have come to dominate the web today. Since its initial release, WebAssembly has made great strides on the web. With the WebAssembly System Interface release in March 2018, which allows WebAssembly to communicate with the operating system, it has become possible to run WebAssembly applications outside web browsers. This paper reviews the current state of WebAssembly and its system interface, describes the costs and benefits of these technologies for applications in different environments, and evaluates performance and portability. Our performance measurements demonstrate that WebAssembly is generally faster than JavaScript and, in some cases, can approach native code performance. Despite its limitations which make WebAssembly useless for specific applications domains, it nevertheless, has the potential to be beneficial in many environments and is likely to grow further even outside its original web environment. Benedikt Spies, Markus Mock |
CLEI | 2 |
| 2019 | Seneca: Fast and Low Cost Hyperparameter Search for Machine Learning ModelsabstractThe goal of our work is to simplify and expedite the construction and evaluation of machine learning models using autoscaled cloud computing resources. To enable this, we develop an open source system called Seneca, which leverages the serverless programming model and its implementation in Amazon Web Services (AWS) Lambda. Seneca takes a machine learning application, dataset, and a list of possible hyperparameter options as input and automatically constructs an AWS Lambda function. The function ingresses and splits the input dataset into training and testing subsets and constructs, tests, and evaluates (i.e. scores) a machine learning model for a given set of hyperparameter values. Seneca concurrently invokes functions for all combinations of the hyperparameters specified. It then returns the configuration (or model) that results in the best score to the user. In this paper, we overview the design and implementation of Seneca, and empirically evaluate its performance for a popular classification application. Chandra Krintz, Markus Mock, Richard Wolski |
CLOUD | 3 |
| 2005 | Program Slicing with Dynamic Points-To SetsabstractProgram slicing is a potentially useful analysis for aiding program understanding. However, in reality even slices of small programs are often too large to be useful. Imprecise pointer analyses have been suggested as one cause of this problem. In this paper, we use dynamic points-to data, which represents optimistic pointer information, to obtain a bound on the best case slice size improvement that can be achieved with improved pointer precision. Our experiments show that slice size can be reduced significantly for programs that make frequent use of calls through function pointers because for them the dynamic pointer data results in a considerably smaller call graph, which leads to fewer data dependences. Programs without or with only few calls through function pointers, however, show considerably less improvement. We discovered that C programs appear to have a significant fraction of direct and nonspurious pointer data dependences so that reducing spurious dependences via pointers is only of limited benefit. Consequently, to make slicing useful in general for such programs, improvements beyond better pointer analyses are necessary. On the other hand, since we show that collecting dynamic function pointer information can be performed with little overhead (average slowdown of 10 percent for our benchmarks), dynamic pointer information may be a practical approach to making slicing of programs with frequent function pointer use more successful in practice. Markus Mock, Darren C. Atkinson, Craig Chambers, Susan J. Eggers |
IEEE Trans. Software Eng. | 1 |
| 2002 | Improving program slicing with dynamic points-to dataabstractProgram slicing is a potentially useful analysis for aiding program understanding. However, slices of even small programs are often too large to be generally useful. Imprecise pointer analyses have been suggested as one cause of this problem. In this paper, we use dynamic points-to data, which represents optimal or optimistic pointer information, to obtain a bound on the best case slice size improvement that can be achieved with improved pointer precision. Our experiments show that slice size can be reduced significantly for programs that make frequent use of calls through function pointers because for them the dynamic pointer data results in a considerably smaller call graph, which leads to fewer data dependences. Programs without or with only few calls through function pointers, however, show only insignificant improvement. We identified Amdahl's law as the reason for this behavior: C programs appear to have a large fraction of direct data dependences so that reducing spurious dependences via pointers is only of limited benefit. Consequently, to make slicing useful in general for such programs, improvements beyond better pointer analyses will be necessary. On the other hand, since we show that collecting dynamic function pointer information can be performed with little overhead (average slowdown of 10% for our benchmarks), dynamic pointer information may be a practical approach to making slicing of programs with frequent function pointer use more successful in reality. Markus Mock, Darren C. Atkinson, Craig Chambers, Susan J. Eggers |
SIGSOFT FSE | 1 |
| 2001 | Dynamic points-to sets: a comparison with static analyses and potential applications in program understanding and optimizationabstractIn this paper, we compare the behavior of pointers in C programs, as approximated by static pointer analysis algorithms, with the actual behavior of pointers when these programs are run. In order to perform this comparison, we have implemented several well known pointer analysis algorithms, and we have built an instrumentation infrastructure for tracking pointer values during program execution. Markus Mock, Manuvir Das, Craig Chambers, Susan J. Eggers |
PASTE | 1 |
| 2000 | Calpa: a tool for automating selective dynamic compilationabstractSelective dynamic compilation systems, typically driven by annotations that identify run-time constants, can achieve significant program speedups. However, manually inserting annotations is a tedious and time-consuming process that requires careful inspection of a program's static characteristics and run-time behavior and much trial and error in order to select the most beneficial annotations. Calpa is a system that generates annotations automatically for the DyC dynamic compiler. Calpa combines execution frequency and value profile information with a model of dynamic compilation cost and dynamically generated code benefit to choose run-time constants and other dynamic compilation strategies. For the programs tested so far, Calpa generates annotations of the same or better quality as those found by a human, but in a fraction of the time. The result was equal or-better program speedups from dynamic compilation, but without the need for programmer intervention. Markus Mock, Craig Chambers, Susan J. Eggers |
MICRO | 1 |
| 2000 | DyC: an expressive annotation-directed dynamic compiler for C
Brian Grant, Markus Mock, Matthai Philipose, Craig Chambers, Susan J. Eggers |
Theor. Comput. Sci. | 2 |
| 2000 | The benefits and costs of DyC's run-time optimizationsabstractDyC selectively dynamically compiles programs during their execution, utilizing the run-time-computed values of variables and data structures to apply optimizations that are based on partial evaluation. The dynamic optimizations are preplanned at static compile time in order to reduce their run-time cost; we call this staging . DyC's staged optimizations include (1) an advanced binding-time analysis that supports polyvariant specialization (enabling both single-way and multiway complete loop unrolling), polyvariant division, static loads, and static calls, (2) low-cost, dynamic versions of traditional global optimizations, such as zero and copy propagation and dead-assignment elimination, and (3) dynamic peephole optimizations, such as strength reduction. Because of this large suite of optimizations and its low dynamic compilation overhead, DyC achieves good performance improvements on programs that are larger and more complex than the kernels previously targeted by other dynamic compilation systems. This paper evaluates the benefits and costs of applying DyC's optimizations. We assess their impact on the performance of a variety of small to medium-sized programs, both for the regions of code that are actually transformed and for the entire application as a whole. Our study includes an analysis of the contribution to performance of individual optimizations, the performance effect of changing the applications' inputs, and a detailed accounting of dynamic compilation costs. Brian Grant, Markus Mock, Matthai Philipose, Craig Chambers, Susan J. Eggers |
ACM Trans. Program. Lang. Syst. | 2 |
| 1999 | An Evaluation of Staged Run-Time Optimizations in DyCabstractPrevious selective dynamic compilation systems have demonstrated that dynamic compilation can achieve performance improvements at low cost on small kernels, but they have had difficulty scaling to larger programs. To overcome this limitation, we developed DyC, a selective dynamic compilation system that includes more sophisticated and flexible analyses and transformations. DyC is able to achieve good performance improvements on programs that are much larger and more complex than the kernels. We analyze the individual optimizations of DyC and assess their impact on performance collectively and individually. Brian Grant, Matthai Philipose, Markus Mock, Craig Chambers, Susan J. Eggers |
PLDI | 3 |
| 1997 | Annotation-Directed Run-Time Specialization in CabstractWe present the design of a dynamic compilation system for C. Directed by a few declarative user annotations specifying where and on what dynamic compilation is to take place, a binding time analysis computes the set of run-time constants at each program point in each annotated procedure's control flow graph; the analysis supports program-point-specific polyvariant division and specialization. The analysis results guide the construction of a specialized run-time specializer for each dynamically compiled region; the specializer supports various caching strategies for managing dynamically generated code and supports mixes of speculative and demand-driven specialization of dynamic branch successors. Most of the key cost/benefit trade-offs in the binding time analysis and the run-time specialize are open to user control through declarative policy annotations. Our design is being implemented in the context of art existing optimizing compiler. Brian Grant, Markus Mock, Matthai Philipose, Craig Chambers, Susan J. Eggers |
PEPM | 2 |