Martin Grund

dblp:90/5189 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
0since 2021 · last 2015
0009-0001-1655-0133ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 13 · 3 first-authorArtificial intelligence and machine learning · 4Applied, interdisciplinary, general and emerging computing · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Indexing and storage engines · 40% Database system architecture and tuning · 33% Query processing and optimization · 27%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Database system architecture and tuning
main-memory database
0.322013
CPU and cache efficient management of memory-resident databases · ICDE 2013
A Demonstration of HYRISE - A Main Memory Hybrid Storage Engine · Proc. VLDB Endow. 2011
Indexing and storage engines › storage management
hybrid storage engine
0.222011
A Demonstration of HYRISE - A Main Memory Hybrid Storage Engine · Proc. VLDB Endow. 2011
HYRISE - A Main Memory Hybrid Storage Engine · Proc. VLDB Endow. 2010
Query processing and optimization › query compilation
just-in-time compilation
0.212013
CPU and cache efficient management of memory-resident databases · ICDE 2013
Query processing and optimization
query compilation
0.212013
CPU and cache efficient management of memory-resident databases · ICDE 2013
Indexing and storage engines › column store
main-memory column store
0.112011
Fast Updates on Read-Optimized Databases Using Multi-Core CPUs · Proc. VLDB Endow. 2011
Indexing and storage engines
in-memory storage
0.112010
HYRISE - A Main Memory Hybrid Storage Engine · Proc. VLDB Endow. 2010
Query processing and optimization
cost model
0.012013
CPU and cache efficient management of memory-resident databases · ICDE 2013
Indexing and storage engines › storage management
storage layout optimization
0.012013
CPU and cache efficient management of memory-resident databases · ICDE 2013
Indexing and storage engines › storage model
hybrid row-column storage
0.012011
A Demonstration of HYRISE - A Main Memory Hybrid Storage Engine · Proc. VLDB Endow. 2011
Database system architecture and tuning
hybrid transactional and analytical processing
0.012011
Fast Updates on Read-Optimized Databases Using Multi-Core CPUs · Proc. VLDB Endow. 2011

Methods — techniques the papers use, named apart from their topics

cost modeling · 0.3just-in-time compilation · 0.2parallelization · 0.1architecture-aware optimization · 0.1cache miss modeling · 0.1
YearPublicationVenuePosition
2015 Impala: A Modern, Open-Source SQL Engine for Hadoop
Marcel Kornacker, Alexander Behm, Victor Bittorf, Taras Bobrovytsky, Casey Ching, Alan Choi, Justin Erickson, Martin Grund, Daniel Hecht, Matthew Jacobs, Ishaan Joshi, Lenni Kuff, Alex Leblang, Nong Li, Ippokratis Pandis, Henry Robinson, David Rorke, Silvius Rus, Dimitris Tsirogiannis, Skye Wanderman-Milne, Michael Yoder 0002
CIDR8
2014 TRISTAN: Real-time analytics on massive time series using sparse dictionary compression
abstract
Large-scale critical infrastructures such as transportation, energy, or water distribution networks are increasingly equipped with smart sensor technologies. Low-latency analytics on the resulting times series would open the door to many exciting opportunities to improve our grasp on complex urban systems. However, sensor-generated time series often turn out to be noisy, non-uniformly sampled, and misaligned in practice, making them ill-suited for traditional data processing. In this paper, we introduce TRISTAN (massive TRIckletS Time series ANalysis), a new data management system for efficient storage and real-time processing of fine-grained time series data. TRISTAN relies on a dedicated, compressed sparse representation of the time series using a dictionary. In contrast to previous approaches, TRISTAN is able to execute most analytics queries on the compressed data directly, and supports efficient and approximate query answering based on the most significant atoms of the dictionary only. We present the overall architecture of our system and discuss its performance on several smarter city datasets, showing that TRISTAN can achieve up to 20:1 compression ratios and 250x speedup compared to a state-of-the-art system.
Alice Marascu, Pascal Pompey, Eric Bouillet, Michael Wurst, Olivier Verscheure, Martin Grund, Philippe Cudré-Mauroux
IEEE BigData6
2014 Correct Me If I'm Wrong: Fixing Grammatical Errors by Preposition Ranking
abstract
The detection and correction of grammatical errors still represent very hard problems for modern error-correction systems. As an example, the top-performing systems at the preposition correction challenge CoNLL-2013 only achieved a F1 score of 17%. In this paper, we propose and extensively evaluate a series of approaches for correcting prepositions, analyzing a large body of high-quality textual content to capture language usage. Leveraging n-gram statistics, association measures, and machine learning techniques, our system is able to learn which words or phrases govern the usage of a specific preposition. Our approach makes heavy use of n-gram statistics generated from very large textual corpora. In particular, one of our key features is the use of n-gram association measures (e.g., Pointwise Mutual Information) between words and prepositions to generate better aggregated preposition rankings for the individual n-grams. We evaluate the effectiveness of our approach using cross-validation with different feature combinations and on two test collections created from a set of English language exams and StackExchange forums. We also compare against state-of-the-art supervised methods. Experimental results from the CoNLL-2013 test collection show that our approach to preposition correction achieves ∼30% in F1 score which results in 13% absolute improvement over the best performing approach at that challenge.
Roman Prokofyev, Ruslan Mavlyutov, Martin Grund, Gianluca Demartini, Philippe Cudré-Mauroux
CIKM3
2014 Concurrent Execution of Mixed Enterprise Workloads on In-Memory Databases
Johannes Wust, Martin Grund, Kai Hoewelmeyer, David Schwalb, Hasso Plattner
DASFAA (1)2
2013 Big data analytics on high Velocity streams: A case study
abstract
Big data management is often characterized by three Vs: Volume, Velocity and Variety. While traditional batch-oriented systems such as MapReduce are able to scale-out and process very large volumes of data in parallel, they also introduce some significant latency. In this paper, we focus on the second V (Velocity) of the Big Data triad; We present a case-study where we use a popular open-source stream processing engine (Storm) to perform real-time integration and trend detection on Twitter and Bitly streams. We describe our trend detection solution below and experimentally demonstrate that our architecture can effectively process data in real-time - even for high-velocity streams.
Thibaud Chardonnens, Philippe Cudré-Mauroux, Martin Grund, Benoit Perroud
IEEE BigData3
2013 MiSTRAL: An architecture for low-latency analytics on MasSive time series
abstract
Smart sensors are increasingly being used to manage and monitor critical urban infrastructures, e.g., for telecommunication, transport, water, or energy networks, as well as for healthcare or smart buildings. Sensor-based monitoring systems offer ways of continuously monitoring low frequency activities, and open the door to new analytic and predictive applications in Smarter Cities. Such sensors generate “tricklets”, i.e., noisy and continuous time series. Tricklets are typically misaligned, non-uniformly sampled, and comprise low frequency activities and recurring patterns. Storing and making sense of such data in a typical database management system is difficult, due to the impedance mismatch between classical (e.g., relational) data and tricklets. In this paper, we investigate the management of large amounts of tricklets from an architectural perspective, and propose MiSTRAL (MaSsive TRicklets anALysis), an architecture designed for executing low-latency analytics on time series warehouses. MiSTRAL uses a dictionary based representation for tricklets that allows queries to be run natively on compressed representations and thus to achieve the low-latency goal. The architecture of MiSTRAL is presented in detail in the following, along with early experimental results on several Smarter Cities datasets.
Alice Marascu, Pascal Pompey, Eric Bouillet, Olivier Verscheure, Michael Wurst, Martin Grund, Philippe Cudré-Mauroux
IEEE BigData6
2013 Elastic online analytical processing on RAMCloud
abstract
A shared-nothing architecture is state-of-the-art for deploying a distributed analytical in-memory database management system: it preserves the in-memory performance advantage by processing data locally on each node but is difficult to scale out. Modern switched fabric communication links such as InfiniBand narrow the performance gap between local and remote DRAM data access to a single order of magnitude. Based on these premises, we introduce a distributed in-memory database architecture that separates the query execution engine and data access: this enables a) the usage of a large-scale DRAM-based storage system such as Stanford's RAMCloud and b) the push-down of bandwidth-intensive database operators into the storage system. We address the resulting challenges such as finding the optimal operator execution strategy and partitioning scheme. We demonstrate that such an architecture delivers both: the elasticity of a shared-storage approach and the performance characteristics of operating on local DRAM.
Christian Tinnefeld, Donald Kossmann, Martin Grund, Joos-Hendrik Böse, Frank Renkes, Vishal Sikka, Hasso Plattner
EDBT3
2013 CPU and cache efficient management of memory-resident databases
abstract
Memory-Resident Database Management Systems (MRDBMS) have to be optimized for two resources: CPU cycles and memory bandwidth. To optimize for bandwidth in mixed OLTP/OLAP scenarios, the hybrid or Partially Decomposed Storage Model (PDSM) has been proposed. However, in current implementations, bandwidth savings achieved by partial decomposition come at increased CPU costs. To achieve the aspired bandwidth savings without sacrificing CPU efficiency, we combine partially decomposed storage with Just-in-Time (JiT) compilation of queries, thus eliminating CPU inefficient function calls. Since existing cost based optimization components are not designed for JiT-compiled query execution, we also develop a novel approach to cost modeling and subsequent storage layout optimization. Our evaluation shows that the JiT-based processor maintains the bandwidth savings of previously presented hybrid query processors but outperforms them by two orders of magnitude due to increased CPU efficiency.
Holger Pirk, Florian Funke 0001, Martin Grund, Thomas Neumann 0001, Ulf Leser, Stefan Manegold, Alfons Kemper, Martin L. Kersten
ICDE3
2011 A Demonstration of HYRISE - A Main Memory Hybrid Storage Engine
Martin Grund, Philippe Cudré-Mauroux, Samuel Madden 0001
Proc. VLDB Endow.1
2011 Fast Updates on Read-Optimized Databases Using Multi-Core CPUs
abstract
Read-optimized columnar databases use differential updates to handle writes by maintaining a separate write-optimized delta partition which is periodically merged with the read-optimized and compressed main partition. This merge process introduces significant overheads and unacceptable downtimes in update intensive systems, aspiring to combine transactional and analytical workloads into one system. In the first part of the paper, we report data analyses of 12 SAP Business Suite customer systems. In the second half, we present an optimized merge process reducing the merge overhead of current systems by a factor of 30. Our linear-time merge algorithm exploits the underlying high compute and bandwidth resources of modern multi-core CPUs with architecture-aware optimizations and efficient parallelization. This enables compressed in-memory column stores to handle the transactional update rate required by enterprise applications, while keeping properties of read-optimized databases for analytic-style queries.
Jens Krüger 0003, Changkyu Kim, Martin Grund, Nadathur Satish, David Schwalb, Jatin Chhugani, Hasso Plattner, Pradeep Dubey, Alexander Zeier
Proc. VLDB Endow.3
2010 The effects of virtualization on main memory systems
abstract
Virtualization is mainly employed for increasing the utilization of a lightly-loaded system by consolidation, but also to ease the administration based on the possibility to rapidly provision or migrate virtual machines. These facilities are crucial for efficiently managing large data centers. At the same time, modern hardware --- such as Intel's Nehalem microarchitecure --- change critical assumptions about performance bottlenecks and software systems explicitly exploiting the underlying hardware --- such as main memory databases --- gain increasing momentum.
Martin Grund, Jan Schaffner, Jens Krüger 0003, Jan Brunnert, Alexander Zeier
DaMoN1
2010 Optimizing Write Performance for Read Optimized Databases
Jens Krüger 0003, Martin Grund, Christian Tinnefeld, Hasso Plattner, Alexander Zeier, Franz Färber
DASFAA (2)2
2010 Enterprise Application-Specific Data Management
abstract
Enterprise applications are presently built on a 20-year old data management infrastructure that was designed to meet a specific set of requirements for OLTP systems. In the meantime, enterprise applications have become more sophisticated, data set sizes have increased, requirements on the freshness of input data have been strengthened, and the time allotted for completing business processes has been reduced. To meet these challenges, enterprise applications have become increasingly complicated to make up for short-comings in the data management infrastructure. To address this issue we investigate recent trends in data management such as main memory databases, column stores and compression techniques with regards to the workload requirements and data characteristics derived from actual customer systems. We show that a main memory column store is better suited for to days enterprise systems, which we validate by using SAP's Net Weaver Business Warehouse Accelerator and a realistic set of data from an inventory management application.
Jens Krüger 0003, Martin Grund, Alexander Zeier, Hasso Plattner
EDOC2
2010 HYRISE - A Main Memory Hybrid Storage Engine
abstract
In this paper, we describe a main memory hybrid database system called HYRISE, which automatically partitions tables into vertical partitions of varying widths depending on how the columns of the table are accessed. For columns accessed as a part of analytical queries (e.g., via sequential scans), narrow partitions perform better, because, when scanning a single column, cache locality is improved if the values of that column are stored contiguously. In contrast, for columns accessed as a part of OLTP-style queries, wider partitions perform better, because such transactions frequently insert, delete, update, or access many of the fields of a row, and co-locating those fields leads to better cache locality. Using a highly accurate model of cache misses, HYRISE is able to predict the performance of different partitionings, and to automatically select the best partitioning using an automated database design algorithm. We show that, on a realistic workload derived from customer applications, HYRISE can achieve a 20% to 400% performance improvement over pure all-column or all-row designs, and that it is both more scalable and produces better designs than previous vertical partitioning approaches for main memory systems.
Martin Grund, Jens Krüger 0003, Hasso Plattner, Alexander Zeier, Philippe Cudré-Mauroux, Samuel Madden 0001
Proc. VLDB Endow.1