EDBT 2026 Demo / reviewers in the wild / expert
Christian A. Lang
dblp:l/ChristianALang
· DBLP profile ↗
18ranked-venue papers
5as first author
0since 2021 · last 2013
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 18 · 5 first-authorArtificial intelligence and machine learning · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
10 papers |
Query processing and optimization · 37% Database system architecture and tuning · 22% Indexing and storage engines · 18% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational finance and economics · 100% |
Topics — the 20 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Indexing and storage engines
buffer management |
0.2 | 3 | 2010 | SSD Bufferpool Extensions for Database Systems · Proc. VLDB Endow. 2010 Increasing Buffer-Locality for Multiple Relational Table Scans through Grouping and Throttling · ICDE 2007 Joining Massive High-Dimensional Datasets · ICDE 2003 |
Storage systems
flash and SSD |
0.2 | 2 | 2010 | SSD Bufferpool Extensions for Database Systems · Proc. VLDB Endow. 2010 An Object Placement Advisor for DB2 Using Solid State Storage · Proc. VLDB Endow. 2009 |
Query processing and optimization
top-k query processing |
0.1 | 2 | 2006 | Boolean + ranking: querying a database by k-constrained optimization · SIGMOD Conference 2006 Making the Threshold Algorithm Access Cost Aware · IEEE Trans. Knowl. Data Eng. 2004 |
Distributed and cloud data management › data placement
object placement |
0.1 | 1 | 2009 | An Object Placement Advisor for DB2 Using Solid State Storage · Proc. VLDB Endow. 2009 |
Database system architecture and tuning › database design
physical database design |
0.1 | 1 | 2009 | An Object Placement Advisor for DB2 Using Solid State Storage · Proc. VLDB Endow. 2009 |
Database system architecture and tuning
database tuning |
0.1 | 1 | 2008 | QueryScope: visualizing queries for repeatable database tuning · Proc. VLDB Endow. 2008 |
Data stream processing
distributed stream processing |
0.1 | 1 | 2007 | WhiteWater: Distributed Processing of Fast Streams · IEEE Trans. Knowl. Data Eng. 2007 |
Query processing and optimization › query execution › scan processing
index scan |
0.1 | 1 | 2007 | Increasing Buffer-Locality for Multiple Index Based Scans through Intelligent Placement and Index Scan Speed Control · VLDB 2007 |
Query processing and optimization › query optimization › cost-based optimization
access cost optimization |
0.0 | 1 | 2004 | Making the Threshold Algorithm Access Cost Aware · IEEE Trans. Knowl. Data Eng. 2004 |
Query processing and optimization › top-k query processing
threshold algorithm |
0.0 | 1 | 2004 | Making the Threshold Algorithm Access Cost Aware · IEEE Trans. Knowl. Data Eng. 2004 |
Query processing and optimization
join processing |
0.0 | 1 | 2003 | Joining Massive High-Dimensional Datasets · ICDE 2003 |
Spatial and temporal data management › spatial query processing
spatial join |
0.0 | 1 | 2003 | Joining Massive High-Dimensional Datasets · ICDE 2003 |
Storage systems
storage hierarchy |
0.0 | 1 | 2010 | SSD Bufferpool Extensions for Database Systems · Proc. VLDB Endow. 2010 |
Query processing and optimization
cardinality estimation |
0.0 | 1 | 2001 | Modeling High-Dimensional Index Structures using Sampling · SIGMOD Conference 2001 |
Database system architecture and tuning
workload characterization |
0.0 | 1 | 2009 | An Object Placement Advisor for DB2 Using Solid State Storage · Proc. VLDB Endow. 2009 |
Visualization and visual analytics › visual analytics
query visualization |
0.0 | 1 | 2008 | QueryScope: visualizing queries for repeatable database tuning · Proc. VLDB Endow. 2008 |
Storage systems
buffer management |
0.0 | 1 | 2007 | Increasing Buffer-Locality for Multiple Index Based Scans through Intelligent Placement and Index Scan Speed Control · VLDB 2007 |
Information retrieval › retrieval models
ranked retrieval |
0.0 | 1 | 2004 | Making the Threshold Algorithm Access Cost Aware · IEEE Trans. Knowl. Data Eng. 2004 |
Information retrieval
retrieval models |
0.0 | 1 | 2004 | Making the Threshold Algorithm Access Cost Aware · IEEE Trans. Knowl. Data Eng. 2004 |
Indexing and storage engines › multidimensional indexing
high-dimensional indexing |
0.0 | 1 | 2001 | Modeling High-Dimensional Index Structures using Sampling · SIGMOD Conference 2001 |
Methods — techniques the papers use, named apart from their topics
extent-level temperature statistics · 0.2aging mechanism · 0.2greedy knapsack · 0.2dynamic programming · 0.2graph-based query representation · 0.2supervised learning · 0.1simulation · 0.1hill-climbing heuristic · 0.1dynamic grouping · 0.1adaptive throttling · 0.1NP-hardness analysis · 0.1a* search · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2013 | Query-aware compression of join resultsabstractClient-server database query processing has become an important paradigm in many data processing applications today. In cloud-based data services, for example, queries over structured data are sent to cloud-based servers for processing and the results relayed back to the client devices. Network bandwidth between client devices and cloud-based servers is often a limited resource and the use of data compression to reduce the amount of query result data transmitted would not only conserve bandwidth but also help with battery lifetime in the case of mobile client devices. For query result compression, current data compression methods do not exploit redundancy information that can be inferred from the query structure itself for greater compression. In this paper we propose a novel query-aware compression method for compressing query results sent from database servers to client applications. Our method is based on two key ideas. We exploit redundancy information obtained from the query plan and possibly from the database schema to achieve better compression than standard non-query aware compressors. We use a collection of memory-limited dictionaries to encode attribute values in a lightweight and efficient manner. Each dictionary in the collection of dictionaries are also dynamically resized to adapt to changing temporal access characteristics. We evaluated our method empirically using the TPC-H benchmark show that this technique is effective especially when used in conjunction with standard compressors. Our results show that compression ratios of up to twice that of gzip are possible. Christopher M. Mullins, Lipyeow Lim, Christian A. Lang |
EDBT | 3 |
| 2011 | Enhancing recovery using an SSD buffer pool extensionabstractRecent advances in solid state technology have led to the introduction of solid state drives (SSDs). Today's SSDs store data persistently using NAND flash memory and support good random IO performance. Current work in exploiting flash in database systems has primarily focused on using its random IO capability for second level bufferpools below main memory. There has not been much emphasis on exploiting its persistence. Bishwaranjan Bhattacharjee, Kenneth A. Ross, Christian A. Lang, George A. Mihaila, Mohammad Banikazemi |
DaMoN | 3 |
| 2011 | Column-oriented query processing for row storesabstractColumn-oriented DBMSs have gained increasing interest due to their superior performance for analytical workloads. Prior efforts tried to determine the possibility of simulating the query processing techniques of column-oriented systems in row-oriented databases, in a hope to improve their performance, especially for OLAP and data warehousing applications. In this paper, we show that column-oriented query processing can significantly improve the performance of row-oriented DBMSs. We introduce new operators that take into account the unique characteristics of data obtained from indexes, and exploit new technologies such as flash SSDs and multi-core processors to boost the performance. We demonstrate our approach with an experimental study using a prototype built on a commercial row-oriented DBMS. Amr El-Helw, Kenneth A. Ross, Bishwaranjan Bhattacharjee, Christian A. Lang, George A. Mihaila |
DOLAP | 4 |
| 2010 | SSD Bufferpool Extensions for Database SystemsabstractHigh-end solid state disks (SSDs) provide much faster access to data compared to conventional hard disk drives. We present a technique for using solid-state storage as a caching layer between RAM and hard disks in database management systems. By caching data that is accessed frequently, disk I/O is reduced. For random I/O, the potential performance gains are particularly significant. Our system continuously monitors the disk access patterns to identify hot regions of the disk. Temperature statistics are maintained at the granularity of an extent, i.e., 32 pages, and are kept current through an aging mechanism. Unlike prior caching methods, once the SSD is populated with pages from warm regions cold pages are not admitted into the cache, leading to low levels of cache pollution. Simulations based on DB2 I/O traces, and a prototype implementation within DB2 both show substantial performance improvements. Mustafa Canim, George A. Mihaila, Bishwaranjan Bhattacharjee, Kenneth A. Ross, Christian A. Lang |
Proc. VLDB Endow. | 5 |
| 2009 | An Object Placement Advisor for DB2 Using Solid State StorageabstractSolid state disks (SSDs) provide much faster random access to data compared to conventional hard disk drives. Therefore, the response time of a database engine could be improved by moving the objects that are frequently accessed in a random fashion to the SSD. Considering the price and limited storage capacity of solid state disks, the database administrator needs to determine which objects (tables, indexes, materialized views, etc.), if placed on the SSD, would most improve the performance of the system. In this paper we propose a tool called "Object Placement Advisor" for making a wise decision for the object placement problem. By collecting profile inputs from workload runs, the advisor utility provides a list of objects to be placed on the SSD by applying heuristics like the greedy knapsack technique or dynamic programming. To show that the proposed approach is effective in conventional database management systems, we have conducted experiments on IBM DB2 with queries and schemas based on the TPC-H and TPC-C benchmarks. The results indicate that using a relatively small amount of SSD storage, the response time of the system can be reduced significantly by considering the recommendation of the advisor. Mustafa Canim, Bishwaranjan Bhattacharjee, George A. Mihaila, Christian A. Lang, Kenneth A. Ross |
Proc. VLDB Endow. | 4 |
| 2008 | Anomaly-free incremental output in stream processingabstractContinuous queries enable alerts, predictions, and early warning in various domains such as health care, business process monitoring, financial applications, and environment protection. Currently, the consistency of the result cannot be assessed by the application, since only the query processor has enough internal information to determine when the output has reached a consistent state. To our knowledge, this is the first paper that addresses the problem of consistency under the assumptions and constraints of a continuous query model. In addition to defining an appropriate consistency notion, we propose techniques for guaranteeing consistency. We implemented the proposed techniques in our existing stream engine, and we report on the characteristics of the observed performance. As we show, these methods are practical as they impose only a small overhead on the system. George A. Mihaila, Ioana Stanoi, Christian A. Lang |
CIKM | 3 |
| 2008 | Using predictive analysis to improve invoice-to-cash collectionabstractIt is commonly agreed that accounts receivable (AR) can be a source of financial difficulty for firms when they are not efficiently managed and are underperforming. Experience across multiple industries shows that effective management of AR and overall financial performance of firms are positively correlated. In this paper we address the problem of reducing outstanding receivables through improvements in the collections strategy. Specifically, we demonstrate how supervised learning can be used to build models for predicting the payment outcomes of newly-created invoices, thus enabling customized collection actions tailored for each invoice or customer. Our models can predict with high accuracy if an invoice will be paid on time or not and can provide estimates of the magnitude of the delay. We illustrate our techniques in the context of real-world transaction data from multiple firms. Finally, simulation results show that our approach can reduce collection time up to a factor of four compared to a baseline that is not model-driven. Sai Zeng, Prem Melville, Christian A. Lang, Ioana M. Boier-Martin, Conrad Murphy |
KDD | 3 |
| 2008 | QueryScope: visualizing queries for repeatable database tuningabstractReading and perceiving complex SQL queries has been a time consuming task in traditional database applications for decades. When it comes to decision support systems with automatically generated and sometimes highly nested SQL queries, human understanding or tuning of these workloads becomes even more challenging. This demonstration explores visualization methods to represent queries as graphs. We developed the QueryScope tool to help visualize and understand critical elements of a query, thereby cutting down the learning curve. We show how the tool allows the user to drill down on particular queries or to find similarly structured queries that may exhibit similar tuning opportunities. The queries shown in the demonstration are taken from real tuning engagements. Kenneth A. Ross, Yuan-Chi Chang, Christian A. Lang |
Proc. VLDB Endow. | 4 |
| 2007 | Increasing Buffer-Locality for Multiple Relational Table Scans through Grouping and ThrottlingabstractDecision support (DSS) workloads generally contain multiple large concurrent scan operations. These are often executed as relational table scans which can take up a lot of I/O bandwidth. This is especially true for ad-hoc queries where the workload is not known in advance. Common database management systems have only limited ability to reuse memory buffer content across multiple running queries due to their treatment of queries in isolation. Previous attempts to coordinate scans for better buffer reuse were less than satisfactory due to drifting between scans and the required radical DBMS architecture changes. In this paper, we describe a new mechanism to keep similar table scans closer together during scanning. This is achieved via dynamic grouping and regrouping of scans based on their runtime behavior and via adaptive throttling of scan speeds based on scan group characteristics. The required memory footprint is very small and the effort required to extend existing database management systems is minimal, as shown in our DB2 UDB prototype. Our experiments show significant gains in end-to-end response times as well as average response times for TPC-H workloads. Christian A. Lang, Bishwaranjan Bhattacharjee, Timothy Malkemus, Sriram Padmanabhan, Kwai Wong |
ICDE | 1 |
| 2007 | Increasing Buffer-Locality for Multiple Index Based Scans through Intelligent Placement and Index Scan Speed Control
Christian A. Lang, Bishwaranjan Bhattacharjee, Timothy Malkemus, Kwai Wong |
VLDB | 1 |
| 2007 | WhiteWater: Distributed Processing of Fast StreamsabstractMonitoring systems today often involve continuous queries over streaming data in a distributed collaborative fashion. The distribution of query operators over a network of processors, as well as their processing sequence, form a query configuration with inherent constraints on the throughput that it can support. In this paper, we discuss the implications of measuring and optimizing for output throughput, as well as its limitations. We propose to use instead the more granular input throughput and a version of throughput measure, the profiled input throughput, that is focused on matching the expected behavior of the input streams. We show how we can evaluate a query configuration based on profiled input throughput and that the problem of finding the optimal configuration is NP-hard. Furthermore, we describe how we can overcome the complexity limitation by adapting hill-climbing heuristics to reduce the search space of configurations. We show experimentally that the approach used is not only efficient but also effective. Ioana Stanoi, George A. Mihaila, Themis Palpanas, Christian A. Lang |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2006 | Maximizing the sustained throughput of distributed continuous queriesabstractMonitoring systems today often involve continuous queries over streaming data, in a distributed collaborative system. The distribution of query operators over a network of processors, and their processing sequence, form a query configuration with inherent constraints on the throughput it can support. In this paper we propose to optimize stream queries with respect to a version of throughput measure, the profiled input throughput. This measure is focused on matching the expected behavior of the input streams. To prune the search space we used hill-climbing techniques that proved to be efficient and effective. Ioana Stanoi, George A. Mihaila, Themis Palpanas, Christian A. Lang |
CIKM | 4 |
| 2006 | Boolean + ranking: querying a database by k-constrained optimizationabstractThe wide spread of databases for managing structured data, compounded with the expanded reach of the Internet, has brought forward interesting data retrieval and analysis scenarios to RDBMS. In such settings, queries often take the form of k-constrained optimization, with a Boolean constraint and a numeric optimization expression as the goal function, retrieving only the top-k tuples. This paper proposes the concept of supporting such queries, as their nature implies, by a functional optimization machinery over the search space of multiple indices. To realize this concept, we combine the dual perspectives of discrete state search (from the view of indices) and continuous function optimization (from the view of goal functions). We present, as the marriage of the two perspectives, the OPT* framework, which encodes k-constrained optimization as an A* search over the composite space of multiple indices, driven by functional optimization for providing tight heuristics. By processing queries as optimization, OPT* significantly outperforms baseline approaches, with up to 3 orders of magnitude margins. Zhen Zhang 0001, Seung-won Hwang, Kevin Chen-Chuan Chang, Min Wang 0001, Christian A. Lang, Yuan-Chi Chang |
SIGMOD Conference | 5 |
| 2005 | Hint and Run: Accelerating XPath QueriesabstractXML documents are often represented as DOM structures or trees. In some instances, due to the complexity of queries, XPath queries are better evaluated by traversing these structures rather than using summarization indexes. A requirement for the optimization of document navigation, is to efficiently decrease the number of traversed nodes. The optimization task is more difficult when the query framework allows a syntax that enlarges the search space, such as wildcards and descendant queries. To reduce the overhead of query processing, many database systems supporting XML rely on indexes. Such indexes are typically not able to adapt to memory restrictions or workload changes. A secondary data structure that uses little storage space and tunes itself to address hot spots in processing, can therefore be beneficial. In this paper, we propose a first method for creating, using, and maintaining selective signatures called hints to aid in the navigation of XML documents. Hints form a flexible data structure for pruning the search space, that can be used on its own or it can complement existing indexes. The amount of hints used is variable, and it depends on the storage limitation set a priori, and on the efficiency of the hints. Our experiments show that hints can improve the efficiency of navigational XML query processing by a large margin while using only little extra memory. Ioana Stanoi, Christian A. Lang, Sriram Padmanabhan |
IDEAS | 2 |
| 2004 | Making the Threshold Algorithm Access Cost AwareabstractAssume a database storing N objects with d numerical attributes or feature values. All objects in the database can be assigned an overall score that is derived from their single feature values (and the feature values of a user-defined query). The problem considered here is then to efficiently retrieve the k objects with minimum (or maximum) overall score. The well-known threshold algorithm (TA) was proposed as a solution to this problem. TA views the database as a set of d sorted lists storing the feature values. Even though TA is optimal with regard to the number of accesses, its overall access cost can be high since, in practice, some list accesses may be more expensive than others. We therefore propose to make TA access cost aware by choosing the next list to access such that the overall cost is minimized. Our experimental results show that this overall cost is close to the optimal cost and significantly lower than the cost of prior approaches. Christian A. Lang, Yuan-Chi Chang, John R. Smith |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2003 | Joining Massive High-Dimensional DatasetsabstractWe consider the problem of joining massive datasets. We propose two techniques for minimizing disk I/O cost of join operations for both spatial and sequence data. Our techniques optimize the available buffer space using a global view of the datasets. We build a boolean matrix on the pages of the given datasets using a lower bounding distance predictor. The marked entries of this matrix represent candidate page pairs to be joined. Our first technique joins the marked pages iteratively. Our second technique clusters the marked entries using rectangular dense regions that have minimal perimeter and fit into buffer. These clusters are then ordered so that the total number of common pages between consecutive clusters is maximal. The clusters are then read from disk and joined. Our experimental results on various real datasets show that our techniques are 2 to 86 times faster than the competing techniques for spatial datasets, and 13 to 133 times faster than the competing techniques for sequence datasets. Tamer Kahveci, Christian A. Lang, Ambuj K. Singh |
ICDE | 2 |
| 2002 | Accelerating High-Dimensional Nearest Neighbor QueriesabstractThe performance of nearest neighbor (NN) queries degrades noticeably with increasing dimensionality of the data due to reduced selectivity of high-dimensional data and an increased number of seek operations during NN-query execution. If the NN-radii were known in advance, the disk accesses could be reordered such that seek operations are minimized. We therefore propose a new way of estimating the NN-radius based on the fractal dimensionality and sampling. It is applicable to any page-based index structure. We show that the estimation error is considerably lower than for previous approaches. In the second part of the paper, we present two applications of this technique. We show how the radius estimations can be used to transform k-NN queries into at most two range queries, and how it can be used to reduce the number of page reads during all-NN queries. In both cases, we observe significant speedups over traditional techniques for synthetic and real-world data. Christian A. Lang, Ambuj K. Singh |
SSDBM | 1 |
| 2001 | Modeling High-Dimensional Index Structures using SamplingabstractA large number of index structures for high-dimensional data have been proposed previously. In order to tune and compare such index structures, it is vital to have efficient cost prediction techniques for these structures. Previous techniques either assume uniformity of the data or are not applicable to high-dimensional data. We propose the use of sampling to predict the number of accessed index pages during a query execution. Sampling is independent of the dimensionality and preserves clusters which is important for representing skewed data. We present a general model for estimating the index page layout using sampling and show how to compensate for errors. We then give an implementation of our model under restricted memory assumptions and show that it performs well even under these constraints. Errors are minimal and the overall prediction time is up to two orders of magnitude below the time for building and probing the full index without sampling. 1. Christian A. Lang, Ambuj K. Singh |
SIGMOD Conference | 1 |