Tomas Talius

dblp:46/9582 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
4 papers
Indexing and storage engines · 28% Query processing and optimization · 28% Data stream processing · 25%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 41% High-performance computing · 26% Storage systems · 17%

Topics — the 9 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › query execution
index-based query processing
0.512021
Hyperspace: The Indexing Subsystem of Azure Synapse · Proc. VLDB Endow. 2021
Indexing and storage engines
index management
0.512021
Hyperspace: The Indexing Subsystem of Azure Synapse · Proc. VLDB Endow. 2021
Distributed systems › distributed data processing
distributed indexing
0.412020
Helios: Hyperscale Indexing for the Cloud & Edge · Proc. VLDB Endow. 2020
High-performance computing › data-intensive computing
large-scale data processing
0.412020
Helios: Hyperscale Indexing for the Cloud & Edge · Proc. VLDB Endow. 2020
Storage systems
database recovery
0.112012
Transaction Log Based Application Error Recovery and Point In-Time Query · Proc. VLDB Endow. 2012
Storage systems › database recovery
point-in-time recovery
0.112012
Transaction Log Based Application Error Recovery and Point In-Time Query · Proc. VLDB Endow. 2012
Cloud and datacenter computing
database-as-a-service
0.112011
Adapting microsoft SQL server for cloud computing · ICDE 2011
Distributed systems › replication
primary-backup replication
0.112011
Adapting microsoft SQL server for cloud computing · ICDE 2011
Distributed systems
replication
0.112011
Adapting microsoft SQL server for cloud computing · ICDE 2011

Methods — techniques the papers use, named apart from their topics

distributed systems design · 0.9concurrency control · 0.5snapshot integration · 0.3log-based rewind · 0.3shared-nothing partitioning · 0.2
YearPublicationVenuePosition
2021 Hyperspace: The Indexing Subsystem of Azure Synapse
abstract
Microsoft recently introduced Azure Synapse Analytics, which offers an integrated experience across data ingestion, storage, and querying in Apache Spark and T-SQL over data in the lake, including files and warehouse tables. In this paper, we present our experiences with designing and implementing Hyperspace, the indexing subsystem underlying Synapse. Hyperspace enables users to build multiple types of secondary indexes on their data, maintain them through a multi-user concurrency model, and leverage them automatically---without any change to their application code---for query/workload acceleration. Many requirements of Hyperspace are based on feedback from several enterprise customers. We present the details of Hyperspace's underlying design, the user-facing APIs, its concurrency control protocol for index access, its index-aware query processing techniques, and its maintenance mechanisms for handling index updates. Evaluations over standard industry benchmarks and real customer workloads show that Hyperspace can accelerate query execution by up to 10x and in certain real-world workloads, even up to two orders of magnitude.
Rahul Potharaju, Terry Kim, Eunjin Song, Wentao Wu 0001, Lev Novik, Apoorve Dave, Pouria Pirzadeh, Andrew Fogarty, Gurleen Dhody, Jiying Li, Vidip Acharya, Sinduja Ramanujam, Nicolas Bruno, César A. Galindo-Legaria, Vivek R. Narasayya, Surajit Chaudhuri, Anil K. Nori, Tomas Talius, Raghu Ramakrishnan 0001
Proc. VLDB Endow.18
2020 Helios: Hyperscale Indexing for the Cloud & Edge
abstract
Helios is a distributed, highly-scalable system used at Microsoft for flexible ingestion, indexing, and aggregation of large streams of real-time data that is designed to plug into relational engines. The system collects close to a quadrillion events indexing approximately 16 trillion search keys per day from hundreds of thousands of machines across tens of data centers around the world. Helios use cases within Microsoft include debugging/diagnostics in both public and government clouds, workload characterization, cluster health monitoring, deriving business insights and performing impact analysis of incidents in other large-scale systems such as Azure Data Lake and Cosmos. Helios also serves as a reference blueprint for other large-scale systems within Microsoft. We present the simple data model behind Helios, which offers great flexibility and control over costs, and enables the system to asynchronously index massive streams of data. We also present our experiences in building and operating Helios over the last five years at Microsoft.
Rahul Potharaju, Terry Kim, Wentao Wu 0001, Vidip Acharya, Steve Suh, Andrew Fogarty, Apoorve Dave, Sinduja Ramanujam, Tomas Talius, Lev Novik, Raghu Ramakrishnan 0001
Proc. VLDB Endow.9
2012 Transaction Log Based Application Error Recovery and Point In-Time Query
abstract
Database backups have traditionally been used as the primary mechanism to recover from hardware and user errors. High availability solutions maintain redundant copies of data that can be used to recover from most failures except user or application errors. Database backups are neither space nor time efficient for recovering from user errors which typically occur in the recent past and affect a small portion of the database. Moreover periodic full backups impact user workload and increase storage costs. In this paper we present a scheme that can be used for both user and application error recovery starting from the current state and rewinding the database back in time using the transaction log. While we provide a consistent view of the entire database as of a point in time in the past, the actual prior versions are produced only for data that is accessed. We make the as of data accessible to arbitrary point in time queries by integrating with the database snapshot feature in Microsoft SQL Server.
Tomas Talius, Robin Dhamankar, Andrei Dumitrache, Hanuma Kodavalla
Proc. VLDB Endow.1
2011 Adapting microsoft SQL server for cloud computing
abstract
Cloud SQL Server is a relational database system designed to scale-out to cloud computing workloads. It uses Microsoft SQL Server as its core. To scale out, it uses a partitioned database on a shared-nothing system architecture. Transactions are constrained to execute on one partition, to avoid the need for two-phase commit. The database is replicated for high availability using a custom primary-copy replication scheme. It currently serves as the storage engine for Microsoft's Exchange Hosted Archive and SQL Azure.
Philip A. Bernstein, Istvan Cseri, Nishant Dani, Nigel Ellis, Ajay Kalhan, Gopal Kakivaya, David B. Lomet, Ramesh Manne, Lev Novik, Tomas Talius
ICDE10