Joseph A. Tucek

dblp:78/7202 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 3 first-authorSoftware engineering, systems software and programming languages · 8 · 2 first-authorDatabases, data management, data science and information retrieval · 4Artificial intelligence and machine learning · 1Security and privacy · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Storage systems · 60% Hardware reliability and fault tolerance · 26% High-performance computing · 12%
Databases, data mining, and information retrieval
3 papers
Transaction processing and concurrency control · 39% Information retrieval · 32% Database system architecture and tuning · 19%
Software engineering, system software, and programming languages
5 papers
Debugging and program repair · 47% Operating systems · 17% Concurrent programming · 14%

Topics — the 28 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › flash and SSD
solid-state drive
0.522017
Reliability Analysis of SSDs Under Power Fault · ACM Trans. Comput. Syst. 2017
Understanding the robustness of SSDS under power fault · FAST 2013
Hardware reliability and fault tolerance
fault injection
0.312017
Reliability Analysis of SSDs Under Power Fault · ACM Trans. Comput. Syst. 2017
Storage systems
storage reliability
0.322017
Torturing Databases for Fun and Profit · OSDI 2014
Reliability Analysis of SSDs Under Power Fault · ACM Trans. Comput. Syst. 2017
Database system architecture and tuning
pointer swizzling
0.212014
In-Memory Performance for Big Data · Proc. VLDB Endow. 2014
Transaction processing and concurrency control › concurrency control › locking
early lock release
0.212013
Controlled lock violation · SIGMOD Conference 2013
Transaction processing and concurrency control › concurrency control
locking
0.212013
Controlled lock violation · SIGMOD Conference 2013
Operating systems › fault tolerance
checkpoint and rollback
0.122007
Rx: Treating bugs as allergies - a safe method to survive software failures · ACM Trans. Comput. Syst. 2007
Rx: treating bugs as allergies - a safe method to survive software failures · SOSP 2005
Debugging and program repair
failure recovery
0.122007
Rx: Treating bugs as allergies - a safe method to survive software failures · ACM Trans. Comput. Syst. 2007
Rx: treating bugs as allergies - a safe method to survive software failures · SOSP 2005
High-performance computing
performance optimization at scale
0.112011
Applying idealized lower-bound runtime models to understand inefficiencies in data-intensive computing · SIGMETRICS 2011
Information retrieval › text matching
document matching
0.112009
Applying syntactic similarity algorithms for enterprise information management · KDD 2009
Information retrieval › similarity measure
document similarity
0.112009
Applying syntactic similarity algorithms for enterprise information management · KDD 2009
Data integration and cleaning
entity resolution
0.112009
Applying syntactic similarity algorithms for enterprise information management · KDD 2009
Information retrieval › similarity search
near-duplicate detection
0.112009
Applying syntactic similarity algorithms for enterprise information management · KDD 2009
Debugging and program repair › automated program repair
patch validation
0.112009
Efficient online validation with delta execution · ASPLOS 2009
Software testing
software validation
0.112009
Efficient online validation with delta execution · ASPLOS 2009
Debugging and program repair › fault localization
bug diagnosis
0.122007
Rx: Treating bugs as allergies - a safe method to survive software failures · ACM Trans. Comput. Syst. 2007
Rx: treating bugs as allergies - a safe method to survive software failures · SOSP 2005
Malware analysis › malware defense
worm defense
0.112007
Sweeper: a lightweight end-to-end system for defending against fast worms · EuroSys 2007
Debugging and program repair › fault localization
production-run failure diagnosis
0.112007
Triage: diagnosing production run failures at the user's site · SOSP 2007
Concurrent programming › concurrency bugs
atomicity violation
0.112006
AVIO: detecting atomicity violations via access interleaving invariants · ASPLOS 2006
Concurrent programming
concurrency bugs
0.112006
AVIO: detecting atomicity violations via access interleaving invariants · ASPLOS 2006
Program analysis
dynamic analysis
0.112006
AVIO: detecting atomicity violations via access interleaving invariants · ASPLOS 2006
Transaction processing and concurrency control › contention management
lock contention
0.012013
Controlled lock violation · SIGMOD Conference 2013
High-performance computing
data-intensive computing
0.012011
Applying idealized lower-bound runtime models to understand inefficiencies in data-intensive computing · SIGMETRICS 2011
Information retrieval
retrieval models
0.012009
Applying syntactic similarity algorithms for enterprise information management · KDD 2009
Software testing
regression testing
0.012009
Efficient online validation with delta execution · ASPLOS 2009
Network security › intrusion detection and prevention
intrusion detection
0.012007
Sweeper: a lightweight end-to-end system for defending against fast worms · EuroSys 2007
Debugging and program repair
bug reproduction
0.012007
Triage: diagnosing production run failures at the user's site · SOSP 2007
Memory systems
cache coherence
0.012006
AVIO: detecting atomicity violations via access interleaving invariants · ASPLOS 2006

Methods — techniques the papers use, named apart from their topics

experimental evaluation · 0.4workload stress testing · 0.3power fault injection · 0.3failure detection · 0.3environment modification · 0.1checkpointing · 0.1runtime monitoring · 0.1access interleaving invariants · 0.1shingling · 0.1fingerprint sampling · 0.1delta execution · 0.1content-based chunking · 0.1signature generation · 0.1rollback and reexecution · 0.1antibody dissemination · 0.1rollback · 0.1
YearPublicationVenuePosition
2017 Sparkle: optimizing spark for large memory machines and analytics
abstract
Given the growing availability of affordable scale-up servers, our goal is to bring the performance benefits of in-memory processing on scale-up servers to an increasingly common class of data analytics applications that process small to medium size datasets (up to a few 100GBs) that can easily fit in the memory of a typical scale-up server To achieve this goal, we leverage Spark, an existing memory-centric data analytics framework with wide-spread adoption among data scientists. Bringing Spark's data analytic capabilities to a scale-up system requires rethinking the original design assumptions, which, although effective for a scale-out system, are a poor match to a scale-up system resulting in unnecessary communication and memory inefficiencies.
Mijung Kim, Jun Li 0008, Haris Volos 0001, Manish Marwah, Alexander Ulanov, Kimberly Keeton, Joseph A. Tucek, Ludmila Cherkasova, Pradeep Fernando
SoCC7
2017 Reliability Analysis of SSDs Under Power Fault
abstract
Modern storage technology (solid-state disks (SSDs), NoSQL databases, commoditized RAID hardware, etc.) brings new reliability challenges to the already-complicated storage stack. Among other things, the behavior of these new components during power faults—which happen relatively frequently in data centers—is an important yet mostly ignored issue in this dependability-critical area. Understanding how new storage components behave under power fault is the first step towards designing new robust storage systems. In this article, we propose a new methodology to expose reliability issues in block devices under power faults. Our framework includes specially designed hardware to inject power faults directly to devices, workloads to stress storage components, and techniques to detect various types of failures. Applying our testing framework, we test 17 commodity SSDs from six different vendors using more than three thousand fault injection cycles in total. Our experimental results reveal that 14 of the 17 tested SSD devices exhibit surprising failure behaviors under power faults, including bit corruption, shorn writes, unserializable writes, metadata corruption, and total device failure.
Mai Zheng, Joseph A. Tucek, Mark Lillibridge, Bill W. Zhao, Elizabeth S. Yang
ACM Trans. Comput. Syst.2
2014 Torturing Databases for Fun and Profit
Mai Zheng, Joseph A. Tucek, Dachuan Huang, Mark Lillibridge, Elizabeth S. Yang, Bill W. Zhao, Shashank Singh 0003
OSDI2
2014 In-Memory Performance for Big Data
abstract
When a working set fits into memory, the overhead imposed by the buffer pool renders traditional databases non-competitive with in-memory designs that sacrifice the benefits of a buffer pool. However, despite the large memory available with modern hardware, data skew, shifting workloads, and complex mixed workloads make it difficult to guarantee that a working set will fit in memory. Hence, some recent work has focused on enabling in-memory databases to protect performance when the working data set almost fits in memory. Contrary to those prior efforts, we enable buffer pool designs to match in-memory performance while supporting the "big data" workloads that continue to require secondary storage, thus providing the best of both worlds. We introduce here a novel buffer pool design that adapts pointer swizzling for references between system objects (as opposed to application objects), and uses it to practically eliminate buffer pool overheads for memoryresident data. Our implementation and experimental evaluation demonstrate that we achieve graceful performance degradation when the working set grows to exceed the buffer pool size, and graceful improvement when the working set shrinks towards and below the memory and buffer pool sizes.
Goetz Graefe, Haris Volos 0001, Hideaki Kimura 0001, Harumi A. Kuno, Joseph A. Tucek, Mark Lillibridge, Alistair C. Veitch
Proc. VLDB Endow.5
2013 Understanding the robustness of SSDS under power fault
Mai Zheng, Joseph A. Tucek, Mark Lillibridge
FAST2
2013 Controlled lock violation
abstract
In databases with a large buffer pool, a transaction may run in less time than it takes to log the transaction's commit record on stable storage. Such cases motivate a technique called early lock release: immediately after appending its commit record to the log buffer in memory, a transaction may release its locks. Thus, it cuts overall lock duration to a fraction and reduces lock contention accordingly.
Goetz Graefe, Mark Lillibridge, Harumi A. Kuno, Joseph A. Tucek, Alistair C. Veitch
SIGMOD Conference4
2011 Disks Are Like Snowflakes: No Two Are Alike
Elie Krevat, Joseph A. Tucek, Gregory R. Ganger
HotOS2
2011 Applying idealized lower-bound runtime models to understand inefficiencies in data-intensive computing
abstract
No abstract available.
Elie Krevat, Tomer Shiran, Eric Anderson 0003, Joseph A. Tucek, Jay J. Wylie, Gregory R. Ganger
SIGMETRICS4
2010 Efficient eventual consistency in Pahoehoe, an erasure-coded key-blob archive
abstract
Cloud computing demands cheap, always-on, and reliable storage. We describe Pahoehoe, a key-value cloud storage system we designed to store large objects cost-effectively with high availability. Pahoehoe stores objects across multiple data centers and provides eventual consistency so to be available during network partitions. Pahoehoe uses erasure codes to store objects with high reliability at low cost. Its use of erasure codes distinguishes Pahoehoe from other cloud storage systems, and presents a challenge for efficiently providing eventual consistency. We describe Pahoehoe's put, get, and convergence protocols-convergence being the decentralized protocol that ensures eventual consistency. We use simulated executions of Pahoehoe to evaluate the efficiency of convergence, in terms of message count and message bytes sent, for failure-free and expected failure scenarios (e.g., partitions and server unavailability). We describe and evaluate optimizations to the naïve convergence protocol that reduce the cost of convergence in all scenarios.
Eric Anderson 0003, Xiaozhou Li 0001, Arif Merchant, Mehul A. Shah, Kevin Smathers, Joseph A. Tucek, Mustafa Uysal, Jay J. Wylie
DSN6
2009 Efficient online validation with delta execution
abstract
Software systems are constantly changing. Patches to fix bugs and patches to add features are all too common. Every change risks breaking a previously working system. Hence administrators loathe change, and are willing to delay even critical security patches until after fully validating their correctness. Compared to off-line validation, on-line validation has clear advantages since it tests against real life workloads. Yet unfortunately it imposes restrictive overheads as it requires running the old and new versions side-by-side. Moreover, due to spurious differences (e.g. event timing, random number generation, and thread interleavings), it is difficult to compare the two for validation.
Joseph A. Tucek, Weiwei Xiong, Yuanyuan Zhou 0001
ASPLOS1
2009 Applying syntactic similarity algorithms for enterprise information management
abstract
For implementing content management solutions and enabling new applications associated with data retention, regulatory compliance, and litigation issues, enterprises need to develop advanced analytics to uncover relationships among the documents, e.g., content similarity, provenance, and clustering. In this paper, we evaluate the performance of four syntactic similarity algorithms. Three algorithms are based on Broder's "shingling" technique while the fourth algorithm employs a more recent approach, "content-based chunking". For our experiments, we use a specially designed corpus of documents that includes a set of "similar" documents with a controlled number of modifications. Our performance study reveals that the similarity metric of all four algorithms is highly sensitive to settings of the algorithms' parameters: sliding window size and fingerprint sampling frequency. We identify a useful range of these parameters for achieving good practical results, and compare the performance of the four algorithms in a controlled environment. We validate our results by applying these algorithms to finding near-duplicates in two large collections of HP technical support documents.
Ludmila Cherkasova, Kave Eshghi, Charles B. Morrey III, Joseph A. Tucek, Alistair C. Veitch
KDD4
2009 Efficient tracing and performance analysis for large distributed systems
abstract
Distributed systems are notoriously difficult to implement and debug. One important tool for understanding the behavior of distributed systems is tracing. Unfortunately, effective tracing for modern distributed systems faces several challenges. First, many interesting behaviors in distributed systems only occur rarely, or at full production scale. Hence we need tracing mechanisms which impose minimal overhead, in order to allow always-on tracing of production instances. Second, for high-speed systems, messages can be delivered in significantly less time than the error of traditional time synchronization techniques such as network time protocol (NTP), necessitating time adjustment techniques with much higher precision. Third, distributed systems today may generate millions of events per second systemwide, resulting in traces consisting of billions of events. Such large traces can overwhelm existing trace analysis tools. These challenges make effective tracing difficult. We present techniques that address these three challenges. Our contributions include (1) a low-overhead tracing mechanism, which allows tracing of large systems without impacting their behavior or performance (0.14 ¿s/event), (2) a post hoc technique for producing highly accurate time synchronization across hosts (within 10/ts, compared to between 100 ¿s to 2 ms for NTP), and (3) incremental data processing techniques which facilitate analyzing traces containing billions of trace points on desktop systems. We have successfully applied these techniques to two distributed systems, a cooperative caching system and a distributed storage system, and from our experience, we believe our techniques are applicable to other distributed systems.
Eric Anderson 0003, Christopher Hoover, Xiaozhou Li 0001, Joseph A. Tucek
MASCOTS4
2007 Sweeper: a lightweight end-to-end system for defending against fast worms
abstract
The vulnerabilities that plague computers cause endless grief to users. Slammer compromised millions of hosts in minutes; a hit-list worm would take under a second. Recently proposed techniques respond better than manual approaches, but require expensive instrumentation, which limits deployment. Although spreading "antibodies" (e.g. signatures) ameliorates this limitation, hosts depending on antibodies are defenseless until inoculation; to the fastest hit-list worms this delay is crucial. Additionally, most recently proposed techniques cannot provide recovery to provide continuous service after an attack.
Joseph A. Tucek, James Newsome, Shan Lu 0001, Chengdu Huang, Spiros Xanthos, David Brumley, Yuanyuan Zhou 0001, Dawn Song
EuroSys1
2007 Triage: diagnosing production run failures at the user's site
abstract
Diagnosing production run failures is a challenging yet importanttask. Most previous work focuses on offsite diagnosis, i.e.development site diagnosis with the programmers present. This is insufficient for production-run failures as: (1) it is difficult to reproduce failures offsite for diagnosis; (2) offsite diagnosis cannot provide timely guidance for recovery or security purposes; (3)it is infeasible to provide a programmer to diagnose every production run failure; and (4) privacy concerns limit the release of information(e.g. coredumps) to programmers.
Joseph A. Tucek, Shan Lu 0001, Chengdu Huang, Spiros Xanthos, Yuanyuan Zhou 0001
SOSP1
2007 Rx: Treating bugs as allergies - a safe method to survive software failures
abstract
Many applications demand availability. Unfortunately, software failures greatly reduce system availability. Prior work on surviving software failures suffers from one or more of the following limitations: required application restructuring, inability to address deterministic software bugs, unsafe speculation on program execution, and long recovery time. This paper proposes an innovative safe technique, called Rx, which can quickly recover programs from many types of software bugs, both deterministic and nondeterministic. Our idea, inspired from allergy treatment in real life, is to rollback the program to a recent checkpoint upon a software failure, and then to reexecute the program in a modified environment. We base this idea on the observation that many bugs are correlated with the execution environment, and therefore can be avoided by removing the “allergen” from the environment. Rx requires few to no modifications to applications and provides programmers with additional feedback for bug diagnosis. We have implemented Rx on Linux. Our experiments with five server applications that contain seven bugs of various types show that Rx can survive six out of seven software failures and provide transparent fast recovery within 0.017--0.16 seconds, 21--53 times faster than the whole program restart approach for all but one case (CVS). In contrast, the two tested alternatives, a whole program restart approach and a simple rollback and reexecution without environmental changes, cannot successfully recover the four servers (Squid, Apache, CVS, and ypserv) that contain deterministic bugs, and have only a 40% recovery rate for the server (MySQL) that contains a nondeterministic concurrency bug. Additionally, Rx's checkpointing system is lightweight, imposing small time and space overheads.
Feng Qin 0003, Joseph A. Tucek, Yuanyuan Zhou 0001, Jagadeesan Sundaresan
ACM Trans. Comput. Syst.2
2006 AVIO: detecting atomicity violations via access interleaving invariants
abstract
Concurrency bugs are among the most difficult to test and diagnose of all software bugs. The multicore technology trend worsens this problem. Most previous concurrency bug detection work focuses on one bug subclass, data races, and neglects many other important ones such as atomicity violations, which will soon become increasingly important due to the emerging trend of transactional memory models.This paper proposes an innovative, comprehensive, invariantbased approach called AVIO to detect atomicity violations. Our idea is based on a novel observation called access interleaving invariant, which is a good indication of programmers' assumptions about the atomicity of certain code regions. By automatically extracting such invariants and detecting violations of these invariants at run time, AVIO can detect a variety of atomicity violations.Based on this idea, we have designed and built two implementations of AVIO and evaluated the trade-offs between them. The first implementation, AVIO-S, is purely in software, while the second, AVIO-H, requires some simple extensions to the cache coherence hardware. AVIO-S is cheaper and more accurate but incurs much higher overhead and thus more run-time perturbation than AVIOH. Therefore, AVIO-S is more suitable for in-house bug detection and postmortem bug diagnosis, while AVIO-H can be used for bug detection during production runs.We evaluate both implementations of AVIO using large realworld server applications (Apache and MySQL) with six representative real atomicity violation bugs, and SPLASH-2 benchmarks. Our results show that AVIO detects more tested atomicity violations of various types and has 25 times fewer false positives than previous solutions on average.
Shan Lu 0001, Joseph A. Tucek, Feng Qin 0003, Yuanyuan Zhou 0001
ASPLOS2
2005 Treating Bugs as Allergies: A Safe Method for Surviving Software Failures
Feng Qin 0003, Joseph A. Tucek, Yuanyuan Zhou 0001
HotOS2
2005 Trade-Offs in Protecting Storage: A Meta-Data Comparison of Cryptographic, Backup/Versioning, Immutable/Tamper-Proof, and Redundant Storage Solutions
abstract
Modern storage systems are responsible for increasing amounts of data and the value of the data itself is growing in importance. Several primary storage system solutions have emerged for the protection of data: (1) secure storage through cryptography, (2) backup and versioning systems, (3) immutable and tamper-proof storage, and (4) redundant storage. Using results from published studies, we compare these four solutions against different requirements highlighting trade-offs in performance, space, attack resistance, and cost. We also present a case study of applying these solutions based on design work at NCSA. Lastly, we conclude that while different storage protection solutions may be appropriate for different requirements, some general conclusions can be made about current state-of-the-art storage protection solutions as well as directions for future research.
Joseph A. Tucek, Paul Stanton, Elizabeth Haubert, Ragib Hasan, Larry Brumbaugh, William Yurcik
MSST1
2005 Rx: treating bugs as allergies - a safe method to survive software failures
abstract
Many applications demand availability. Unfortunately, software failures greatly reduce system availability. Prior work on surviving software failures suffers from one or more of the following limitations: Required application restructuring, inability to address deterministic software bugs, unsafe speculation on program execution, and long recovery time.This paper proposes an innovative safe technique, called Rx, which can quickly recover programs from many types of software bugs, both deterministic and non-deterministic. Our idea, inspired from allergy treatment in real life, is to rollback the program to a recent checkpoint upon a software failure, and then to re-execute the program in a modified environment. We base this idea on the observation that many bugs are correlated with the execution environment, and therefore can be avoided by removing the "allergen" from the environment. Rx requires few to no modifications to applications and provides programmers with additional feedback for bug diagnosis.We have implemented RX on Linux. Our experiments with four server applications that contain six bugs of various types show that RX can survive all the six software failures and provide transparent fast recovery within 0.017-0.16 seconds, 21-53 times faster than the whole program restart approach for all but one case (CVS). In contrast, the two tested alternatives, a whole program restart approach and a simple rollback and re-execution without environmental changes, cannot successfully recover the three servers (Squid, Apache, and CVS) that contain deterministic bugs, and have only a 40% recovery rate for the server (MySQL) that contains a non-deterministic concurrency bug. Additionally, RX's checkpointing system is lightweight, imposing small time and space overheads.
Feng Qin 0003, Joseph A. Tucek, Jagadeesan Sundaresan, Yuanyuan Zhou 0001
SOSP2
2004 Who Moved My Data? A Backup Tracking System for Dynamic Workstation Environments
Gregory Pluta, Larry Brumbaugh, William Yurcik, Joseph A. Tucek
LISA4