Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Craig A. N. Soules

dblp:35/680 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 3 first-authorSoftware engineering, systems software and programming languages · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 1 first-authorSecurity and privacy · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
Storage systems · 88% Distributed systems · 11% Reconfigurable computing and FPGAs · 1%
Databases, data mining, and information retrieval
5 papers
Information retrieval · 37% Distributed and cloud data management · 22% Data models and query languages · 17%
Software engineering, system software, and programming languages
4 papers
Operating systems · 75% Software maintenance and evolution · 25%
Network and information security
2 papers
Network security · 60% Systems and software security · 40%

Topics — the 24 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › file systems
distributed file system
0.212014
From research to practice: experiences engineering a production metadata database for a scale out file system · FAST 2014
Data models and query languages
NoSQL database
0.112012
LazyBase: trading freshness for performance in a scalable database · EuroSys 2012
Distributed and cloud data management › large-scale data management
scalable database systems
0.112012
LazyBase: trading freshness for performance in a scalable database · EuroSys 2012
Storage systems
storage reliability
0.132009
SCAN-Lite: enterprise-wide analysis on the cheap · EuroSys 2009
Self-Securing Storage: Protecting Data in Compromised Systems · OSDI 2000
Journaling Versus Soft Updates: Asynchronous Meta-data Protection in File Systems · USENIX ATC, General Track 2000
Information retrieval › search engines
file search
0.122007
Confluence: enhancing contextual desktop search · SIGIR 2007
Connections: using context to enhance file search · SOSP 2005
Storage systems
file systems
0.132003
Metadata Efficiency in Versioning File Systems · FAST 2003
Soft updates: a solution to the metadata update problem in file systems · ACM Trans. Comput. Syst. 2000
Journaling Versus Soft Updates: Asynchronous Meta-data Protection in File Systems · USENIX ATC, General Track 2000
Storage systems › data reduction
data deduplication
0.112009
SCAN-Lite: enterprise-wide analysis on the cheap · EuroSys 2009
Distributed systems
distributed scheduling
0.112009
SCAN-Lite: enterprise-wide analysis on the cheap · EuroSys 2009
Operating systems › resource management › storage management
file systems
0.122007
Connections: using context to enhance file search · SOSP 2005
Using Provenance to Aid in Personal File Search · USENIX ATC 2007
Information retrieval › search engines
desktop search
0.112007
Confluence: enhancing contextual desktop search · SIGIR 2007
Storage systems › file systems
soft updates
0.122000
Soft updates: a solution to the metadata update problem in file systems · ACM Trans. Comput. Syst. 2000
Journaling Versus Soft Updates: Asynchronous Meta-data Protection in File Systems · USENIX ATC, General Track 2000
Network security › intrusion detection and prevention
intrusion detection
0.012003
Storage-based Intrusion Detection: Watching Storage Activity for Suspicious Behavior · USENIX Security Symposium 2003
Software maintenance and evolution › software evolution › software adaptation
dynamic reconfiguration
0.012003
System Support for Online Reconfiguration · USENIX ATC, General Track 2003
Storage systems
metadata management
0.012003
Metadata Efficiency in Versioning File Systems · FAST 2003
Storage systems › file systems › versioning
versioning file system
0.012003
Metadata Efficiency in Versioning File Systems · FAST 2003
Storage systems
crash recovery
0.012000
Soft updates: a solution to the metadata update problem in file systems · ACM Trans. Comput. Syst. 2000
Storage systems › file systems
file system recovery
0.012000
Soft updates: a solution to the metadata update problem in file systems · ACM Trans. Comput. Syst. 2000
Storage systems › storage reliability
file system reliability
0.012000
Soft updates: a solution to the metadata update problem in file systems · ACM Trans. Comput. Syst. 2000
Storage systems › file systems
journaling file system
0.012000
Journaling Versus Soft Updates: Asynchronous Meta-data Protection in File Systems · USENIX ATC, General Track 2000
Operating systems
provenance
0.012007
Using Provenance to Aid in Personal File Search · USENIX ATC 2007
Information retrieval › ranking › result ranking
search result ranking
0.012005
Connections: using context to enhance file search · SOSP 2005
Reconfigurable computing and FPGAs
dynamic reconfiguration
0.012003
System Support for Online Reconfiguration · USENIX ATC, General Track 2003
Storage systems › file systems
file system performance
0.012000
Soft updates: a solution to the metadata update problem in file systems · ACM Trans. Comput. Syst. 2000
Storage systems › metadata management
metadata protection
0.012000
Journaling Versus Soft Updates: Asynchronous Meta-data Protection in File Systems · USENIX ATC, General Track 2000

Methods — techniques the papers use, named apart from their topics

replication-aware scheduling · 0.2idle-time exploitation · 0.2content hashing · 0.2update scanning · 0.1pipelining · 0.1batching · 0.1temporal relationship analysis · 0.1file system call tracing · 0.1storage activity analysis · 0.1window focus event analysis · 0.1user study · 0.1write-ahead logging · 0.0
YearPublicationVenuePosition
2014 From research to practice: experiences engineering a production metadata database for a scale out file system
Kimberly Keeton, Charles B. Morrey III, Craig A. N. Soules, Alistair C. Veitch, Stephen Bacon, Oskar Batuner, Marcelo Condotta, Hamilton Coutinho, Patrick J. Doyle, Rafael Eichelberger, Hugo Kiehl, Guilherme R. Magalhaes, James McEvoy, Padmanabhan Nagarajan, Patrick Osborne, Joaquim Souza, Andy Sparkes, Mike Spitzer, Sébastien Tandel, Lincoln Thomas, Sebastian Zangaro
FAST4
2012 LazyBase: trading freshness for performance in a scalable database
abstract
The LazyBase scalable database system is specialized for the growing class of data analysis applications that extract knowledge from large, rapidly changing data sets. It provides the scalability of popular NoSQL systems without the query-time complexity associated with their eventual consistency models, offering a clear consistency model and explicit per-query control over the trade-off between latency and result freshness. With an architecture designed around batching and pipelining of updates, LazyBase simultaneously ingests atomic batches of updates at a very high throughput and offers quick read queries to a stale-but-consistent version of the data. Although slightly stale results are sufficient for many analysis queries, fully up-to-date results can be obtained when necessary by also scanning updates still in the pipeline. Compared to the Cassandra NoSQL system, LazyBase provides 4X--5X faster update throughput and 4X faster read query throughput for range queries while remaining competitive for point queries. We demonstrate LazyBase's tradeoff between query latency and result freshness as well as the benefits of its consistency model. We also demonstrate specific cases where Cassandra's consistency model is weaker than LazyBase's.
James Cipar, Gregory R. Ganger, Kimberly Keeton, Charles B. Morrey III, Craig A. N. Soules, Alistair C. Veitch
EuroSys5
2009 SCAN-Lite: enterprise-wide analysis on the cheap
abstract
Background data analysis due to virus scanning, backup, and desktop search is increasingly prevalent on client systems. As the number of tools and their resource requirements grow, their impact on foreground workloads can be prohibitive. This creates a tension between users' foreground work and the background work that makes information management possible. We present a system called SCAN-Lite that addresses this tension. SCAN-Lite exploits the fact that data in an enterprise is often replicated to efficiently schedule background data analyses. It uses content hashing to identify duplicate content, and scans each unique piece of content only once. It delays scheduling these scans to increase the likelihood that the content will be replicated on multiple machines, thus providing more choices for where to perform the scan. Furthermore, it prioritizes machines to maximize use of idle time and minimize the impact on foreground activities. We evaluate SCAN-Lite using measurements of enterprise replication behavior. We find that SCAN-Lite significantly improves scanning performance over the naive approach, and that it effectively exploits replication to reduce total work done and the impact on client foreground activity.
Craig A. N. Soules, Kimberly Keeton, Charles B. Morrey III
EuroSys1
2008 Seeing is retrieving: building information context from what the user sees
abstract
As the user's document and application workspace grows more diverse, supporting personal information management becomes increasingly important. This trend toward diversity renders it difficult to implement systems which are tailored to specific applications, file types, or other information sources.
Karl Gyllstrom, Craig A. N. Soules
IUI2
2007 Confluence: enhancing contextual desktop search
abstract
We present Confluence, an enhancement to a desktop file search tool called Confluence which extracts conceptual relationships between files by their temporal access patterns in the file system. A limitation of a purely file-based approach is that as file operations are increasingly abstracted by applications, their correlation to a user's activity weakens and thereby reduces the applicability of their temporal patterns. To deal with this problem, we augment the file event stream with a stream of window focus events from the UI layer. We present 3 algorithms that analyze this new stream, extracting the user's task information which informs the existing Confluence algorithms. We present results and conclusions from a preliminary user study on Confluence.
Karl Gyllstrom, Craig A. N. Soules, Alistair C. Veitch
SIGIR2
2007 Using Provenance to Aid in Personal File Search
Sam Shah, Craig A. N. Soules, Gregory R. Ganger, Brian D. Noble
USENIX ATC2
2005 Connections: using context to enhance file search
abstract
Connections is a file system search tool that combines traditional content-based search with context information gathered from user activity. By tracing file system calls, Connections can identify temporal relationships between files and use them to expand and reorder traditional content search results. Doing so improves both recall (reducing false-positives) and precision (reducing false-negatives). For example, Connections improves the average recall (from 13% to 22%) and precision (from 23% to 29%) on the first ten results. When averaged across all recall levels, Connections improves precision from 17% to 28%. Connections provides these benefits with only modest increases in average query time (2 seconds), indexing time (23 seconds daily), and index size(under 1% of the user's data set).
Craig A. N. Soules, Gregory R. Ganger
SOSP1
2003 Metadata Efficiency in Versioning File Systems
Craig A. N. Soules, Garth R. Goodson, John D. Strunk, Gregory R. Ganger
FAST1
2003 Why Can't I Find My Files? New Methods for Automating Attribute Assignment
Craig A. N. Soules, Gregory R. Ganger
HotOS1
2003 System Support for Online Reconfiguration
Craig A. N. Soules, Jonathan Appavoo, Kevin Hui, Robert W. Wisniewski, Dilma Da Silva, Gregory R. Ganger, Orran Krieger, Michael Stumm, Marc A. Auslander, Michal Ostrowski, Bryan S. Rosenburg, Jimi Xenidis
USENIX ATC, General Track1
2003 Storage-based Intrusion Detection: Watching Storage Activity for Suspicious Behavior
Adam G. Pennington, John D. Strunk, John Linwood Griffin, Craig A. N. Soules, Garth R. Goodson, Gregory R. Ganger
USENIX Security Symposium4
2000 Self-Securing Storage: Protecting Data in Compromised Systems
John D. Strunk, Garth R. Goodson, Michael L. Scheinholtz, Craig A. N. Soules, Gregory R. Ganger
OSDI4
2000 Journaling Versus Soft Updates: Asynchronous Meta-data Protection in File Systems
Margo I. Seltzer, Gregory R. Ganger, Marshall K. McKusick, Keith A. Smith, Craig A. N. Soules, Christopher A. Stein
USENIX ATC, General Track5
2000 Soft updates: a solution to the metadata update problem in file systems
abstract
Metadata updates, such as file creation and block allocation, have consistently been identified as a source of performance, integrity, security, and availability problems for file systems. Soft updates is an implementation technique for low-cost sequencing of fine-grained updates to write-back cache blocks. Using soft updates to track and enforce metadata update dependencies, a file system can safely use delayed writes for almost all file operations. This article describes soft updates, their incorporation into the 4.4BSD fast file system, and the resulting effects on the sytem. We show that a disk-based file system using soft updates achieves memory-based file system performance while providing stronger integrity and security guarantees than most disk-based file systems. For workloads that frequently perform updates on metadata (such as creating and deleting files), this improves performance by more than a factor of two and up to a factor of 20 when compared to the conventional synchronous write approach and by 4-19% when compared to an aggressive write-ahead logging approach. In addition, soft updates can improve file system availablity by relegating crash-recovery assistance (e.g., the fsck utility) to an optional and background role, reducing file system recovery time to less than one second.
Gregory R. Ganger, Marshall K. McKusick, Craig A. N. Soules, Yale N. Patt
ACM Trans. Comput. Syst.3