EDBT 2026 Demo / reviewers in the wild / expert
Fred Douglis
dblp:d/FredDouglis
· DBLP profile ↗
45ranked-venue papers
12as first author
1since 2021 · last 2023
0000-0003-2472-0339ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 6 first-authorDatabases, data management, data science and information retrieval · 10 · 1 first-authorComputer networks · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 5 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 3 first-authorArtificial intelligence and machine learning · 1Security and privacy · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
18 papers |
Storage systems · 71% Memory systems · 14% Cloud and datacenter computing · 8% | |
| Computer networks
5 papers |
Internet architecture and protocols · 32% Content delivery and video streaming · 31% Transport protocols and congestion control · 11% |
Topics — the 30 heaviest of 54, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
key-value storage |
0.8 | 2 | 2020 | Customizable Scale-Out Key-Value Stores · IEEE Trans. Parallel Distributed Syst. 2020 bespoKV: application tailored scale-out key-value stores · SC 2018 |
Storage systems › data reduction
data deduplication |
0.8 | 5 | 2017 | The Logic of Physical Garbage Collection in Deduplicating Storage · FAST 2017 A Comprehensive Study of the Past, Present, and Future of Data Deduplication · Proc. IEEE 2016 Tradeoffs in Scalable Data Routing for Deduplication Clusters · FAST 2011 |
Storage systems
storage reliability |
0.6 | 4 | 2018 | RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk Failures · ACM Trans. Storage 2015 RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk Failures · FAST 2015 Can't We All Get Along? Redesigning Protection Storage for Modern Workloads · USENIX ATC 2018 |
Memory systems › cache management › storage caching
flash cache |
0.5 | 2 | 2017 | Pannier: Design and Analysis of a Container-Based Flash Cache for Compound Objects · ACM Trans. Storage 2017 Erasing Belady's Limitations: In Search of Flash Cache Offline Optimality · USENIX ATC 2016 |
Storage systems › key-value storage
distributed key-value store |
0.4 | 1 | 2020 | Customizable Scale-Out Key-Value Stores · IEEE Trans. Parallel Distributed Syst. 2020 |
Storage systems › key-value storage
scalable key-value store |
0.4 | 1 | 2020 | Customizable Scale-Out Key-Value Stores · IEEE Trans. Parallel Distributed Syst. 2020 |
Storage systems
backup storage |
0.3 | 1 | 2018 | Can't We All Get Along? Redesigning Protection Storage for Modern Workloads · USENIX ATC 2018 |
Cloud and datacenter computing › serverless computing
container caching |
0.3 | 1 | 2017 | Pannier: Design and Analysis of a Container-Based Flash Cache for Compound Objects · ACM Trans. Storage 2017 |
Storage systems
flash and SSD |
0.3 | 1 | 2017 | Pannier: Design and Analysis of a Container-Based Flash Cache for Compound Objects · ACM Trans. Storage 2017 |
Storage systems › flash and SSD › flash memory management
garbage collection |
0.3 | 1 | 2017 | The Logic of Physical Garbage Collection in Deduplicating Storage · FAST 2017 |
Storage systems
data reduction |
0.2 | 1 | 2016 | A Comprehensive Study of the Past, Present, and Future of Data Deduplication · Proc. IEEE 2016 |
Storage systems › storage reliability
disk failure prediction |
0.2 | 1 | 2015 | RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk Failures · ACM Trans. Storage 2015 |
Storage systems
data compression |
0.2 | 2 | 2014 | Migratory compression: coarse-grained data reordering to improve compressibility · FAST 2014 Application-specific Delta-encoding via Resemblance Detection · USENIX ATC, General Track 2003 |
Memory systems › data layout optimization
data reordering |
0.2 | 1 | 2014 | Migratory compression: coarse-grained data reordering to improve compressibility · FAST 2014 |
Storage systems › flash and SSD
SSD cache |
0.2 | 1 | 2014 | Nitro: A Capacity-Optimized SSD Cache for Primary Storage · USENIX ATC 2014 |
Memory systems › cache management
cache replacement |
0.2 | 2 | 2017 | Pannier: Design and Analysis of a Container-Based Flash Cache for Compound Objects · ACM Trans. Storage 2017 Erasing Belady's Limitations: In Search of Flash Cache Offline Optimality · USENIX ATC 2016 |
Storage systems › storage management
backup systems |
0.1 | 1 | 2012 | Characteristics of backup workloads in production systems · FAST 2012 |
Performance modeling and evaluation
workload characterization |
0.1 | 1 | 2012 | Characteristics of backup workloads in production systems · FAST 2012 |
Memory systems
cache |
0.1 | 2 | 2016 | Erasing Belady's Limitations: In Search of Flash Cache Offline Optimality · USENIX ATC 2016 Nitro: A Capacity-Optimized SSD Cache for Primary Storage · USENIX ATC 2014 |
High-performance computing › scientific computing
HPC applications |
0.1 | 1 | 2020 | Customizable Scale-Out Key-Value Stores · IEEE Trans. Parallel Distributed Syst. 2020 |
Electronic design automation › physical design
routing |
0.1 | 1 | 2011 | Tradeoffs in Scalable Data Routing for Deduplication Clusters · FAST 2011 |
Storage systems
file systems |
0.1 | 2 | 2016 | A Comprehensive Study of the Past, Present, and Future of Data Deduplication · Proc. IEEE 2016 Redundancy Elimination Within Large Collections of Files · USENIX ATC, General Track 2004 |
Memory systems › cache management
storage caching |
0.1 | 1 | 2017 | Pannier: Design and Analysis of a Container-Based Flash Cache for Compound Objects · ACM Trans. Storage 2017 |
Storage systems › data placement
data placement optimization |
0.1 | 1 | 2008 | Storage optimization for large-scale distributed stream-processing systems · ACM Trans. Storage 2008 |
Storage systems
distributed storage |
0.1 | 1 | 2008 | Storage optimization for large-scale distributed stream-processing systems · ACM Trans. Storage 2008 |
Storage systems › storage management
storage reclamation |
0.1 | 1 | 2008 | Storage optimization for large-scale distributed stream-processing systems · ACM Trans. Storage 2008 |
Storage systems › storage reliability
RAID |
0.1 | 1 | 2015 | RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk Failures · FAST 2015 |
Distributed and cloud data management
web caching |
0.1 | 1 | 2005 | Automatic Fragment Detection in Dynamic Web Pages and Its Impact on Caching · IEEE Trans. Knowl. Data Eng. 2005 |
Electronic design automation › logic synthesis › logic optimization
redundancy removal |
0.0 | 1 | 2004 | Redundancy Elimination Within Large Collections of Files · USENIX ATC, General Track 2004 |
Distributed systems
web caching |
0.0 | 1 | 2004 | Automatic detection of fragments in dynamically generated web pages · WWW 2004 |
Methods — techniques the papers use, named apart from their topics
feedback control · 0.3cryptographic hashing · 0.2content-defined chunking · 0.2reallocated sector analysis · 0.2medium error analysis · 0.2joint failure probability · 0.2SSD caching · 0.2graph algorithms · 0.1simulation · 0.1optimization · 0.1information sharing analysis · 0.1change pattern analysis · 0.1hierarchical modeling · 0.0proximity evaluation · 0.0packet trace simulation · 0.0trace analysis · 0.0indexing · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Secure NDN Packet EncapsulationabstractPacket encapsulation is a general network technique that provides an essential building block for constructing secure networks. While extensively used in IP networks over the last few decades, secure packet encapsulation remains largely unexplored in the context of Named Data Networking (NDN) networks. NDN represents a radical departure from traditional endpoint-oriented networking by making secured data the centerpiece of communication. This new data-centric design brings both advantages and new challenges for the development of secure packet encapsulation that can preserve essential properties of an NDN network, including in-network data caching and builtin multicast data delivery. In this paper, we first identify the major differences between encapsulation solution designs in IP and NDN, highlighting the ensuing challenges, both inherent and practical. We then present a novel design to achieve secure NDN data packet encapsulation, and showcase an implementation suite that enables efficient fetching of securely encapsulated data.11The views, opinions and/or findings expressed are those of the authors and should not be interpreted as representing the official views or policies of the Department of Defense or the U.S. Government. Distribution Statement “A” (Approved for Public Release, Distribution Unlimited) Daniel Townley, Fred Douglis, Jesse Elwell, Constantin Serban, Alexander Afanasyev, Lixia Zhang 0001 |
ICC | 3 |
| 2020 | Customizable Scale-Out Key-Value StoresabstractEnterprise KV stores are often not well suited for HPC applications, and thus cumbersome end-to-end KV design customization is required to meet the needs of modern HPC applications. To this end, in this article we present bespoKV, an adaptive, extensible, and scale-out KV store framework. bespoKV decouples the KV store design into the control plane for distributed management and the data plane for local data store. For the control plane, bespoKVprovides pre-built modules, called controlets, supporting common distributed functionalities (e.g., replication, consistency, and topology) and their various combinations. This decoupling allows bespoKV to take a user-provided single-server KV store, called a datalet, and transparently enables a scalable and fault-tolerant distributed KV store service. The resulting distributed stores are also adaptive to consistency or topology requirement changes and can be easily extended for new types of services. Such specializations enable innovative uses of KV stores in HPC applications, especially for emerging applications that utilize KV-friendly workloads. We evaluate bespoKV in a local testbed as well as in a public cloud settings. Experiments show that bespoKV-enabled distributed KV stores scale horizontally to a large number of nodes, and performs comparably and sometimes 1.2× to 2.6× better than the state-of-the-art systems. Ali Anwar 0001, Yue Cheng 0001, Hai Huang 0002, Jingoo Han, Hyogi Sim, Fred Douglis, Ali Raza Butt |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2018 | bespoKV: application tailored scale-out key-value stores
Ali Anwar 0001, Yue Cheng 0001, Hai Huang 0002, Jingoo Han, Hyogi Sim, Fred Douglis, Ali Raza Butt |
SC | 7 |
| 2018 | Can't We All Get Along? Redesigning Protection Storage for Modern Workloads
Yamini Allu, Fred Douglis, Mahesh Kamat, Ramya Prabhakar, Philip Shilane, Rahul Ugale |
USENIX ATC | 2 |
| 2017 | The Logic of Physical Garbage Collection in Deduplicating Storage
Fred Douglis, Abhinav Duggal, Philip Shilane, Tony Wong, Shiqin Yan, Fabiano C. Botelho |
FAST | 1 |
| 2017 | Pannier: Design and Analysis of a Container-Based Flash Cache for Compound ObjectsabstractClassic caching algorithms leverage recency, access count, and/or other properties of cached blocks at per-block granularity. However, for media such as flash which have performance and wear penalties for small overwrites, implementing cache policies at a larger granularity is beneficial. Recent research has focused on buffering small blocks and writing in large granularities, sometimes called containers, but it has not explored the ramifications and best strategies for caching compound blocks consisting of logically distinct, but physically co-located, blocks. Containers may have highly diverse blocks, with mixtures of frequently accessed, infrequently accessed, and invalidated blocks. We propose and evaluate Pannier, a flash cache layer that provides high performance while extending flash lifespan. Pannier uses three main techniques: (1) leveraging block access counts to manage cache containers, (2) incorporating block liveness as a property to improve flash cache space efficiency, and (3) designing a multi-step feedback controller to ensure a flash cache reaches its desired lifespan while maintaining performance. Our evaluation shows that Pannier improves flash cache performance and extends lifespan beyond previous per-block and container-aware caching policies. More fundamentally, our investigation highlights the importance of creating new policies for caching compound blocks in flash. Philip Shilane, Fred Douglis, Grant Wallace |
ACM Trans. Storage | 3 |
| 2016 | Erasing Belady's Limitations: In Search of Flash Cache Offline Optimality
Yue Cheng 0001, Fred Douglis, Philip Shilane, Grant Wallace, Peter Desnoyers, Kai Li 0001 |
USENIX ATC | 2 |
| 2016 | A Comprehensive Study of the Past, Present, and Future of Data DeduplicationabstractData deduplication, an efficient approach to data reduction, has gained increasing attention and popularity in large-scale storage systems due to the explosive growth of digital data. It eliminates redundant data at the file or subfile level and identifies duplicate content by its cryptographically secure hash signature (i.e., collision-resistant fingerprint), which is shown to be much more computationally efficient than the traditional compression approaches in large-scale storage systems. In this paper, we first review the background and key features of data deduplication, then summarize and classify the state-of-the-art research in data deduplication according to the key workflow of the data deduplication process. The summary and taxonomy of the state of the art on deduplication help identify and understand the most important design considerations for data deduplication systems. In addition, we discuss the main applications and industry trend of data deduplication, and provide a list of the publicly available sources for deduplication research and studies. Finally, we outline the open problems and future research directions facing deduplication-based storage systems. Wen Xia, Hong Jiang 0001, Dan Feng 0001, Fred Douglis, Philip Shilane, Yu Hua 0001, Min Fu 0002 |
Proc. IEEE | 4 |
| 2016 | State of the JournalabstractDiscusses the current state of the journal, reports on current and future areas of exploration and research, and presents new editors. Paolo Montuschi, Edward J. McCluskey, Samarjit Chakraborty, Jason Cong, Ramón M. Rodríguez-Dagnino, Fred Douglis, Lieven Eeckhout, Gernot Heiser, Sushil Jajodia, Ruby B. Lee, Dinesh Manocha, Tomás F. Pena, Isabelle Puaut, Hanan Samet, Donatella Sciuto |
IEEE Trans. Computers | 6 |
| 2015 | RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk Failures
Fred Douglis, Guanlin Lu, Darren Sawyer, Surendar Chandra, Windsor W. Hsu |
FAST | 2 |
| 2015 | Metadata Considered Harmful...to Deduplication
Fred Douglis, Jim Li, Robert Ricci, Stephen Smaldone, Grant Wallace |
HotStorage | 2 |
| 2015 | Pannier: A Container-based Flash Cache for Compound ObjectsabstractClassic caching algorithms leverage recency, access count, and/or other properties of cached blocks at per-block granularity. However, for media such as flash which have performance and wear penalties for small overwrites, implementing cache policies at a larger granularity is beneficial. Recent research has focused on buffering small blocks and writing in large granularities, called containers, but it has not explored the ramifications and best strategies for caching compound blocks consisting of logically distinct, but physically co-located, blocks. Containers may have highly diverse blocks, with mixtures of frequently accessed, infrequently accessed, and invalidated blocks. Philip Shilane, Fred Douglis, Grant Wallace |
Middleware | 3 |
| 2015 | RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk FailuresabstractModern storage systems orchestrate a group of disks to achieve their performance and reliability goals. Even though such systems are designed to withstand the failure of individual disks, failure of multiple disks poses a unique set of challenges. We empirically investigate disk failure data from a large number of production systems, specifically focusing on the impact of disk failures on RAID storage systems. Our data covers about one million SATA disks from six disk models for periods up to 5 years. We show how observed disk failures weaken the protection provided by RAID. The count ofreallocated sectorscorrelates strongly with impending failures. With these findings we designed RAIDShield, which consists of two components. First, we have built and evaluated an active defense mechanism that monitors the health of each disk and replaces those that are predicted to fail imminently. This proactive protection has been incorporated into our product and is observed to eliminate 88% of triple disk errors, which are 80% of all RAID failures. Second, we have designed and simulated a method of using the joint failure probability to quantify and predict how likely a RAID group is to face multiple simultaneous disk failures, which can identify disks that collectively represent a risk of failure even when no individual disk is flagged in isolation. We find in simulation that RAID-level analysis can effectively identify most vulnerable RAID-6 systems, improving the coverage to 98% of triple errors. We conclude with discussions of operational considerations in deploying RAIDShieldmore broadly and new directions in the analysis of disk errors. One interesting approach is to combine multiple metrics, allowing the values of different indicators to be used for predictions. Using newer field data that reports an additional metric,medium errors, we find that the relative efficacy of reallocated sectors and medium errors varies across disk models, offering an additional way to predict failures. Rachel Traylor, Fred Douglis, Mark Chamness, Guanlin Lu, Darren Sawyer, Surendar Chandra, Windsor W. Hsu |
ACM Trans. Storage | 3 |
| 2014 | Migratory compression: coarse-grained data reordering to improve compressibility
Guanlin Lu, Fred Douglis, Philip Shilane, Grant Wallace |
FAST | 3 |
| 2014 | Assert(!Defined(Sequential I/O))
Philip Shilane, Fred Douglis, Darren Sawyer, Hyong Shim |
HotStorage | 3 |
| 2014 | Nitro: A Capacity-Optimized SSD Cache for Primary Storage
Philip Shilane, Fred Douglis, Hyong Shim, Stephen Smaldone, Grant Wallace |
USENIX ATC | 3 |
| 2012 | Characteristics of backup workloads in production systems
Grant Wallace, Fred Douglis, Hangwei Qian, Philip Shilane, Stephen Smaldone, Mark Chamness, Windsor W. Hsu |
FAST | 2 |
| 2011 | Tradeoffs in Scalable Data Routing for Deduplication Clusters
Wei Dong 0003, Fred Douglis, Kai Li 0001, R. Hugo Patterson, Sazzala Reddy, Philip Shilane |
FAST | 2 |
| 2011 | Content-aware Load Balancing for Distributed Backup
Fred Douglis, Deepti Bhardwaj, Hangwei Qian, Philip Shilane |
LISA | 1 |
| 2009 | Applying Knowledge Sharing for Business Intelligence CollaborationabstractIT services need an automatic and flexible ability to react to dynamic changes in their environment. Managing change effectively and reducing the negative effects of day-today operations has become one of the most important tasks in IT service management, which require hiring highly skilled IT professionals with correspondingly high labor costs. There is a challenge to select, implement and integrate the right resources quickly and effectively. Although there is a growing body of research into IT management, many techniques are either too narrow (focusing on a single component rather than the entire system), or they address only configuration data collection and integration. Instead, one needs to scale or respond to special domain knowledge, collaboration and the right data for helping IT professionals to improve their work. In this paper, we propose a knowledge-sharing based collaborating management system for IT service management. It aims to bridge the gap between domain experts' knowledge and manageable systems. We developed a proof-of-concept of an impact analysis service based on knowledge-sharing; it establishes an IT service management collaboration paradigm and ecosystem, leveraging the expert's rich knowledge and experience to improve the management quality and reduce the cost. A case study driven by customers demonstrates that collaboration with knowledge sharing is effective both at constructing useful system analysis services and in using those services to improve system management. Bo Yang 0013, Hao Wang 0208, Fred Douglis |
ICWS | 3 |
| 2008 | Storage optimization for large-scale distributed stream-processing systemsabstractWe consider storage in an extremely large-scale distributed computer system designed for stream processing applications. In such systems, both incoming data and intermediate results may need to be stored to enable analyses at unknown future times. The quantity of data of potential use would dominate even the largest storage system. Thus, a mechanism is needed to keep the data most likely to be used. One recently introduced approach is to employ retention value functions, which effectively assign each data object a value that changes over time in a prespecified way [Douglis et al.2004]. Storage space for data entering the system is reclaimed automatically by deleting data of the lowest current value. In such large systems, there will naturally be multiple file systems available, each with different properties. Choosing the right file system for a given incoming stream of data presents a challenge. In this article we provide a novel and effective scheme for optimizing the placement of data within a distributed storage subsystem employing retention value functions. The goal is to keep the data of highest overall value, while simultaneously balancing the read load to the file system. The key aspects of such a scheme are quite different from those that arise in traditional file assignment problems. We further motivate this optimization problem and describe a solution, comparing its performance to other reasonable schemes via simulation experiments. Kirsten Hildrum, Fred Douglis, Joel L. Wolf, Philip S. Yu, Lisa Fleischer, Akshay Katta |
ACM Trans. Storage | 2 |
| 2007 | Failure Recovery in Cooperative Data Stream AnalysisabstractWe present a failure recovery framework for System S, a large-scale stream data analysis environment. It is intended to support multiple sites, which have their own local administration and goals. However, it is beneficial for these sites to cooperate with each other, especially in the presence of various failures. Our ultimate goal is to support automatic, timely failure recovery through cooperation among sites. We identify the unique challenges in the context of System S and present our initial design work. In particular, we consider a backup selection problem, specifying where to recover failed jobs, which we formulate as an optimization problem. We present an approximation algorithm together with empirical results obtained through simulations. Our numerical evaluations show that the proposed approximation algorithm is very efficient and effective compared to the optimal solutions. It exhibits a promising empirical performance ratio that is close to the theoretical limit of polynomial approximations of such a problem Bin Rong, Fred Douglis, Cathy H. Xia, Zhen Liu 0001 |
ARES | 2 |
| 2007 | Storage Optimization for Large-Scale Distributed Stream Processing SystemsabstractWe consider storage in an extremely large-scale distributed computer system designed for stream processing applications. In such systems, incoming data and intermediate results may need to be stored to enable future analyses. The quantity of such data would dominate even the largest storage system. Thus, a mechanism is needed to keep the most useful data. One recently introduced approach is to employ retention value functions, which effectively assign each data object a value that changes over time. Storage space is then reclaimed automatically by deleting data of lowest current value. In such large systems, there can naturally be multiple file systems available, each with different properties. Choosing the right file system for a given incoming data stream presents a challenge. In this paper we provide a novel and effective scheme for optimizing the placement of data within a distributed storage subsystem employing retention value functions. The goal is to keep the data of highest overall value, while simultaneously balancing the read load to the file system. Kirsten Hildrum, Fred Douglis, Joel L. Wolf, Philip S. Yu, Lisa Fleischer, Akshay Katta |
IPDPS | 2 |
| 2007 | CLASP: Collaborating, Autonomous Stream Processing Systems
Michael Branson, Fred Douglis, Brad Fawcett, Zhen Liu 0001, Anton Riabov, Fan Ye 0003 |
Middleware | 2 |
| 2006 | Guest Editors' Introduction
Fred Douglis, Prabhakar Raghavan |
World Wide Web | 1 |
| 2005 | Automatic Fragment Detection in Dynamic Web Pages and Its Impact on CachingabstractConstructing Web pages from fragments has been shown to provide significant benefits for both content generation and caching. In order for a Web site to use fragment-based content generation, however, good methods are needed for fragmenting the Web pages. Manual fragmentation of Web pages is expensive, error prone, and unscalable. This paper proposes a novel scheme to automatically detect and flag fragments that are cost-effective cache units in Web sites serving dynamic content. Our approach analyzes Web pages with respect to their information sharing behavior, personalization characteristics, and change patterns. We identify fragments which are shared among multiple documents or have different lifetime or personalization characteristics. Our approach has three unique features. First, we propose a framework for fragment detection, which includes a hierarchical and fragment-aware model for dynamic Web pages and a compact and effective data structure for fragment detection. Second, we present an efficient algorithm to detect maximal fragments that are shared among multiple documents. Third, we develop a practical algorithm that effectively detects fragments based on their lifetime and personalization characteristics. This paper shows the results when the algorithms are applied to real Web sites. We evaluate the proposed scheme through a series of experiments, showing the benefits and costs of the algorithms. We also study the impact of using the fragments detected by our system on key parameters such as disk space utilization, network bandwidth consumption, and load on the origin servers. Lakshmish Ramaswamy, Arun Iyengar, Ling Liu 0001, Fred Douglis |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2004 | Redundancy Elimination Within Large Collections of Files
Purushottam Kulkarni, Fred Douglis, Jason D. LaVoie, John M. Tracey |
USENIX ATC, General Track | 2 |
| 2004 | Automatic detection of fragments in dynamically generated web pagesabstractDividing web pages into fragments has been shown to provide significant benefits for both content generation and caching. In order for a web site to use fragment-based content generation, however, good methods are needed for dividing web pages into fragments. Manual fragmentation of web pages is expensive, error prone, and unscalable. This paper proposes a novel scheme to automatically detect and flag fragments that are cost-effective cache units in web sites serving dynamic content. We consider the fragments to be interesting if they are shared among multiple documents or they have different lifetime or personalization characteristics. Our approach has three unique features. First, we propose a hierarchical and fragment-aware model of the dynamic web pages and a data structure that is compact and effective for fragment detection. Second, we present an efficient algorithm to detect maximal fragments that are shared among multiple documents. Third, we develop a practical algorithm that effectively detects fragments based on their lifetime and personalization characteristics. We evaluate the proposed scheme through a series of experiments, showing the benefits and costs of the algorithms. We also study the impact of adopting the fragments detected by our system on disk space utilization and network bandwidth consumption. Lakshmish Ramaswamy, Arun Iyengar, Ling Liu 0001, Fred Douglis |
WWW | 4 |
| 2003 | Techniques for efficient fragment detection in web pagesabstractThe existing approaches to fragment-based publishing, delivery and caching of web pages assume that the web pages are manually fragmented at their respective web sites. However manual fragmentation of web pages is expensive, error prone, and not scalable. This paper proposes a novel scheme to automatically detect and flag possible fragments in a web site. Our approach is basedonananalysisofthewebpagesdynamicallygeneratedat given web sites with respect to their information sharing behavior, personalization characteristics and change patterns. Categories and Subject Descriptors: H.3.3 [Information Systems- Information storage and retrieval]: Information search and retrieval Lakshmish Ramaswamy, Arun Iyengar, Ling Liu 0001, Fred Douglis |
CIKM | 4 |
| 2003 | Application-specific Delta-encoding via Resemblance Detection
Fred Douglis, Arun Iyengar |
USENIX ATC, General Track | 1 |
| 2002 | ACDN: a content delivery network for applicationsabstractNo abstract available. Pradnya Karbhari, Michael Rabinovich, Fred Douglis |
SIGMOD Conference | 4 |
| 2002 | A Precise and Efficient Evaluation of the Proximity Between Web Clients and Their Local DNS Servers
Z. Morley Mao, Chuck Cranor, Fred Douglis, Michael Rabinovich, Oliver Spatscheck, Jia Wang 0001 |
USENIX ATC, General Track | 3 |
| 2002 | CDN brokering
Alexandros Biliris, Chuck Cranor, Fred Douglis, Michael Rabinovich, Sandeep Sibal, Oliver Spatscheck, Walter Sturm |
Comput. Commun. | 3 |
| 1999 | Performance of Web Proxy Caching in Heterogeneous Bandwidth EnvironmentsabstractMuch work on the performance of Web proxy caching has focused on high-level metrics such as hit rates, but has ignored low level details such as "cookies", aborted connections, and persistent connections between clients and proxies as well as between proxies and servers. These details have a strong impact on performance, particularly in heterogeneous bandwidth environments where network speeds between clients and proxies are significantly different than speeds between proxies and servers. We evaluate through detailed simulations the latency and bandwidth effects of Web proxy caching in such environments. We drive our simulations with packet traces from two scenarios: clients connected through slow dialup modems to a commercial ISP, and clients on a fast LAN in an industrial research lab. We present three main results. First, caching persistent connections at the proxy can improve latency much more than simply caching Web data. Second, aborted connections can waste more bandwidth than that saved by caching data. Third, cookies can dramatically reduce hit rates by making many documents effectively uncacheable. Anja Feldmann, Ramón Cáceres, Fred Douglis, Gideon Glass, Michael Rabinovich |
INFOCOM | 3 |
| 1999 | Adaptive Modem Connection Lifetimes
Fred Douglis, Tom Killian |
USENIX ATC, General Track | 1 |
| 1999 | World Wide Web Characterization and Performance Evaluation - Preface
Fred Douglis |
World Wide Web | 1 |
| 1998 | The AT&T Internet Difference Engine: Tracking and Viewing Changes on the Web
Fred Douglis, Thomas Ball 0001, Yih-Farn Robin Chen, Eleftherios Koutsofios |
World Wide Web | 1 |
| 1997 | Potential Benefits of Delta Encoding and Data Compression for HTTPabstractCaching in the World Wide Web currently follows a naive model, which assumes that resources are referenced many times between changes. The model also provides no way to update a cache entry if a resource does change, except by transferring the resource's entire new value. Several previous papers have proposed updating cache entries by transferring only the differences, or "delta," between the cached entry and the current value.In this paper, we make use of dynamic traces of the full contents of HTTP messages to quantify the potential benefits of delta-encoded responses. We show that delta encoding can provide remarkable improvements in response size and response delay for an important subset of HTTP content types. We also show the added benefit of data compression, and that the combination of delta encoding and data compression yields the best results.We propose specific extensions to the HTTP protocol for delta encoding and data compression. These extensions are compatible with existing implementations and specifications, yet allow efficient use of a variety of encoding techniques. Jeffrey C. Mogul, Fred Douglis, Anja Feldmann, Balachander Krishnamurthy |
SIGCOMM | 2 |
| 1996 | Tracking and Viewing Changes on the Web
Fred Douglis, Thomas Ball 0001 |
USENIX ATC | 1 |
| 1996 | WebGUIDE: Querying and Navigating Changes in Web Repositories
Fred Douglis, Thomas Ball 0001, Yih-Farn Robin Chen, Eleftherios Koutsofios |
Comput. Networks | 1 |
| 1996 | TeleWeb: Loosely Connected Access to the World Wide Web
Bill N. Schilit, Fred Douglis, David M. Kristol, Paul Krzyzanowski, James Sienicki, John A. Trotter |
Comput. Networks | 2 |
| 1994 | Storage Alternatives for Mobile Computers
Fred Douglis, Ramón Cáceres, M. Frans Kaashoek, Kai Li 0001, Brian Marsh, Joshua A. Tauber |
OSDI | 1 |
| 1993 | The Gold MailerabstractThe Gold Mailer, a system that provides users with an integrated way to send and receive messages using different media, efficiently store and retrieve these messages, and access a variety of sources of other useful information, is described. The mailer solves the problems of information overload, organization of messages and multiple interfaces. By providing good storage and retrieval facilities, it can be used as a powerful information processing engine covering a range of useful office information. The Gold Mailer's query language, indexing engine, file organization, data structures, and support of mail message data and multimedia documents are discussed.> Daniel Barbará, Chris Clifton, Fred Douglis, Hector Garcia-Molina, Ben Kao, Sharad Mehrotra, Jens Tellefsen, Rosemary Walsh |
ICDE | 3 |
| 1991 | Transparent Process Migration: Design Alternatives and the Sprite ImplementationabstractAbstract The Sprite operating system allows executing processes to be moved between hosts at any time. We use this process migration mechanism to offload work onto idle machines, and also to evict migrated processes when idle workstations are reclaimed by their owners. Sprite's migration mechanism provides a high degree of transparency both for migrated processes and for users. Idle machines are identified, and eviction is invoked, automatically by daemon processes. On Sprite it takes up to a few hundred milliseconds on SPARCstation 1 workstations to perform a remote exec, whereas evictions typically occur in a few seconds. The pmake program uses remote invocation to invoke tasks concurrently. Compilations commonly obtain speed‐up factors in the range of three to six; they are limited primarily by contention for centralized resources such as file servers. CPU‐bound tasks such as simulations can make more effective use of idle hosts, obtaining as much as eight‐fold speed‐up over a period of hours. Process migration has been in regular service for over two years. Fred Douglis, John K. Ousterhout |
Softw. Pract. Exp. | 1 |
| 1987 | Process Migration in the Sprite Operating System
Fred Douglis, John K. Ousterhout |
ICDCS | 1 |