EDBT 2026 Demo / reviewers in the wild / expert
Dimitre Trendafilov
dblp:98/3215
· DBLP profile ↗
2ranked-venue papers
0as first author
0since 2021 · last 2004
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 50% Storage systems · 50% | |
| Computer networks
1 paper |
Content delivery and video streaming · 100% |
Topics — the 2 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › file systems
file synchronization |
0.0 | 1 | 2004 | Improved File Synchronization Techniques for Maintaining Large Replicated Collections over Slow Networks · ICDE 2004 |
Distributed systems
replication |
0.0 | 1 | 2004 | Improved File Synchronization Techniques for Maintaining Large Replicated Collections over Slow Networks · ICDE 2004 |
Methods — techniques the papers use, named apart from their topics
rsync · 0.1delta encoding · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2004 | Improved File Synchronization Techniques for Maintaining Large Replicated Collections over Slow NetworksabstractWe study the problem of maintaining large replicated collections of files or documents in a distributed environment with limited bandwidth. This problem arises in a number of important applications, such as synchronization of data between accounts or devices, content distribution and Web caching networks, Web site mirroring, storage networks, and large scale Web search and mining. At the core of the problem lies the following challenge, called the file synchronization problem: given two versions of a file on different machines, say an outdated and a current one, how can we update the outdated version with minimum communication cost, by exploiting the significant similarity between the versions? While a popular open source tool for this problem called rsync is used in hundreds of thousands of installations, there have been only very few attempts to improve upon this tool in practice. We propose a framework for remote file synchronization and describe several new techniques that result in significant bandwidth savings. Our focus is on applications where very large collections have to be maintained over slow connections. We show that a prototype implementation of our framework and techniques achieves significant improvements over rsync. As an example application, we focus on the efficient synchronization of very large Web page collections for the purpose of search, mining, and content distribution. Torsten Suel, Patrick Noel, Dimitre Trendafilov |
ICDE | 3 |
| 2002 | Cluster-Based Delta Compression of a Collection of FilesabstractDelta compression techniques are commonly used to succinctly represent an updated version of a file with respect to an earlier one. We study the use of delta compression in a somewhat different scenario, where we wish to compress a large collection of (more or less) related files by performing a sequence of pairwise delta compressions. The problem of finding an optimal delta encoding for a collection of files by taking pairwise deltas can be reduced to the problem of computing a branching of maximum weight in a weighted directed graph, but this solution is inefficient and thus does not scale to larger file collections. This motivates us to propose a framework for cluster-based delta compression that uses text clustering techniques to prune the graph of possible pairwise delta encodings. To demonstrate the efficacy of our approach, we present experimental results on collections of Web pages. Our experiments show that cluster-based delta compression of collections provides significant improvements in compression ratio as compared to individually compressing each file or using tar+gzip, at a moderate cost in efficiency. Zan Ouyang, Nasir Memon, Torsten Suel, Dimitre Trendafilov |
WISE | 4 |