EDBT 2026 Demo / reviewers in the wild / expert
Hong Tang 0004
dblp:00/2111-4
· DBLP profile ↗
15ranked-venue papers
6as first author
0since 2021 · last 2015
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 5 first-authorComputer networks · 2Software engineering, systems software and programming languages · 2 · 1 first-authorSecurity and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
10 papers |
Parallel and multicore computing · 30% Storage systems · 28% Distributed systems · 22% | |
| Software engineering, system software, and programming languages
3 papers |
Compilers and program optimization · 56% Operating systems · 44% |
Topics — the 29 heaviest of 32, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
distributed storage |
0.1 | 3 | 2004 | The Panasas ActiveScale Storage Cluster - Delivering Scalable High Bandwidth Storage · SC 2004 A Self-Organizing Storage Cluster for Parallel Data-Intensive Applications · SC 2004 An Efficient Data Location Protocol for Self.organizing Storage Clusters · SC 2003 |
Distributed systems › peer-to-peer systems
consistent hashing |
0.1 | 2 | 2004 | A Self-Organizing Storage Cluster for Parallel Data-Intensive Applications · SC 2004 An Efficient Data Location Protocol for Self.organizing Storage Clusters · SC 2003 |
Storage systems
data placement |
0.1 | 2 | 2004 | A Self-Organizing Storage Cluster for Parallel Data-Intensive Applications · SC 2004 An Efficient Data Location Protocol for Self.organizing Storage Clusters · SC 2003 |
Distributed systems
fault tolerance |
0.1 | 2 | 2005 | Dependency isolation for thread-based multi-tier Internet services · INFOCOM 2005 Optimizing data aggregation for cluster-based internet services · PPoPP 2003 |
Cloud and datacenter computing › datacenter services › online service systems › internet services
cluster-based internet services |
0.1 | 2 | 2003 | Optimizing data aggregation for cluster-based internet services · PPoPP 2003 Integrated Resource Management for Cluster-based Internet Services · OSDI 2002 |
Parallel and multicore computing
MPI |
0.1 | 2 | 2000 | Program transformation and runtime support for threaded MPI execution on shared-memory machines · ACM Trans. Program. Lang. Syst. 2000 Compile/Run-Time Support for Threaded MPI Execution on Multiprogrammed Shared Memory Machines · PPoPP 1999 |
Parallel and multicore computing
parallel programming models |
0.1 | 2 | 2000 | Program transformation and runtime support for threaded MPI execution on shared-memory machines · ACM Trans. Program. Lang. Syst. 2000 Compile/Run-Time Support for Threaded MPI Execution on Multiprogrammed Shared Memory Machines · PPoPP 1999 |
Storage systems
object storage |
0.0 | 1 | 2004 | The Panasas ActiveScale Storage Cluster - Delivering Scalable High Bandwidth Storage · SC 2004 |
Distributed systems
data aggregation |
0.0 | 1 | 2003 | Optimizing data aggregation for cluster-based internet services · PPoPP 2003 |
Parallel and multicore computing › parallel programming models
partitioned parallelism |
0.0 | 1 | 2003 | Optimizing data aggregation for cluster-based internet services · PPoPP 2003 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management |
0.0 | 1 | 2002 | Integrated Resource Management for Cluster-based Internet Services · OSDI 2002 |
Embedded and real-time systems › real-time scheduling
admission control |
0.0 | 1 | 2001 | Demand-driven Service Differentiation in Cluster-based Network Servers · INFOCOM 2001 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.0 | 1 | 2001 | Demand-driven Service Differentiation in Cluster-based Network Servers · INFOCOM 2001 |
Cloud and datacenter computing › quality of service
differentiated service |
0.0 | 1 | 2001 | Demand-driven Service Differentiation in Cluster-based Network Servers · INFOCOM 2001 |
Compilers and program optimization › program transformation
compile-time transformation |
0.0 | 1 | 2000 | Program transformation and runtime support for threaded MPI execution on shared-memory machines · ACM Trans. Program. Lang. Syst. 2000 |
Parallel and multicore computing
parallel programming runtimes |
0.0 | 1 | 2000 | Program transformation and runtime support for threaded MPI execution on shared-memory machines · ACM Trans. Program. Lang. Syst. 2000 |
Parallel and multicore computing › parallel programming models
message passing |
0.0 | 1 | 1999 | Compile/Run-Time Support for Threaded MPI Execution on Multiprogrammed Shared Memory Machines · PPoPP 1999 |
Parallel and multicore computing › parallel programming models › message passing
MPI runtime |
0.0 | 1 | 1999 | Adaptive Two-level Thread Management for Fast MPI Execution on Shared Memory Machines · SC 1999 |
Parallel and multicore computing
parallel programming models and runtimes |
0.0 | 1 | 1999 | Adaptive Two-level Thread Management for Fast MPI Execution on Shared Memory Machines · SC 1999 |
Parallel and multicore computing › parallel scheduling
runtime scheduling |
0.0 | 1 | 1999 | Adaptive Two-level Thread Management for Fast MPI Execution on Shared Memory Machines · SC 1999 |
Parallel and multicore computing › parallel programming runtimes
runtime systems and scheduling |
0.0 | 1 | 1999 | Compile/Run-Time Support for Threaded MPI Execution on Multiprogrammed Shared Memory Machines · PPoPP 1999 |
Parallel and multicore computing › parallel programming runtimes
thread management |
0.0 | 1 | 1999 | Adaptive Two-level Thread Management for Fast MPI Execution on Shared Memory Machines · SC 1999 |
Distributed systems
replication and consistency |
0.0 | 1 | 2004 | A Self-Organizing Storage Cluster for Parallel Data-Intensive Applications · SC 2004 |
Storage systems › networked storage
storage networking |
0.0 | 1 | 2004 | The Panasas ActiveScale Storage Cluster - Delivering Scalable High Bandwidth Storage · SC 2004 |
Distributed systems › replication
replication and fault tolerance |
0.0 | 1 | 2003 | An Efficient Data Location Protocol for Self.organizing Storage Clusters · SC 2003 |
Performance modeling and evaluation
queueing models |
0.0 | 1 | 2001 | Demand-driven Service Differentiation in Cluster-based Network Servers · INFOCOM 2001 |
Memory systems
shared memory |
0.0 | 1 | 2000 | Program transformation and runtime support for threaded MPI execution on shared-memory machines · ACM Trans. Program. Lang. Syst. 2000 |
Operating systems › resource management › process management › CPU scheduling
thread scheduling |
0.0 | 1 | 1999 | Adaptive Two-level Thread Management for Fast MPI Execution on Shared Memory Machines · SC 1999 |
Operating systems › resource management › process management
user-level threads |
0.0 | 1 | 1999 | Adaptive Two-level Thread Management for Fast MPI Execution on Shared Memory Machines · SC 1999 |
Methods — techniques the papers use, named apart from their topics
compile-time transformation · 0.1consistent hashing · 0.1thread management · 0.1system call interception · 0.1versioning · 0.0object-based storage device · 0.0aggregation algorithms · 0.0staged timeout · 0.0event-driven thread pool · 0.0bloom filter · 0.0runtime support · 0.0lock-free queue management · 0.0cache affinity · 0.0adaptive event waiting · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | VM-centric snapshot deduplication for cloud data backupabstractData deduplication is important for snapshot backup of virtual machines (VMs) because of excessive redundant content. Fingerprint search for source-side duplicate detection is resource intensive when the backup service for VMs is co-located with other cloud services. This paper presents the design and analysis of a fast VM-centric backup service with a tradeoff for a competitive deduplication efficiency while using small computing resources, suitable for running on a converged cloud architecture that cohosts many other services. The design consideration includes VM-centric file system block management for the increased VM snapshot availability. This paper describes an evaluation of this VM-centric scheme to assess its deduplication efficiency, resource usage, and fault tolerance. Wei Zhang 0118, Daniel Agun, Tao Yang 0009, Richard Wolski, Hong Tang 0004 |
MSST | 5 |
| 2013 | Low-Cost Data Deduplication for Virtual Machine Backup in Cloud Storage
Wei Zhang 0118, Tao Yang 0009, Gautham Narayanasamy, Hong Tang 0004 |
HotStorage | 4 |
| 2012 | Multi-level Selective Deduplication for VM Snapshots in Cloud StorageabstractIn a virtualized cloud computing environment, frequent snapshot backup of virtual disks improves hosting reliability but storage demand of such operations is huge. While dirty bit-based technique can identify unmodified data between versions, full deduplication with fingerprint comparison can remove more redundant content at the cost of computing resources. This paper presents a multi-level selective deduplication scheme which integrates inner-VM and cross-VM duplicate elimination under a stringent resource requirement. This scheme uses popular common data to facilitate fingerprint comparison while reducing the cost and it strikes a balance between local and global deduplication to increase parallelism and improve reliability. Experimental results show the proposed scheme can achieve high deduplication ratio while using a small amount of cloud resources. Wei Zhang 0118, Hong Tang 0004, Tao Yang 0009, Xiaogang Li 0003 |
IEEE CLOUD | 2 |
| 2010 | Programming support and adaptive checkpointing for high-throughput data services with log-based recoveryabstractMany applications in large-scale data mining and offline processing are organized as network services, running continuously or for a long period of time. To sustain high-throughput, these services often keep their data in memory, thus susceptible to failures. On the other hand, the availability requirement for these services is not as stringent as online services exposed to millions of users. But those data-intensive offline or mining applications do require data persistence to survive failures. This paper presents programming and runtime support called SLACH for building multi-threaded high-throughput persistent services. To keep in-memory objects persistent, SLACH employs application-assisted logging and checkpointing for log-based recovery while maximizing throughput and concurrency. SLACH adaptively adjusts checkpointing frequency based on log growth and throughput demand to balance between runtime overhead and recovery speed. This paper describes the design and API of SLACH, adaptive checkpoint control, and our experiences and experiments in using SLACH at Ask.com. Jingyu Zhou, Caijie Zhang, Hong Tang 0004, Jiesheng Wu, Tao Yang 0009 |
DSN | 3 |
| 2005 | Dependency isolation for thread-based multi-tier Internet servicesabstractMulti-tier Internet service clusters often contain complex calling dependencies among service components spreading across cluster nodes. Without proper handling, partial failure or overload at one component can cause cascading performance degradation in the entire system. While dependency management may not present significant challenges for even-driven services (particularly in the context of staged event-driven architecture), there is a lack of system support for thread-based online services to achieve dependency isolation automatically. To this end, we propose dependency capsule, a new mechanism that supports automatic recognition of dependency states and per-dependency management for thread-based services. Our design employs a number of dependency capsules at each service node: one for each remote service component. Dependency capsules monitor and manage threads that block on supporting services and isolate their performance impact on the capsule host and the rest of the system. In addition to the failure and overload isolation, each capsule can also maintain dependency-specific feedback information to adjust control strategies for better availability and performance. In our implementation, dependency capsules are transparent to application-level services and clustering middleware, which is achieved by intercepting dependency-induced system calls. Additionally, we employ two-level thread management so that only light-weight user-level threads block in dependency capsules. Using four applications on two different clustering middleware platforms, we demonstrate the effectiveness of dependency capsules in improving service availability and throughput during component failures and overload. Lingkun Chu, Hong Tang 0004, Tao Yang 0009, Jingyu Zhou |
INFOCOM | 3 |
| 2004 | A Self-Organizing Storage Cluster for Parallel Data-Intensive ApplicationsabstractCluster-based storage systems are popular for data-intensive applications and it is desirable yet challenging to provide incremental expansion and high availability while achieving scalability and strong consistency. This paper presents the design and implementation of a self-organizing storage cluster called Sorrento, which targets data-intensive workload with highly parallel requests and low write-sharing patterns. Sorrento automatically adapts to storage node joins and departures, and the system can be configured and maintained incrementally without interrupting its normal operation. Data location information is distributed across storage nodes using consistent hashing and the location protocol differentiates small and large data objects for access efficiency. It adopts versioning to achieve single-file serializability and replication consistency. In this paper, we present experimental results to demonstrate features and performance of Sorrento using microbenchmarks, application benchmarks, and application trace replay. Hong Tang 0004, Aziz Gulbeden, Jingyu Zhou, William Strathearn, Tao Yang 0009, Lingkun Chu |
SC | 1 |
| 2004 | The Panasas ActiveScale Storage Cluster - Delivering Scalable High Bandwidth StorageabstractFundamental advances in high-level storage architectures and low-level storage-device interfaces greatly improve the performance and scalability of storage systems. Specifically, the decoupling of storage control (i.e., file system policy) from datapath operations (i.e., read, write) allows client applications to leverage the readily available bandwidth of storage devices while continuing to rely on the rich semantics of today’s file systems. Further, the evolution of storage interfaces from block-based devices with no protection to object-based devices with per-command access control enables storage to become secure, first-class IP-based network citizens. This paper examines how the Panasas ActiveScale Storage Cluster leverages distributed storage and object-based devices to achieve linear scalability of storage bandwidth. Specifically, we focus on implementation issues with our Object-based Storage Device, aggregation algorithms across the collection of OSDs, and the close coupling of networking and storage to achieve scalability. Hong Tang 0004, Aziz Gulbeden, Jingyu Zhou, William Strathearn, Tao Yang 0009, Lingkun Chu |
SC | 1 |
| 2003 | Optimizing data aggregation for cluster-based internet servicesabstractLarge-scale cluster-based Internet services often host partitioned datasets to provide incremental scalability. The aggregation of results produced from multiple partitions is a fundamental building block for the delivery of these services. This paper presents the design and implementation of a programming primitive -- Data Aggregation Call (DAC) -- to exploit partition parallelism for cluster-based Internet services. A DAC request specifies a local processing operator and a global reduction operator, and it aggregates the local processing results from participating nodes through the global reduction operator. Applications may allow a DAC request to return partial aggregation results as a tradeoff between quality and availability. Our architecture design aims at improving interactive responses with sustained throughput for typical cluster environments where platform heterogeneity and software/hardware failures are common. At the cluster level, our load-adaptive reduction tree construction algorithm balances processing and aggregation load across servers while exploiting partition parallelism. Inside each node, we employ an event-driven thread pool design that prevents slow nodes from adversely affecting system throughput under highly concurrent workload. We further devise a staged timeout scheme that eagerly prunes slow or unresponsive servers from the reduction tree to meet soft deadlines. We have used the DAC primitive to implement several applications: a search engine document retriever, a parallel protein sequence matcher, and an online parallel facial recognizer. Our experimental and simulation results validate the effectiveness of the proposed optimization techniques for reducing response time, improving throughput, and gracefully handling server unresponsiveness. We also demonstrate the ease-of use of the DAC primitive and the scalability of our architecture design. Lingkun Chu, Hong Tang 0004, Tao Yang 0009 |
PPoPP | 2 |
| 2003 | An Efficient Data Location Protocol for Self.organizing Storage ClustersabstractComponent additions and failures are common for large-scale storage clusters in production environments. To improve availability and manageability, we investigate and compare data location schemes for a large self-organizing storage cluster that can quickly adapt to the additions or departures of storage nodes. We further present an efficient location scheme that differentiates between small and large file blocks for reduced management overhead compared to uniform strategies. In our protocol, small blocks, which are typically in large quantities, are placed through consistent hashing. Large blocks, much fewer in practice, are placed through a usage-based policy, and their locations are tracked by Bloom filters. The proposed scheme results in improved storage utilization even with non-uniform cluster nodes. To achieve high scalability and fault resilience, this protocol is fully distributed, relies only on soft states, and supports data replication. We demonstrate the effectiveness and efficiency of this protocol through trace-driven simulation. Hong Tang 0004, Tao Yang 0009 |
SC | 1 |
| 2002 | Integrated Resource Management for Cluster-based Internet Services
Hong Tang 0004, Tao Yang 0009, Lingkun Chu |
OSDI | 2 |
| 2001 | Optimizing threaded MPI execution on SMP clustersabstractOur previous work has shown that using threads to execute MPI programs can yield great performance gain on multiprogrammed shared-memory machines. This paper investigates the design and implementation of a thread-based MPI system on SMP clusters. Our study indicates that with a proper design for threaded MPI execution, both point-to-point and collective communication performance can be improved substantially, compared to a process-based MPI implementation in a cluster environment. Our contribution includes a hierarchy-aware and adaptive communication scheme for threaded MPI execution and a thread-safe network device abstraction that uses event-driven synchronization and provides separated collective and point-to-point communication channels. This paper describes the implementation of our design and illustrates its performance advantage on a Linux SMP cluster. Hong Tang 0004, Tao Yang 0009 |
ICS | 1 |
| 2001 | Demand-driven Service Differentiation in Cluster-based Network ServersabstractService differentiation that provides prioritized service qualities to multiple classes of client requests can effectively utilize available server resources. This paper studies how demand-driven service differentiation in terms of end-user performance can be supported in cluster-based network servers. Our objective is to deliver better services to high priority request classes without over-sacrificing low priority classes. To achieve this objective, we propose a dynamic scheduling scheme, called DDSD that adapts to fluctuating request resource demands by periodically repartitioning servers. This scheme also employs priority-based admission control to drop excessive user requests and achieve soft performance guarantees. For each scheduling period, our scheme monitors the system status and uses a queuing model to approximate server behaviors and guide resource allocation. Our experiments show that the proposed technique achieves demand-driven service differentiation while maximizing resource utilization and that it can substantially outperform static server partitioning. Huican Zhu, Hong Tang 0004, Tao Yang 0009 |
INFOCOM | 2 |
| 2000 | Program transformation and runtime support for threaded MPI execution on shared-memory machinesabstractParallel programs written in MPI have been widely used for developing high-performance applications on various platforms. Because of a restriction of the MPI computation model, conventional MPI implementations on shared-memory machines map each MPI node to an OS process, which can suffer serious performance degradation in the presence of multiprogramming. This paper studies compile-time and runtime techniques for enhancing performance portability of MPI code running on multiprogrammed shared-memory machines. The proposed techniques allow MPI nodes to be executed safety and efficiently as threads. Compile-time transformation eliminates global and static variables in C code using node-specific data. The runtime support includes an efficient and provably correct communication protocol that uses lock-free data structure and takes advantage of address space sharing among threads. The experiments on SGI Origin 2000 show that our MPI prototype called TMPI using the proposed techniques is competitive with SGI's native MPI implementation in a dedicated environment, and that it has significant performance advantages in a multiprogrammed environment. Hong Tang 0004, Tao Yang 0009 |
ACM Trans. Program. Lang. Syst. | 1 |
| 1999 | Compile/Run-Time Support for Threaded MPI Execution on Multiprogrammed Shared Memory MachinesabstractMPI is a message-passing standard widely used for developing high-performance parallel applications. Because of the restriction in the MPI computation model, conventional implementations on shared memory machines map each MPI node to an OS process, which suffers serious performance degradation in the presence of multiprogramming, especially when a space/time sharing policy is employed in OS job scheduling. In this paper, we study compile-time and run-time support for MPI by using threads and demonstrate our optimization techniques for executing a large class of MPI programs written in C. The compile-time transformation adopts thread-specific data structures to eliminate the use of global and static variables in C code. The runtime support includes an efficient point-to-point communication protocol based on a novel lock-free queue management scheme. Our experiments on an SGI Origin 2000 show that our MPI prototype called TMPI using the proposed techniques is competitive with SGI's native MPI implementation in a dedicated environment, and it has significant performance advantages with up to a 23-fold improvement in a multiprogrammed environment. Hong Tang 0004, Tao Yang 0009 |
PPoPP | 1 |
| 1999 | Adaptive Two-level Thread Management for Fast MPI Execution on Shared Memory MachinesabstractThis paper addresses performance portability of MPI code on multiprogrammed shared memory machines. Conventional MPI implementations map each MPI node to an OS process, which suffers severe performance degradation in multiprogrammed environments. Our previous work (TMPI) has developed compile/run-time techniques to support threaded MPI execution by mapping each MPI node to a kernel thread. However, kernel threads have context switch cost higher than user-level threads and this leads to longer spinning time requirement during MPI synchronization. This paper presents an adaptive two-level thread scheme for MPI to reduce context switch and synchronization cost. This scheme also exposes thread scheduling information at user-level, which allows us to design an adaptive event waiting strategy to minimize CPU spinning and exploit cache affinity. Our experiments show that the MPI system based on the proposed techniques has great performance advantages over the previous version of TMPI and the SGI MPI implementation in multiprogrammed environments. The improvement ratio can reach as much as 161 % or even more depending on the degree of multiprogramming. 1 Hong Tang 0004, Tao Yang 0009 |
SC | 2 |