VLDB 2026 Research / reviewers in the wild / expert
Bing Zhu 0003
dblp:71/4410-3
· DBLP profile ↗
22ranked-venue papers
11as first author
7since 2021 · last 2023
0000-0002-5164-9550ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-authorArtificial intelligence and machine learning · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 5 · 1 first-authorTheory of computation · 3 · 3 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Secure Fractional Repetition Codes for Distributed Storage SystemsabstractFractional Repetition (FR) codes are a class of exact-repair regenerating codes known for their ability to minimize repair bandwidth and provide uncoded repair capability. However, the table-based repair mechanism employed by FR codes brings both advantages and challenges. While it enables FR codes to exceed the storage-bandwidth trade-off, it also exposes them to potential information leakage and makes traditional security methods less effective. In this paper, we address the issue of securing FR codes against eavesdroppers who can access the content of a subset of storage nodes. To mitigate the risk of information leakage, we propose a novel code construction that enhances the security of FR codes in a flexible manner. Zhihang Deng, Bing Zhu 0003, Kenneth W. Shum, Weiping Wang 0003 |
GLOBECOM | 2 |
| 2023 | ISAA: Boost Repair Process by Constructing the Degree Constrained Optimal Repair Tree for Erasure-coded SystemsabstractTo ensure data reliability, large-scale distributed systems usually adopt erasure codes to restore failed nodes. However, existing erasure-coded repair strategies will cause heavy network traffics, which will increase the repair time. In order to boost the repair process, we consider optimizing the repair path which can be abstracted to a repair tree. Moreover, we add a degree constraint to each node to avoid local congestion. In this paper, we study the degree constrained optimal repair tree, which is an NP-hard problem. Current methods cannot find the optimal solution in a short time in complex non-uniform bandwidth networks. To obtain the optimal repair tree, an improved simulated annealing algorithm (ISAA) based on the Prufer code representation is proposed in this paper. In addition, we simulate the repair process of erasure codes in a non-uniform bandwidth network and experiments show that the repair time reduction can reach up to 66.4% and 88.6% with ISAA over Repair Pipelining and Partial-Parallel-Repair. Xianzhi Du, Bing Zhu 0003, Zhihang Deng, Kenneth W. Shum, Weiping Wang 0003 |
ICPADS | 2 |
| 2022 | A Genetic Algorithm-based Construction of Fractional Repetition CodesabstractFractional repetition (FR) codes form a special family of minimum bandwidth regenerating codes, characterized by an uncoded exact repair process. This low-complexity repair process benefits from the two-layer encoding structure consisting of an outer maximum distance separable code and an inner repetition code. However, it sacrifices certain storage efficiency, or supported file size. In order to improve the supported file size, we present in this paper a genetic algorithm-based construction of FR codes. In particular, we first review the existence of FR codes and introduce a general construction of FR codes. We implement an iterative optimization process of mimicking biological evolution on constructed FR codes. Through simulations, it is shown that the supported file size can be improved for different network scales. Zhihang Deng, Bing Zhu 0003, Xianzhi Du, Kenneth W. Shum |
GLOBECOM | 2 |
| 2022 | High-Rate Constructions of Exact-Repair Regenerating CodesabstractRegenerating codes are a class of distributed storage codes proposed to reduce the bandwidth consumption in the node repair process. In this paper, we present explicitly a construction of exact-repair regenerating codes, which is a two-layer encoding structure that consists of concatenating an outer scalar maximum distance separable (MDS) code followed by an inner tailor-made MDS array code. These coded symbols are distributed across the storage nodes based on a family of combinatorial objects termed t-designs, and this sophisticated symbol placement enables that a failed node can be repaired by simple data transfer. Furthermore, our proposed regenerating codes generally have a high code rate and extend the parameter values of existing constructions. Bing Zhu 0003, Xuyu Zhao, Weiping Wang 0003 |
WCNC | 2 |
| 2022 | An Improved Bound and Singleton-Optimal Constructions of Fractional Repetition CodesabstractFractional repetition (FR) codes are a class of repair efficient erasure codes that can recover a failed storage node with both optimal repair bandwidth and complexity. In this paper, we study the minimum distance of FR codes, which is the smallest number of nodes whose failure leads to the unrecoverable loss of the stored file. We derive a new upper bound on the minimum distance of FR codes, which is tighter than the Singleton bound and a Singleton-like bound that takes locality into account. Based on regular graphs and combinatorial designs, several families of FR codes with optimal minimum distance are obtained. Bing Zhu 0003, Kenneth W. Shum, Weiping Wang 0003, Jianxin Wang 0001 |
IEEE Trans. Commun. | 1 |
| 2021 | Square Fractional Repetition Codes for Distributed Storage Systems
Bing Zhu 0003, Shigeng Zhang, Weiping Wang 0003 |
ICA3PP (2) | 1 |
| 2021 | Expandable Fractional Repetition Codes for Distributed Storage SystemsabstractModern distributed storage systems are increasingly implementing erasure codes to obtain better storage performance. In such systems, it is desirable to regenerate a failed storage node in a cost-effective manner since node failures occur frequently in real-world storage networks. Fractional repetition (FR) codes are a special class of regenerating codes that enable efficient recovery of failed storage nodes. In this paper, we introduce expandable FR codes, wherein both the number of storage nodes and the capacity of each node in the storage systems can be readily expanded. We present explicit constructions of expandable FR codes by applying two families of combinatorial structures called embeddable quasi-residual designs and extendible t-designs. Moreover, we study the property of constructed codes for some special scenarios. Bing Zhu 0003, Shigeng Zhang, Weiping Wang 0003 |
ITW | 1 |
| 2020 | On the Optimal Minimum Distance of Fractional Repetition CodesabstractFractional repetition (FR) codes are a class of repair efficient erasure codes that can recover a failed storage node with both optimal repair bandwidth and complexity. In this paper, we focus on the minimum distance of FR codes, which is the smallest number of nodes whose failure leads to the unrecoverable loss of stored files. We consider upper bounds on the minimum distance and present several families of FR codes attaining these bounds. The optimal constructions are derived from regular graphs and combinatorial designs, respectively. Bing Zhu 0003, Kenneth W. Shum, Weiping Wang 0003, Jianxin Wang 0001 |
GLOBECOM | 1 |
| 2020 | Accurate human activity recognition with multi-task learning
Yinggang Li, Shigeng Zhang, Bing Zhu 0003, Weiping Wang 0003 |
CCF Trans. Pervasive Comput. Interact. | 3 |
| 2020 | Fractional Repetition Codes With Optimal Reconstruction DegreeabstractFractional repetition (FR) codes form a special class of minimum bandwidth regenerating codes by providing uncoded repairs via a table-based repair model. For a given file size, it is desirable to design FR codes that minimize the reconstruction degree, which is defined as the number of storage nodes required for data retrieval. In this paper, we first consider a lower bound on the reconstruction degree of FR codes obtained by Silberstein and Etzion. We present several families of FR codes that attain this lower bound, which are derived from combinatorial designs and regular graphs. We further provide a new lower bound on the reconstruction degree of FR codes, which is tighter than the existing one. Moreover, we show that for an FR code with reconstruction degree achieving the lower bounds, the corresponding dual code attains upper bounds on the supported file size. Bing Zhu 0003, Kenneth W. Shum, Hui Li 0022 |
IEEE Trans. Inf. Theory | 1 |
| 2019 | On the Optimal Reconstruction Degree of Fractional Repetition CodesabstractFractional repetition (FR) codes form a special class of minimum bandwidth regenerating codes by providing uncoded repairs with a table-based repair model. In this paper, we focus on a lower bound on the reconstruction degree of FR codes, which is the smallest number of storage nodes required for data retrieval. We show that for an FR code with reconstruction degree attaining this lower bound, the corresponding dual FR code is optimal with respect to an upper bound on the file size, and vice versa. Using this duality relationship, we present several families of FR codes with optimal reconstruction degree. Bing Zhu 0003, Kenneth W. Shum, Hui Li 0022, Weiping Wang 0003 |
ISIT | 1 |
| 2019 | On the Duality and File Size Hierarchy of Fractional Repetition CodesabstractDistributed storage systems that deploy erasure codes can provide better features such as lower storage overhead and higher data reliability. In this paper, we focus on fractional repetition (FR) codes, which are a class of storage codes characterized by the features of uncoded exact repair and minimum repair bandwidth. We study the duality of FR codes and investigate the relationship between the supported file size of an FR code and its dual code. Based on the established relationship, we derive an improved dual bound on the supported file size of FR codes. We further show that FR codes constructed from t-designs are optimal when the size of the stored file is sufficiently large. Moreover, we present the tensor product technique for combining FR codes and elaborate on the file size hierarchy of resulting codes. Bing Zhu 0003, Kenneth W. Shum, Hui Li 0022 |
Comput. J. | 1 |
| 2017 | On the duality of fractional repetition codesabstractErasure codes have emerged as an efficient technology for providing data redundancy in distributed storage systems. However, it is a challenging task to repair the failed storage nodes in erasure-coded storage systems, which requires large quantities of network resources. In this paper, we study fractional repetition (FR) codes, which enable the minimal repair complexity and also minimum repair bandwidth during node repair. We focus on the duality of FR codes, and investigate the relationship between the supported file size of an FR code and its dual code. Furthermore, we present a dual bound on the supported file size of FR codes. Bing Zhu 0003, Kenneth W. Shum, Hui Li 0022 |
ITW | 1 |
| 2016 | EStore: An effective optimized data placement structure for HiveabstractThe data warehouse system Hive has emerged as an important facility for supporting data computing and storage. In particular, RCFile is a tailor-made data placement structure implemented in Hive, which is designed for the data processing efficiency. In this paper, we propose several optimized schemes based on RCFile and introduce EStore, which is an optimized data placement structure that is able to improve the query rate and reduce storage space for Hive. Specifically, it adopts both row-store and column-store in blocks, and further classifies the columns by the frequency of each table-column. Moreover, we also employ the classic RDP code to store files of the data table. We conduct experiments on a real cluster, and the results show that EStore has better features in terms of data query rate and storage space compared with RCFile. Hui Li 0022, Bing Zhu 0003, Jiawei Cai |
IEEE BigData | 4 |
| 2015 | On the implementation of Zigzag codes for distributed storage systemabstractErasure codes such as Reed-Solomon (RS) codes are widely used to improve data reliability in distributed storage systems. Although erasure codes indeed greatly reduce the storage overhead compared to the replication schemes, it is still very costly in terms of network bandwidth when repairing a failed node. To address such problem, we employ the Zigzag code, a MDS array code with optimal repair property, in the practical system. Specifically, we first build a general system on Hadoop to evaluate the encoding, decoding and repair performance of different codes, and then implement Zigzag codes on our system. The experimental results show that the Zigzag codes coincide with the theoretical findings and has certain advantages. Compared to current HDFS modules that use RS codes, our Zigzag based HDFS implementation shows significant reduction of repair disk I/O and repair bandwidth with the same computation complexity. Lijia Lu, Hui Li 0022, Bing Zhu 0003, Weijuan Yin |
IEEE BigData | 4 |
| 2015 | HFR code: a flexible replication scheme for cloud storage systemsabstractFractional repetition (FR) codes are a family of repair‐efficient storage codes that provide exact and uncoded node repair at the minimum bandwidth regenerating point. The advantageous repair properties are achieved by a tailor‐made two‐layer encoding scheme which concatenates an outer maximum‐distance‐separable (MDS) code and an inner repetition code. In this study, the authors generalise the application of FR codes and propose heterogeneous fractional repetition (HFR) code, which is adaptable to the scenario where the repetition degrees of coded packets are different. The authors provide explicit code constructions by utilising group divisible designs, which allow the design of HFR codes over a large range of parameters. The constructed codes achieve the system storage capacity under random access repair and have multiple repair alternatives for node failures. Further, the authors take advantage of the systematic feature of MDS codes and present a novel design framework of HFR codes, in which storage nodes can be wisely partitioned into clusters such that data reconstruction time can be reduced when contacting nodes in the same cluster. Bing Zhu 0003, Hui Li 0022, Kenneth W. Shum, Shuo-Yen Robert Li |
IET Commun. | 1 |
| 2014 | A new Zigzag MDS code with optimal encoding and efficient decodingabstractDistributed file system has emerged in recent years as an efficient solution to store the large amount of data produced anytime and anywhere. In order to guarantee data reliability, it is necessary to introduce redundancy to the storage systems. Compared to simple replication, practical systems are increasingly adopting erasure codes for better storage efficiency. However, traditional erasure codes such as maximum-distance-separable (MDS) codes, are designed over a large finite field, which inevitably hinders the wide implementation of erasure codes. In this paper, we propose a new family of MDS codes with high computation efficiency. More specifically, only XOR operation is included in the encoding process to generate parity blocks. Upon failure of a storage node, we use the efficient Zigzag decoding method to recover the failed blocks, which achieves the optimal encoding and an efficient decoding. Furthermore, we implement the proposed codes in a distributed file system, and the results show the high performance of the new codes. Hui Li 0022, Hanxu Hou, Bing Zhu 0003, Tai Zhou, Lijia Lu |
IEEE BigData | 4 |
| 2014 | STORE: Data recovery with approximate minimum network bandwidth and disk I/O in distributed storage systemsabstractRecently, traditional erasure codes such as Reed-Solomon (RS) codes have been increasingly deployed in many distributed storage systems to reduce the large storage overhead incurred by the widely adopted replication scheme. However, these codes require significantly high resources with respect to network bandwidth and disk I/O during recovery of missing or unavailable data. It is referred as the recovery problem. In this paper, we dedicate to integrating exact minimum bandwidth regenerating codes into practical systems to solve the recovery problem. We design an implementation friendly storage code with the recently proposed BASIC framework and ZigZag decodable code for saving recovery bandwidth and disk I/O. We build a system called STORE based on this code and evaluate our prototype atop a HDFS cluster testbed with 21 nodes. As shown in this paper, the recovery bandwidth achieves minimum approximately during recovery of both data block and parity block with STORE. Another attractive result is that the recovery disk I/O also achieves minimum approximately during recovery of data block. Due to the reduction of recovery bandwidth and disk I/O, the degraded read throughput is boosted notably. Tai Zhou, Hui Li 0022, Bing Zhu 0003, Hanxu Hou |
IEEE BigData | 3 |
| 2014 | Repair efficient storage codes via combinatorial configurationsabstractFractional repetition (FR) codes are a special class of regenerating codes characterized by the exact and uncoded repair property. In this work, we propose an explicit method to construct FR codes from combinatorial configurations. The proposed construction gives FR codes with parameters that are not covered by prior approaches. Bing Zhu 0003, Hui Li 0022, Kenneth W. Shum |
IEEE BigData | 1 |
| 2014 | Dictionary construction for sparse representation classification: A novel cluster-based approachabstractThere has been a rapid development in sparse representation classification (SRC) since it came out. Most previous work on dictionary improvement was to enhance the classification performance by modifying the dictionary representation structure while this paper concentrates on the reduction of dictionary length with nearly no sacrifice in classification accuracy. A novel cluster-based dictionary construction approach for SRC is proposed in this paper. Both cluster technique and clustering evaluation index are introduced to help construct an optimal dictionary for better classification performance. Results of experiments have verified that the new dictionary does not lose discrimination ability while its running time is greatly reduced. Most importantly, its robustness is also preserved. Weiyang Liu, Yandong Wen, Hui Li 0022, Bing Zhu 0003 |
ISCC | 4 |
| 2014 | On low repair complexity storage codes via group divisible designsabstractFractional repetition (FR) codes are a family of storage codes that provide efficient node repair at the minimum bandwidth regenerating point. Specifically, the repair process is exact and uncoded, but table-based. Existing constructions of FR codes are primarily based on combinatorial designs such as Steiner systems, resolvable designs, etc. In this paper, we present a new explicit construction of FR codes, which adopts the theory of uniform group divisible designs, termed GDDFR codes. Our codes achieve the storage capacity of random access and are available for a wide range of parameters. In addition, our techniques allow for constructing FR codes with parameters that are not covered by Steiner systems, which answers an open question put forward in prior work. Bing Zhu 0003, Kenneth W. Shum, Hui Li 0022, Shuo-Yen Robert Li |
ISCC | 1 |
| 2014 | Fast Repair for Single Failure in Erasure Coding-Based Distributed Storage SystemsabstractIn order to guarantee data reliability in distributed storage systems, erasure codes are widely used for the desirable storage properties. Nevertheless, the codes have one drawback that overmuch data are needed to repair a failure, resulting in both large bandwidth consuming in the network and high calculation pressure on the replacement node. For repair bandwidth problem, researchers derive the tradeoffs between storage and repair traffic from network coding and propose regenerating codes. However, the constructions of regenerating codes complicate the systems as well as recovery calculation. Hence, this paper proposes a distributed repair method based on general erasure codes to mitigate the burden of both recovery computation and network traffic. We observe that distributing recovery computation among helpers can distract the whole calculation procedure and accelerate repair speed in practical systems. Furthermore, by combining this technique with network topology, we introduce a novel repair tree to minimize repair traffic. Repair tree is also derived from network coding. The performance of the repair tree is preliminarily analyzed and evaluated, which infers that the storage-bandwidth bound of regenerating codes can be broken under this model. Hui Li 0022, Bing Zhu 0003 |
SRDS | 3 |