EDBT 2026 Demo / reviewers in the wild / expert
Juntao Fang
dblp:70/8586
· DBLP profile ↗
7ranked-venue papers
3as first author
4since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Storage systems · 78% Distributed systems · 18% Performance modeling and evaluation · 4% | |
| Theoretical computer science
1 paper |
Graph algorithms and graph theory · 87% Mathematical optimization · 13% | |
| Network and information security
1 paper |
Blockchain and cryptocurrency security · 100% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
storage reliability |
1.2 | 3 | 2021 | Design and Evaluation of a Risk-Aware Failure Identification Scheme for Improved RAS in Erasure-Coded Data Centers · IEEE Trans. Parallel Distributed Syst. 2021 Early Identification of Critical Blocks: Making Replicated Distributed Storage Systems Reliable Against Node Failures · IEEE Trans. Parallel Distributed Syst. 2018 RAFI: Risk-Aware Failure Identification to Improve the RAS in Erasure-coded Data Centers · USENIX ATC 2018 |
Blockchain and cryptocurrency security
blockchain scalability |
0.8 | 1 | 2024 | Data Deduplication Based on Content Locality of Transactions to Enhance Blockchain Scalability · ACM Trans. Archit. Code Optim. 2024 |
Storage systems › data reduction
data deduplication |
0.8 | 1 | 2024 | Data Deduplication Based on Content Locality of Transactions to Enhance Blockchain Scalability · ACM Trans. Archit. Code Optim. 2024 |
Graph algorithms and graph theory › network analysis
network robustness |
0.8 | 1 | 2024 | Optimizing Network Resilience via Vertex Anchoring · WWW 2024 |
Distributed systems › fault tolerance
failure diagnosis |
0.6 | 2 | 2021 | Design and Evaluation of a Risk-Aware Failure Identification Scheme for Improved RAS in Erasure-Coded Data Centers · IEEE Trans. Parallel Distributed Syst. 2021 RAFI: Risk-Aware Failure Identification to Improve the RAS in Erasure-coded Data Centers · USENIX ATC 2018 |
Storage systems › storage reliability › data recovery
data repair |
0.5 | 1 | 2021 | Design and Evaluation of a Risk-Aware Failure Identification Scheme for Improved RAS in Erasure-Coded Data Centers · IEEE Trans. Parallel Distributed Syst. 2021 |
Storage systems › storage reliability
erasure coding |
0.5 | 1 | 2021 | Design and Evaluation of a Risk-Aware Failure Identification Scheme for Improved RAS in Erasure-Coded Data Centers · IEEE Trans. Parallel Distributed Syst. 2021 |
Mathematical optimization
combinatorial optimization |
0.2 | 1 | 2024 | Optimizing Network Resilience via Vertex Anchoring · WWW 2024 |
Performance modeling and evaluation
simulation |
0.1 | 1 | 2021 | Design and Evaluation of a Risk-Aware Failure Identification Scheme for Improved RAS in Erasure-Coded Data Centers · IEEE Trans. Parallel Distributed Syst. 2021 |
Distributed systems
fault tolerance |
0.1 | 1 | 2018 | Early Identification of Critical Blocks: Making Replicated Distributed Storage Systems Reliable Against Node Failures · IEEE Trans. Parallel Distributed Syst. 2018 |
Methods — techniques the papers use, named apart from their topics
content locality analysis · 1.5alias generation · 1.5time-dependent framework · 0.8greedy algorithm · 0.8simulation · 0.5prototyping · 0.5emulation · 0.5replica state management · 0.3adaptive time-out · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Optimizing Network Resilience via Vertex AnchoringabstractNetwork resilience is a critical ability of a network to maintain its functionality against disturbances. A network is resilient/robust when a large portion of the nodes are to be better engaged in the network, i.e., they are less likely to leave given the changes on the network. Existing studies validate that the engagement of a node can be well captured by its coreness on network topology. Therefore, it is promising to maximize the number of nodes with increasing coreness values. In this paper, we propose and study thefollower maximization problem: maximizing the resilience gain (the number of coreness-increased vertices) via anchoring a set of vertices within a given budget. We prove that the problem is NP-hard and W[2]-hard, and it is NP-hard to approximate within an O(n^1-ε ) factor. We first propose an advanced greedy approach, followed by a time-dependent framework designed to quickly find high-quality results. The framework is initialized by the advanced greedy algorithm and incorporates novel techniques for optimizing the search space. The effectiveness and efficiency of our solution are verified with extensive experiments on 8 real-life datasets. Our source codes are available at https://github.com/Tsyxxxka/Follower-Maximization. Siyi Teng, Jiadong Xie 0002, Fan Zhang 0036, Juntao Fang, Kai Wang 0037 |
WWW | 5 |
| 2024 | Data Deduplication Based on Content Locality of Transactions to Enhance Blockchain ScalabilityabstractBlockchain is a promising infrastructure for the internet and digital economy, but it has serious scalability problems, that is, long block synchronization time and high storage cost. Conventional coarse-grained data deduplication schemes (block or file level) are proved to be ineffective on improving the scalability of blockchains. Based on comprehensive analysis on typical blockchain workloads, we propose two new locality concepts (economic and argument locality) and a novel fine-grained data deduplication scheme (transaction level) named Alias-Chain. Specifically, Alias-Chain replaces frequently used data, for example, smart contract arguments, with much shorter aliases to reduce the block sizes, which results in both shorter synchronization time and lower storage cost. Furthermore, to solve the potential consistency issue in Alias-Chain, we propose two complementary techniques: one is generating aliases from history blocks with high consistency, and the other is speeding up the generation of aliases via a specific algorithm. Our simulation results show: (1) the average transfer and SC-call transaction (a transaction used to call the smart contracts in the blockchain) sizes can be significantly reduced by up to 11.03% and 79.44% in native Ethereum, and up to 39.29% and 81.84% in Ethereum optimized by state-of-the-art techniques; and (2) the two complementary techniques well address the inconsistency risk with very limited impact on the benefit of Alias-Chain. Prototyping-based experiments are further conducted on a testbed consisting of up to 3200 miners. The results demonstrate the effectiveness and efficiency of Alias-Chain on reducing block synchronization time and storage cost under typical real-world workloads. Chenglong Yi, Shenggang Wan, Juntao Fang, Liqiang Zhang 0010 |
ACM Trans. Archit. Code Optim. | 4 |
| 2023 | Predicting miRNA-Disease Associations via Node-Level Attention Graph Auto-EncoderabstractPrevious studies have confirmed microRNA (miRNA), small single-stranded non-coding RNA, participates in various biological processes and plays vital roles in many complex human diseases. Therefore, developing an efficient method to infer potential miRNA disease associations could greatly help understand operational mechanisms for diseases at the molecular level. However, during these early stages for miRNA disease prediction, traditional biological experiments are laborious and expensive. Therefore, this study proposes a novel method called AGAEMD (node-level Attention Graph Auto-Encoder to predict potential MiRNA Disease associations). We first create a heterogeneous matrix incorporating miRNA similarity, disease similarity, and known miRNA-disease associations. Then these matrixes are input into a node-level attention encoder-decoder network which utilizes low dimensional dense embeddings to represent nodes and calculate association scores. To verify the effectiveness of the proposed method, we conduct a series of experiments on two benchmark datasets (the Human MicroRNA Disease Database v2.0 and v3.2) and report the averages over 10 runs in comparison with several state-of-the-art methods. Experimental results have demonstrated the excellent performance of AGAEMD in comparison with other methods. Three important diseases (Colon Neoplasms, Lung Neoplasms, Lupus Vulgaris) were applied in case studies. The results comfirm the reliable predictive performance of AGAEMD. Huizhe Zhang, Juntao Fang, Yuping Sun, Guobo Xie, Zhiyi Lin 0001, Guosheng Gu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | Design and Evaluation of a Risk-Aware Failure Identification Scheme for Improved RAS in Erasure-Coded Data CentersabstractData reliability and availability, and serviceability (RAS) of erasure-coded data centers are highly affected by data repair induced by node failures. In a traditional failure identification scheme, all chunks share the same identification time threshold, thus losing opportunities to further improve the RAS. To solve this problem, we propose RAFI, a novel risk-aware failure identification scheme. In RAFI, chunk failures in stripes experiencing different numbers of failed chunks are identified using different time thresholds. For those chunks in a high-risk stripe, a shorter identification time is adopted, thus improving the overall data reliability and availability. For those chunks in a low-risk stripe, a longer identification time is adopted, thus reducing the repair network traffic. Therefore, RAS can be improved simultaneously. We also propose three optimization techniques to reduce the additional overhead that RAFI imposes on management nodes and to ensure that RAFI can work properly under large-scale clusters. We use simulation, emulation, and prototyping implementation to evaluate RAFI from multiple aspects. Simulation and prototype results prove the effectiveness and correctness of RAFI, and the performance improvement of the optimization techniques on RAFI is demonstrated by running the emulator. Weichen Huang, Juntao Fang, Shenggang Wan, Changsheng Xie 0001, Xubin He |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2018 | RAFI: Risk-Aware Failure Identification to Improve the RAS in Erasure-coded Data Centers
Juntao Fang, Shenggang Wan, Xubin He |
USENIX ATC | 1 |
| 2018 | Early Identification of Critical Blocks: Making Replicated Distributed Storage Systems Reliable Against Node FailuresabstractIn large-scale replicated distributed storage systems consisting of hundreds to thousands of nodes, node failures are not rare and can cause data blocks to lose their replicas and become faulty. A simple but effective approach to prevent data loss from the node failures, i.e., ensuring reliability, is to shorten the identification time of the node failures and faulty blocks, which is determined by both timeouts and check intervals for node states. However, to maintain low repair network traffic, the identification time is actually relatively long and even dominates repair processes of critical blocks. In this paper, we propose a novel scheme, named RICK, to explore the potential in the identification time, and thus improve data reliability of replicated distributed storage systems while maintaining a low repair cost. First, by introducing an additional replica state, critical blocks (with two or more lost replicas) have individual short timeouts while sick blocks (with only one lost replica) preserve the long timeouts. Second, by replacing the static check intervals for node states with adaptive ones, the check intervals and the identification time of critical blocks are further shortened, which improves data reliability. Meanwhile, due to the low ratio of critical blocks in all faulty blocks, the repair network traffic remains low. The results from our simulation and prototype implementation show that RICK improves data reliability of replicated distributed storage systems by a factor of up to 14 in terms of mean time to data loss. Meanwhile, the extra repair network traffic caused by RICK is less than 1.5 percent of the total network traffic for data repairs. Juntao Fang, Shenggang Wan, Ping Huang 0001, Changsheng Xie 0001, Xubin He |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Achieving High Reliability via Expediting the Repair of Critical Blocks in Replicated Storage SystemsabstractHigh reliability is critical to large data centers consisting of hundreds to thousands of storage nodes where node failures are not rare. Data replication is a typical technique deployed to achieve high reliability. When a node failure is detected, blocks with lost replicas are identified and recovered. Long timeouts are usually used for node failure detection. For blocks with one lost replica, the long timeouts can significantly reduce network traffic induced by data recovery. However, for blocks with two or more lost replicas, which can be caused by concurrent node failures that are not rare in large data centers, the long timeouts will result in a high risk of loss of these blocks. In this paper, we propose MFR to separate the identification of the blocks with two or more lost replicas from that of the blocks with one lost replica in a way that the identification of the blocks with two or more replicas can be accelerated while that of the blocks with one lost replica stays the same. Consequently, MFR can significantly improve data reliability while keeping the network traffic induced by data recovery stable. The results from our simulation and prototype implementation show that MFR improves the reliability of storage systems by a factor of up to 4.0 in terms of mean time to data loss. As blocks with two or more lost replicas are far fewer than blocks with one lost replica, the extra network traffic caused by MFR is less than 0.54% of total network traffic for data recovery. Juntao Fang, Shenggang Wan, Ping Huang 0001, Xubin He, Changsheng Xie 0001 |
SRDS | 1 |