EDBT 2026 Demo / reviewers in the wild / expert
Yunfeng Zhu
dblp:85/1368
· DBLP profile ↗
14ranked-venue papers
5as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Distributed systems · 66% Storage systems · 34% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 15 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Distributed systems
distributed database |
0.8 | 1 | 2024 | Timestamp as a Service, not an Oracle · Proc. VLDB Endow. 2024 |
Distributed systems › concurrency control
timestamping |
0.8 | 1 | 2024 | Timestamp as a Service, not an Oracle · Proc. VLDB Endow. 2024 |
Distributed systems › consensus
transaction ordering |
0.8 | 1 | 2024 | Timestamp as a Service, not an Oracle · Proc. VLDB Endow. 2024 |
Storage systems
erasure-coded storage |
0.4 | 2 | 2015 | Boosting Degraded Reads in Heterogeneous Erasure-Coded Storage Systems · IEEE Trans. Computers 2015 On the Speedup of Recovery in Large-Scale Erasure-Coded Storage Systems · IEEE Trans. Parallel Distributed Syst. 2014 |
Storage systems
storage reliability |
0.4 | 2 | 2014 | On the Speedup of Recovery in Large-Scale Erasure-Coded Storage Systems · IEEE Trans. Parallel Distributed Syst. 2014 Single Disk Failure Recovery forX-Code-Based Parallel Storage Systems · IEEE Trans. Computers 2014 |
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis |
0.3 | 1 | 2018 | Ms2lda.org: web-based topic modelling for substructure discovery in mass spectrometry · Bioinform. 2018 |
Bioinformatics and computational biology
metabolomics |
0.3 | 1 | 2018 | Ms2lda.org: web-based topic modelling for substructure discovery in mass spectrometry · Bioinform. 2018 |
Distributed systems
fault tolerance |
0.3 | 2 | 2024 | Timestamp as a Service, not an Oracle · Proc. VLDB Endow. 2024 On the Speedup of Recovery in Large-Scale Erasure-Coded Storage Systems · IEEE Trans. Parallel Distributed Syst. 2014 |
Distributed systems › fault tolerance
high availability |
0.2 | 1 | 2024 | Timestamp as a Service, not an Oracle · Proc. VLDB Endow. 2024 |
Storage systems
distributed storage |
0.2 | 1 | 2015 | Boosting Degraded Reads in Heterogeneous Erasure-Coded Storage Systems · IEEE Trans. Computers 2015 |
Storage systems › storage reliability
erasure coding |
0.2 | 1 | 2014 | Single Disk Failure Recovery forX-Code-Based Parallel Storage Systems · IEEE Trans. Computers 2014 |
Distributed systems › fault tolerance
failure recovery |
0.2 | 1 | 2014 | On the Speedup of Recovery in Large-Scale Erasure-Coded Storage Systems · IEEE Trans. Parallel Distributed Syst. 2014 |
Storage systems › distributed storage
parallel storage system |
0.2 | 1 | 2014 | Single Disk Failure Recovery forX-Code-Based Parallel Storage Systems · IEEE Trans. Computers 2014 |
Storage systems › erasure-coded storage
XOR-based erasure coding |
0.2 | 1 | 2014 | On the Speedup of Recovery in Large-Scale Erasure-Coded Storage Systems · IEEE Trans. Parallel Distributed Syst. 2014 |
Distributed systems › fault tolerance › resilience
node failure resilience |
0.1 | 1 | 2014 | On the Speedup of Recovery in Large-Scale Erasure-Coded Storage Systems · IEEE Trans. Parallel Distributed Syst. 2014 |
Methods — techniques the papers use, named apart from their topics
logical timestamping · 0.8bipartite architecture · 0.8topic modelling · 0.3latent dirichlet allocation · 0.3testbed experiments · 0.2simulation · 0.2greedy algorithm · 0.2trace-driven simulation · 0.2replace recovery algorithm · 0.2parallelized architecture · 0.2integer linear programming · 0.2hill-climbing · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Task Learning for OSA Detection and Sleep Staging via Multi-Scale ModelingabstractObstructive sleep apnea (OSA) and sleep fragmentation are closely linked physiological phenomena that play crucial roles in the diagnosis and management of sleep disorders. While numerous deep learning models have been developed for either OSA detection or sleep stage classification, few attempts have been made to address both tasks simultaneously. To this end, we propose MT-TASPPNet (Multi-Task Triple Atrous Spatial Pyramid Pooling Network), a unified multi-modal multi-task network that jointly performs automatic OSA event detection and sleep staging. The model integrates modality-specific feature extractors for EEG, ECG, and airflow signals, and employs Atrous Spatial Pyramid Pooling modules in both the modality-specific and shared representation pathways to capture multi-scale temporal-frequency patterns. Additionally, an EOG-guided prior mechanism is incorporated to enhance the discrimination of subtle sleep stages. We use a 3-min input window (1-min target with $\pm$ 1-min context) and evaluate our method on three large-scale datasets: SHHS1, SHHS2, and Sydney Sleep Biobank. The model achieves OSA detection accuracy between 0.798 and 0.884 (MF1: 0.772 to 0.821), and sleep staging accuracy between 0.776 and 0.834 (MF1: 0.735 to 0.749, $\mathcal {K}$: 0.697 to 0.77). Notably, the model maintains consistent performance despite data heterogeneity and individual variability. These results validate the stability and adaptability of MT-TASPPNet in clinical settings, paving the way for efficient and scalable multi-task sleep analysis systems. Zhiya Wang, Yunfeng Zhu, Jia Liu 0092, Peter A. Cistulli, Wei Chen 0015 |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Timestamp as a Service, not an OracleabstractWe present a logical timestamping mechanism for ordering transactions in distributed databases, eliminating the single point of failure (SPoF) that bother existing timestamp "oracles". The main innovation is a bipartite client-server architecture, where the servers do not communicate with each other. The result is a highly available timestamping "service" that guarantees the availability of time-stamps, unless half the servers are down at the same time. We study the fundamental needs of timestamping, and formalize its availability and correctness properties in a distributed setting. We then introduce the TaaS (timestamp as a service) algorithm, which defines a monotonic spacetime over multiple server clocks. We prove, mathematically: (i) Availability that the timestamps are always computable, provided any majority of the server clocks being observable; and (ii) Correctness that all the computed time-stamps must increase monotonically over time, even if some clocks become unobservable. We evaluate our algorithm by prototyping TaaS and benchmarking it against state of the art timestamp oracle in TiDB. Our experiment shows that TaaS is indeed immune to SPoF (as we have proven mathematically), while exhibiting a reasonable performance at the same order of magnitude with TiDB. We also demonstrate the stability of our bipartite architecture, by deploying TaaS across datacenters and showing its resilience to datacenter-level failures. Yishuai Li, Yunfeng Zhu |
Proc. VLDB Endow. | 2 |
| 2024 | Unsupervised Transfer Learning Approach With Adaptive Reweighting and Resampling Strategy for Inter-Subject EOG-Based Gaze Angle EstimationabstractGaze estimation based on electrooculograms (EOGs) has been widely explored. However, the inter-subject variability of EOGs still leaves a significant challenge for practical applications. It contributes to performance degradation when handling inter-subject issues. In this paper, an unsupervised transfer learning approach with an adaptive reweighting and resampling (ARR) strategy to fully consider individual variability is proposed for EOG-based gaze angle estimation. It allows quantifying domain shifts by leveraging the source-target similarities, reweighting and resampling the source data to retain relevant instances and disregard irrelevant instances during adaptation. Specifically, our proposed methodology first assesses the domain shifts via decomposing transformation matrices, which are estimated between the training subjects (denoted as multi-source domains) and the test subject (denoted as target domain). Then, the multi-domain shifts are assigned as weighted indicators to resample the multi-source domains for model training. Comparative experiments with several prevailing transfer learning methods including CORrelation ALignment (CORAL), Geodesic Flow Kernel (GFK), Joint Distribution Adaptation (JDA), Transfer component analysis (TCA), and Balanced distribution adaption (BDA) using two different normalization processes were conducted on a realistic scenario across 18 subjects. Experimental results demonstrate that the ARR strategy can significantly improve performance (mean absolute error (MAE) reduction: 7.0%, root mean square error (RMSE) reduction: 6.3%), outperforming the prevailing methods. Besides, the impacts of data diversity and data size on ARR strategy are further investigated. It exhibits that data size is more important than data diversity for EOG-based gaze angle estimation, and also presents the benefits of the ARR strategy for dealing with practical scenarios. Linkai Tao, Ruizhi Su, Yunfeng Zhu, Long Meng, Adili Tuheti, Feng Shu 0001, Wei Chen 0015, Chen Chen 0039 |
IEEE J. Biomed. Health Informatics | 4 |
| 2018 | Ms2lda.org: web-based topic modelling for substructure discovery in mass spectrometryabstractMOTIVATION: We recently published MS2LDA, a method for the decomposition of sets of molecular fragment data derived from large metabolomics experiments. To make the method more widely available to the community, here we present ms2lda.org, a web application that allows users to upload their data, run MS2LDA analyses and explore the results through interactive visualizations. RESULTS: Ms2lda.org takes tandem mass spectrometry data in many standard formats and allows the user to infer the sets of fragment and neutral loss features that co-occur together (Mass2Motifs). As an alternative workflow, the user can also decompose a data set onto predefined Mass2Motifs. This is accomplished through the web interface or programmatically from our web service. AVAILABILITY AND IMPLEMENTATION: The website can be found at http://ms2lda.org, while the source code is available at https://github.com/sdrogers/ms2ldaviz under the MIT license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Joe Wandy, Yunfeng Zhu, Justin J. J. van der Hooft, Rónán Daly, Michael P. Barrett, Simon Rogers |
Bioinform. | 2 |
| 2015 | POS: A Popularity-based Online Scaling scheme for RAID-structured storage systemsabstractThe ever-increasing demand of storage capability leads to scaling requirement in RAID-structured storage systems. Previous approaches to RAID scaling mainly focus on minimizing data migration, without considering the user-level application accesses. However, the mixed scaling I/Os and user accesses in practical systems will interfere with each other, which results in significant performance degradation of both the data migration time and the user response time. In this paper, we divide the whole storage space into multiple zones and measure the popularity (mainly using the metric of access frequency) of each zone. Based on the measured popularity, we propose an online scheme, namely Popularity-based Online Scaling (POS), to scale RAID-structured storage systems. The main idea of POS is to scale storage areas with high popularity first so as to better exploit workload locality. POS can efficiently alleviate the performance degradation of user response time and data migration time during the scaling process. It can be readily deployed atop various conventional RAID scaling approaches to improve their performance. To evaluate the performance of POS, we implement FastScale and FastScale with POS (POS-FS) in the same system. Through extensive benchmark studies on real-system workloads, we show that POS can efficiently reduce the response time to user requests and scaling I/Os and improve the sequentiality of data accesses. Si Wu 0003, Yinlong Xu 0001, Yongkun Li 0001, Yunfeng Zhu |
ICCD | 4 |
| 2015 | Even data placement for load balance in reliable distributed deduplication storage systemsabstractModern distributed storage systems often deploy deduplication to remove content-level redundancy and hence improve storage efficiency. However, deduplication inevitably leads to unbalanced data placement across storage nodes, thereby degrading read performance. This paper studies the load balance problem in the setting of a reliable distributed deduplication storage system, which deploys deduplication for storage efficiency and erasure coding for reliability. We argue that in such a setting, it is generally challenging to find a data placement that simultaneously achieves both read balance and storage balance objectives. To this end, we formulate a combinatorial optimization problem, and propose a greedy, polynomial-time Even Data Placement (EDP) algorithm, which identifies a data placement that effectively achieves read balance while maintaining storage balance. We further extend our EDP algorithm to heterogeneous environments. We demonstrate the effectiveness of our EDP algorithm under real-world workloads using both extensive simulations and prototype testbed experiments. In particular, our testbed experiments show that our EDP algorithm reduces the file read time by 37.41% compared to the baseline round-robin placement, and the reduction can further reach 52.11% in a heterogeneous setting. Yunfeng Zhu, Patrick P. C. Lee, Yinlong Xu 0001 |
IWQoS | 2 |
| 2015 | Boosting Degraded Reads in Heterogeneous Erasure-Coded Storage SystemsabstractDistributed storage systems provide large-scale data storage services, yet they are confronted with frequent node failures. To ensure data availability, a storage system often introduces data redundancy via replication or erasure coding. As erasure coding incurs significantly less redundancy overhead than replication under the same fault tolerance, it has been increasingly adopted in large-scale storage systems. In erasure-coded storage systems, degraded reads to temporarily unavailable data are very common, and hence boosting the performance of degraded reads becomes important. One challenge is that storage nodes tend to be heterogeneous with different storage capacities and I/O bandwidths. To this end, we propose FastDR, a system that addresses node heterogeneity and exploits I/O parallelism, so as to boost the performance of degraded reads to temporarily unavailable data. FastDR incorporates a greedy algorithm that seeks to reduce the data transfer cost of reading surviving data for degraded reads, while allowing the search of the efficient degraded read solution to be completed in a timely manner. We implement a FastDR prototype, and conduct extensive evaluation through simulation studies as well as testbed experiments on a Hadoop cluster with 10 storage nodes. We demonstrate that our FastDR achieves efficient degraded reads compared to existing approaches. Yunfeng Zhu, Patrick P. C. Lee, Yinlong Xu 0001 |
IEEE Trans. Computers | 1 |
| 2014 | Enhancing scalability in distributed storage systems with Cauchy Reed-Solomon codesabstractSystem scaling becomes essential and indispensable for distributed storage systems due to the explosive growth of data volume. As fault-protection is also a necessity in large-scale distributed storage systems, and Cauchy Reed-Solomon (CRS) codes are widely deployed to tolerate multiple simultaneous node failures, this paper studies the scaling of distributed storage systems with CRS codes. In particular, we formulate the scaling problem with an optimization model in which both the post-scaling encoding matrix and the data migration policy are assumed to be unknown in advance. To minimize the I/O overhead for CRS scaling, we first derive the optimal post-scaling encoding matrix under a given data migration policy, and then optimize the data migration process using the selected postscaling encoding matrix. Our scaling scheme requires the minimal data movement while achieving uniform data distribution. To validate the efficiency of our scheme, we implement it atop a networked file system. Extensive experiments show that our scaling scheme reduces 7.94% to 58.87%, and 39.52% on average, of the scaling time over the basic scheme. Si Wu 0003, Yinlong Xu 0001, Yongkun Li 0001, Yunfeng Zhu |
ICPADS | 4 |
| 2014 | Single Disk Failure Recovery forX-Code-Based Parallel Storage SystemsabstractIn modern parallel storage systems (e.g., cloud storage and data centers), it is important to provide data availability guarantees against disk (or storage node) failures via redundancy coding schemes. One coding scheme is X-code, which is double-fault tolerant while achieving the optimal update complexity. When a disk/node fails, recovery must be carried out to reduce the possibility of data unavailability. We propose an X-code-based optimal recovery scheme called minimum-disk-read-recovery (MDRR), which minimizes the number of disk reads for single-disk failure recovery. We make several contributions. First, we show that MDRR provides optimal single-disk failure recovery and reduces about 25 percent of disk reads compared to the conventional recovery approach. Second, we prove that any optimal recovery scheme for X-code cannot balance disk reads among different disks within a single stripe in general cases. Third, we propose an efficient logical encoding scheme that issues balanced disk read in a group of stripes for any recovery algorithm (including the MDRR scheme). Finally, we implement our proposed recovery schemes and conduct extensive testbed experiments in a networked storage system prototype. Experiments indicate that MDRR reduces around 20 percent of recovery time of the conventional approach, showing that our theoretical findings are applicable in practice. Silei Xu, Runhui Li, Patrick P. C. Lee, Yunfeng Zhu, Liping Xiang, Yinlong Xu 0001, John C. S. Lui |
IEEE Trans. Computers | 4 |
| 2014 | On the Speedup of Recovery in Large-Scale Erasure-Coded Storage SystemsabstractModern storage systems stripe redundant data across multiple nodes to provide availability guarantees against node failures. One form of data redundancy is based on XOR-based erasure codes, which use only XOR operations for encoding and decoding. In addition to tolerating failures, a storage system must also provide fast failure recovery to reduce the window of vulnerability. This work addresses the problem of speeding up the recovery of a single-node failure for general XOR-based erasure codes. We propose a replace recovery algorithm, which uses a hill-climbing technique to search for a fast recovery solution, such that the solution search can be completed within a short time period. We further extend the algorithm to adapt to the scenario where nodes have heterogeneous capabilities (e.g., processing power and transmission bandwidth). We implement our replace recovery algorithm atop a parallelized architecture to demonstrate its feasibility. We conduct experiments on a networked storage system testbed, and show that our replace recovery algorithm uses less recovery time than the conventional recovery approach. Yunfeng Zhu, Patrick P. C. Lee, Yinlong Xu 0001, Yuchong Hu, Liping Xiang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2013 | PHR: A Pipelined Heterogeneous Recovery for RAID6-Coded Storage SystemsabstractWith the rapid growth of data and the growing demand from users on the system performance, data availability has become the most important issue in large-scale storage systems. Due to the ability to provide space-optimal data redundancy to protect against node failures, erasure codes have seen widely deployment. To ensure data availability, it is crucial to recover node failures quickly. In this paper, we propose PHR, a pipelined heterogeneous recovery optimization for failure recovery in RAID-6 coded storage systems. Our PHR takes into account both I/O parallelism and node heterogeneity in practical storage systems, and returns an efficient recovery solution timely. We parallelize our PHR algorithm in a pipelined manner, so as to further improve failure recovery performance. With the quantitative simulation studies and extensive test bed experiments, we show our PHR significantly reduces recovery time and also user response time. Fang Niu, Yinlong Xu 0001, Yunfeng Zhu |
PDCAT | 3 |
| 2012 | A cost-based heterogeneous recovery scheme for distributed storage systems with RAID-6 codesabstractModern distributed storage systems provide large-scale, fault-tolerant data storage. To reduce the probability of data unavailability, it is important to recover the lost data of any failed storage node efficiently. In practice, storage nodes are of heterogeneous types and have different transmission bandwidths. Thus, traditional recovery solutions that simply minimize the number of data blocks being read may no longer be optimal in a heterogeneous environment. We propose a cost-based heterogeneous recovery (CHR) algorithm for RAID-6-coded storage systems. We formulate the recovery problem as an optimization model in which storage nodes are associated with generic costs. We narrow down the solution space of the model to make it practically tractable, while still achieving the global optimal solution in most cases. We implement different recovery algorithms and conduct testbed experiments on a real networked storage system with heterogeneous storage devices. We show that our CHR algorithm reduces the total recovery time of existing recovery solutions in various scenarios. Yunfeng Zhu, Patrick P. C. Lee, Liping Xiang, Yinlong Xu 0001, Lingling Gao |
DSN | 1 |
| 2012 | On the speedup of single-disk failure recovery in XOR-coded storage systems: Theory and practiceabstractModern storage systems stripe redundant data across multiple disks to provide availability guarantees against disk failures. One form of data redundancy is based on XOR-based erasure codes, which use only XOR operations for encoding and decoding. In addition to providing failure tolerance, a storage system must also provide fast failure recovery to avoid data unavailability. We consider the problem of speeding up the recovery of a single-disk failure for arbitrary XOR-based erasure codes. We address this problem from both theoretical and practical perspectives. We propose a replace recovery algorithm, which uses a hill-climbing technique to search for a fast recovery solution, such that the solution search can be completed within a short time period. We further implement our replace recovery algorithm atop a parallelized architecture to justify its practicality. We experiment our replace recovery algorithm and its parallelized implementation on a networked storage system testbed, and demonstrate that our replace recovery algorithm uses less recovery time than the conventional approach. Yunfeng Zhu, Patrick P. C. Lee, Yuchong Hu, Liping Xiang, Yinlong Xu 0001 |
MSST | 1 |
| 2011 | Dynamic Cascades with Bidirectional Bootstrapping for Action Unit Detection in Spontaneous Facial BehaviorabstractAutomatic facial action unit detection from video is a long-standing problem in facial expression analysis. Research has focused on registration, choice of features, and classifiers. A relatively neglected problem is the choice of training images. Nearly all previous work uses one or the other of two standard approaches. One approach assigns peak frames to the positive class and frames associated with other actions to the negative class. This approach maximizes differences between positive and negative classes, but results in a large imbalance between them, especially for infrequent AUs. The other approach reduces imbalance in class membership by including all target frames from onsets to offsets in the positive class. However, because frames near onsets and offsets often differ little from those that precede them, this approach can dramatically increase false positives. We propose a novel alternative, dynamic cascades with bidirectional bootstrapping (DCBB), to select training samples. Using an iterative approach, DCBB optimally selects positive and negative samples in the training data. Using Cascade Adaboost as basic classifier, DCBB exploits the advantages of feature selection, efficiency, and robustness of Cascade Adaboost. To provide a real-world test, we used the RU-FACS (a.k.a. M3) database of nonposed behavior recorded during interviews. For most tested action units, DCBB improved AU detection relative to alternative approaches. Yunfeng Zhu, Fernando De la Torre, Jeffrey F. Cohn, Yu-Jin Zhang |
IEEE Trans. Affect. Comput. | 1 |