VLDB 2026 Research / reviewers in the wild / expert
Zhiguang Chen 0001
dblp:75/7449-1
· DBLP profile ↗
94ranked-venue papers
11as first author
50since 2021 · last 2026
0000-0002-9318-5715ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 58 · 7 first-author · 30 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Software engineering, systems software and programming languages · 5 · 4 since 2021Computer networks · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bullet: Boosting GPU Utilization for LLM Serving via Dynamic Spatial-Temporal OrchestrationabstractModern large language model (LLM) serving systems confront inefficient GPU utilization due to the fundamental mismatch between compute-intensive prefill phase and memory-bound decode phase. While current practices attempt to address this by organizing these phases into hybrid batches, such solutions create an inefficient tradeoff that sacrifices either throughput or latency, leaving substantial GPU resources underutilized. For this, we identify two key root causes: 1) the prefill phase suffers from suboptimal compute utilization due to wave quantization and attention bottlenecks, and 2) hybrid batching disproportionately prioritizes latency over throughput, wasting both compute resources and memory bandwidth. To mitigate the issues, we present Bullet, a novel spatial-temporal orchestration system that eliminates these inefficiencies through fine-grained phase coordination. Bullet enables concurrent execution of prefill and decode requests, while dynamically provisioning GPU resources based on real-time performance modeling. By integrating SLO-aware scheduling and adaptive resource allocation, Bullet maximizes GPU utilization without compromising latency targets. Experimental evaluations on real-world workloads demonstrate that Bullet delivers 1.26× average throughput gains (up to 1.55×) over state-of-the-arts, while consistently meeting latency constraints. Zejia Lin 0001, Hongxin Xu, Guanyi Chen, Zhiguang Chen 0001, Yutong Lu, Xianwei Zhang 0001 |
ASPLOS (2) | 4 |
| 2026 | KirbyMM: Outer-Product Based Matrix Multiplication on ARMv9 Processor
Lanshu Huang, Zhiguang Chen 0001, Yutong Lu |
DATE | 3 |
| 2026 | RecDB: An LSM-Tree based Storage System for Training Large Recommendation Model in Low-Resource Scenarios
Qingyin Lin, Zhitao Chen, Yunling Chen, Zhiguang Chen 0001 |
EDBT | 5 |
| 2026 | Matrix‑PIC: Harnessing Matrix Outer-product for High‑Performance Particle‑in‑Cell SimulationsabstractParticle-in-Cell (PIC) simulations devote most cycles to particle-grid interactions, and their fine-grained atomic updates become a severe bottleneck on traditional many-core CPUs. The evolution of CPU architectures, particularly the integration of specialized Matrix Processing Units (MPUs) designed for efficient matrix outer-product operations, presents a paradigm shift and an opportunity to alleviate these bottlenecks. Capitalizing on this architectural advancement, this work focuses on adapting the critical current deposition step in PIC simulations to this new matrix-centric computational model. Yizhuo Rao, Xingjian Cui, Jiabin Xie, Shangzhi Pang, Guangnan Feng, Jinhui Wei, Zhiguang Chen 0001, Yutong Lu |
EuroSys | 7 |
| 2026 | POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and CommunicationabstractParticle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle–grid interaction bottlenecks and particle redistribution costs. Specifically, the particle–grid interaction computations have not taken full advantage of the emerging Matrix Processing Units (MPUs), the particle motion introduces irregular memory accesses, and the bulk-synchronous redistribution further destroys long-term data locality thereby limiting parallel efficiency. To address these inefficiencies, we present POLAR-PIC, a co-designed framework for large-scale PIC simulations that (i) reformulates Field Interpolation into an MPU-friendly outer-product form, (ii) maintains a physically ordered particle layout to preserve memory contiguity, and (iii) overlaps particle communication with Deposition to hide redistribution overhead. The evaluation on the pilot system of an Exascale supercomputer demonstrates that POLAR-PIC accelerates the entire particle-processing phase by up to 10.9 × in uniform plasma and 4.4 × in real-world laser-ion acceleration scenarios compared to the native WarpX reference pipeline on LX2. Ablation studies reveal that the speedups achieved by Interpolation and Deposition are 8.0 × and 13.2 × , respectively, and the asynchronous communication design sustains a \(99.1\%\) overlap ratio. In cross-platform comparisons, POLAR-PIC achieves \(13.2\%\) of theoretical peak efficiency on the CPU-based LS system, while WarpX reaches \(9.6\%\) on NVIDIA A800 GPUs. Notably, the scalability evaluation demonstrates that POLAR-PIC maintains \(67.5\%\) weak scaling efficiency on over 2 million cores under high-migration dynamic workloads, highlighting the importance of holistic co-design for future matrix-centric HPC systems. Yizhuo Rao, Xingjian Cui, Shangzhi Pang, Jiabin Xie, Guangnan Feng, Jinhui Wei, Languang Gao, Zhiguang Chen 0001, Yutong Lu |
HPDC | 10 |
| 2026 | TADS: Trend-Aware Dynamic Load Balancing for Large-Scale SNN Simulations with Delay-Sharded Graph InfrastructureabstractLarge-scale simulation of Spiking Neural Networks (SNNs) on supercomputers is pivotal for unraveling the mechanisms of brain function and advancing brain-inspired intelligence. However, efficiently mapping billions of neurons onto distributed nodes presents a significant challenge due to the heterogeneity of neuronal activities and complex, irregular network connectivity. While static partitioning strategies perform well in stable states, they often fail under metastable neurodynamics where theoretical models cannot accurately predict neuronal firing rates, leading to severe load imbalance. To address this, we propose TADS (Trend-Aware Dynamic load balancing with Delay-Sharded graph infrastructure), a framework tailored for large-scale SNN simulations. First, we identify the specific failure modes of static partitioning under metastable dynamics, establishing the necessity for runtime intervention. Second, we introduce a trend-aware dynamic load balancing strategy. By analyzing the temporal evolution of loads, this approach effectively distinguishes persistent imbalance from transient fluctuations, thereby avoiding unnecessary migrations caused by momentary jitter. Third, we design a delay-sharded graph infrastructure that leverages synaptic delays to parallelize graph modifications, significantly reducing the overhead associated with dynamic load balancing. Experimental results on the Tianhe-Xingyi supercomputer demonstrate that TADS effectively handles metastable scenarios, achieving up to 2.13 × speedup over state-of-the-art static partitioning methods, while sustaining 1.58 × performance improvement at the largest evaluated scale of 192 nodes. Shangzhi Pang, Yangle Zeng, Guangnan Feng, Zhiguang Chen 0001, Yutong Lu |
ICS | 5 |
| 2026 | Sin-PPI: Sub-linear Proteome-Wide PPI Screening via Orthogonal Manifold Learning
Xinqi Zeng, Huajian Mao, Zhiguang Chen 0001, Nong Xiao 0001 |
ISBRA (2) | 5 |
| 2026 | Superior F1-score: I/O feature driven algorithms for stream computing systems workload identification
Yuxiao Han, Zhiguang Chen 0001, Nong Xiao 0001 |
Frontiers Comput. Sci. | 5 |
| 2026 | Adaptive load balance scheme for the distributed control plane in SDN
Yuwen Zhou, Bangbang Ren, Zhi Zhou 0006, Xu Chen 0004, Zhiguang Chen 0001, Deke Guo |
Frontiers Comput. Sci. | 6 |
| 2025 | DynoInfer: Adaptive Resource Orchestration for LLM Inference on Resource-Constrained PCs
Yunling Chen, Qingyin Lin, Zhitao Chen, Zhiguang Chen 0001 |
Euro-Par (1) | 5 |
| 2025 | AuLoRA: Fine-Grained Loading and Computation Orchestration for Efficient LoRA LLM ServingabstractLoRA is a widely used Parameter-Efficient FineTuning (PEFT) technique for customizing pre-trained backbone models to specific tasks. Serving a backbone model with numerous LoRA adapters, known as multi-tenant LoRA serving, is a common scenario where different users utilize distinct LoRA adapters while sharing the same backbone model. To support more LoRA adapters simultaneously and improve efficiency, existing solutions dynamically load adapters from host memory and separate workloads into batched backbone model computation and adapter computation. However, they introduce complex data dependencies and necessitate careful coordination of loading and computation to enhance efficiency. We introduce AuLoRA, a multi-tenant LoRA serving system that achieves fine-grained orchestration of adapter loading and computation alongside backbone model execution. It optimizes both Time-to-First-Token (TTFT) and throughput by: 1) layer-wise-priority LoRA adapter loading, which reorganizes adapter loading by layer, to perform inference before adapters are fully loaded and overlaps adapter loading with backbone computation. 2) Intra-layer pipelined LoRA adapter execution, which loads LoRA adapters and performs computation in a pipelined manner, further hiding the adapter loading overhead. 3) Dynamic LoRA adapter batching, which explores the optimal LoRA adapter batching plan by comprehensively considering kernel launch overhead and modern hardware parallelism, improving computational efficiency. We compare AuLoRA with S-LoRA, a state-of-the-art multi-tenant LoRA serving system, and the results show that AuLoRA can achieve up to$3.03 \times$TTFT reduction and$1.27 \times$throughput improvement. Jiangsu Du, Zhiguang Chen 0001, Yutong Lu |
ICCD | 3 |
| 2025 | TD-Pipe: Temporally-Disaggregated Pipeline Parallelism Architecture for High-Throughput LLM InferenceabstractAs the model size continuously increases, pipeline parallelism shows great promise in throughput-oriented LLM inference due to its low demand on communications. However, imbalanced pipeline workloads and complex data dependencies in the prefill and decode phases result in massive pipeline bubbles and further severe performance reduction. Hongbin Zhang 0006, Taosheng Wei, Zhenyi Zheng, Jiangsu Du, Zhiguang Chen 0001, Yutong Lu |
ICPP | 5 |
| 2025 | LEAF: Latent Diffusion with Efficient Encoder Distillation for Aligned Features in Medical Image Segmentation
Zhiguang Chen 0001, Fudan Zheng |
MICCAI (6) | 3 |
| 2025 | StrategyAdapter: One-Shot Learning for Unseen-Domain Procedural Sequence Generation
Zhiguang Chen 0001, Nong Xiao 0001 |
PRCV (6) | 2 |
| 2025 | gLLM: Global Balanced Pipeline Parallelism Systems for Distributed LLMs Serving with Token ThrottlingabstractPipeline parallelism has emerged as a predominant approach for deploying large language models (LLMs) across distributed nodes, owing to its lower communication overhead compared to tensor parallelism. While demonstrating high throughput in request serving, pipeline parallelism often faces performance limitations caused by pipeline bubbles, which are primarily resulted from imbalanced computation delays across batches. Existing methods like Sarathi-Serve attempt to address this through hybrid scheduling of chunked prefill and decode tokens with a fixed token budget. However, such methods may still experience significant fluctuations, arising either from insufficient prefill tokens or uneven distribution of decode tokens, ultimately leading to computational imbalance. Tianyu Guo 0009, Xianwei Zhang 0001, Jiangsu Du, Zhiguang Chen 0001, Nong Xiao 0001, Yutong Lu |
SC | 4 |
| 2025 | coMtainer: Compilation-assisted HPC Container Images with Enhanced AdaptabilityabstractThe increasing interconnectivity of HPC systems has highlighted the need for efficient application migration across different environments. Containers, widely adopted for this purpose, simplify deployment but often fail to deliver optimal performance due to the separated build and execution container workflow. This leads to generic container images that miss out on system-specific software stack advantages, a challenge we define as the adaptability issue. Yuhao Gu, Haoquan Chen, Xianjie Chen, Jiangsu Du, Zhiguang Chen 0001, Nong Xiao 0001, Xianwei Zhang 0001, Yutong Lu |
SC | 5 |
| 2025 | HStencil: Matrix-Vector Stencil Computation with Interleaved Outer Product and MLAabstractStencil computations are fundamental to various HPC and intelligent computing applications, often consuming significant execution time. The emergence of specialized matrix units presents new opportunities to accelerate stencil computations. While scalable matrix compute units provide substantial computing horsepower, prior efforts fail to fully utilize the computing capabilities for stencils due to suboptimal matrix-unit utilization, limited instruction-level parallelism, and low cache hit rates. This paper introduces HStencil, a novel stencil computing framework utilizing matrix and vector units. HStencil addresses these challenges through three contributions: 1) microkernels that jointly leverage matrix and vector units to enhance hardware utilization; 2) fine-grained instruction scheduling with interleaved execution to enhance instruction-level parallelism; and 3) spatial prefetch to sustain high performance when working sets exceed cache capacity. Evaluations on representative benchmarks demonstrate that HStencil achieves maximum speedups of 1.81x – 5.76x over auto-vectorization across different CPU platforms, delivers 31% - 91% higher performance versus state-of-the-art methods. Jiabin Xie, Guangnan Feng, Xianwei Zhang 0001, Dan Huang 0001, Zhiguang Chen 0001, Yutong Lu |
SC | 6 |
| 2025 | Learning-based parallel acceleration for HaplotypeCallerabstractIn the genome analysis workflow, Genome Analysis Toolkit (GATK) HaplotypeCaller is a widely used variant calling tool designed to accurately identify single nucleotide polymorphisms (SNPs) and insertions/deletions (Indels) in samples. However, when processing large-scale datasets, HaplotypeCaller often faces the challenge of excessively long runtime. Parallelizing GATK HaplotypeCaller with data segmentation is an effective solution, but existing methods struggle to accurately estimate the computational complexity of each data block, leading to severe computational skew. This paper introduces a learning-based framework LPA (learning-based parallel acceleration), leveraging model to accurately predict the computational complexity of data. By employing adaptive data segmentation algorithms and Multi-Knapsack Problem (MKP) based task scheduling, LPA significantly alleviates computational skew. We evaluated LPA in multiple datasets, demonstrating that its execution speed is 30x–40x faster than HaplotypeCaller and 2x–5x faster than HaplotypeCallerSpark. LPA achieves a speedup of 1.3x–2x compared to similar methods. And LPA maintaining a high accuracy with over 99.9%, enhancing the efficiency and reliability of variant calling. The source code of LPA is publicly available at https://github.com/laixx9/LPA . Xiangxing Lai, Minguang Xiao, Lingling Weng, Zhiguang Chen 0001 |
BMC Bioinform. | 4 |
| 2025 | GPU acceleration for DNA sequence alignment algorithm and its application
Heming Zhong, Xiaojian Pan, Zengquang He, Haoling Wang, Dan Huang 0001, Zhiguang Chen 0001 |
CCF Trans. High Perform. Comput. | 6 |
| 2025 | WALSH: Write-Aggregating Log-Structured Hashing for Hybrid MemoryabstractPersistent memory (PM) brings important opportunities for improving data storage including the widely used hash tables. However, PM is not friendly to small writes, which causes existing PM hashes to suffer from high hardware write amplification. Hybrid memory offers the performance and concurrency of DRAM and the durability and capacity of PM, but existing hybrid memory hashes cannot deliver high performance, low DRAM footprint, and fast recovery at the same time. This paper proposes WALSH, a flat hash with novel log-structured separate chaining designs to optimize the performance while ensuring low DRAM footprint and fast recovery. To address the overhead of hash resizing and garbage collection (GC), WALSH further proposes partial resizing/GC mechanisms and a 4-phase protocol for concurrent hash operations. As a result, WALSH is the first flat index for hybrid memory with embedded write aggregation ability. A comprehensive evaluation shows that WALSH substantially outperforms state-of-the-art hybrid memory hashes; e.g., its insert throughput is up to 2.4X that of related works while saving more than 87% of DRAM. WALSH also provides efficient recovery; e.g., it can recover a dataset with 1 billion objects in just a few seconds. Yongfeng Wang, Zhiguang Chen 0001, Yutong Lu, Ming Zhao 0002 |
ACM Trans. Storage | 3 |
| 2025 | Critique of "Productivity, Portability, Performance Data-Centric Python" by SCC Team From Sun Yat-sen UniversityabstractIn SC21, Ziogas et al. proposed Data-Centric (DaCe) Python. It attains high performance and portability, and further extends the original productivity of Python. This paper analyzes the reproducibility of the DaCe paper as part of the SC22 Student Cluster Competition (SCC). The reproduction experiments are conducted on the Azure CycleCloud. Different from the DaCe paper, we use AMD EPYC 7V73X processors for CPU-based experiments. We successfully reproduce most of the results of the DaCe paper. The remaining results are also explainable. Tengyang Zheng, Tianxing Yang, Siran Liu, Shengyou Lu, Guangnan Feng, Zhiguang Chen 0001, Dan Huang 0001 |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2024 | SparkTH: NUMA-Aware and I/O-Efficient Spark for ARM-Based Many-Core Supercomputing SystemabstractARM-based many-core processors in supercomputers enable exascale-level data analysis by leveraging the massive number of cores within a single chip. However, most existing big data frameworks, designed for distributed environments, fail to effectively harness the full potential of advanced supercomputing systems. Two significant issues impeding data processing efficiency are the remote memory access between NUMA nodes and the disregard for architectures within compute nodes. Furthermore, the explosive growth of intermediate data results in a severe mismatch between I/O capabilities and computation performance due to frequent read and write operations. In this paper, we present SparkTH, a NUMA-aware and I/O-efficient framework for big data processing on many-core supercomputing systems. SparkTH incorporates a NUMA resource management layer to fully utilize many-core resources and cuts the number of I/O operations in half in parallel file systems. We evaluated SparkTH on two ARM-based many-core systems, and it outperformed the original Spark by up to 2.1× on typical big data benchmarks and 8.7× on scientific computing applications. Minguang Xiao, Zhiguang Chen 0001, Yutong Lu |
ISPA | 3 |
| 2024 | ATM: Area-based Partition and Topology-aware Mapping for Large-scale SNN SimulationabstractSpiking Neural Network (SNN) is an effective tool for the simulation of neuronal dynamics as well as the understanding of brain structure and functions. However, scaling up SNN for large-scale simulations poses significant computational demands that necessitate the supercomputers. The advent of distributed simulation introduces the requirement of SNN partition and process mapping, which becomes a critical challenge in the context of large-scale distributed SNN simulations. In this paper, we propose an Area-based partition and Topology-aware process Mapping (ATM) strategy to balance the computation workload while coping with the heterogeneity of communication interconnect. We first model the computation workload and communication volume of the SNN simulation according to its biological features. Based on this model, we design an area-based SNN partition strategy to balance the computation workload. Subsequently, we introduce a topology-aware strategy for process mapping, Bottleneck Fulfilling (BF), tailored specifically for collective communication paradigms. Experiments are conducted on an HPC cluster with a multi-area model of the marmoset brain. The results demonstrate that the proposed approach achieves up to 2.2x speedup compared with the baseline on 290 compute nodes. Yangle Zeng, Guangnan Feng, Zhiguang Chen 0001, Yutong Lu, Nong Xiao 0001 |
ISPA | 3 |
| 2024 | Stable Diffusion Segmentation for Biomedical Images with Single-Step Reverse Process
Zhiguang Chen 0001, Zhonghao Yan, Weijiang Yu, Fudan Zheng |
MICCAI (8) | 2 |
| 2024 | Liger: Interleaving Intra- and Inter-Operator Parallelism for Distributed Large Model InferenceabstractDistributed large model inference is still in a dilemma where balancing cost and effect. The online scenarios demand intraoperator parallelism to achieve low latency and intensive communications makes it costly. Conversely, the inter-operator parallelism can achieve high throughput with much fewer communications, but it fails to enhance the effectiveness. Jiangsu Du, Jinhui Wei, Jiazhi Jiang, Shenggan Cheng, Dan Huang 0001, Zhiguang Chen 0001, Yutong Lu |
PPoPP | 6 |
| 2024 | Extreme-scale Direct Numerical Simulation of Incompressible Turbulence on the Heterogeneous Many-core SystemabstractDirect numerical simulation (DNS) is a technique that directly solves the fluid Navier-Stokes equations with high spatial and temporal resolutions, which has driven much research regarding the nature of turbulence. For high-Reynolds number (Re) incompressible turbulence of particular interest, where the nondimensional Re characterizes the flow regime, the application of DNS is hindered by the fact that the numerical grid size (i.e., the memory requirement) scales with Re3, while the overall computational cost scales with Re4. Recent studies have shown that developing efficient parallel methods for heterogeneous many-core systems is promising to solve this computational challenge. Jiabin Xie, Guangnan Feng, Junxuan Feng, Zhiguang Chen 0001, Yutong Lu |
PPoPP | 5 |
| 2024 | Topo: Towards a fine-grained topological data processing framework on Tianhe-3 supercomputer
Yutong Lu, Zhuo Tang, Dan Huang 0001, Zhiguang Chen 0001 |
J. Parallel Distributed Comput. | 6 |
| 2024 | Exploring low-resource medical image classification with weakly supervised prompt learning
Fudan Zheng, Jindong Cao, Weijiang Yu, Zhiguang Chen 0001, Nong Xiao 0001, Yutong Lu |
Pattern Recognit. | 4 |
| 2024 | IncrCP: Decomposing and Orchestrating Incremental Checkpoints for Effective Recommendation Model TrainingabstractTraining large models for modern recommendation systems requires a substantial number of computational devices and extended periods. Since it is essential to store model checkpoints throughout the training progress for accuracy debugging or mitigating potential failures, checkpointing systems are widely used. However, given that recommendation models can scale to hundreds of gigabytes or more, existing solutions often introduce significant overhead in terms of both storage and I/O. In this paper, we present IncrCP, a checkpointing system specifically designed for recommendation models. Given that only a small fraction of model parameters are modified in each iteration, IncrCP creatively leverages the incremental checkpointing strategy and overcomes the inherent slow recovery problem. To support recovering all states throughout the training process, while also ensuring efficient storage utilization and rapid recovery, IncrCP proposes the 2-D chunk approach. It proactively records changed parameters in the training process as well as their indexes, extracts parameters according to duplicated indexes as independent chunk files, and then orchestrates these chunks in the 2-dimensional linked list. In this way, IncrCP achieves fast recovery by loading less unnecessary parameters and performing less deduplication during recovery. Furthermore, IncrCP includes a selective extraction approach to reduce I/O by avoiding worthless extractions and a concatenate approach to reduce random disk access when recovery. Evaluations show that IncrCP achieves up to 6.6× recovery speedup compared to the naive incremental strategy and saves storage space by 60.4% with slight overhead compared to another recovery-friendly strategy. Qingyin Lin, Jiangsu Du, Zhiguang Chen 0001, Nong Xiao 0001 |
Proc. VLDB Endow. | 4 |
| 2024 | DumpKV: Learning based lifetime aware garbage collection for key value separation in LSM-treeabstractKey-value separation is used in LSM-tree to store large values in separate log files to reduce write amplification but requires garbage collection to recycle invalid values. Existing LSM-tree typically adopts a static policy to recycle obsolete values, struggling to achieve low write amplification as it is challenging to predefine the static parameters for garbage collection. In this work we propose DumpKV, a learning-based lifetime-aware garbage collection mechanism which achieves lower write amplification. DumpKV trains a machine learning model based on the access history of keys and accordingly uses the lightweight model to predict the lifetime of each key, where the predicted lifetime can be used to guide the garbage collection. To reduce the interference to write throughput introduced by garbage collection, DumpKV conducts feature collection during L0-L1 compaction, leveraging the fact that LSM-tree is small under KV separation. Experimental results show that DumpKV reduces GC write size by 25.7%-53.3% in real-world workloads and 19%-65% in synthetic workloads compared to baseline key-value separation LSM-tree KV stores with small feature storage overhead. Zhutao Zhuang, Zhiguang Chen 0001, Xinqi Zeng |
Proc. VLDB Endow. | 2 |
| 2024 | Contrastive Transformer Learning With Proximity Data Generation for Text-Based Person SearchabstractGiven a descriptive text query, text-based person search (TBPS) aims to retrieve the best matched target person from an image gallery. Such a cross-modal retrieval task is quite challenging due to significant modality gap, fine-grained differences and insufficiency of annotated data. To better align the two modalities, most existing works focus on introducing sophisticated network structures and auxiliary tasks, which are complex and hard to implement. In this paper, we propose a simple yet effective dual Transformer model for text-based person search. By exploiting a hardness-aware contrastive learning strategy, our model achieves state-of-the-art performance without any special design for local feature alignment or side information. Moreover, we propose a proximity data generation (PDG) module to automatically produce more diverse data for cross-modal training. The PDG module first introduces an automatic generation algorithm based on a text-to-image diffusion model, which generates new text-image pair samples in the proximity space of original ones. Then it combines approximate text generation and feature-level mixup during training to further strengthen the data diversity. The PDG module can largely guarantee the reasonability of the generated samples that are directly used for training without any human inspection for noise rejection. It improves the performance of our model significantly, providing a feasible solution to the data insufficiency problem faced by such fine-grained visual-linguistic tasks. Extensive experiments on two popular datasets of the TBPS task (i.e., CUHK-PEDES and ICFG-PEDES) show that the proposed approach outperforms state-of-the-art approaches evidently, e.g., improving by 3.88%, 4.02%, 2.92% in terms of Top1, Top5, Top10 on CUHK-PEDES. Hefeng Wu, Tianshui Chen, Zhiguang Chen 0001, Liang Lin 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Accurately Identifying Muscle-Invasive Bladder Cancer from MRI via Weakly Supervised LearningabstractBladder cancer (BCa) is one of the most common malignancies in the world, which can be categorized into muscleinvasive (MIBC) and non-muscle-invasive (NMIBC). These two types of BCa must be treated differently, and thus it is essential to correctly distinguish MIBC and NMIBC patients preoperatively for adopting different treatment methods accordingly. Currently, the two types can be distinguished through MRI images by radiologists, but manual inspection is time and labor-consuming. Existing machine learning based methods attempt to free radiologists from manual inspection. However, they fail to take full advantage of image features and always require extra laborious refined manual labeling in addition to the classification labels. In this study, we propose a Tumor Staging and Localization Network (TSLNet) to perform preoperative non-invasive assessment of muscle invasion of BCa, which can automatically distinguish MIBC patients from NMIBC patients based on MRI T2-weighted images of BCa. The model adopts the weakly supervised learning method. Specifically, self-produced guidance is used as pixellevel segmentation pseudo labels for auxiliary supervision to extract basic features, and location-recognition based fine-grained image classification technology and inexact consistency labels are used for auxiliary supervision to extract fine-grained features. Moreover, the model can visualize the critical regions of the lesions, which can provide practical reference and a basis for clinicians’ clinical diagnosis. Experimental results show that the model achieves high AUC, accuracy, specificity, sensitivity, and F1-score, which is comparable to experienced clinicians. Fudan Zheng, Yuedong Yang, Tianxin Lin, Shaoxu Wu, Yutong Lu, Zhiguang Chen 0001, Huiying Zhao |
BIBM | 8 |
| 2023 | LazySort: A customized sorting algorithm for non-volatile memory
Yang Liu 0259, Zhiguang Chen 0001, Nong Xiao 0001 |
Inf. Sci. | 4 |
| 2023 | Securing the Ethereum from Smart Ponzi Schemes: Identification Using Static FeaturesabstractMalware detection approaches have been extensively studied for traditional software systems. However, the development of blockchain technology has promoted the birth of a new type of software system–decentralized applications. Composed of smart contracts, a type of application that implements the Ponzi scheme logic (called smart Ponzi schemes) has caused irreversible loss and hindered the development of blockchain technology. These smart contracts generally had a short life but involved a large amount of money. Whereas identification of these Ponzi schemes before causing financial loss has been significantly important, existing methods suffer from three main deficiencies, i.e., the insufficient dataset, the reliance on the transaction records, and the low accuracy. In this study, we first build a larger dataset. Then, a large number of features from multiple views, including bytecode, semantic, and developers, are extracted. These features are independent of the transaction records. Furthermore, we leveraged machine learning methods to build our identification model, i.e., Mul ti-view Cas cade Ensemble model (MulCas). The experiment results show that MulCas can achieve higher performance and robustness in the scope of our dataset. Most importantly, the proposed method can identify smart Ponzi scheme at the creation time. Zibin Zheng, Weili Chen, Zhiguang Chen 0001, Yutong Lu |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | Full-Stack Optimizing Transformer Inference on ARM Many-Core CPUabstractThe past several years have witnessed tremendous success of transformer models in natural language processing (NLP), and their current landscape is increasingly diverse. Although GPU gradually becomes the dominating workhorse and de facto standard for deep learning, there are still many scenarios where using CPU remains a prevalent choice.Recently, ARM many-core processor starts emigrating to cloud computing and high-performance computing, which is promising to deploy transformer inference. In this paper, we identify several performance bottlenecks of existing inference runtime on many-core CPU including low-core usage, isolated thread configuration, inappropriate implementation of general matrix multiply (GEMM), and redundant computations for variable-length inputs. To tackle these problems, full-stack optimizations are conducted for these challenges from service level to operator level. We explore multi-instance parallelization at the service level to improve CPU core usage. To improve parallel efficiency of the inference runtime, we design NUMA-aware thread scheduling and a look-up table for optimal parallel configurations. The GEMM implementation is tailored for some critical modules to exploit the characteristics of transformer workload. To eliminate redundant computations, a novel storage format is designed and implemented to pack sparse data and a load balancing strategy is proposed for tasks with different sparsity. Experiments show that our implementation can outperform existing solutions by 1.1x to 6x with fixed-length inputs. For variable-length inputs, it achieves 1.9x to 8x speedups on different ARM many-core processors. Jiazhi Jiang, Jiangsu Du, Dan Huang 0001, Zhiguang Chen 0001, Yutong Lu, Xiangke Liao |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | An SIMD-Accelerated Metadata Management Scheme for Persistent Memory File SystemsabstractPersistent memory (PM) offers byte-addressable persistence with high random access performance close to DRAM. The special characteristics of PM have brought new opportunities and challenges to design file systems. Intuitively, file systems are IO-intensive and the computation overhead is negligible. Whereas, PM dramatically improves IO performance and we observe that the computation overhead of metadata operations in PM file systems is becoming increasingly non-negligible. Furthermore, with the heavy computation overhead in metadata operations, the CPU is easy to be saturated under a highly concurrent workload. Fortunately, due to abundant computation resources, SIMD technology provides potential opportunities to accelerate metadata operations for PM file systems. In this paper, we present an SIMD-accelerated metadata management scheme for PM file systems. Specifically, we design the SIMD-aware data structures and algorithms involved in metadata operations for PM file systems to accelerate metadata operations. In addition, to take the full performance of SIMD and leverage the compatibility of SIMD instructions and PM, we perform operations on PM directly to eliminate the overhead of data interaction between PM and DRAM. We implement a prototype called SPFS, and our evaluation demonstrates that SPFS can outperform other tested PM file systems in a variety of test scenarios. Zejie Hu, Jarvan Law, Zhiguang Chen 0001, Nong Xiao 0001 |
CCGRID | 3 |
| 2022 | Characterizing and Optimizing Hybrid DRAM-PM Main Memory System with Application AwarenessabstractPersistent memory (PM) has always been used in combination with DRAM to configure hybrid main memory systems that can obtain both the high performance of DRAM and large capacity of PM. There are critical management challenges in data placement, memory concurrency and workload scheduling for the concurrent execution of multiple application workloads. But the non-negligible performance gap between DRAM and PM makes the existing application-agnostic management strategies inefficient in reaching the full potential of hybrid memory. In this paper, we propose a series of application aware optimization strategies, including application aware data placement, adaptive thread allocation and inter-application interference avoiding, to improve the concurrent performance of different application workloads on hybrid memory. Finally, we provide the performance evaluation for our application aware solutions on real hybrid memory hardware with some comprehensive benchmark suites. Our experimental results show that the duration of multi-application concurrent execution on hybrid memory can be reduced by at most 60.7% for application aware data placement, 37.7% for adaptive thread allocation and 34.8% for workload scheduling with inter-application interference avoiding, respectively. And the additive effects of all these three optimization methods can reach 62.8% performance improvement with negligible overheads. Yongfeng Wang, Yinjin Fu, Zhiguang Chen 0001, Nong Xiao 0001 |
DATE | 4 |
| 2022 | SpacKV: A Pmem-Aware Key-Value Separation Store Based on LSM-Tree
Xuran Ge, Yang Liu 0259, Lizhou Wu, Zhutao Zhuang, Zhiguang Chen 0001, Nong Xiao 0001 |
NPC | 7 |
| 2022 | TopKmer: Parallel High Frequency K-mer Counting on Distributed Memory
Mocheng Li, Zhiguang Chen 0001, Nong Xiao 0001, Luo Xi, Tao Chen 0013 |
NPC | 2 |
| 2022 | A tail-tolerant cloud storage scheduling based on precise periodicity detectionabstractAbstract Cloud storage is a fundamental component of the cloud computing system, which significantly affects the overall performance and quality of service of the cloud. Cloud storage servers face the challenge of imbalanced workloads. According to our observations on the time series generated by cloud storage, we found that the imbalance workloads will dramatically increase the tail latency of data access in the multi-tenant scenario. The intuitive solution is to periodicity detect the imbalance storage nodes and re-balance the loads. However, there are four challenges to accurately detect load of storage in the cloud with multiple tenants since the load may change frequently in cloud. This paper proposes PrecisePeriod, a precise periodicity detection algorithm customized for multi-tenant cloud storage. It removes outliers through data preprocessing, employs the discrete wavelet transform to remove high-frequency noise while keeping frequency domain information, computes the candidate periodicity queue using the autocorrelation function, and determines precise period through periodicity verification. Then, we design a cloud storage load balancing scheduling strategy based on PrecisePeriod, and the evaluation shows that the PrecisePeriod scheduling significantly reduces tail latency while only bringing $$1-2\%$$ 1-2% overhead. Yuxiao Han, Jia Ma, Nong Xiao 0001, Yutong Lu, Zhiguang Chen 0001 |
CCF Trans. High Perform. Comput. | 7 |
| 2022 | Optimizing data query performance of Bi-cluster for large-scale scientific data in supercomputers
Xia Liao, Yixian Shen, Shengguo Li, Yutong Lu, Yufei Du, Zhiguang Chen 0001 |
J. Supercomput. | 6 |
| 2022 | Design and Simulation of Content-Aware Hybrid DRAM-PCM Memory SystemabstractPhase Change Memory (PCM) can directly connect persistent memory to main memory bus, while it achieves high read throughput and low standby power, the critical concerns are its poor write performance and limited durability. A naturally in-spired design is the hybrid memory architecture that fuses DRAM and PCM, so as to exploit the positive aspects of both types of memory. Unfortunately, existing solutions are seriously challenged by the limited main memory size, which is the primary bottleneck of in-memory computing. In this paper, we introduce a novel Content Aware Hybrid DRAM-PCM memory system framework—CAHRAM, which exploits deduplication to improve line sharing with high memory efficiency. It reduces write traffic to hybrid memory by removing unnecessary duplicate line writes, thereby further enhancing the write endurance of PCM. And it also substantially extends available free memory space by coalescing redundant lines in hybrid memory. We also design a reference-based page migration technique to minimize the access overheads caused by the performance gap between DRAM and PCM. Compared with the state-of-the-art in a hybrid memory simulator, our experiment results show that CAHRAM can achieve the highest I/O performance and the longest PCM lifetime with the competitive efficiencies in space and energy. Yinjin Fu, Yutong Lu, Zhiguang Chen 0001, Nong Xiao 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | Multi-Layer Networks for Ensemble Precipitation Forecasts PostprocessingabstractThe postprocessing method of ensemble forecasts is usually used to find a more precise estimate of future precipitation, because dynamic meteorology models have limitations in fitting fine-grained atmospheric processes and precipitation is driven more often by smaller-scale processes, while ensemble forecasts can hit this precipitation at times. However, the pattern of these hits cannot be easily summarized. The existing objective postprocessing methods tend to extend the rain area or false alarm the precipitation intensity categories. In this work, we introduce a multi-layer structure to simultaneously reduce the bias in forecast ensembles output by meteorology models and merge them to a quality deterministic (single-valued) forecast using cross-grid information, which differs quite dramatically from the previous statistical postprocessing method. The multi-layer network is designed to model the spatial distribution of future precipitation of different intensity categories(IC-MLNet). We provide a comparison of IC-MLNet to simple average as well as another two state-of-the-art ensemble quantitative precipitation forecasts (QPFs) postprocessing approaches over both single-model and multi-model ensemble forecasts datasets from TIGGE. The experimental results indicate that our model achieves superior performance over the compared baselines in precipitation amount prediction as well as precipitation intensities categories prediction. Fengyang Xu, Guanbin Li, Yunfei Du 0001, Zhiguang Chen 0001, Yutong Lu |
AAAI | 4 |
| 2021 | Modality-shared MRI Image Translation Based on Conditional GANabstractMultimodal MRI images are often necessary for precise clinical diagnosis and the development of high-performance intelligent medical image systems. However, it is expensive and often difficult to obtain sufficient registered multimodal MRI images. One way to alleviate this issue is MRI image translation which aims to generate images of another modality based on images of one modality. This paper proposes a novel image translation framework that can translate images of any MRI modality to any other one. The core part of the framework is a conditional generative adversarial network (CGAN), where unlike commonly used conditions such as one-hot vectors, an image-specific condition can be generated based on the gradient-weighted class activation mapping (Grad-CAM), which can highlight the parts of the input image that need more attention during the translation process. In addition, to translate the lesion information in the MRI images, a modality-shared segmentation model is proposed to extract the lesion information, which is then novelly embedded into the image translation process. Comprehensive experiments on different multi-modality MRI datasets demonstrated the effectiveness of the proposed translation approach, achieving better performance compared to state-of-the-art methods. Chufu Deng, Zhiguang Chen 0001, Wanqi Su, Yili Qu |
BIBM | 2 |
| 2021 | A NUMA-Aware Parallel Truss Decomposition Algorithm for Large Scale Graphs
Zhebin Mou, Nong Xiao 0001, Zhiguang Chen 0001 |
ICA3PP (2) | 3 |
| 2021 | Optimizing Massively Parallel Winograd Convolution on ARM ProcessorabstractConvolution Neural Network (CNN) has gained a great success in deep learning applications and been accelerated by dedicated convolutional algorithms. Winograd-based algorithm can greatly reduce the number of arithmetic operations required in convolution. However, our experiments show that existing implementations in deep learning libraries cannot achieve expected parallel performance on ARM manycore CPUs with last-level cache (LLC). Compared to multicore processor, ARM manycore CPUs have more cores, more NUMA nodes and the parallel performance is more easily restricted by memory bandwidth, cache contention, NUMA configuration and etc. In this paper, we propose an optimized implementation for single-precision Winograd-based algorithm on ARM manycore CPUs. Our algorithm adjusts the data layout according to the input shape and is optimized for the characteristics of ARM processor, thus reducing the matrix transformation overhead and achieving high arithmetic intensity. We redesign the parallel algorithm for Winograd-based convolution to achieve a more efficient implementation for manycore CPUs. The experimental results with 32 cores show that for modern ConvNets, our implementation achieves speedups ranging from 3 × to 5 × over the state-of-the-art Winograd-based convolution on ARM processor. Even conducted on a set of convolutional benchmarks executing on a 128-core system with 4 NUMA nodes, the results show that our implementation can also achieve better performance than existing implementations on ARM processor. Dan Huang 0001, Zhiguang Chen 0001, Yutong Lu |
ICPP | 3 |
| 2021 | BDCuckoo: an Efficient Cuckoo Hash for Block Device
Xianqi Zheng, Jia Ma, Zhiguang Chen 0001 |
NPC | 4 |
| 2021 | Unsupervised Domain Adaptation for 3D Medical Image with High Efficiency
Chufu Deng, Kuilin Li, Zhiguang Chen 0001 |
PAKDD (1) | 3 |
| 2021 | DeepPE: Emulating Parameterization in Numerical Weather Forecast Model Through Bidirectional Network
Fengyang Xu, Wencheng Shi, Yunfei Du 0001, Zhiguang Chen 0001, Yutong Lu |
ECML/PKDD (5) | 4 |
| 2021 | A GPU-Accelerated In-Memory Metadata Management Scheme for Large-Scale Parallel File Systems
Zhiguang Chen 0001, Yongfeng Wang, Yutong Lu |
J. Comput. Sci. Technol. | 1 |
| 2020 | An End-to-end Oxford Nanopore Basecaller Using Convolution-augmented TransformerabstractThe following topics are dealt with: learning (artificial intelligence); diseases; medical image processing; molecular biophysics; genetics; medical computing; feature extraction; cancer; genomics; proteins. Xuan Lv, Zhiguang Chen 0001, Yutong Lu, Yuedong Yang |
BIBM | 2 |
| 2020 | GramFS: The Graph Model-based Namespace Management of Large-scale Distributed File Systems
Hongbo Li 0007, Zhiguang Chen 0001, Nong Xiao 0001 |
HotStorage | 3 |
| 2020 | Synthesis of Registered Multimodal Medical Images with Lesions
Yili Qu, Wanqi Su, Xuan Lv, Chufu Deng, Yutong Lu, Zhiguang Chen 0001, Nong Xiao 0001 |
ICANN (1) | 7 |
| 2020 | CP-GAN: Context Pyramid Generative Adversarial Network for Speech EnhancementabstractThe topic of speech enhancement has been largely improved recently, especially with the development of generative adversarial networks (GANs). However prior methods simply follow the GAN architectures from computer vision tasks without specific designs for the speech enhancement according to the audio characteristics (i.e., different granularity context), which may leave noise points in some segments or disturb the contents of the original audio. In this work, we make the first attempt to explore the global and local speech features for coarse-to-fine speech enhancement and introduce a Context Pyramid Generative Adversarial Network (CPGAN), which contains a densely-connected feature pyramid generator and a dynamic context granularity discriminator to better eliminate audio noise hierarchically. Extensive experiments demonstrate that our CP-GAN effectively achieves state-of-the-art speech enhancement results and boosts the performance of more high-level speech tasks including automatic speech recognition and speaker recognition. Xiaodan Liang, Zhiguang Chen 0001 |
ICASSP | 4 |
| 2020 | Enhance Generative Adversarial Networks By Wavelet Transform To Denoise Low-Dose Ct ImagesabstractComputed Tomography (CT) has been widely used in clinical diagnosis, while its potential risk of X-ray radiation has attracted serious public concerns. Reconstructing high-quality images from low-dose CT devices is a promising solution. Whereas, existing methods mostly relied on the raw data of devices, and cannot be shared among different device suppliers. Inspired by the powerful learning ability of GAN and the structural information extraction ability of wavelet transform, we propose to combine the two together and design the WT-GAN, which extracts structure and noise information by wavelet transform and generates high-quality images by GAN. The two technologies are incorporated with each other by our well-designed loss functions. Experimental results show that the proposed WT-GAN achieves superior performance and can efficiently extract the noise while retaining the texture details. Furthermore, the WT-GAN is a postprocessing method imposed on full-size images, thus it is easy to integrate into any CT systems. Wanqi Su, Yili Qu, Chufu Deng, Fudan Zheng, Zhiguang Chen 0001 |
ICIP | 6 |
| 2020 | Phishing Scam Detection on Ethereum: Towards Financial Security for Blockchain EcosystemabstractIn recent years, blockchain technology has created a new cryptocurrency world and has attracted a lot of attention. It also is rampant with various scams. For example, phishing scams have grabbed a lot of money and has become an important threat to users' financial security in the blockchain ecosystem. To help deal with this issue, this paper proposes a systematic approach to detect phishing accounts based on blockchain transactions and take Ethereum as an example to verify its effectiveness. Specifically, we propose a graph-based cascade feature extraction method based on transaction records and a lightGBM-based Dual-sampling Ensemble algorithm to build the identification model. Extensive experiments show that the proposed algorithm can effectively identify phishing scams. Weili Chen, Xiongfeng Guo, Zhiguang Chen 0001, Zibin Zheng, Yutong Lu |
IJCAI | 3 |
| 2020 | Pacon: Improving Scalability and Efficiency of Metadata Service through Partial ConsistencyabstractTraditional distributed file systems (DFS) use centralized service to manage metadata. Many studies based on this centralized architecture enhanced metadata processing capability by scaling the metadata server cluster, which is however still difficult to keep up with the growing number of clients and the increasingly metadata-intensive applications. Some solutions abandoned the centralized metadata service and improved scalability by embedding a private metadata service in an HPC application, but these solutions are suitable for only some specific applications and the absence of global namespace makes data sharing and management difficult. This paper addresses the shortcomings of existing studies by optimizing the consistency model of client- side metadata cache for the HPC scenario using a novel partial consistency model. It provides the application with strong consistency guarantee for only its workspace, thus improving metadata scalability without adding hardware or sacrificing the versatility and manageability of DFSes. In addition, the paper proposes batch permission management to reduce path traversal overhead, thereby improving metadata processing efficiency. The result is a library (Pacon) that allows existing DFSes to achieve partial consistency for scalable and efficient metadata management. The paper also presents a comprehensive evaluation using intensive benchmarks and representative application. For example, in file creation, Pacon improves the performance of BeeGFS by more than 76.4 times, and outperforms the state-of-the-art metadata management solution (IndexFS) by more than 4.6 times. Yutong Lu, Zhiguang Chen 0001, Ming Zhao 0002 |
IPDPS | 3 |
| 2020 | Honeypot Contract Risk Warning on Ethereum Smart ContractsabstractAs Ethereum's smart contracts have boomed, it has become an integral part of the blockchain ecosystem. Unfortunately, some malicious users also find the opportunity to use fraudulent means to profit. A new reported approach is to lure new users or other attackers into the contract in an attempt to make a profit by exposing seemingly obvious flaws in the contract. But in fact, the contract contains a hidden trap that ultimately benefits the creator of the contract. Such contracts are known as honeypot contracts in the blockchain ecosystem. Previous studies proposed two methods to identify such smart contracts by using symbolic execution and contract behaviors. However, these methods either make it difficult to discover new categories or fail to warn users before they lose money. To solve this problem, we propose a machine learning model to detect honeypot contracts based on N-gram features and LightGBM. Extensive experiments show that our proposed model performs well in different conditions. Weili Chen, Xiongfeng Guo, Zhiguang Chen 0001, Zibin Zheng, Yutong Lu |
JCC | 3 |
| 2020 | Traveling the token world: A graph analysis of Ethereum ERC20 token ecosystemabstractThe birth of Bitcoin ushered in the era of cryptocurrency, which has now become a financial market attracted extensive attention worldwide. The phenomenon of startups launching Initial Coin Offerings (ICOs) to raise capital led to thousands of tokens being distributed on blockchains. Many studies have analyzed this phenomenon from an economic perspective. However, little is know about the characteristics of participants in the ecosystem. To fill this gap and considering over 80% of ICOs launched based on ERC20 token on Ethereum, in this paper, we conduct a systematic investigation on the whole Ethereum ERC20 token ecosystem to characterize the token creator, holder, and transfer activity. By downloading the whole blockchain and parsing the transaction records and event logs, we construct three graphs, namely token creator graph, token holder graph, and token transfer graph. We obtain many observations and findings by analyzing these graphs. Besides, we propose an algorithm to discover potential relationships between tokens and other accounts. The reported case shows that our algorithm can effectively reveal entities and the complex relationship between various accounts in the token ecosystem. Weili Chen, Zhiguang Chen 0001, Zibin Zheng, Yutong Lu |
WWW | 3 |
| 2020 | UniIndex: An index and query middleware for parallel file systemsabstractSummary As data analysis scenarios keep increasing on high‐performance computing systems, the ability to select a small fraction of data from a large volume of scientific data sets is vital to accelerate scientific discovery. However, parallel file systems lack the ability to provide efficient data locating services at the granularity of both a file and a record. Existing methods for identifying and indexing data are often domain‐specific and do not scale to large scientific data sets. In this paper, we describe the design and implementation of UniIndex framework, which combines our proposed techniques for user‐annotation extraction, in‐memory cache layer, in‐situ indexing, and parallel query processing. Acting as middleware on top of production file systems, UniIndex enables efficient data locating services with minimal user effort. Our evaluations show that UniIndex can locate target files from directories containing millions of files in microseconds. By applying in situ indexing and the lightweight range‐bitmap index, record‐level index building time can be dramatically reduced while maintaining up to two orders of magnitude query speedup than scanning the entire data set. Peng Cheng 0012, Yutong Lu, Yunfei Du 0001, Zhiguang Chen 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2020 | Design and Implementation of the Tianhe-2 Data Storage and Management System
Yutong Lu, Peng Cheng 0012, Zhiguang Chen 0001 |
J. Comput. Sci. Technol. | 3 |
| 2019 | SMR-X: Flexible Parallel State Machine Replication for Cloud ComputingabstractState Machine Replication (SMR) is a fundamental fault tolerant technique for distributed systems. SMR traditionally requires sequential execution of commands at each replica node, so as to guarantee strong consistency among replicas. To achieve high performance at large scale cloud datacenters, SMR has been parallelized by employing multiple threads at each replica. In this paper, we propose SMR-X, a novel parallel SMR scheme, which realizes flexible mapping of commands for parallel executing at each replica. The mapping between clients' requests and work threads is dynamically adjusted according to the load level of work threads. Therefore, workloads of different threads can be well balanced and high system throughput can be achieved. The major challenge in our work lies in the inconsistency problem caused by dynamic changes in request-thread mapping. To cope with this, we design delicate mechanisms to synchronize mapping function, so that strong consistency among replicas can be guaranteed. The correctness of the proposed scheme is rigorously proved and its performance is evaluated via simulations. Simulation results show that SMR-X can achieve better load balance and lower access latency than existing parallel SMR schemes. Weigang Wu, Zhiguang Chen 0001, Nong Xiao 0001 |
CCGRID | 3 |
| 2019 | EC-ARR: Using Active Reconstruction to Optimize SSD Read Performance
Shuo Li 0007, Mingzhu Deng, Fang Liu 0002, Zhiguang Chen 0001, Nong Xiao 0001 |
ICA3PP (2) | 4 |
| 2019 | An Efficient and Flexible Metadata Management Layer for Local File SystemsabstractThe efficiency of metadata processing affects the file system performance significantly. There are two bottlenecks in metadata management in existing local file systems: 1) Path lookup is costly because it causes a lot of disk I/Os, which makes metadata operations inefficient. 2) Existing file systems have deep I/O stack in metadata management, resulting in additional processing overhead. To solve these two bottlenecks, we decoupled data and metadata management and proposed a metadata management layer for local file systems. First, we separated the metadata based on their locations in the namespace tree and aggregated the metadata into fixed-size metadata buckets (MDBs). This design fully utilizes the metadata locality and improves the efficiency of disk I/O in the path lookup. Second, we customized an efficient MDB storage system on the raw storage device. This design simplifies the file system I/O stack in the metadata management and allows metadata lookup to be completed with constant time complexity. Finally, this metadata management layer gives users the flexibility to choose metadata storage devices. We implemented a prototype called Otter. Our evaluation demonstrated that Otter outperforms native EXT4, XFS, Btrfs, BetrFS and TableFS in many metadata operations. For instance, Otter has 1.2 times to 9.6 times performance improvement over other tested file systems in file opening. Hongbo Li 0007, Yutong Lu, Zhiguang Chen 0001, Ming Zhao 0002 |
ICCD | 4 |
| 2019 | COMBFT: Conflicting-Order-Match based Byzantine Fault Tolerance Protocol with High Efficiency and RobustnessabstractByzantine Fault-Tolerant (BFT) state machine replication protocol is an important building block for highly available distributed computing. This paper presents COMBFT, a BFT protocol that achieves both efficiency and robustness simultaneously. The major novelty of COMBFT lies in Conflicting-Order-Match (COM), a new request ordering mechanism that uses a new way to select the available sequence number for requests, and detects the possible malicious primary early. COM assigns sequence number based on request interference, and requires both primary and backup nodes to conduct request ordering, which can greatly reduce the impact of malicious primary and clients. When the backup suspects the primary may be malicious, it triggers an efficient commit protocol with two phases (i.e., suspect phase and commit phase) to further confirm whether the primary is malicious, and commit the request. The performance of COMBFT is evaluated via simulations and the results illustrate the outstanding performance of COMBFT in terms of throughput, latency and fault scalability. Yingyao Rong, Weigang Wu, Zhiguang Chen 0001 |
ICPP | 3 |
| 2019 | Optimizing Data Placement on Hierarchical Storage Architecture via Machine Learning
Peng Cheng 0012, Yutong Lu, Yunfei Du 0001, Zhiguang Chen 0001, Yang Liu 0005 |
NPC | 4 |
| 2019 | Tiered data management system: Accelerating data processing on HPC systems
Peng Cheng 0012, Yutong Lu, Yunfei Du 0001, Zhiguang Chen 0001 |
Future Gener. Comput. Syst. | 4 |
| 2018 | PDFE: Flexible Parallel State Machine Replication for Cloud ComputingabstractState machine replication (SMR) is a fundamental fault tolerant technique for distributed systems to guarantee consistency among replicas via sequential execution of commands. With the development of cloud computing, parallel SMR has been recently proposed for large scale cloud datacenters. In this paper, we propose PDFE, a novel parallel SMR scheme, which realizes flexible dispatch of parallel ordered commands for parallel executing. In PDFE, the mapping/binding between ordering threads and work threads becomes dynamic, and commands can be dispatched according to the work load level of different threads. Such flexibility can help achieve two levels of load balancing: load balancing between ordering threads and work threads, and load balancing among work threads. The major challenge in our work lies in the inconsistency problem caused by dynamic changes in command dis-patch, and it is addressed by a specially designed mechanism. Compared with existing parallel SMR schemes, PDFE can achieve better load balancing and higher system efficiency. Such advantages are validated by experimental performance evaluation. Lihui Wu, Weigang Wu, Zhiguang Chen 0001 |
CLUSTER | 4 |
| 2018 | DA Placement: A Dual-Aware Data Placement in a Deduplicated and Erasure-Coded Storage System
Mingzhu Deng, Ming Zhao 0002, Fang Liu 0002, Zhiguang Chen 0001, Nong Xiao 0001 |
ICA3PP (1) | 4 |
| 2018 | Path Prefetching: Accelerating Index Searches for In-Memory DatabasesabstractIn-memory databases (IMDBs) store all working data in main memory, which makes memory accesses become the dominant factor of the whole system performance. Micro-architectural studies of mainstream in-memory on-line transaction processing (OLTP) systems show that more than half of the execution time goes to memory stalls. Moreover, for IMDBs that adopt aggressive transaction compilation optimizations, data misses from the last-level cache (LLC) are responsible for the majority of the overall stall time. In this paper, through profiling analysis of IMDBs we observe that index access misses dominate LLC data misses. Based on the key observation that adjacent keys tend to follow similar traversal paths in ordered index searches, we propose the path prefetching to mitigate LLC misses induced by ordered index searches, which records mappings between keys and their traversal paths and then generate prefetches for future same/adjacent keys. Experimental results show that for ordered index searches the proposed path prefetcher provides an average speedup of 27.4% over the baseline with no prefetching. Shuo Li 0007, Zhiguang Chen 0001, Nong Xiao 0001, Guangyu Sun 0003 |
ICCD | 2 |
| 2018 | RM-KVStore: New MXNet KVStore to Accelerate Transfer Performancewith RDMA
Baocai Lv, Fang Liu 0002, Nong Xiao 0001, Zhiguang Chen 0001 |
ISCC | 5 |
| 2018 | Accelerating Spark Shuffle with RDMAabstractApache Spark is a lightning-fast unified analytics engine for large-scale data processing. When executing an application with Spark, it runs many jobs in parallel. These jobs are divided into stages based on the shuffle boundary. However, shuffling data across the stages in a cluster is time-consuming because it will place significant burden on operating system on both the source and the destination by requiring many remote files and network I/Os. Meanwhile, the latest Spark is based on Netty which is written with Java Sockets and will produce a large number of data copies during the shuffle phase. This has become the major bottleneck for Apache Spark and motivates us to use RDMA technology to accelerate data shuffle. RDMA, with the function of zero-copy transfers, reducing latency and CPU overhead, can reduce stress on operating system during the shuffle phase and improve the throughput of the whole system. In this paper, we present a high-performance RDMA-based design for accelerating data shuffle in Apache Spark framework by providing tiering memory pool and different mechanisms to transfer messages of different sizes. The experimental results show that compared to the default Spark running with IP over InfiniBand (IPoIB), our proposed design can achieve up to 89.8% performance improvement for Spark RDD operation benchmarks (e.g., GroupBy and SortBy), up to 49% performance improvement for iterative algorithms (e.g., TriangleCount and SVM in SparkBench). And the evaluation results also show that our RDMA-based design slightly outperforms Crail-Spark-IO, a recent open-source Spark shuffle plugin from IBM. Fang Liu 0002, Nong Xiao 0001, Zhiguang Chen 0001 |
NAS | 4 |
| 2018 | Mimir+: An Optimized Framework of MapReduce on Heterogeneous High-Performance Computing System
Zhiguang Chen 0001, Yunfei Du 0001, Yutong Lu |
NPC | 2 |
| 2017 | KV-FTL: A novel key-value based FTL scheme for large scale SSDsabstractBoth traditional coarse-grained and fine-grained Flash Translation Layer schemes are unsuitable for ultra-large SSDs. They produce overmuch mapping entries which fail to be kept in embedded DRAM completely and can suffer severely from low spatial and temporal localities. In this paper, we propose a novel KV-FTL for ultra-large SSDs, which mostly maps logical addresses to physical addresses via a simple hash function, while handles hash collisions and out-of-place data updates by the traditional manner, i.e., the mapping table. Our KV-FTL can accelerate address translation by avoiding loading mapping table from flash memory to DRAM, thus improve performance; as well as reduce the write-traffic incurred by the mapping table, thus extend the lifespan of SSDs. Experimental results show that our KV-FTL facilitates SSDs to survive longer lifespan by a factor of up to 18.7% with an average of 13.6%; improves read performance ranging from 18.4% to 50.7% with an average of 39% with optimization, and in the case of extremely intensive requests, improves the access performance for requests with an average of 47%. Zhengguo Chen, Zhiguang Chen 0001, Nong Xiao 0001, Fang Liu 0002 |
ASAP | 3 |
| 2016 | Leader: Accelerating ReRAM-based main memory by leveraging access latency discrepancy in crossbar arrays
Nong Xiao 0001, Fang Liu 0002, Zhiguang Chen 0001 |
DATE | 4 |
| 2016 | Red-Shield: Shielding Read Disturbance for STT-RAM Based Register Files on GPUsabstractTo address the high energy consumption issue of SRAM on GPUs, emerging Spin-Transfer Torque (STT-RAM) memory technology has been intensively studied to build GPU register files for better energy-efficiency, thanks to its benefits of low leakage power, high density, and good scalability. However, STT-RAM suffers from a reliability issue, read disturbance, which stems from the fact that the voltage difference between read current and write current becomes smaller as technology scales. The read disturbance leads to high error rates for read operations, which cannot be effectively protected by SECDEC ECC on large-capacity register files of GPUs. Xuhao Chen 0001, Nong Xiao 0001, Fang Liu 0002, Zhiguang Chen 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2016 | Shielding STT-RAM Based Register Files on GPUs against Read DisturbanceabstractTo address the high energy consumption issue of SRAM on GPUs, emerging Spin-Transfer Torque (STT-RAM) memory technology has been intensively studied to build GPU register files for better energy-efficiency, thanks to its benefits of low leakage power, high density, and good scalability. However, STT-RAM suffers from the read disturbance issue, which stems from the fact that the voltage difference between read current and write current becomes smaller as technology scales. The read disturbance leads to high error rates for read operations, which cannot be effectively protected by the SEC-DED ECC on large-capacity register files of GPUs. Prior schemes (e.g., read-restore) to mitigate the read disturbance usually incur either non-trivial performance loss or excessive energy overhead, thus not applicable for the GPU register file design that aims to achieve both high performance and energy-efficiency. To combat the read disturbance, we propose a novel software-hardware co-designed solution (i.e., Red-Shield ), which consists of three optimizations to overcome the limitations of the existing solutions. First, we identify dead reads at compiling stage and augment instructions to avoid unnecessary restores. Second, we employ a small read buffer to accommodate register reads with high-access locality to further reduce restores. Third, we propose an adaptive restore mechanism to selectively pick the suitable restore scheme, according to the busy status of corresponding register banks. Experimental results show that our proposed design can effectively mitigate the performance loss and energy overhead caused by restore operations while still maintaining the reliability of reads. Xuhao Chen 0001, Nong Xiao 0001, Lei Wang 0011, Fang Liu 0002, Wei Chen 0009, Zhiguang Chen 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 7 |
| 2016 | Me-CLOCK: A Memory-Efficient Framework to Implement Replacement Policies for Large CachesabstractSolid State Drives (SSDs) have been extensively deployed as the cache of hard disk-based storage systems. The SSD-based cache generally supplies ultra-large capacity, whereas managing so large a cache introduces excessive memory overhead, which in turn makes the SSD-based cache neither cost-effective nor energy-efficient. This work targets to reduce the memory overhead introduced by the replacement policy of SSD-based cache. Traditionally, data structures involved in cache replacement policy reside in main memory. While these in-memory data structures are not suitable for SSD-based cache any more since the cache is much larger than ever. We propose a memory-efficient framework which keeps most data structures in SSD while just leaving the memory-efficient data structure (i.e., a new bloom proposed in this work) in main memory. Our framework can be used to implement any LRU-based replacement policies under negligible memory overhead. We evaluate our proposals via theoretical analysis and prototype implementation. Experimental results demonstrate that, our framework is practical to implement most replacement policies for large caches, and is able to reduce the memory overhead by about$10 \times$. Zhiguang Chen 0001, Nong Xiao 0001, Yutong Lu, Fang Liu 0002 |
IEEE Trans. Computers | 1 |
| 2015 | RAID-6Plus: A Fast and Reliable Coding Scheme Aided by Multi-failure Degradation
Mingzhu Deng, Nong Xiao 0001, Songping Yu, Wei Chen 0009, Zhiguang Chen 0001, Fang Liu 0002 |
APSCC | 6 |
| 2015 | NF-Dedupe: A novel no-fingerprint deduplication scheme for flash-based SSDsabstractNAND flash-based Solid State Drives (SSDs) have been widely deployed in data centers of cloud computing due to their high performance compared with hard disks, while the limited lifespan of flash memory makes SSDs not very suitable for write-intensive applications. Deduplication is an effective method used to reduce the write traffic of applications thus can be used to extend the lifespan of SSDs. However, traditional deduplication schemes rely on the time-consuming fingerprint computing process to find duplicated data, which may impair the write performance of SSDs. Accordingly, Pre-hashing was proposed to reduce the chances of fingerprint computing thus improving the performance of SSDs with deduplication, but at the cost of degrading deduplication rate. In this paper, we propose NF-Dedupe, a new deduplication scheme that needs no fingerprint computing for flash-based SSDs. NF-Dedupe determines whether a write page is duplicated or not by comparing the write page with its potential duplicated page read from underlying flash chips byte by byte, rather than relying on the comparison of fingerprints. As flash memory is known for its high parallelism and low read latency, reading a page from flash chip and comparing two pages byte by byte introduce lower overhead than the fingerprint computing does. We evaluate the NF-Dedupe via trace-driven simulations. Experimental results have shown that NF-Dedupe outperforms the other approaches and can achieve the deduplication rate ranging from 5.3% to 29.9% and the write latency is improved by a factor of up to 21% with an average of 12%. Zhengguo Chen, Zhiguang Chen 0001, Nong Xiao 0001, Fang Liu 0002 |
ISCC | 2 |
| 2015 | A theoretical analysis of lifespan impact on flash memory imposed by erasure codeabstractEach cell of flash memory only survives a nominally given number of write/erasure cycles. Beyond the nominal lifespan, flash memory can still record digital information but the bit error rate increases rapidly with the increment of write/erase cycles. Erasure code is a conventional method used to recover corrupted data, but its redundant data produce a large number of additional writes, making the erasure code seem to be unsuitable for the write-sensitive flash memory. We argue that, erasure code influences the lifespan of flash memory in two conflicting directions: its inherent error correction capability enables the flash memory to survive beyond the nominal lifespan, while its redundant data wear out the lifespan of flash memory by increasing the write/erase cycles. This paper builds a theoretical model to analyze both the two aspects and demonstrates that the erasure code is able to extend the nominal lifespan of flash memory by as many as 30×. Enqiang Zhou, Yutong Lu, Nong Xiao 0001, Zhiguang Chen 0001 |
NAS | 5 |
| 2014 | A hybrid memory built by SSD and DRAM to support in-memory Big Data analytics
Zhiguang Chen 0001, Yutong Lu, Nong Xiao 0001, Fang Liu 0002 |
Knowl. Inf. Syst. | 1 |
| 2013 | Reorder Write Sequence by Hetero-Buffer to Extend SSD's Lifespan
Zhiguang Chen 0001, Nong Xiao 0001, Fang Liu 0002, Yimo Du |
J. Comput. Sci. Technol. | 1 |
| 2013 | CSWL: Cross-SSD Wear-Leveling Method in SSD-Based RAID Systems for System Endurance and Performance
Yimo Du, Nong Xiao 0001, Fang Liu 0002, Zhiguang Chen 0001 |
J. Comput. Sci. Technol. | 4 |
| 2013 | An SSD-based accelerator for directory parsing in storage systems containing massive files
Zhiguang Chen 0001, Nong Xiao 0001, Fang Liu 0002 |
Peer-to-Peer Netw. Appl. | 1 |
| 2012 | SAC: rethinking the cache replacement policy for SSD-based storage systemsabstractSolid-state drives (SSDs) are widely used in storage systems. However, algorithms adopted by existing operating systems generally consider the underlying devices as hard disks, and thus are rarely optimized for SSDs. In this paper, we focus on a classical research issue, the cache replacement policy, and design a new policy by taking the parallelism of SSDs into account. Zhiguang Chen 0001, Nong Xiao 0001, Fang Liu 0002 |
SYSTOR | 1 |
| 2012 | Dual queues cache replacement algorithm based on sequentiality detection
Nong Xiao 0001, Yingjie Zhao, Fang Liu 0002, Zhiguang Chen 0001 |
Sci. China Inf. Sci. | 4 |
| 2011 | PBFTL: The Page to Block Mapping FTL with Low Response TimeabstractNAND flash has some inherent peculiarities which increase the access delay seriously. We propose the Page to Block mapping Flash Translation Layer (PBFTL). Solid State Drives (SSDs) adopting PBFTL have lower response time. To achieve low response time for read requests, PBFTL adopts hybrid-level mapping scheme. But, hybrid-level FTL behaves awkwardly for write due to the high overhead of garbage collection. PBFTL takes two measures to optimize garbage collection. The first is to direct hot and cold data to separate blocks, which mitigates write amplification significantly. The second is to reduce the latency of reclaiming a block, which enables PBFTL to spend less time on garbage collection. User's requests are unlikely to be congested for a long time. Trace-driven simulations show that, PBFTL achieves low response for both read- and write-intensive workloads. Zhiguang Chen 0001, Nong Xiao 0001, Fang Liu 0002, Yimo Du |
MASCOTS | 1 |
| 2011 | Reorder the Write Sequence by Virtual Write Buffer to Extend SSD's Lifespan
Zhiguang Chen 0001, Fang Liu 0002, Yimo Du |
NPC | 1 |
| 2011 | WeLe-RAID: A SSD-Based RAID for System Endurance and Performance
Yimo Du, Fang Liu 0002, Zhiguang Chen 0001 |
NPC | 3 |
| 2011 | P3Stor: A parallel, durable flash-based SSD for enterprise-scale storage systems
Nong Xiao 0001, Zhiguang Chen 0001, Fang Liu 0002, Longfei An |
Sci. China Inf. Sci. | 2 |
| 2010 | Hot Data-Aware FTL Based on Page-Level Address MappingabstractThe development of flash memory drives flash based SSD to enter into large-scale storage systems. The performance of SSD is highly dependent on the design of FTL. For the last few years, several FTL schemes have been proposed. Such as FAST, BAST, SAST etc. we design a novel FTL based on page-level mapping scheme. Since one of the major troubles of page-level mapping FTL is the unendurable memory consuming of the fine-grained mapping table. We propose a dedicated cache replacement policy called SRC to mitigate the memory pressure. Our FTL based on SRC is able to distinguish hot data from the cold. This capability highlights the garbage collection efficiency of page-level mapping schemes. As a result, the hot data-aware FTL reduces extra read/write operations by 10 times or more compared with FAST and BAST. As our FTL erases less blocks, the lifetime of SSD is extended by more than 30%. Further experiment shows that the hot data-aware FTL outperforms hybrid-level FTLs on workloads with varied read/write ratios. Zhiguang Chen 0001, Nong Xiao 0001, Fang Liu 0002, Yimo Du |
HPCC | 1 |
| 2010 | 2F: A Special Cache for Mapping Table of Page-Level Flash Translation LayerabstractThe development of flash memory drives flash based SSDs to enter into enterprise-scale storage systems. As the kernel of SSD, flash translation layer (FTL) attracts many attentions. Generally, there are two types of FTLs according to the granularity of address mapping: block-level and page-level mapping FTLs. We focus on the latter one. Typically, page-level mapping scheme must employ a cache to alleviate the memory pressure introduced by the big mapping table. We argue that classic cache replacement policies aren't competent for the page table cache of FTLs. The major contribution of this work is to design a dedicated cache replacement policy called Two Filters (abbreviated as 2F) for page-level mapping FTLs. 2F aims at two goals. The first is higher hit ratio as all the replacement policies pursue. As 2F not only protects frequently accessed pages, but also protects sequentially accessed pages at little cost, it does achieve a higher hit ratio. The second goal is to distinguish hot pages from the cold. This goal is special for page table of FTLs. If hot and cold pages are directed to separate blocks, garbage collection will be more efficient. In order to achieve this goal, 2F employs two filters. One is used for containing sequentially accessed pages. Another is used for selecting hot pages. Trace driven simulations present that 2F outperforms classic replacement policies in both hit ratio and data classification. Zhiguang Chen 0001, Nong Xiao 0001, Fang Liu 0002, Yimo Du |
ICPADS | 1 |
| 2009 | SSARC: The Short-Sighted Adaptive Replacement CacheabstractAs the performance gap between disks and processors continues to increase, dozens of cache replacement policies come up to handle the problem. Unfortunately, most of the policies are static. Nimrod Megiddo etc put forward a low overhead adaptive policy called ARC. It outperforms most of the static policies in most situations. But, ARC adapts itself to the workloads by the feedback of the missed pages. It hasn 't carried out the adaption before missed pages are discovered. We propose a high performance adaptive replacement policy. It adapts itself to the workloads by the feedback of the hit pages, so, it is more sensitive to the changes of the workloads than ARC. As the policy stares at the tails of the queues regardless of other pages, we name the policy as short-sighted adaptive replacement policy. The ARC usually regrets for the missed pages and wishes to rescue the neighborhood of them. However, SSARC endeavors to protect the would-be-reused pages from being replaced aggressively. So, it outperforms ARC in most situations. We compared SSARC with LRU, 2Q and ARC. The trace-driven experiments represent that SSARC gains higher performance. Zhiguang Chen 0001, Nong Xiao 0001, Fang Liu 0002, Yingjie Zhao |
HPCC | 1 |