EDBT 2026 Demo / reviewers in the wild / expert
Jin Xiong
dblp:47/6433
· DBLP profile ↗
51ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 38 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Computer networks · 1Security and privacy · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CETOFS: A High-Performance File System with Host-Server Collaboration for Remote Storage
Wenqing Jia, Dejun Jiang 0001, Jin Xiong |
FAST | 3 |
| 2025 | On the minimum spectral radius of graphs with given order and dissociation number
Huiqing Liu, Jin Xiong |
Discret. Appl. Math. | 3 |
| 2025 | NapFS: A High-Performance Persistent Memory File System for Non-Uniform Memory Access Architectures
Wen-Qing Jia, De-Jun Jiang, Jin Xiong |
J. Comput. Sci. Technol. | 3 |
| 2025 | A High-Performance and Scalable Userspace Log-Structured File System for Modern SSDsabstractWe present AugeFS , a scalable userspace log-structured file system for modern SSDs. AugeFS re-architects the file system stack to address three critical challenges: inefficient control plane, limited metadata scalability, and underutilized device bandwidth. First, we propose a shared and protected address space within the userspace of accessing applications to run AugeFS , which enables the high-performance data plane and efficient control plane. Second, we design a scalable LSM-tree based key-value store called MetaDB to organize small-sized metadata in AugeFS . To improve the metadata scalability, MetaDB employs parallel request processing to reduce thread synchronization overhead and fine-grained parallel write-ahead log to eliminate false sharing in metadata persistence. Finally, AugeFS distributes files into different domains. To reduce contention, we maintain space management metadata for each domain independently, which helps scale data performance and improve the device IO utilization. Moreover, AugeFS designs an asynchronous IO stack for fsync to reduce the latency of synchronous writes. The evaluation results show that AugeFS significantly improves both metadata scalability and data scalability. Wenqing Jia, Dejun Jiang 0001, Jin Xiong |
ACM Trans. Storage | 3 |
| 2024 | SuperMap: High-Performance and Flexible Memory-Mapped IO for Fast Storage DeviceabstractMemory-mapped IO offers several advantages over explicit read/write IO. It requires no system call, incurs minimal overhead in case of cache hits, and avoids extra data copies between user and kernel space. However, we still identify inefficiencies in current memory-mapped IO designs when meeting fast storage devices: i) the heavy IO stack in the page fault handler, ii) the suboptimal prefetching design, and iii) the inefficient eviction policy. To address these limitations, we present SuperMap, an alternative design for the memory-mapped IO in Linux, which specifically brings high performance and flexibility for fast devices. First, SuperMap designs a lightweight and asynchronous IO stack by directly accessing device, reducing software overhead significantly. Second, SuperMap introduces a fine-grained and application-customized prefetcher framework based on eBPF, further improving performance. Third, SuperMap proposes a hotness-aware eviction policy with the hardware assistance, trying to keep frequently accessed data in memory. Through evaluations using benchmarks and real-world applications, we demonstrate that SuperMap outperforms the state-of-the-art memory-mapped IO design (FastMap) up to 67%. Wenqing Jia, Dejun Jiang 0001, Jin Xiong |
ICCD | 3 |
| 2024 | zQoS: Unleashing full performance capabilities of NVMe SSDs while enforcing SLOs in distributed storage systemsabstractNowadays, data centers consolidate latency-critical (LC) tenants and best-effort (BE) tenants on the same cloud platform to increase resource utilization and reduce costs. In such a scenario, the underlying distributed storage systems are responsible for guaranteeing SLOs for LC tenants while maximizing bandwidth for BE tenants. As high-performance NVMe SSDs are widely deployed, how to make full use of their performance capabilities and guarantee SLOs has become an urgent problem. However, current methods restrict the performance capabilities of NVMe SSDs based on a conservative offline model, and also ignore runtime changes in tenant loads and device states, which definitely affect the performance capabilities. Liuying Ma, Zhenqing Liu, Jin Xiong, Renhai Chen, Xi Peng 0006, Gong Zhang 0001, Dejun Jiang 0001 |
ICPP | 3 |
| 2024 | DiStore: A Fully Memory Disaggregation Friendly Key-Value Store with Improved Tail Latency and Space EfficiencyabstractMemory disaggregation decouples CPUs and memory in monolithic servers to form compute nodes (CNs) and memory nodes (MNs) for elastic and efficient memory scaling. Memory disaggregation benefits in-memory key-value stores (KVSs) that demand large memory capacity. Building a KVS with full functionality, low latency, and high space efficiency under memory disaggregation is critically required by real-world applications. However, we observe that existing disaggregated KVSs fail to achieve all the above demands simultaneously. In this paper, we present DiStore, a full-disaggregation-friendly KVS that accomplishes the above goals. We carefully divide the responsibilities of CNs and MNs when involved in index traversing, concurrency control, and memory management. We first design a two-layer indexing structure, separable adaptive linked array, to reduce network RTTs for index traversing and improve space efficiency. Then, DiStore introduces thread-context-based concurrency control to enable inter-thread collaborating to reduce stall time under multi-thread contention on CNs. Finally, we propose cachable disaggregated memory management to allow CNs to manage remote memory locally with marginal space overhead. We implement DiStore and the evaluation shows that DiStore can reduce P999 tail latency by up to 72.1% and improve space efficiency by up to 33.6%. Meanwhile, DiStore achieves comparable throughput as state-of-the-art disaggregated KVSs. Ziwei Xiong, Dejun Jiang 0001, Jin Xiong |
ICPP | 3 |
| 2023 | FastStore: A High-Performance RDMA-Enabled Distributed Key-Value Store with Persistent MemoryabstractDistributed persistent key-value store (KVS) plays an important role in today's storage infrastructure. The development of persistent memory (PM) and remote direct memory access (RDMA) allows to build distributed persistent KVS to provide fast data access. However, prior works focus on either PM-oriented or RDMA-oriented optimizations for key-value stores. We find these optimizations disallow a simple porting of RDMA-enabled KVS to PM or vice versa. This paper proposes FastStore, a high-performance distributed persistent KVS, by fully exploiting RDMA features and PM-friendly optimizations. First, FastStore utilizes RDMA-enabled PM exposure to establish direct indexing at the client side to reduce RTTs for reading values. Meanwhile, PM exposure allows PM sharing among cluster nodes, which helps to mitigate attribute-value skewness. Then, FastStore designs PM-friendly ownership transferring log and failure-atomic slotted-page allocator to achieve highly efficient PM management without PM leakage. Finally, FastStore proposes volatile search key to its B+tree indexing to reduce excessive PM accesses. We implement FastStore and the evaluation shows that FastStore outperforms the state-of-the-art ordered KVS Sherman by 2.8× higher throughput and 71.5% fewer RTTs. Ziwei Xiong, Dejun Jiang 0001, Jin Xiong |
ICDCS | 3 |
| 2023 | Exploiting Hybrid Index Scheme for RDMA-based Key-Value StoresabstractRDMA (Remote Direct Memory Access) is widely studied in building key-value stores to achieve ultra-low latency. In RDMA-based key-value stores, the indexing time takes a large fraction of the overall operation latency as RDMA enables fast data access. However, the single index structure used in existing RDMA-based key-value stores, either hash-based or sorted index, fails to support range queries efficiently while achieving high performance for singlepoint operations. In this paper, we explore the adoption of a hybrid index in the key-value stores based on RDMA, especially under the memory disaggregation architecture, to combine the benefits of a hash table and a sorted index. We propose HStore, an RDMA-based key-value store that uses a hash table for single-point lookups and leverages a skiplist for range queries to index the values stored in the memory pool. Guided by previous work on using RDMA for key-value services, HStore dedicatedly chooses different RDMA verbs to optimize the read and write performance. To efficiently keep the index structures within a hybrid index consistent, HStore asynchronously applies the updates to the sorted index by shipping the update log via two-sided verbs. Compared to state-of-the-art Sherman and Clover, HStore improves the throughput by up to 54.5% and 38.5% respectively under the YCSB benchmark. Shukai Han, Mi Zhang 0007, Dejun Jiang 0001, Jin Xiong |
SYSTOR | 4 |
| 2023 | A Survey of Non-Volatile Main Memory File Systems
Ying Wang 0001, Wenqing Jia, Dejun Jiang 0001, Jin Xiong |
J. Comput. Sci. Technol. | 4 |
| 2023 | Dalea: A Persistent Multi-Level Extendible Hashing with Improved Tail Performance
Ziwei Xiong, Dejun Jiang 0001, Jin Xiong |
J. Comput. Sci. Technol. | 3 |
| 2023 | Design and application of new storage systemsabstract存储系统是计算机的核心,在人工智能、大数据、云计算和物联网等新兴战略产业的可持续发展中起着重要作用。随着处理器和网络设备性能不断提高,存储软件栈成为限制数据密集型系统性能的主要因素。近年来,新型存储设备因其打破“内存墙”的能力而受到广泛关注。这些设备包括支持块寻址的闪存设备、支持字节寻址的非易失性存储器、存算一体化设备以及大容量光存储。构建高吞量、低延迟和高可靠性的大规模存储系统,需要对算法、软件设计和硬件的持续创新。这些创新可以应对大规模、高性能复杂结构系统构建中存在的挑战,还可以增加相关系统的构建和应用经验,加快大数据处理系统的开发速度。 研究人员一直致力于解决“内存墙”问题,并改进相关软硬件生态系统,从而在新型存储系统设计和应用方面取得很大进展,包括但不限于以下方面: Guangyan Zhang, Keqin Li 0001, Zili Shao, Nong Xiao 0001, Jin Xiong |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2022 | NapFS: A High-Performance NUMA-Aware PM File SystemabstractPersistent memory (PM) allows file systems to directly persist data on the memory bus. In order to expand the capacity of PM file system, building a file system across sockets with each attached PMs is attractive. However, accessing data across sockets incurs impacts of non-uniform memory access (NUMA) architecture, which degrade PM file system performance. In this paper, we first conduct experiments to understand the NUMA impacts on building PM file systems. We then propose four design principles for building a high-performance NUMA-aware PM file system NapFS. We architect NapFS with per-socket local PM file systems and per-socket dedicated IO thread pools. This not only allows applications to delegate data access to IO threads for avoiding remote PM access, but also fully reuses existing single-socket PM file systems to reduce implementation complexity. In addition, NapFS utilizes fast DRAM to accelerate performance by adding a global cache. We evaluate NapFS against other multi-socket PM file systems. The evaluation results show that NapFS achieves 2.2× and 1.0× throughput improvement for Filebench and RocksDB. Wenqing Jia, Dejun Jiang 0001, Jin Xiong |
ICCD | 3 |
| 2022 | QWin: Core Allocation for Enforcing Differentiated Tail Latency SLOs at Shared Storage BackendabstractDistributed storage systems consolidate latency-critical (LC) tenants and best-effort (BE) tenants together to increase resources utilization. As a result, storage server processes require to guarantee differentiated tail latency SLOs for LC tenants, and meanwhile provide sustainable high bandwidth for BE tenants. With the adoption of high-performance NVMe SSDs, queuing time on storage server process contributes a large part to tail latency of LC tenant. Partitioning request queues, worker threads, and CPU cores on storage server processes among tenants helps to reduce interference and thus queuing time. This in turn requires careful core allocation to guarantee tail latency SLOs. In this paper, we argue that accurate core allocation is necessary for storage server processes to allocate the actual cores required by LC tenants. Thus, we propose QWin, a tail latency SLO-aware core allocation to enforce differentiated tail latency SLOs for multiple LC tenants. QWin first designs an SLO-to-core calculation model to accurately calculate the number of cores required by LC tenant. Bursty loads or fluctuated I/O latencies of storage devices can change core requirements of LC tenants. Thus, we design three core policies in QWin to adapt to the changing core requirements. We evaluate QWin by consolidating multiple LC and BE tenants together. The experiment results show that QWin outperforms the-state-of-the-art approaches in guaranteeing differentiated tail latency SLOs for LC tenants and meanwhile increasing bandwidth for BE tenants by 4x~21x. Liuying Ma, Zhenqing Liu, Jin Xiong, Dejun Jiang 0001 |
ICDCS | 3 |
| 2021 | Using Vectorized Execution to Improve SQL Query Performance on SparkabstractMapReduce-based SQL processing frameworks, such as Hive and Spark SQL, are widely used to support big data analytics. Currently these systems mainly adopt the record-at-a-time execution model, which is less efficient in terms of CPU utilization. In contrast, vectorized execution is able to make better use of CPU cache by bulk processing a record batch at a time. However, simply applying vectorized execution to MapReduce-based frameworks results in low efficient vectorized shuffle. Moreover, existing vectorized execution donot make full use of CPU cache for complex operators (e.g. Sort and Aggregation). In this paper, we present VEE, a thorough vectorized execution engine designed for SQL query processing on Spark. First, VEE designs compact in-memory data layout and serialization-aware assembling for vectorized shuffle to expedites shuffle execution, since they reduce shuffle data footprint and related computations. Secondly, VEE applies in-memory record batch rearrangement for Sort and Aggregation to greatly reduce random memory access and increase query performance. Thirdly, VEE carefully designs operator-aware batch length when handling different operators, which makes better utilization of CPU cache and increases query performance. We conduct extensive performance evaluations. The experiment results show that the performance speedup of VEE against Spark is up to 72.7% and 25.0% on average for OLAP workloads (TPC-H). The vectorized execution technologies in VEE are also applicable to other MapReduce-based data analytic frameworks to improve their query performance. Yijie Shen, Jin Xiong, Dejun Jiang 0001 |
ICPP | 2 |
| 2021 | Preface
Xian-He Sun, Dong Li 0001, Wen-Guang Chen, Tao Li 0006, Jiwu Shu, Bo Wu 0002, Jin Xiong, Jinging Xue, Feng Zhang 0007, Jidong Zhai, Zhiia Zhao |
J. Comput. Sci. Technol. | 7 |
| 2020 | HiLSM: an LSM-based key-value store for hybrid NVM-SSD storage systemsabstractIn order to ensure data durability and crash consistency, the LSM-tree based key-value stores suffer from high WAL synchronization overhead. Fortunately, the advent of NVM offers an opportunity to address this issue. However, NVM is currently too expensive to meet the demand of massive storage systems. Therefore, the hybrid NVM and SSD storage system provides a more cost-efficient solution. This paper proposes HiLSM, a key-value store for hybrid NVM-SSD storage systems. According to the characteristics of hybrid storage mediums, HiLSM adopts hybrid data structures consisting of the log-structured memory and the LSM-tree. Aiming at the issue of write stalls in write intensive scenario, a fine-grained data migration strategy is proposed to make the data migration start as early as possible. Aiming at the performance gap between NVM and SSD, a multi-threaded data migration strategy is proposed to make the data migration complete as soon as possible. Aiming at the LSM-tree's inherent issue of write amplification, a data filtering strategy is proposed to make data updates be absorbed in NVM as much as possible. We compare HiLSM with the state-of-the-art key-value stores via extensive experiments and the results show that HiLSM achieves 1.3x higher throughput for write, 10x higher throughput for read and 79% less write traffic under the skewed workload. Wen-Jie Li, Dejun Jiang 0001, Jin Xiong, Yungang Bao |
CF | 3 |
| 2020 | SplitKV: Splitting IO Paths for Different Sized Key-Value Items with Advanced Storage Devices
Shukai Han, Dejun Jiang 0001, Jin Xiong |
HotStorage | 3 |
| 2020 | Gecko: Guaranteeing Latency SLO on a Multi-Tenant Distributed Storage SystemabstractMeeting tail latency Service Level Objective (SLO) as well as achieving high resource utilization is important to distributed storage systems. Recent works adopt strict priority scheduling or constant rate limiting to provide SLO guarantee but cause under-utilization resources. To address this issue, we first analyze the relationship between workload burst and latency SLO. Based on burst patterns and latency SLOs, we classify tenants into two categories: Postponement-Tolerable tenant and Postponement-Intolerable tenant. We then explore the opportunity to improve resource utilization by carefully allocating resources to each tenant type. We design Rate-Limiting-Priority scheduling algorithm to limit the impact of high priority tenants on low priority ones. Meanwhile, we propose Postponement-Aware scheduling algorithm which allows Postponement-Intolerable tenants to preempt system capacity from Postponement-Tolerable tenants. This helps to increase resource utilization. We propose a latency SLO guarantee framework Gecko. Gecko guarantees multi-tenant latency SLOs via combining the two proposed scheduling algorithms together with an admission control strategy. We evaluate Gecko with real production traces and the results show that Gecko admits 44% more tenants on average than state-of-the-art techniques meanwhile guaranteeing latency SLO. Zhenyu Leng, Dejun Jiang 0001, Liuying Ma, Jin Xiong |
ICPADS | 4 |
| 2020 | SrSpark: Skew-resilient Spark based on Adaptive Parallel ProcessingabstractMapReduce-based SQL processing systems, e.g., Hive and Spark SQL, are widely used for big data analytic applications due to automatic parallel processing on large-scale machines. They provide high processing performance when loads are balanced across the machines. However, skew loads are not rare in real applications. Although many efforts have been made to address the skew issue in MapReduce-based systems, they can neither fully exploit all available computing resources nor handle skews in SQL processing. Moreover, none of them can expedite the processing of skew partitions in case of failures. In this paper, we present SrSpark, a MapReduce-based SQL processing system that can make full use of all computing resources for both non-skew loads and skew loads. To achieve this goal, SrSpark introduces fine-grained processing and work-stealing into the MapReduce framework. More specifically, SrSpark is implemented based on Spark SQL. In SrSpark, partitions are further divided into sub-partitions and processed in sub-partition granularity. Moreover, SrSpark adaptively uses both intra-node and inter-node parallel processing for skew loads according to available computing resources in realtime. Such adaptive parallel processing increases the degree of parallelism and reduces the interaction overheads among the cooperative worker threads. In addition, SrSpark checkpoints sub-partition's processing results periodically to ensure fast recovery from failures during skew partition processing. Our experiment results show that for skew loads, SrSpark outperforms Spark SQL by up to 3.5x, and 2.2x on average, while the performance overhead is only about 4% under non-skew loads. Yijie Shen, Jin Xiong, Dejun Jiang 0001 |
ICPADS | 2 |
| 2018 | Caching or Not: Rethinking Virtual File System for Non-Volatile Main Memory
Ying Wang 0001, Dejun Jiang 0001, Jin Xiong |
HotStorage | 3 |
| 2018 | H-Scheduler: Storage-Aware Task Scheduling for Heterogeneous-Storage Spark ClustersabstractA trend in nowadays data centers is that heterogeneous storage devices are deployed to meet different storage demands of various big data workloads. For example, many nodes are equipped with both SSDs and HDDs. And HDFS has introduced the heterogeneous-storage-aware feature to adapt to such hybrid storage clusters. However, current task scheduler on big data processing platforms (such as Hadoop and Spark) only considers the overhead of network data transmission by exploiting the data locality principle. On heterogeneous storage clusters, task completion time is also affected by the speed of storage devices (SSDs and HDDs) where the data are stored. Ignoring the different speed of storage devices results in poor utilization of high speed devices such as SSD. In this paper, we propose a task scheduling strategy for heterogeneous storage clusters called H-Scheduler. The key idea of H-Scheduler is to differentiate speeds of storage devices by storage types. It classifies the tasks by both data locality and storage types, and redefines the priorities of different classes of tasks by both storage device speed and data locality to reduce job execution time. We implemented H-Scheduler in Spark, and the experiment results show that H-Scheduler can reduce job execution time by up to 73.6 %, depending on the workload characteristics and data distribution among different types of storage devices. Fengfeng Pan, Jin Xiong, Yijie Shen, Tianshi Wang 0002, Dejun Jiang 0001 |
ICPADS | 2 |
| 2017 | HiKV: A Hybrid Index Key-Value Store for DRAM-NVM Memory Systems
Dejun Jiang 0001, Jin Xiong, Ninghui Sun |
USENIX ATC | 3 |
| 2017 | dCompaction: Speeding up Compaction of the LSM-Tree via Delayed Compaction
Fengfeng Pan, Yinliang Yue, Jin Xiong |
J. Comput. Sci. Technol. | 3 |
| 2017 | HAP: Hybrid-Memory-Aware Partition in Shared Last-Level CacheabstractData-center servers benefit from large-capacity memory systems to run multiple processes simultaneously. Hybrid DRAM-NVM memory is attractive for increasing memory capacity by exploiting the scalability of Non-Volatile Memory (NVM). However, current LLC policies are unaware of hybrid memory. Cache misses to NVM introduce high cost due to long NVM latency. Moreover, evicting dirty NVM data suffer from long write latency. We propose hybrid memory aware cache partitioning to dynamically adjust cache spaces and give NVM dirty data more chances to reside in LLC. Experimental results show Hybrid-memory-Aware Partition (HAP) improves performance by 46.7% and reduces energy consumption by 21.9% on average against LRU management. Moreover, HAP averagely improves performance by 9.3% and reduces energy consumption by 6.4% against a state-of-the-art cache mechanism. Dejun Jiang 0001, Jin Xiong, Mingyu Chen 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2016 | Workload Shifting: Contention-Insular Disk Arrays for Big Data SystemsabstractIt is well known that in-place update index, unordered log structured index and ordered log structured index are three typical data organizations which are designed to meet different workload requirements respectively and wildly used in big data storage systems. Differentiated workload requirements in different phase of the data lifecycle, e.g. various types of data are injected into the big data storage systems in the write optimized manner, then they are needed to be read in the read optimized manner for analysis, lead to data organization transformation(data transformation for short). However, the simple mixture of foreground data injection and background data transformation causes serious disk contention. Frequent disk head seeks result in low disk throughput, and not only prolong the data transformation process, but also increase foreground data injection latency. In this paper, we propose \emph{Workload Shifting}, a novel log- structured design that shifts background data transformation away from the foreground data injection. Compared with conventional RAID0 disk array, \emph{Workload Shifting} effectively isolates background data transformation and foreground data injections, avoids the disk contention between them to boost their performance. We have implemented \emph{Workload Shifting} prototype on one multiple disks based disk array. Extensive experimental evaluation results show that compared with conventional RAID0 disk arrays, \emph{Workload Shifting} can avoid disk contention and speed up both data injection and data transformation significantly. Fengfeng Pan, Yinliang Yue, Jin Xiong |
NAS | 3 |
| 2016 | S-RAC: SSD Friendly Caching for Data Center WorkloadsabstractCurrent data-center applications tend to process increasingly large volume of data sets. The caching effect of page cache is reduced by its limited capacity. Emerging flash-based solid state drives (SSD) have latency and price advantages compared to hard disk and DRAM. Thus, SSD-based caching is widely deployed in data centers. However, SSD caching faces two challenges. First, SSD has limited write endurance, which requires cache manager to reduce write amount to SSD. Second, data-center workloads exhibit a diverse I/O access patterns, which requires one to figure out SSD caching friendly access patterns. This paper first classifies 6 I/O access patterns among 32 data-center workloads using a cost-benefit analysis. We derive implications for designing SSD cache from analyzing the access patterns. We then propose an SSD cache manager S-RAC with re-adding blocks and ghost cache adaptation to retain SSD friendly blocks in SSD. The experimental evaluation shows the efficiency of S-RAC in reducing SSD write amount while improving/maintaining cache hit ratio. Yuanjiang Ni, Ji Jiang, Dejun Jiang 0001, Xiaosong Ma, Jin Xiong, Yuangang Wang |
SYSTOR | 5 |
| 2015 | Exploiting Program Semantics to Place Data in Hybrid MemoryabstractLarge-memory applications like data analytics and graph processing benefit from extended memory hierarchies, and hybrid DRAM/NVM (non-volatile memory) systems represent an attractive means by which to increase capacity at reasonable performance/energy tradeoffs. Compared to DRAM, NVMs generally have longer latencies and higher energies for writes, which makes careful data placement essential for efficient system operation. Data placement strategies that resort to monitoring all data accesses and migrating objects to dynamically adjust data locations incur high monitoring overhead and unnecessary memory copies due to mispredicted migrations. We find that program semantics (specifically, global access characteristics) can effectively guide initial data placement with respect to memory types, which, in turn, makes run-time migration more efficient. We study a combined offline/online placement scheme that uses access profiling information to place objects statically and then selectively monitors run-time behaviors to optimize placements dynamically. We present a software/hardware cooperative framework, 2PP, and evaluate it with respect to state-of-the-art migratory placement, finding that it improves performance by an average of 12.1%. Furthermore, 2PP improves energy efficiency by up to 51.8%, and by an average of 18.4%. It does so by reducing run-time monitoring and migration overheads. Dejun Jiang 0001, Sally A. McKee, Jin Xiong, Mingyu Chen 0001 |
PACT | 4 |
| 2015 | A Survey of Phase Change Memory Systems
Dejun Jiang 0001, Jin Xiong, Ninghui Sun |
J. Comput. Sci. Technol. | 3 |
| 2014 | HAP: Hybrid-memory-Aware Partition in shared Last-Level CacheabstractData-center servers require large capacity main memory to run multiple workloads simultaneously. However, the scalability and power consumption of DRAM limit its capability of constructing large capacity memory. Emerging non-volatile memories (e.g. PCM and STT-RAM) provide better scalability and lower power leakage than DRAM. Especially, hybrid memory consisting of DRAM and NVM is able to exploit advantages of different memory medias. However, NVMs have a few drawbacks, such as relatively longer read and write latency. Cache miss at the shared last level cache (LLC) suffers from longer latency if the missing data resides in NVM. Current LLC policies manage the cache space without being aware of the underlying heterogeneous medias. This results in cache performance degradation if a large number of missing data come from NVM. Taking the asymmetric cache miss cost into account, we first propose a new performance metric -TMPKI, which can exactly reflect the LLC performance on the top of hybrid memories. Then we propose a hybrid memory aware cache partitioning technique (HAP) to dynamically adjust the cache spaces for DRAM and NVM data based on TMPKI. Experimental results show that HAP improves performance against the traditional LRU policy by up to 54.3% (19.6% on average) while it incurs a little storage overhead (0.2%). Dejun Jiang 0001, Jin Xiong, Mingyu Chen 0001 |
ICCD | 3 |
| 2014 | Write-aware random page initialization for non-volatile memory systemsabstractDue to the high scalability and low power leakage, emerging non-volatile memories (NVMs) are promising to be integrated into memory hierarchy. However, NVMs have write issues, such as limited write endurance and high write energy. This paper observes that a large fraction of writes are caused by memory page initialization in the OS kernel stack. Thus, this paper proposes a write-aware random page initialization technique (WRPI) to reduce writes without sacrificing system security. Instead of initializing all bits of an allocated page, WRPI randomly initializes part bits. Moreover, WRPI sets the values of initialized bits to zeros or ones that require writing the least number of bits. The evaluation results show that WRPI can reduce writing bits to NVM memory by up to 21.0% and 11.3% on average, compared to the conventional page initialization. WRPI can also reduce the write energy consumption and the total energy consumption of NVM memory system by 14.0% and 5.8% on average, respectively. Dejun Jiang 0001, Jin Xiong, Ninghui Sun |
ICCD | 3 |
| 2014 | DWC: dynamic write consolidation for phase change memory systemsabstractPhase change memory (PCM) is promising to become an alternative main memory thanks to its better scalability and lower leakage than DRAM. However, the long write latency of PCM puts it at a severe disadvantage against DRAM. In this paper, we propose a Dynamic Write Consolidation (DWC) scheme to improve PCM memory system performance while reducing energy consumption. This paper is motivated by the observation that a large fraction of a cache line being written back to memory is not actually modified. DWC exploits the unnecessary burst writes of unmodified data to consolidate multiple writes targeting the same row into one write. By doing so, DWC enables multiple writes to be send within one. DWC incurs low implementation overhead and shows significant efficiency. The evaluation results show that DWC achieves up to 35.7% performance improvement, and 17.9% on average. The effective write latency are reduced by up to 27.7%, and 16.0% on average. Moreover, DWC reduces the energy consumption by up to 35.3%, and 13.9% on average. Dejun Jiang 0001, Jin Xiong, Mingyu Chen 0001, Lixin Zhang 0002, Ninghui Sun |
ICS | 3 |
| 2014 | Pipelined Compaction for the LSM-TreeabstractWrite-optimized data structures like Log-Structured Merge-tree (LSM-tree) and its variants are widely used in key-value storage systems like Big Table and Cassandra. Due to deferral and batching, the LSM-tree based storage systems need background compactions to merge key-value entries and keep them sorted for future queries and scans. Background compactions play a key role on the performance of the LSM-tree based storage systems. Existing studies about the background compaction focus on decreasing the compaction frequency, reducing I/Os or confining compactions on hot data key-ranges. They do not pay much attention to the computation time in background compactions. However, the computation time is no longer negligible, and even the computation takes more than 60% of the total compaction time in storage systems using flash based SSDs. Therefore, an alternative method to speedup the compaction is to make good use of the parallelism of underlying hardware including CPUs and I/O devices. In this paper, we analyze the compaction procedure, recognize the performance bottleneck, and propose the Pipelined Compaction Procedure (PCP) to better utilize the parallelism of CPUs and I/O devices. Theoretical analysis proves that PCP can improve the compaction bandwidth. Furthermore, we implement PCP in real system and conduct extensive experiments. The experimental results show that the pipelined compaction procedure can increase the compaction bandwidth and storage system throughput by 77% and 62% respectively. Zigang Zhang, Yinliang Yue, Bingsheng He, Jin Xiong, Mingyu Chen 0001, Lixin Zhang 0002, Ninghui Sun |
IPDPS | 4 |
| 2013 | Performance analysis and resource allocation of heterogeneous cognitive gaussian relay channelsabstractMotivated by the deployment of cognitive radio (CR) based relays in cellular networks, this paper studies the fundamental limits of heterogeneous cognitive Gaussian relay channels (HCGRCs). Unlike conventional relay channels, in HCGRC a source transmits to a relay in a licensed band, while the relay transmits to a destination in an unlicensed cognitive spectrum band. The licensed and unlicensed bands are characterized by different power, bandwidth and reliability constraints. Taking an information-theoretic perspective, the fundamental properties of the HCGRC are analyzed thoroughly in terms of capacity, spectral efficiency (SE), and energy efficiency (EE). With regard to each metric, we derive the optimal resource allocation strategy and discuss the impacts of CR spectrum reliability and relay location on the metric. We find that in HCGRC, improving the SE and EE are not necessarily conflicting objectives. Instead, both metrics can be optimized simultaneously with proper resource allocation. Xuemin Hong, Jin Xiong, Jianghong Shi, Cheng-Xiang Wang 0001 |
GLOBECOM | 4 |
| 2012 | Mastiff: A MapReduce-based System for Time-Based Big Data AnalyticsabstractExisting MapReduce-based warehousing systems are not specially optimized for time-based big data analysis applications. Such applications have two characteristics: 1) data are continuously generated and are required to be stored persistently for a long period of time, 2) applications usually process data in some time period so that typical queries use time-related predicates. Time-based big data analytics requires both high data loading speed and high query execution performance. However, existing systems including current MapReduce-based solutions do not solve this problem well because the two requirements are contradictory. We have implemented a MapReduce-based system, called Mastiff, which provides a solution to achieve both high data loading speed and high query performance. Mastiff exploits a systematic combination of a column group store structure and a lightweight helper structure. Furthermore, Mastiff uses an optimized table scan method and a column-based query execution engine to boost query performance. Based on extensive experiments results with diverse workloads, we will show that Mastiff can significantly outperform existing systems including Hive, HadoopDB, and GridSQL. Sijie Guo, Jin Xiong, Rubao Lee |
CLUSTER | 2 |
| 2011 | HR-NET: A Highly Reliable Message-Passing Mechanism for Cluster File SystemabstractAs PC clusters increase in popularity and quantity, message-passing between nodes has been an important issue for high failure rate in the network. File access in a cluster file system often contains several sub-operations, each includes one or more network transmissions. Any network failures will cause the file system service unavailable. In this paper, we describe a highly reliable message-passing mechanism (HRNET), which tolerates both software and hardware network failures. HR-NET provides fine-grained, connection-level fail over across communication path redundancy. With it the file system can keep passing messages until it either recovers from network failures or it is failed over to a backup. Load balance for messages is also achieved to relieve network traffic. For transmission timeout, HR-NET proposes the message priority scheduling which dynamically manages messages in an appropriate order to tolerate request-response failures between clients and servers. As HR-NET is completely independent, there are neither any changes to standard protocol stacks nor modifications at upper file system. Performance results show that HR-NET takes full advantage of network bandwidth with average 6.17% throughput loss and provides a fast recovery. Experiments with cluster file system dispose that the overall performance degradation is below 8% due to failover of HR-NET while the reliability is highly enhanced. Can Ma, Jin Xiong |
NAS | 3 |
| 2011 | A Load-Aware Data Placement Policy on Cluster File System
Jin Xiong |
NPC | 3 |
| 2011 | Dawning Nebulae: A PetaFLOPS Supercomputer with a Heterogeneous Structure
Ninghui Sun, Zhigang Huo, Guangming Tan, Jin Xiong, Bo Li 0009, Can Ma |
J. Comput. Sci. Technol. | 5 |
| 2011 | Metadata Distribution and Consistency Techniques for Large-Scale Cluster File SystemsabstractMost supercomputers nowadays are based on large clusters, which call for sophisticated, scalable, and decentralized metadata processing techniques. From the perspective of maximizing metadata throughput, an ideal metadata distribution policy should automatically balance the namespace locality and even distribution without manual intervention. None of existing metadata distribution schemes is designed to make such a balance. We propose a novel metadata distribution policy, Dynamic Dir-Grain (DDG), which seeks to balance the requirements of keeping namespace locality and even distribution of the load by dynamic partitioning of the namespace into size-adjustable hierarchical units. Extensive simulation and measurement results show that DDG policies with a proper granularity significantly outperform traditional techniques such as the Random policy and the Subtree policy by 40 percent to 62 times. In addition, from the perspective of file system reliability, metadata consistency is an equally important issue. However, it is complicated by dynamic metadata distribution. Metadata consistency of cross-metadata server operations cannot be solved by traditional metadata journaling on each server. While traditional two-phase commit (2PC) algorithm can be used, it is too costly for distributed file systems. We proposed a consistent metadata processing protocol, S2PC-MP, which combines the two-phase commit algorithm with metadata processing to reduce overheads. Our measurement results show that S2PC-MP not only ensures fast recovery, but also greatly reduces fail-free execution overheads. Jin Xiong, Yiming Hu, Guojie Li, Rongfeng Tang, Zhihua Fan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2010 | Replication-Based Highly Available Metadata Management for Cluster File SystemsabstractIn cluster file systems, the metadata management is critical to the whole system. Past researches mainly focus on journaling which alone is not enough to provide high-available metadata service. Some others try to use replication, but the extra latency accompanied is a main problem. To guarantee both availability and efficiency, we propose a mechanism for building highly available metadata servers based on replication, which integrates Paxos algorithm effectively into metadata service. The Packed Multi-Paxos is proposed to reduce the latency brought by replication, which is self-adaptive and can make the replication to achieve high throughput under heavy client load and low latency under light client load. By designing efficient architecture and coordination mechanism, all replica server nodes simultaneously provide metadata read-access service. This high-available mechanism could decrease the impact of server failures and there is no interruption of service. The performance results show that the latency caused by replication and redundancy is well under control, and the performance of metadata read operation gains improvement. Zhuan Chen, Jin Xiong |
CLUSTER | 2 |
| 2009 | HMF: High-available Message-passing Framework for Cluster File SystemabstractIn large-scale cluster systems, the failure rate of network connection is non-negligibly high. A cluster file system must have the ability to handle network failures in order to provide high-available data accesses service. Traditionally, network failure handling is only guaranteed by network protocol, or implemented within the file system semantic layer. We present the high-available message-passing framework which is called HMF. Based on the operation hierarchy in cluster file system, HMF guarantees the availability of each pair of network transmissions and their interaction with the file system sub-operations. It separates the network fault-tolerance design from the file system and keeps a simple interface between them. HMF could handle a lot of network failures internally, which greatly simplifies the implementation of file system semantic layer. Performance results show that HMF can increase the availability of message passing and reduce the cost of recovery from network failures. When there are two network channels, HMF also improves aggregate I/O bandwidth by 80% in normal condition while the performance degradation due to recovery is below 10%. Zhuan Chen, Rongfeng Tang, Jin Xiong |
NAS | 4 |
| 2009 | Building Highly Available Cluster File System Based on ReplicationabstractIn order to gain better cost-effectiveness, current large-scale storage systems are typically built up by thousands of individual components. As systems scale up, the probability of the failure of multiple components increases. And for large-scale storage system, failures are normal rather than exception. How to build file systems providing both high throughput and highly available service under such circumstances is a big challenge. We have designed and implemented HA-DCFS3, a highly available cluster file system prototype. It uses a scalable replication algorithm called asynchronous primary copy protocol (APCP). Unlike traditional primary copy protocol that must synchronize updates to all replicas, APCP introduces an asynchronous approach where write operation is permitted to be synchronized to a subset of replicas. This flexible approach greatly improves the write performance. Furthermore, HA-DCFS3 also introduces a fine-grained failure detection called ¿ data path detection¿, which is integrated into the fault-tolerant framework based on data replication. Hence, HA-DCFS3 can provide continuous service even when component failures occur. And finally, HA-DCFS3 adopts a two-level data recovery strategy that handles transient failures with reintegration and persistent failures with re-replication respectively to reduce the cost of data repair. Our performance results show that HA-DCFS3 can deliver high and scalable aggregate performance and provide highly available service at very low cost. Jin Xiong |
PDCAT | 3 |
| 2009 | Adaptive and scalable metadata management to support a trillion filesabstractNowadays more and more applications require file systems to efficiently maintain million or more files. How to provide high access performance with such a huge number of files and such large directories is a big challenge for cluster file systems. Limited by static directory structures, existing file systems will be prohibitively inefficient for this use. To address this problem, we present a scalable and adaptive metadata management system which aims to maintain a trillion files efficiently. Firstly, our system exploits an adaptive two-level directory partitioning based on extendible hashing to manage very large directories. Secondly, our system utilizes fine-grained parallel processing within a directory and greatly improves performance of file creation or deletion. Thirdly, our system uses multiple-layered metadata cache management which improves memory utilization on the servers. And finally, our system uses a dynamic loadbalance mechanism based on consistent hashing which enables our system to scale up and down easily. Jin Xiong, Ninghui Sun |
SC | 2 |
| 2008 | A novel hint-based I/O mechanism for centralized file server of clusterabstractIn the medium and small cluster systems, the centralized file server such as NFS is the main approach to provide the storage service with low cost and easy management. However, when multiple parallel applications access the shared storage at the same time, the I/O performance decreases much because of the interference of the I/O requests coming from the different clients. In this paper, a hint-based I/O mechanism is proposed and implemented in the United-FS. By analyzing the hint information of the I/O requests, the related requests are grouped, sorted and scheduled by our hint-based I/O scheduler. The experiments show that our hint-based I/O mechanism nearly doubles the read performance compared with NFS, and has better scalability. Huan Chen 0003, Jin Xiong, Ninghui Sun |
CLUSTER | 2 |
| 2008 | Improving data availability for a cluster file system through replicationabstractData availability is a challenging issue for large-scale cluster file systems built upon thousands of individual storage devices. Replication is a well-known solution used to improve data availability. However, how to efficiently guarantee replicas consistency under concurrent conflict mutations remains a challenge. Moreover, how to quickly recover replica consistency from a storage server crash or storage device failure is also a tough problem. In this paper, we present a replication-based data availability mechanism designed for a large-scale cluster file system prototype named LionFS. Unlike other replicated storage systems that serialize replica updates, LionFS introduces a relaxed consistency model to enable concurrent updating all replicas for a mutation operation, greatly reducing the latency of operations. LionFS ensures replica consistency if applications use file locks to synchronize the concurrent conflict mutations. Another novelty of this mechanism is its light-weight log, which only records failed mutations and imposes no overhead on failure-free execution and low overhead when some storage devices are unavailable. Furthermore, recovery of replica consistency needs not stop the file system services and running applications. Performance evaluation shows that our solution achieves 50–70% higher write performance than serial replica updates. The logging overhead is shown to be low, and the recovery time is proportional to the amount of data written during the failure. Jin Xiong, Rongfeng Tang, Yiming Hu |
IPDPS | 1 |
| 2007 | United-FS: A Logical File System Providing a Single Image of Multiple Physical File Systems on NFS ServerabstractNFS is considered to be the bottleneck in cluster computing environment because of its limited resources and centralized data management. With the development of hardware, NFS server has more than one I/O channel, more storage space and more powerful CPU. In this paper, we describe the design and the implementation of a new logical file system called United-FS. It can make storage devices connected to multiple I/O channels work concurrently and cooperatively. It can be exported by NFS server to provide a single file system image to clients by hiding a variety of native file systems built on different type of storage devices. This paper also compares the United-FS with the software RAID system both from theoretical analysis and experiments. The results show that United-FS is much more flexible and its performance is better than software RAID in most cases. Huan Chen 0003, Yi Zhao 0013, Jin Xiong, Ninghui Sun |
IPDPS | 3 |
| 2006 | Trojan Horse Attack Strategy on Quantum Private Communication
Guangqiang He, Jin Xiong, Guihua Zeng |
ISPEC | 3 |
| 2005 | An Efficient Metadata Distribution Policy for Cluster File SystemsabstractHow to distribute the items in the file system hierarchy across a group of metadata servers is an important issue that determines the holistic metadata processing performance (HMPP) of a cluster file system which manages its metadata by a group of metadata servers. The HMPP is affected by two factors: balance degree of metadata distribution and number of branch points. Two types of well-used metadata distribution policies are the dynamic subtree policy and the random policy. Both of them emphasize one factor and neglect the other factor. As a result, their HMPP is low. In order to make good use of processing capacity of all metadata servers, we present a novel metadata distribution policy, called dynamic dir-grain (DDG) policy, which takes both factors into account. Our performance results show that this policy is potentially more efficient than the other two types of policies under real environments, as well as the conditions of creation or removal of a large hierarchy Jin Xiong, Rongfeng Tang, Sining Wu, Dan Meng 0002, Ninghui Sun |
CLUSTER | 1 |
| 2005 | A New Way to High Performance NFS for ClustersabstractFor its simplicity, reliability and maturity, NFS is widely-used in clusters. However, due to its high overheads and implementation limitations, the standard NFS cannot fully exert the potential abilities provided by multiple network channels and multiple SCSI channels on the server. In this paper, we present a new efficient way to high performance NFS implementation for cluster applications. By adding mechanisms to make good use of NFS server's multiple communication channels and multiple I/O channels, CluserNFS can potentially provide better I/O performance than standard NFS, as illustrated by our simulation experiment results. Rongfeng Tang, Jin Xiong |
PDCAT | 2 |
| 2004 | SuperNBD: An Efficient Network Storage Software for Cluster
Rongfeng Tang, Jin Xiong |
NPC | 3 |
| 2003 | Design and Performance of the Dawning Cluster File SystemabstractCluster file system is a key component of system software of clusters. It attracts more and more attention in recent years. In this paper, we introduce the design and implementation of DCFS (the Dawning Cluster File System) - a cluster file system developed for Dawning4000-L. DCFS is a global file system sharing among all cluster nodes. Applications see a single uniform name space, and can use system calls to access DCFS files. The features of DCFS include its scalable architecture, metadata policy, server-side optimization, flexible communication mechanism and easy management. Performance tests of DCFS on Dawning4000-L show that DCFS can provide high aggregate bandwidth and throughput. Jin Xiong, Sining Wu, Dan Meng 0002, Ninghui Sun, Guojie Li |
CLUSTER | 1 |