Jinglei Ren

dblp:67/8972 · DBLP profile ↗
← Back
14ranked-venue papers
4as first author
0since 2021 · last 2019
0000-0001-8386-117XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 3 first-authorSoftware engineering, systems software and programming languages · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Memory systems · 38% Storage systems · 31% Distributed systems · 19%
Software engineering, system software, and programming languages
4 papers
Concurrent programming · 51% Software maintenance and evolution · 22% Empirical software engineering · 13%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
non-volatile memory
1.662019
Dual-Page Checkpointing: An Architectural Approach to Efficient Data Persistence for In-Memory Applications · ACM Trans. Archit. Code Optim. 2019
DudeTx: Durable Transactions Made Decoupled · ACM Trans. Storage 2018
Crash Consistency in Encrypted Non-volatile Main Memory Systems · HPCA 2018
Storage systems
crash consistency
0.932018
DudeTx: Durable Transactions Made Decoupled · ACM Trans. Storage 2018
Crash Consistency in Encrypted Non-volatile Main Memory Systems · HPCA 2018
ThyNVM: enabling software-transparent crash consistency in persistent memory systems · MICRO 2015
Memory systems › non-volatile memory
persistent memory
0.832018
DudeTx: Durable Transactions Made Decoupled · ACM Trans. Storage 2018
DudeTM: Building Durable Transactions with Decoupling for Persistent Memory · ASPLOS 2017
ThyNVM: enabling software-transparent crash consistency in persistent memory systems · MICRO 2015
Storage systems › transaction support
durable transactions
0.622018
DudeTx: Durable Transactions Made Decoupled · ACM Trans. Storage 2018
DudeTM: Building Durable Transactions with Decoupling for Persistent Memory · ASPLOS 2017
Distributed systems
fault tolerance
0.622018
Crash Consistency in Encrypted Non-volatile Main Memory Systems · HPCA 2018
Realizing the Fault-Tolerance Promise of Cloud Storage Using Locks with Intent · OSDI 2016
Concurrent programming
concurrency bugs
0.522016
A Lightweight System for Detecting and Tolerating Concurrency Bugs · IEEE Trans. Software Eng. 2016
Fixing, preventing, and recovering from concurrency bugs · Sci. China Inf. Sci. 2015
Concurrent programming › concurrency bugs
atomicity violation
0.422016
A Lightweight System for Detecting and Tolerating Concurrency Bugs · IEEE Trans. Software Eng. 2016
AI: a lightweight system for tolerating concurrency bugs · SIGSOFT FSE 2014
Distributed systems › fault tolerance
checkpointing
0.412019
Dual-Page Checkpointing: An Architectural Approach to Efficient Data Persistence for In-Memory Applications · ACM Trans. Archit. Code Optim. 2019
Empirical software engineering
mining software repositories
0.312018
Towards quantifying the development value of code contributions · ESEC/SIGSOFT FSE 2018
Memory systems › memory encryption
counter-mode encryption
0.312018
Crash Consistency in Encrypted Non-volatile Main Memory Systems · HPCA 2018
Storage systems
crash recovery
0.312018
Crash Consistency in Encrypted Non-volatile Main Memory Systems · HPCA 2018
Storage systems
logging
0.312018
DudeTx: Durable Transactions Made Decoupled · ACM Trans. Storage 2018
Memory systems
memory encryption
0.312018
Crash Consistency in Encrypted Non-volatile Main Memory Systems · HPCA 2018
Cloud and datacenter computing
cloud storage
0.212016
Realizing the Fault-Tolerance Promise of Cloud Storage Using Locks with Intent · OSDI 2016
Software maintenance and evolution › software maintenance
bug fixing
0.212015
Fixing, preventing, and recovering from concurrency bugs · Sci. China Inf. Sci. 2015
Storage systems › storage devices › storage media
mobile storage
0.212015
Memory-Centric Data Storage for Mobile Systems · USENIX ATC 2015
Concurrent programming
concurrency bug detection
0.212014
AI: a lightweight system for tolerating concurrency bugs · SIGSOFT FSE 2014
Distributed systems
data aggregation
0.212014
Quatrain: Accelerating Data Aggregation between Multiple Layers · IEEE Trans. Computers 2014
Distributed systems › operating system support
interprocess communication
0.212014
Quatrain: Accelerating Data Aggregation between Multiple Layers · IEEE Trans. Computers 2014
Performance modeling and evaluation › statistical analysis
outlier detection
0.212014
NO2: Speeding up Parallel Processing of Massive Compute-Intensive Tasks · IEEE Trans. Computers 2014
Distributed systems
remote procedure call
0.212014
Quatrain: Accelerating Data Aggregation between Multiple Layers · IEEE Trans. Computers 2014
Processor architecture and microarchitecture
speculative execution
0.212014
NO2: Speeding up Parallel Processing of Massive Compute-Intensive Tasks · IEEE Trans. Computers 2014
Storage systems › storage reliability
durability
0.112019
Dual-Page Checkpointing: An Architectural Approach to Efficient Data Persistence for In-Memory Applications · ACM Trans. Archit. Code Optim. 2019
Hardware security and side channels
memory encryption
0.112018
Crash Consistency in Encrypted Non-volatile Main Memory Systems · HPCA 2018
Parallel and multicore computing
transactional memory
0.112018
DudeTx: Durable Transactions Made Decoupled · ACM Trans. Storage 2018
Memory systems › non-volatile memory › persistent memory
byte-addressable persistent memory
0.112017
DudeTM: Building Durable Transactions with Decoupling for Persistent Memory · ASPLOS 2017
Storage systems › non-volatile memory storage
persistent memory systems
0.112017
Log-Structured Non-Volatile Main Memory · USENIX ATC 2017
Program analysis › static analysis
bug detection
0.112016
A Lightweight System for Detecting and Tolerating Concurrency Bugs · IEEE Trans. Software Eng. 2016
Distributed systems
replication
0.112016
Realizing the Fault-Tolerance Promise of Cloud Storage Using Locks with Intent · OSDI 2016
Embedded and real-time systems
mobile computing
0.112015
Memory-Centric Data Storage for Mobile Systems · USENIX ATC 2015

Methods — techniques the papers use, named apart from their topics

undo logging · 0.9redo logging · 0.9versioning · 0.7selective counter-atomicity · 0.7decoupling · 0.6thread stalling · 0.4anticipating invariant · 0.4logging · 0.4dual-page checkpointing · 0.4copy-on-write · 0.4shadow DRAM · 0.3learning-to-rank · 0.3hardware transactional memory · 0.3
YearPublicationVenuePosition
2019 Dual-Page Checkpointing: An Architectural Approach to Efficient Data Persistence for In-Memory Applications
abstract
Data persistence is necessary for many in-memory applications. However, the disk-based data persistence largely slows down in-memory applications. Emerging non-volatile memory (NVM) offers an opportunity to achieve in-memory data persistence at the DRAM-level performance. Nevertheless, NVM typically requires a software library to operate NVM data, which brings significant overhead. This article demonstrates that a hardware-based high-frequency checkpointing mechanism can be used to achieve efficient in-memory data persistence on NVM. To maintain checkpoint consistency, traditional logging and copy-on-write techniques incur excessive NVM writes that impair both performance and endurance of NVM; recent work attempts to solve the issue but requires a large amount of metadata in the memory controller. Hence, we design a new dual-page checkpointing system, which achieves low metadata cost and eliminates most excessive NVM writes at the same time. It breaks the traditional trade-off between metadata space cost and extra data writes. Our solution outperforms the state-of-the-art NVM software libraries by 13.6× in throughput, and leads to 34% less NVM wear-out and 1.28× higher throughput than state-of-the-art hardware checkpointing solutions, according to our evaluation with OLTP, graph computing, and machine-learning workloads.
Song Wu 0001, Hai Jin 0001, Jinglei Ren
ACM Trans. Archit. Code Optim.5
2018 Crash Consistency in Encrypted Non-volatile Main Memory Systems
abstract
Non-Volatile Main Memory (NVMM) systems provide high performance by directly manipulating persistent data in-memory, but require crash consistency support to recover data in a consistent state in case of a power failure or system crash. In this work, we focus on the interplay between the crash consistency mechanisms and memory encryption. Memory encryption is necessary for these systems to protect data against the attackers with physical access to the persistent main memory. As decrypting data at every memory read access can significantly degrade the performance, prior works propose to use a memory encryption technique, counter-mode encryption, that reduces the decryption overhead by performing a memory read access in parallel with the decryption process using a counter associated with each cache line. Therefore, a pair of data and counter value is needed to correctly decrypt data after a system crash. We demonstrate that counter-mode encryption does not readily extend to crash consistent NVMM systems as the system will fail to recover data in a consistent state if the encrypted data and associated counter are not written back to memory atomically, a requirement we refer to as counter-atomicity. We show that näıvely enforcing counter-atomicity for all NVMM writes can serialize memory accesses and results in a significant performance degradation. In order to improve the performance, we make an observation that not all writes to NVMM need to be counter-atomic. The crash consistency mechanisms rely on versioning to keep one consistent copy of data intact while manipulating another version directly in-memory. As the recovery process only relies on the unmodified consistent version, it is not necessary to strictly enforce counter-atomicity for the writes that do not affect data recovery. Based on this insight, we propose selective counter-atomicity that allows reordering of writes to data and associated counters when the writes to persistent memory do not alter the recoverable consistent state. We propose efficient software and hardware support to enforce selective counter-atomicity. Our evaluation demonstrates that in a 1/2/4/8- core system, selective counter-atomicity improves performance by 6/11/22/40% compared to a system that enforces counter-atomicity for all NVMM writes. The performance of our selective counter-atomicity design comes within 5% of an ideal NVMM system that provides crash consistency of encrypted data at no cost.
Sihang Liu 0001, Aasheesh Kolli, Jinglei Ren, Samira Manabi Khan
HPCA3
2018 Towards quantifying the development value of code contributions
abstract
Quantifying the value of developers’ code contributions to a software project requires more than simply counting lines of code or commits. We define the development value of code as a combination of its structural value (the effect of code reuse) and its non-structural value (the impact on development). We propose techniques to automatically calculate both components of development value and combine them using Learning to Rank. Our preliminary empirical study shows that our analysis yields richer results than those obtained by human assessment or simple counting methods and demonstrates the potential of our approach.
Jinglei Ren, Hezheng Yin, Qingda Hu, Armando Fox, Wojciech Koszek
ESEC/SIGSOFT FSE1
2018 DudeTx: Durable Transactions Made Decoupled
abstract
Emerging non-volatile memory (NVM) offers non-volatility, byte-addressability, and fast access at the same time. It is suggested that programs should access NVM directly through CPU load and store instructions. To guarantee crash consistency, durable transactions are regarded as a common choice of applications for accessing persistent memory data. However, existing durable transaction systems employ either undo logging , which requires a fence for every memory write, or redo logging , which requires intercepting all memory reads within transactions. Both approaches incur significant overhead. This article presents D ude T x , a crash-consistent durable transaction system that avoids the drawbacks of both undo and redo logging. D ude T x uses shadow DRAM to decouple the execution of a durable transaction into three fully asynchronous steps. The advantage is that only minimal fences and no memory read instrumentation are required. This design enables an out-of-the-box concurrency control mechanism, transactional memory or fine-grained locks, to be used as an independent component. The evaluation results show that D ude T x adds durability to a software transactional memory system with only 7.4%--24.6% throughput degradation. Compared to typical existing durable transaction systems, D ude T x provides 1.7× --4.4× higher throughput. Moreover, D ude T x can be implemented with hardware transactional memory or lock-based concurrency control, leading to a further 1.7× and 3.3× speedup, respectively.
Mengxing Liu, Kang Chen 0001, Xuehai Qian, Yongwei Wu 0001, Jinglei Ren
ACM Trans. Storage7
2017 DudeTM: Building Durable Transactions with Decoupling for Persistent Memory
abstract
Emerging non-volatile memory (NVM) offers non-volatility, byte-addressability and fast access at the same time. To make the best use of these properties, it has been shown by empirical evidence that programs should access NVM directly through CPU load and store instructions, so that the overhead of a traditional file system or database can be avoided. Thus, durable transactions become a common choice of applications for accessing persistent memory data in a crash consistent manner. However, existing durable transaction systems employ either undo logging, which requires a fence for every memory write, or redo logging, which requires intercepting all memory reads within transactions.
Mengxing Liu, Kang Chen 0001, Xuehai Qian, Yongwei Wu 0001, Jinglei Ren
ASPLOS7
2017 Log-Structured Non-Volatile Main Memory
Qingda Hu, Jinglei Ren, Anirudh Badam, Jiwu Shu, Thomas Moscibroda
USENIX ATC2
2016 Realizing the Fault-Tolerance Promise of Cloud Storage Using Locks with Intent
Srinath Setty, Chunzhi Su, Jacob R. Lorch, Lidong Zhou, Hao Chen 0030, Parveen Patel, Jinglei Ren
OSDI7
2016 A Lightweight System for Detecting and Tolerating Concurrency Bugs
abstract
Along with the prevalence of multi-threaded programs, concurrency bugs have become one of the most important sources of software bugs. Even worse, due to the non-deterministic nature of concurrency bugs, these bugs are both difficult to detect and fix even after the detection. As a result, it is highly desired to develop an all-around approach that is able to not only detect them during the testing phase but also tolerate undetected bugs during production runs. However, existing bug-detecting and bug-tolerating tools are usually either1)constrained in types of bugs they can handle or2)requiring specific hardware supports for achieving an acceptable overhead. In this paper, we present a novel program invariant, name Anticipating Invariant (Ai), that can detect most types of concurrency bugs. More importantly,Aican be used to anticipate many concurrency bugs before any irreversible changes have been made. Thus it enables us to develop a software-only system that is able to forestall failures with a simple thread stalling technique, which does not rely on execution roll-back and hence has good performance. Experiments with 35 real-world concurrency bugs demonstrate thatAiis capable of detecting and tolerating many important types of concurrency bugs, including both atomicity and order violations. It has also exposed two new bugs (confirmed by developers) that were never reported before in the literature. Performance evaluation with 6 representative parallel programs shows thatAiincurs negligible overhead ($ < 1\%$) for many nontrivial desktop and server applications.
Yongwei Wu 0001, Shan Lu 0001, Shanxiang Qi, Jinglei Ren
IEEE Trans. Software Eng.5
2015 ThyNVM: enabling software-transparent crash consistency in persistent memory systems
abstract
Emerging byte-addressable nonvolatile memories (NVMs) promise persistent memory, which allows processors to directly access persistent data in main memory. Yet, persistent memory systems need to guarantee a consistent memory state in the event of power loss or a system crash (i.e., crash consistency). To guarantee crash consistency, most prior works rely on programmers to (1) partition persistent and transient memory data and (2) use specialized software interfaces when updating persistent memory data. As a result, taking advantage of persistent memory requires significant programmer effort, e.g., to implement new programs as well as modify legacy programs. Use cases and adoption of persistent memory can therefore be largely limited.
Jinglei Ren, Jishen Zhao, Samira Manabi Khan, Jongmoo Choi, Yongwei Wu 0001, Onur Mutlu
MICRO1
2015 Memory-Centric Data Storage for Mobile Systems
Jinglei Ren, Chieh-Jan Mike Liang, Yongwei Wu 0001, Thomas Moscibroda
USENIX ATC1
2015 Fixing, preventing, and recovering from concurrency bugs
Dongdong Deng, Guoliang Jin, Marc de Kruijf, Ben Liblit, Shan Lu 0001, Shanxiang Qi, Jinglei Ren, Karthikeyan Sankaralingam, Linhai Song, Yongwei Wu 0001, Wei Zhang 0022
Sci. China Inf. Sci.8
2014 AI: a lightweight system for tolerating concurrency bugs
abstract
Concurrency bugs are notoriously difficult to eradicate during software testing because of their non-deterministic nature. Moreover, fixing concurrency bugs is time-consuming and error-prone. Thus, tolerating concurrency bugs during production runs is an attractive complementary approach to bug detection and testing. Unfortunately, existing bug-tolerating tools are usually either 1) constrained in types of bugs they can handle or 2) requiring roll-back mechanism, which can hitherto not be fully achieved efficiently without hardware supports. This paper presents a novel program invariant, called Anticipating Invariant (AI), which can help anticipate bugs before any irreversible changes are made. Benefiting from this ability of anticipating bugs beforehand, our software-only system is able to forestall the failures with a simple thread stalling technique, which does not rely on execution roll-back and hence has good performance Experiments with 35 real-world concurrency bugs demonstrate that AI is capable of detecting and tolerating most types of concurrency bugs, including both atomicity and order violations. Two new bugs have been detected and confirmed by the corresponding developers. Performance evaluation with 6 representative parallel programs shows that AI incurs negligible overhead (<1%) for many nontrivial desktop and server applications.
Yongwei Wu 0001, Shan Lu 0001, Shanxiang Qi, Jinglei Ren
SIGSOFT FSE5
2014 Quatrain: Accelerating Data Aggregation between Multiple Layers
abstract
Composition of multiple layers (or components/services) has been a dominant practice in building distributed systems, meanwhile aggregation has become a typical pattern of data flows nowadays. However, the efficiency of data aggregation is usually impaired by multiple layers due to amplified delay. Current solutions based on data/execution flow optimization mostly counteract flexibility, reusability, and isolation of layers abstraction. Otherwise, programmers have to do much error-prone manual programming to optimize communication, and it is complicated in a multithreaded environment. To resolve the dilemma, we propose a new style of inter-process communication that not only optimizes data aggregation but also retains the advantages of layered (or component-based/service-oriented) architecture. Our approach relaxes the traditional definition of procedure and allows a procedure to return multiple times. Specifically, we implement an extended remote procedure calling framework Quatrain to support the new multireturn paradigm. In this paper, we establish the importance of multiple returns, introduce our very simple semantics, and present a new synchronization protocol that frees programmers from multireturn-related thread coordination. Several practical applications are constructed with Quatrain, and the evaluation shows an average of 56% reduction of response time, compared with the traditional calling paradigm, in realistic environments.
Jinglei Ren, Yongwei Wu 0001
IEEE Trans. Computers1
2014 NO2: Speeding up Parallel Processing of Massive Compute-Intensive Tasks
abstract
Large-scale computing frameworks, either tenanted on the cloud or deployed in the high-end local cluster, have become an indispensable software infrastructure to support numerous enterprise and scientific applications. Tasks executed on these frameworks are generally classified into data-intensive and compute-intensive ones. However, most existing frameworks, led by MapReduce, are mainly suitable for data-intensive tasks. Their task schedulers assume that the proportion of data I/O reflects the task progress and state. Unfortunately, this assumption does not apply to most compute-intensive tasks. Due to biased estimation of task progress, traditional frameworks cannot timely cut off outliers and therefore largely prolong execution time when performing compute-intensive tasks. We propose a new framework designed for compute-intensive tasks. By using instrumentation and automatic instrument point selector, our framework estimates the compute-intensive task progress without resorting to data I/O. We employ a clustering method to identify outliers at runtime and perform speculative execution/aborting, speeding up task execution by up to 25%. Moreover, our improvement to bare instrumentation limits overhead within 0.1%, and the aborting-based execution only introduces 10% more average CPU usage. Low overhead and resource consumption make our framework practically usable in the production environment.
Yongwei Wu 0001, Weichao Guo, Jinglei Ren
IEEE Trans. Computers3