VLDB 2026 Research / reviewers in the wild / expert
Zehan Cui
dblp:21/11279
· DBLP profile ↗
12ranked-venue papers
3as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 70% Performance modeling and evaluation · 20% Storage systems · 10% | |
| Software engineering, system software, and programming languages
4 papers |
Operating systems · 75% Program analysis · 25% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Memory systems › virtual memory management
page coloring |
0.2 | 2 | 2014 | BPM/BPM+: Software-based dynamic memory partitioning mechanisms for mitigating DRAM bank-/channel-level interferences in multicore systems · ACM Trans. Archit. Code Optim. 2014 A Swap-based Cache Set Index Scheme to Leverage both Superpage and Page Coloring Optimizations · DAC 2014 |
Memory systems
cache |
0.2 | 1 | 2014 | A Swap-based Cache Set Index Scheme to Leverage both Superpage and Page Coloring Optimizations · DAC 2014 |
Memory systems › cache design
cache indexing |
0.2 | 1 | 2014 | A Swap-based Cache Set Index Scheme to Leverage both Superpage and Page Coloring Optimizations · DAC 2014 |
Performance modeling and evaluation
memory access traces |
0.2 | 1 | 2014 | HMTT: A hybrid hardware/software tracing system for bridging the DRAM access trace's semantic gap · ACM Trans. Archit. Code Optim. 2014 |
Memory systems
memory management |
0.2 | 1 | 2014 | Going vertical in memory management: Handling multiplicity by multi-policy · ISCA 2014 |
Memory systems › memory controller
memory scheduling |
0.2 | 1 | 2014 | BPM/BPM+: Software-based dynamic memory partitioning mechanisms for mitigating DRAM bank-/channel-level interferences in multicore systems · ACM Trans. Archit. Code Optim. 2014 |
Storage systems › data placement
vertical partitioning |
0.2 | 1 | 2014 | Going vertical in memory management: Handling multiplicity by multi-policy · ISCA 2014 |
Performance modeling and evaluation
workload characterization |
0.2 | 1 | 2014 | HMTT: A hybrid hardware/software tracing system for bridging the DRAM access trace's semantic gap · ACM Trans. Archit. Code Optim. 2014 |
Operating systems › resource management
memory management |
0.2 | 3 | 2014 | BPM/BPM+: Software-based dynamic memory partitioning mechanisms for mitigating DRAM bank-/channel-level interferences in multicore systems · ACM Trans. Archit. Code Optim. 2014 Going vertical in memory management: Handling multiplicity by multi-policy · ISCA 2014 A Swap-based Cache Set Index Scheme to Leverage both Superpage and Page Coloring Optimizations · DAC 2014 |
Program analysis › dynamic analysis
profiling |
0.1 | 1 | 2014 | HMTT: A hybrid hardware/software tracing system for bridging the DRAM access trace's semantic gap · ACM Trans. Archit. Code Optim. 2014 |
Memory systems › memory management › virtual memory › address translation
TLB |
0.1 | 1 | 2014 | A Swap-based Cache Set Index Scheme to Leverage both Superpage and Page Coloring Optimizations · DAC 2014 |
Memory systems › memory management
virtual memory |
0.1 | 1 | 2014 | A Swap-based Cache Set Index Scheme to Leverage both Superpage and Page Coloring Optimizations · DAC 2014 |
Methods — techniques the papers use, named apart from their topics
workload characterization · 0.4performance-monitoring unit monitoring · 0.4hardware snooping · 0.4event encoding · 0.4data mining · 0.4address space indirection · 0.4kernel modules · 0.2kernel module · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | A Swap-based Cache Set Index Scheme to Leverage both Superpage and Page Coloring OptimizationsabstractWe propose a novel cache set index scheme called SWAP (swap-based cache set index). SWAP introduces a pseudo-physical address space that is used by the operating system. The real physical address used for cache and main memory access is obtained by simply swapping some of superpage number bits with cache set index bits from the pseudo-physical address. By adding a level of indirection to the physical memory management, we simultaneously support both page coloring and superpage optimizations. These work together to improve TLB and shared LLC performance with negligible cost. Our results show that SWAP can improve performance by an average of 15.1% (by up to 25.2%) compared to 7.34% and 8.26% for superpage and page coloring, respectively. Zehan Cui, Licheng Chen, Yungang Bao, Mingyu Chen 0001 |
DAC | 1 |
| 2014 | DTail: a flexible approach to DRAM refresh managementabstractDRAM cells must be refreshed (or rewritten) periodically to maintain data integrity, and as DRAM density grows, so does the refresh time and energy. Not all data need to be refreshed with the same frequency, though, and thus some refresh operations can safely be delayed. Tracking such information allows the memory controller to reduce refresh costs by judiciously choosing when to refresh different rows Zehan Cui, Sally A. McKee, Zhongbin Zha, Yungang Bao, Mingyu Chen 0001 |
ICS | 1 |
| 2014 | Going vertical in memory management: Handling multiplicity by multi-policyabstractMany emerging applications from various domains often exhibit heterogeneous memory characteristics. When running in combination on parallel platforms, these applications present a daunting variety of workload behaviors that challenge the effectiveness of any memory allocation strategy. Prior partitioning-based or random memory allocation schemes typically manage only one level of the memory hierarchy and often target specific workloads. To handle diverse and dynamically changing memory and cache allocation needs, we augment existing “horizontal” cache/DRAM bank partitioning with vertical partitioning and explore the resulting multi-policy space. We study the performance of these policies for over 2000 workloads and correlate the results with application characteristics via a data mining approach. Based on this correlation we derive several practical memory allocation rules that we integrate into a unified multi-policy framework to guide resources partitioning and coalescing for dynamic and diverse multi-programmed/threaded workloads. We implement our approach in Linux kernel 2.6.32 as a restructured page indexing system plus a series of kernel modules. Extensive experiments show that, in practice, our framework can select proper memory allocation policy and consistently outperforms the unmodified Linux kernel, achieving up to 11% performance gains compared to prior techniques. Lei Liu 0030, Zehan Cui, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
ISCA | 3 |
| 2014 | CMD: classification-based memory deduplication through page access characteristicsabstractLimited main memory size is considered as one of the major bottlenecks in virtualization environments. Content-Based Page Sharing (CBPS) is an efficient memory deduplication technique to reduce server memory requirements, in which pages with same content are detected and shared into a single copy. As the widely used implementation of CBPS, Kernel Samepage Merging (KSM) maintains the whole memory pages into two global comparison trees (a stable tree and an unstable tree). To detect page sharing opportunities, each tracked page needs to be compared with pages already in these two large global trees. However since the vast majority of compared pages have different content with it, that will induce massive futility comparisons and thus heavy overhead. Licheng Chen, Zehan Cui, Mingyu Chen 0001, Haiyang Pan, Yungang Bao |
VEE | 3 |
| 2014 | MIMS: Towards a Message Interface Based Memory System
Licheng Chen, Mingyu Chen 0001, Yuan Ruan, Yongbing Huang, Zehan Cui, Tianyue Lu, Yungang Bao |
J. Comput. Sci. Technol. | 5 |
| 2014 | HMTT: A hybrid hardware/software tracing system for bridging the DRAM access trace's semantic gapabstractDRAM access traces (i.e., off-chip memory references) can be extremely valuable for the design of memory subsystems and performance tuning of software. Hardware snooping on the off-chip memory interface is an effective and nonintrusive approach to monitoring and collecting real-life DRAM accesses. However, compared with software-based approaches, hardware snooping approaches typically lack semantic information, such as process/function/object identifiers, virtual addresses, and lock contexts, that is essential to the complete understanding of the systems and software under investigation. In this article, we propose a hybrid hardware/software mechanism that is able to collect off-chip memory reference traces with semantic information. We have designed and implemented a prototype system called HMTT (Hybrid Memory Trace Tool), which uses a custom-made DIMM connector to collect off-chip memory references and a high-level event-encoding scheme to correlate semantic information with memory references. In addition to providing complete, undistorted DRAM access traces, the proposed system is also able to perform various types of low-overhead profiling, such as object-relative accesses and multithread lock accesses. Yongbing Huang, Licheng Chen, Zehan Cui, Yuan Ruan, Yungang Bao, Mingyu Chen 0001, Ninghui Sun |
ACM Trans. Archit. Code Optim. | 3 |
| 2014 | BPM/BPM+: Software-based dynamic memory partitioning mechanisms for mitigating DRAM bank-/channel-level interferences in multicore systemsabstractThe main memory system is a shared resource in modern multicore machines that can result in serious interference leading to reduced throughput and unfairness. Many new memory scheduling mechanisms have been proposed to address the interference problem. However, these mechanisms usually employ relative complex scheduling logic and need modifications to Memory Controllers (MCs), which incur expensive hardware design and manufacturing overheads. This article presents a practical software approach to effectively eliminate the interference without any hardware modifications. The key idea is to modify the OS memory management system and adopt a page-coloring-based Bank-level Partitioning Mechanism (BPM) that allocates dedicated DRAM banks to each core (or thread). By using BPM, memory requests from distinct programs are segregated across multiple memory banks to promote locality/fairness and reduce interference. We further extend BPM to BPM+ by incorporating channel-level partitioning, on which we demonstrate additional gain over BPM in many cases. To achieve benefits in the presence of diverse application memory needs and avoid performance degradation due to resource underutilization, we propose a dynamic mechanism upon BPM/BPM+ that assigns appropriate bank/channel resources based on application memory/bandwidth demands monitored through PMU (performance-monitoring unit) and a low-overhead OS page table scanning process. We implement BPM/BPM+ in Linux 2.6.32.15 kernel and evaluate the technique on four-core and eight-core real machines by running a large amount of randomly generated multiprogrammed and multithreaded workloads. Experimental results show that BPM/BPM+ can improve the overall system throughput by 4.7%/5.9%, on average, (up to 8.6%/9.5%) and reduce the unfairness by an average of 4.2%/6.1% (up to 15.8%/13.9%). Lei Liu 0030, Zehan Cui, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
ACM Trans. Archit. Code Optim. | 2 |
| 2013 | Scattered superpage: A case for bridging the gap between superpage and page coloringabstractSuperpage and page coloring are two important practical techniques to improve the performance of Translation Lookaside Buffers (TLBs) and shared Last Level Cache (LLC) respectively. However, there exists a gap between these two techniques in current hardware-architecture design, resulting in the contradiction in adopting these two optimizations simultaneously: a superpage requires hundreds of contiguous (e.g. a power of two) base pages in both virtual and physical memory, which would compulsorily occupy all available page colors (or cache sets), thus making page coloring failed to work. This is because most contemporary architecture adopts the design with cache set indexes placed in the least significant part of block address. In this paper, we propose a lightweight approach named Scattered Superpage to bridge this gap. Scattered Superpage decouples a superpage from the limitation of occupying multiple contiguous physical base pages. A superpage is still contiguous in virtual memory, but it is scattered mapping into multiple physical superpages, and it just occupies specified partial page colors in each physical superpage, thus it allows us to configure page color for each superpage. The huge TLB is slightly modified to store page color configuration for each superpage and to calculate target physical address based on this configuration when doing address translation. The experimental results show that the Scattered Superpage can improve system performance by 20.51% and reduce unfairness by 27.77% in our 4-core simulation system (with multi-program memory-intensive workloads). It achieves this by reducing last level cache miss by 17.05% and reducing TLB miss by 86.02% simultaneously. Licheng Chen, Zehan Cui, Yongbing Huang, Yungang Bao, Mingyu Chen 0001 |
ICCD | 3 |
| 2012 | HaLock: hardware-assisted lock contention detection in multithreaded applicationsabstractMultithreaded programming relies on locks to ensure the consistency of shared data. Lock contention is the main reason of low parallel efficiency and poor scalability of multithreaded programs. Lock profiling is the primary approach to detect lock contention. Prior lock profiling tools are able to track lock behaviors but directly store profiling data into local memory regardless of the memory interference on targeted programs. Yongbing Huang, Zehan Cui, Licheng Chen, Yungang Bao, Mingyu Chen 0001 |
PACT | 2 |
| 2012 | A software memory partition approach for eliminating bank-level interference in multicore systemsabstractMain memory system is a shared resource in modern multicore machines, resulting in serious interference, which causes performance degradation in terms of throughput slowdown and unfairness. Numerous new memory scheduling algorithms have been proposed to address the interference problem. However, these algorithms usually employ complex scheduling logic and need hardware modification to memory controllers, as a result, industrial venders seem to have some hesitation in adopting them. Lei Liu 0030, Zehan Cui, Mingjie Xing, Yungang Bao, Mingyu Chen 0001, Chengyong Wu |
PACT | 2 |
| 2012 | Evaluation and Optimization of Breadth-First Search on NUMA ClusterabstractGraph is widely used in many areas. Breadth-First Search (BFS), a key subroutine for many graph analysis algorithms, has become the primary benchmark for Graph500 ranking. Due to the high communication cost of BFS, multi-socket nodes with large memory capacity (NUMA) are supposed to reduce network pressure. However, the longer latency to remote memory may cause problem if not treated well. In this work, we first demonstrate that simply spawning and binding one MPI process for each socket can achieve the best performance for MPI/OpenMP hybrid programmed BFS algorithm, resulting in 1.53X of performance on 16 nodes. Nevertheless, we notice that one MPI process per socket may exacerbate the communication cost. We propose to share some communication data structure among the processes inside the same node, to eliminate most of the intra-node communication. To fully utilize the network bandwidth, we make all the processes in a node to perform communication simultaneously. We further adjust the granularity of a key bitmap for better cache locality to speed up the computation. With all the optimizations for NUMA, communication and computation together, 2.44X of performance is achieved on 16 nodes, which is 39.2 Billion Traversed Edges per Second for an R-MAT graph of scale 32 (4 billion vertices and 64 billion edges). Zehan Cui, Licheng Chen, Mingyu Chen 0001, Yungang Bao, Yongbing Huang, Huiwei Lv |
CLUSTER | 1 |
| 2012 | A lightweight hybrid hardware/software approach for object-relative memory profilingabstractMemory profiling is the process of collecting memory address traces during the execution of a program, then analyzing and characterizing the memory behavior of the program offline. With the trend that there will be more and more cores integrated in a processor chip, the “Memory Wall” problem will become more serious in the chip multiprocessor (CMP) system. Thus accurate and effective memory profiling is becoming one of the keys to identify the source of memory system bottlenecks. A large body of work has been contributed to memory profiling, however, most adopts instrumentation, simulator which suffers heavy overhead, or hardware performance counter which is lack of detail trace information. Furthermore, correlating the raw memory address traces with object-relative information allows us to separate regular pattern for certain object from the irregular mixed, thus helps the optimization. In this paper, we propose a lightweight hybrid hardware/software approach for object-relative memory profiling. We monitor physical memory addresses through hardware snooping with negligible overhead; meanwhile we dump Linux kernel page tables of processes, as well as object-relative memory allocation information. Our approach supports not only to collect applications' full memory traces with detail object relative information, but also to identify hardware-generated memory accesses such as page memory walks due to TLB miss at object level. The experimental results on real system show that our approach is highly accurate (the largest error is 2.04%) and low overhead (the average overhead is 1.60%). Furthermore, we profile two multi-thread applications in detail, and successfully identity hot TLB-miss objects. With object-targeted optimization, we can improve applications' performance by nearly 6.86%. Licheng Chen, Zehan Cui, Yungang Bao, Mingyu Chen 0001, Yongbing Huang, Guangming Tan |
ISPASS | 2 |