VLDB 2026 Research / reviewers in the wild / expert
Yu Chen 0004
dblp:87/1254-4
· DBLP profile ↗
42ranked-venue papers
0as first author
10since 2021 · last 2026
0009-0007-8455-3663ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 6 since 2021Software engineering, systems software and programming languages · 10 · 4 since 2021Security and privacy · 6 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SKernel: An Elastic and Efficient Secure Container System at Scale with a Split-Kernel ArchitectureabstractSecure containers leverage hardware virtualization to isolate container sandboxes, enabling dedicated guest kernels to mitigate shared kernel attacks prevalent in traditional systems. However, existing approaches struggle with a fundamental trade-off: VM-based solutions (e.g., Kata) prioritize performance but lack elasticity and on-demand usage for volatile and bursty workloads, while lightweight methods (e.g., gVisor) rely on the host kernel for dynamic resource management at the cost of significant performance degradation due to guest-host dependencies. Xiaohu Chai, Keyang Hu, Jianfeng Tan, Tiwei Bie, Guotao Tan, Anqi Shen, Dawei Shen, Xinyao Yang, Zhengyu He, Dong Du 0003, Yubin Xia, Kang Chen 0001, Yu Chen 0004 |
EuroSys | 17 |
| 2025 | Fork in the Road: Reflections and Optimizations for Cold Start Latency in Production Serverless Systems
Xiaohu Chai, Keyang Hu, Jianfeng Tan, Tiwei Bie, Anqi Shen, Dawei Shen, Qi Xing, Shun Song, Tongkai Yang, Zhengyu He, Dong Du 0003, Yubin Xia, Kang Chen 0001, Yu Chen 0004 |
OSDI | 17 |
| 2024 | Verifying Rust Implementation of Page Tables in a Software Enclave HypervisorabstractAs trusted execution environments (TEE) have become the corner stone for secure cloud computing, it is critical that they are reliable and enforce proper isolation, of which a key ingredient is spatial isolation. Many TEEs are implemented in software such as hypervisors for flexibility, and in a memory-safe language, namely Rust to alleviate potential memory bugs. Still, even if memory bugs are absent from the TEE, it may contain semantic errors such as mis-configurations in its memory subsystem which breaks spatial isolation. Zhenyang Dai, Vilhelm Sjöberg, Xupeng Li, Yu Chen 0004, Wenhao Wang 0001, Yuekai Jia, Sean Noble Anderson, Laila Elbeheiry, Shubham Sondhi, Yu Zhang 0313, Zhaozhong Ni, Shoumeng Yan, Ronghui Gu, Zhengyu He |
ASPLOS (2) | 5 |
| 2024 | Intelligent Hybrid Memory Scheduling Based on Page Pattern RecognitionabstractHybrid memory systems exhibit disparities in their heterogeneous memory components' access speeds. Dynamic page scheduling to ensure memory access predominantly occurs in the faster memory components is essential for optimizing the performance of hybrid memory systems. Recent works attempt to optimize page scheduling by predicting their hotness using neural network models. However, they face two crucial challenges: the page explosion problem and the new pages problem. We propose an intelligent hybrid memory scheduler driven by page pattern recognition to address these two challenges. Experimental results demonstrate that our approach outperforms state-of-the-art intelligent schedulers regarding effectiveness and cost. Yanjie Zhen, Weining Chen, Wei Gao 0006, Ju Ren 0001, Kang Chen 0001, Yu Chen 0004 |
DATE | 6 |
| 2024 | Skyloft: A General High-Efficient Scheduling Framework in User SpaceabstractSkyloft is a general and highly efficient user-space scheduling framework. It leverages user-mode interrupt to deliver and process hardware timers directly in user space. This capability enables Skyloft to achieve μs-scale preemption. Skyloft offers a set of scheduling interfaces that supports different scheduling policies, including both preemptive and nonpreemptive ones. Operating as a user-space scheduling framework, Skyloft is compatible with Linux and integrates seamlessly with high-performance I/O frameworks like DPDK. Yuekai Jia, Kaifu Tian, Yuyang You, Yu Chen 0004, Kang Chen 0001 |
SOSP | 4 |
| 2024 | Unishyper: A Rust-based unikernel enhancing reliability and efficiency of embedded systems
Keyang Hu, Wang Huang, Lei Wang 0126, Ce Mo, Runxiang Wang, Yu Chen 0004, Ju Ren 0001, Bo Jiang 0001 |
J. Syst. Archit. | 6 |
| 2024 | PatternS: An intelligent hybrid memory scheduler driven by page pattern recognition
Yanjie Zhen, Weining Chen, Wei Gao 0006, Ju Ren 0001, Kang Chen 0001, Yu Chen 0004 |
J. Syst. Archit. | 6 |
| 2023 | A Low-Cost and Pages-Interrelation-Aware Attention Model for Hybrid Memory SchedulingabstractHybrid memory architecture has become an important solution to address the increasing demand for the main memory capacity of big data applications. Due to the varying properties of different components in hybrid memory, accurately predicting the hotness of pages and timely scheduling hot pages to fast memory becomes crucial for optimal performance. However, existing hybrid memory schedulers using non-intelligent policy exhibit low performance. Although schedulers employing neural models can improve performance, they suffer limitations such as long inference time and loss of interrelation between pages. This paper presents PI-Attention, a low-cost and pages-interrelation-aware attention model for hybrid memory scheduling. It addresses the limitations above by utilizing two attention modules in the page and time sequence dimensions. Our experiments show that PI-Attention brings 11.14% performance improvement and a 3.75x reduction in inference time. Yanjie Zhen, Yu Chen 0004 |
SMC | 2 |
| 2022 | HyperEnclave: An Open and Cross-platform Trusted Execution Environment
Yuekai Jia, Wenhao Wang 0001, Yu Chen 0004, Zhengde Zhai, Shoumeng Yan, Zhengyu He |
USENIX ATC | 4 |
| 2021 | Scaling Camouflage: Content Disguising Attack Against Computer Vision ApplicationsabstractRecently, deep neural networks have achieved state-of-the-art performance in multiple computer vision tasks, and become core parts of computer vision applications. In most of their implementations, a standard input preprocessing component called image scaling is embedded, in order to resize the original data to match the input size of pre-trained neural networks. This article demonstrates content disguising attacks by exploiting the image scaling procedure, which cause machine's extracted content to be dramatically dissimilar with that before scaled. Different from previous adversarial attacks, our attacks happen in the data preprocessing stage, and hence they are not subject to specific machine learning models. To achieve a better deceiving and disguising effect, we propose and implement three feasible attack approaches with L0- and L∞-norm distance metrics. We have conducted a comprehensive evaluation on various image classification applications, including three local demos and two remote proprietary services. We also investigate the attack effects on a YOLO-v3 object detection demo. Our experimental results demonstrate successful content disguising against all of them, which validate our approaches are practical. Yufei Chen 0001, Chao Shen 0001, Cong Wang 0001, Qixue Xiao, Kang Li 0001, Yu Chen 0004 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2020 | Occlum: Secure and Efficient Multitasking Inside a Single Enclave of Intel SGXabstractIntel Software Guard Extensions (SGX) enables user-level code to create private memory regions called enclaves, whose code and data are protected by the CPU from software and hardware attacks outside the enclaves. Recent work introduces library operating systems (LibOSes) to SGX so that legacy applications can run inside enclaves with few or even no modifications. As virtually any non-trivial application demands multiple processes, it is essential for LibOSes to support multitasking. However, none of the existing SGX LibOSes support multitasking both securely and efficiently. Youren Shen, Hongliang Tian, Yu Chen 0004, Kang Chen 0001, Runji Wang, Yubin Xia, Shoumeng Yan |
ASPLOS | 3 |
| 2020 | Automatic kernel code synthesis and verification
Qiang Zhang 0032, Jianzhong Qiao, Qingyang Meng, Yu Chen 0004 |
Comput. Secur. | 4 |
| 2019 | Seeing is Not Believing: Camouflage Attacks on Image Scaling Algorithms
Qixue Xiao, Yufei Chen 0001, Chao Shen 0001, Yu Chen 0004, Kang Li 0001 |
USENIX Security Symposium | 4 |
| 2017 | pbSE: Phase-Based Symbolic ExecutionabstractThe study of software bugs has long been a key area in software security. Dynamic symbolic execution, in exploring the program's execution paths, finds bugs by analyzing all potential dangerous operations. Due to its high coverage and abilities to generate effective testcases, dynamic symbolic execution has attracted wide attention in the research community. However, the success of dynamic symbolic execution is limited due to complex program logic and its difficulty to handle large symbolic data. In our experiments we found that phase-related features of a program often prevents dynamic symbolic execution from exploring deep paths. On the basis of this discovery, we proposed a novel symbolic execution technology guided by program phase characteristics. Compared to KLEE, the most well-known symbolic execution approach, our method is capable of covering more code and discovering more bugs. We designed and implemented pbSE system, which was used to test several commonly used tools and libraries in Linux. Our results showed that pbSE on average covers code twice as much as what KLEE does, and we discovered 21 previously unknown vulnerabilities by using pbSE, out of which 7 are assigned CVE IDs. Qixue Xiao, Yu Chen 0004, Chengang Wu, Kang Li 0001, Junjie Mao, Shize Guo, Yuanchun Shi |
DSN | 2 |
| 2016 | Scalable Kernel TCP Design and Implementation for Short-Lived ConnectionsabstractWith the rapid growth of network bandwidth, increases in CPU cores on a single machine, and application API models demanding more short-lived connections, a scalable TCP stack is performance-critical. Although many clean-state designs have been proposed, production environments still call for a bottom-up parallel TCP stack design that is backward-compatible with existing applications. Yu Chen 0004, Junjie Mao, Jiaquan He, Wei Xu 0005, Yuanchun Shi |
ASPLOS | 2 |
| 2016 | RID: Finding Reference Count Bugs with Inconsistent Path Pair CheckingabstractReference counts are widely used in OS kernels for resource management. However, reference counts are not trivial to be used correctly in large scale programs because it is left to developers to make sure that an increment to a reference count is always paired with a decrement. This paper proposes inconsistent path pair checking, a novel technique that can statically discover bugs related to reference counts without knowing how reference counts should be changed in a function. A prototype called RID is implemented and evaluations show that RID can discover more than 80 bugs which were confirmed by the developers in the latest Linux kernel. The results also show that RID tends to reveal bugs caused by developers' misunderstanding on API specifications or error conditions that are not handled properly. Junjie Mao, Yu Chen 0004, Qixue Xiao, Yuanchun Shi |
ASPLOS | 2 |
| 2015 | Mitigating Code-Reuse Attacks on CISC Architectures in a Hardware Approach
Zhijiao Zhang, Ya-Shuai Lü, Yu Chen 0004, Yongqiang Lyu 0001, Yuanchun Shi |
SEC | 3 |
| 2015 | Requester-Based Spin Lock: A Scalable and Energy Efficient Locking Scheme on Multicore SystemsabstractIn response to the increasing ubiquity of multicore processors, applications are usually designed or deployed to make each core busy. Unfortunately, lock contention within operating systems can limit the scalability of multicore systems so severely that an increase in the number of cores can actually lead to reduced performance (i.e., scalability collapse). Existing lock implementations have disadvantages in scalability, power consumption, and energy efficiency. In this paper, we observe that the number of tasks requesting a lock has a significant correlation with the occurrence of scalability collapse. Based on this observation, a lock implementation that allows tasks waiting for a lock to either spin or enter a power-saving state based on the number of requesters is proposed. Our lock protocol is called requester-based lock and is implemented in the Linux kernel to replace its default spin lock. Based on the results of a sensitivity analysis, we find that the best policy, in practice, for a task waiting for a lock to be granted is to enter the power-saving state immediately after noticing the lock cannot be acquired. Our requester-based lock scheme is evaluated using intensive benchmarking on AMD 32-core and Intel 40-core systems. Experimental results suggest that our lock avoids scalability collapse completely for most applications and shows better scalability, power consumption, and energy efficiency than previous work. Besides, the requester-based lock is extensible, which means using together with other kinds of spin locks can provide better scalability and energy efficiency. Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi |
IEEE Trans. Computers | 3 |
| 2015 | LockSim: An Event-Driven Simulator for Modeling Spin Lock ContentionabstractSpin lock contention in operating systems can limit scalability on multicore systems so significantly that an increase in the number of cores actually leads to reduced speedup (i.e., scalability collapse). Modeling spin lock contention is an effective way to understand the scalability collapse phenomenon and explore collapse avoidance schemes. However, previous spin lock models have disadvantages in accuracy and efficiency. To overcome these drawbacks, this paper proposes LockSim, an event-driven simulator which models both the sequential execution in lock-protected codes (i.e., critical sections) and shared hardware resource contention caused by the cache coherence protocol. Our simulator is verified against real-world workloads with different degrees of spin lock contention. Experimental results suggest that LockSim can reproduce the scalability collapse phenomenon with better accuracy than previous work. Besides, several metrics are also used to characterize this phenomenon and collapse avoidance methods are investigated. Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2014 | Running Multiple Androids on One ARM Platform
Zhijiao Zhang, Lei Zhang 0060, Yu Chen 0004, Yuanchun Shi |
ACISP | 3 |
| 2014 | Mitigating Resource Contention on Multicore Systems via SchedulingabstractThis paper addresses the resource contention issue caused by the sharing of the last level caches by introducing a novel contention-aware scheduler. To accurately determine a task's resource requirements, i.e. the unique input of our scheduler, we develop a methodology to select the best heuristic metric from five candidates to represent a task's resource requirements. Based on the heuristic of each task acquired by exploiting the performance monitor unit, our scheduler co-schedules tasks with complementary resource requirements by combining scheduling order adjustments with task-to-core reassignments. The proposed scheduler has been implemented in the completely fair scheduler, rotating staircase deadline scheduler and O(1) schedulers. Using eight workloads constructed from nine NASA advanced supercomputing serial benchmarks on an Intel dual-core platform, the execution time of an individual task is reduced by up to 21%, system scalability and performance of a workload are improved by up to 13%, and the full potential of the contention-aware scheduling can be achieved if the time slice length and the period of executing each task once are short enough. In addition, our proposal exhibits benefits in the reduction of the execution time fluctuation of individual tasks due to an enforcement of reasonable usage of shared resources. Finally, we demonstrate an expected performance improvement on an Intel eight-core platform in order to suggest the broad applicability of our protocol. Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi |
Comput. J. | 3 |
| 2014 | Towards scalability collapse behavior on multicoresabstractSUMMARY Multicore processor systems have become mainstream. To release the full potential of multiple cores, applications are programmed to be parallel to keep every core busy. Unfortunately, lock contention within operating systems can limit the scalability so seriously that use of more cores leads to reduced throughput (scalability collapse). To understand and characterize the collapse behavior easily, a discrete‐event simulation model, which considers both the sequential execution of critical sections and the overhead of hardware resource contention, is designed and implemented. By the use of the model, we observe that the percentage of time used to wait for locks and the number of tasks requesting for a lock have a significant correlation with the occurrence of scalability collapse. On the basis of these observations, two new techniques (lock contention aware scheduler and requester‐based adaptive lock) are proposed to remove the scalability collapse on multicores. The proposed methods are implemented in the Linux kernel 2.6.29.4 and evaluated on an AMD 32‐core system to verify their effectiveness. By using micro‐benchmarks and macro‐benchmarks, we find that these methods can remove scalability collapse totally for four of five workloads exhibiting the collapse behavior. For one workload that does not suffer scalability collapse, these proposed methods only introduce negligible overhead. Copyright © 2012 John Wiley & Sons, Ltd. Yan Cui 0002, Yu Chen 0004, Yuanchun Shi |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | A multi-communication-fusion based mobile monitoring system for maternal and fetal informationabstractMeasurements of vital signs can be translated into accurate predictors of pregnant diseases, even at an early stage. They can also be combined with alarm-triggering systems to initiate the appropriate actions. Because of the emphasis on healthcare awareness and their specific needs, gravidae prefer regular vital signs monitoring in a flexible manner. Thus, different types of sensors are used that involve complex operations, networks, and results. To enhance usability and feasibility, we propose a mobile vital signs monitoring system based on multi-communication fusion and the Android OS so pregnant women can monitor maternal and fetal information anywhere they want. They can also access comprehensive care by transferring data to the server for further processing and remote diagnosis. The accuracy of remote diagnosis is also improved. Pei Lyu, Manman Peng, Yongqiang Lyu 0001, Yu Chen 0004 |
Healthcom | 4 |
| 2013 | Lock-contention-aware scheduler: A scalable and energy-efficient method for addressing scalability collapse on multicore systemsabstractIn response to the increasing ubiquity of multicore processors, there has been widespread development of multithreaded applications that strive to realize their full potential. Unfortunately, lock contention within operating systems can limit the scalability of multicore systems so severely that an increase in the number of cores can actually lead to reduced performance (i.e., scalability collapse). Existing efforts of solving scalability collapse mainly focus on making critical sections of kernel code fine-grained or designing new synchronization primitives. However, these methods have disadvantages in scalability or energy efficiency. In this article, we observe that the percentage of lock-waiting time over the total execution time for a lock intensive task has a significant correlation with the occurrence of scalability collapse. Based on this observation, a lock-contention-aware scheduler is proposed. Specifically, each task in the scheduler monitors its percentage of lock waiting time continuously. If the percentage exceeds a predefined threshold, this task is considered as lock intensive and migrated to a Special Set of Cores (i.e., SSC). In this way, the number of concurrently running lock-intensive tasks is limited to the number of cores in the SSC, and therefore, the degree of lock contention is controlled. A central challenge of using this scheme is how many cores should be allocated in the SSC to handle lock-intensive tasks. In our scheduler, the optimal number of cores is determined online by the model-driven search. The proposed scheduler is implemented in the recent Linux kernel and evaluated using micro- and macrobenchmarks on AMD and Intel 32-core systems. Experimental results suggest that our proposal is able to remove scalability collapse completely and sustains the maximal throughput of the spin-lock-based system for most applications. Furthermore, the percentage of lock-waiting time can be reduced by up to 84%. When compared with scalability collapse reduction methods such as requester-based locking scheme and sleeping-based synchronization primitives, our scheme exhibits significant advantages in scalability, power consumption, and energy efficiency. Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi |
ACM Trans. Archit. Code Optim. | 3 |
| 2012 | Built-in Device Simulator for OS Performance EvaluationabstractI/O devices are evolving rapidly, while OS optimization is always slower because of its dependence on physical devices. This inevitably prevents latest devices from working with their rating performance, which remains a big problem for performance-critical applications. Though I/O device simulators can help carry out performance evaluation before physical devices are ready, the existing simulator implementations are still unsatisfactory, either having too big overhead or requiring too much extra work. In this paper, we propose kernel built-in device simulation to provide accurate real time evaluations with acceptable extra effort. With the work of simulation well isolated, the overhead is reasonable compared to native environment. A bonding Ethernet interface is implemented in this way and experiments on it confirm the close-to-native performance of the idea. Junjie Mao, Yu Chen 0004, Yaozu Dong |
CLUSTER | 2 |
| 2012 | Lock-Visor: An Efficient Transitory Co-scheduling for MP GuestabstractMultiprocessor (MP) virtual machines (VMs) are widely used in cloud environments. However, MP VMs suffer from lock holder preemption (LHP) issue. This causes a tremendous waste of CPU cycles, leading to deteriorated synchronization latency and a significant degradation in system performance. Previous works have addressed the problem with software co-scheduling or lock waiter yielding. However, co-scheduling suffers from CPU utility fragmentation, priority inversion and loss of the flexibility of hyper visor scheduler, which causes inefficiency in CPU usage. Lock waiter yielding, another solution, suffers from a large impact on hyper visor scheduler and issues with response latency. In this paper, we propose Lock-visor, an efficient transitory co-scheduling algorithm, to bypass the guest spin lock loop effectively. Our protocol has little to no impact on the flexibility of hyper visor scheduler, and achieves better system performance. Multiple policies are explored on top of transitory co-scheduling to maximize the efficiency of Lock-visor, i.e. instant transitory, selective instant transitory and deferred transitory co-scheduling. Comprehensive experiments are conducted using CPU-intensive, I/O-intensive and lock-intensive workloads. Our experimental results show that Lock-visor can significantly improve system performance (e.g. Lock-visor has up to 341.3% performance advantage over original Linux kernel 2.6.38 in Sys Bench 4-VM case), while at the same time improve system latency with little to no effect on scheduling fairness. Lei Zhang 0060, Yu Chen 0004, Yaozu Dong |
ICPP | 2 |
| 2012 | Reducing Scalability Collapse via Requester-Based Locking on Multicore SystemsabstractIn response to the increasing ubiquity of multicore processors, there has been widespread development of multithreaded applications that strive to realize their full potential. Unfortunately, lock contention within operating systems can limit the scalability of multicore systems so severely that an increase in the number of cores can actually lead to reduced performance (i.e. scalability collapse). Existing lock implementations have disadvantages in scalability, resource utilization and energy efficiency. In this work, we observe that the number of tasks requesting a lock has a significant correlation with the occurrence of scalability collapse. Based on this observation, we propose a novel lock implementation that allows tasks blocked on a lock to either spin or maintain a power-saving state according to the number of lock requesters. We call our lock implementation protocol a requester-based lock and implement it in the Linux kernel to replace its default spin lock. Based on the results of an analysis, we find that the best policy for a task waiting for a lock to become free is to enter the power saving state immediately after noticing that the lock cannot be acquired. Our lock-requester based lock scheme is evaluated using micro- and macro-benchmarks on AMD 32-core and Intel 40-core systems. Experimental results indicate our lock scheme removes scalability collapse completely for most applications. Furthermore, our method shows better scalability and energy efficiency than mutex locks and adaptive locks. Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi |
MASCOTS | 3 |
| 2012 | CompSC: live migration with pass-through devicesabstractLive migration is one of the most important features of virtualization technology. With regard to recent virtualization techniques, performance of network I/O is critical. Current network I/O virtualization (e.g. Para-virtualized I/O, VMDq) has a significant performance gap with native network I/O. Pass-through network devices have near native performance, however, they have thus far prevented live migration. No existing methods solve the problem of live migration with pass-through devices perfectly. Zhenhao Pan, Yaozu Dong, Yu Chen 0004, Lei Zhang 0060, Zhijiao Zhang |
VEE | 3 |
| 2011 | Experience on Comparison of Operating Systems Scalability on the Multi-core ArchitectureabstractMulti-core processor architectures have become ubiquitous in today's computing platforms, especially in parallel computing installations, with their power and cost advantages. While the technology trend continues towards having hundreds of cores on a chip in the foreseeable future, an urgent question posed to system designers as well as application users is whether applications can receive sufficient support on today's operating systems for them to scale to many cores. To this end, people need to understand the strengths and weaknesses on their support on scalability and to identify major bottlenecks limiting the scalability, if any. As open-source operating systems are of particular interests in the research and industry communities, in this paper we choose three operating systems (Linux, Solaris and FreeBSD) to systematically evaluate and compare their scalability by using a set of highly-focused microbenchmarks for broad and detailed understanding their scalability on an AMD 32-core system. We use system profiling tools and analyze kernel source codes to find out the root cause of each observed scalability bottleneck. Our results reveal that there is no single operating system among the three standing out on all system aspects, though some system(s) can prevail on some of the system aspects. For example, Linux outperforms Solaris and FreeBSD significantly for file-descriptor and process-intensive operations. For applications with intensive sockets creation and deletion operations, Solaris leads FreeBSD, which scales better than Linux. With the help of performance tools and source code instrumentation and analysis, we find that synchronization primitives protecting shared data structures in the kernels are the major bottleneck limiting system scalability. Empowered by the knowledge obtained through targeted experiments and analysis on a small-scale system, we are able to project the scalability of an application on any of the investigated operating systems running on a system of a larger number of cores. Yan Cui 0002, Yingxin Wang, Yu Chen 0004, Yuanchun Shi |
CLUSTER | 3 |
| 2011 | S-FTL: An efficient address translation for flash memory by exploiting spatial localityabstractThe solid-state disk (SSD) is becoming increasingly popular, especially among users whose workloads exhibit substantial random access patterns. As SSD competes with the hard disk, whose per-GB cost keeps dramatically falling, the SSD must retain its performance advantages even with low-cost configurations, such as those with a small built-in DRAM cache for mapping table and using MLC NAND. To this end, we need to make the limited cache space efficiently used to support fast logical-to-physical address translation in the flash translation layer (FTL) with minimal access of flash memory and minimal merge operations. Existing schemes usually require a large number of overhead accesses, either for accessing uncached entries of the mapping table or for the merge operation, and achieve suboptimal performance when the cache space is limited. In this paper we take into account spatial locality exhibited in the workloads to obtain a highly efficient FTL even with a relatively small cache, named as S-FTL. Specifically, we identify three access patterns related to spatial locality, including sequential writes, clustered access, and sparse writes. Accordingly we propose designs to take advantage of these patterns to reduce mapping table size, increase hit ratio for in-cache address translation, and minimize expensive writes to flash memory. We have conducted extensive trace-driven simulations to evaluate S-FTL and compared it with other state-of-the-art FTL schemes. Our experiments show that S-FTL can reduce accesses to the flash for address translation by up to 70% and reduce response time of SSD by up to 25%, compared with the state-of-the-art FTL strategies such as FAST and DFTL. Song Jiang 0001, Lei Zhang 0060, XinHao Yuan, Yu Chen 0004 |
MSST | 5 |
| 2010 | A Discrete Event Simulation Model for Understanding Kernel Lock Thrashing on Multi-core ArchitecturesabstractMulti-core architectures have become mainstream. Trends suggest that the number of cores integrated on a single chip will increase continuously. However, lock contention in operating systems can limit the parallel scalability on multi-cores so significantly that the speedup decreases with the increasing number of cores (thrashing). Although the phenomenon can be easily reproduced experimentally, most existing lock models are not able to do so. To overcome this challenge, this paper develops a discrete event simulation model which has the capability of capturing both the sequential execution in critical sections and the contention for shared hardware resources. The model is evaluated using a series of typical parameter configurations which can represent different degrees of lock contention. Experimental results suggest that the thrashing phenomenon can be observed when the model parameters are selected properly. To further understand this phenomenon, statistics such as the percentage of time spent waiting for locks and the number of cores waiting for a lock are exploited to characterize the lock thrashing. In addition, the model sensitivity to changes in memory latency and hardware architectures are also examined. Finally, we use this model to compare three methods which are proposed for preventing the lock thrashing. Yan Cui 0002, Weiyi Wu, Yingxin Wang, Xufeng Guo, Yu Chen 0004, Yuanchun Shi |
ICPADS | 5 |
| 2010 | A Scheduling Method for Avoiding Kernel Lock Thrashing on Multi-coresabstractMulti-core architectures have been adopted in various computing environments. Predictions based on Moore's Law state that thousands of cores can be integrated on a single chip within 10 years. To achieve better performance and scalability on multi-cores, applications should be multi-threaded, and therefore threads assigned on different cores can execute concurrently. However, lock contention in kernels can affect the scalability so significantly that the speedup decreases with the increasing number of cores (thrashing). Existing efforts to address this problem mainly focus on deferring lock thrashing, and therefore these techniques cannot prevent thrashing fundamentally. In this paper, we propose to use lock-aware scheduling to avoid thrashing. Our method detects thrashing on a per-thread basis and migrates contended threads to a smaller set of cores. The optimal number of cores is determined by maximizing the proposed normalized throughput model of migrated threads. The proposed method is implemented in Linux 2.6.29.4 and evaluated on a 32-core system. Experimental results on a series of lock-intensive micro- and macro-benchmarks show the effectiveness: for 3 of 5 workloads exhibiting thrashing behaviour, lock-aware scheduling can detect the speedup decrease accurately and sustain the maximal speedup, for the remaining 2 workloads, the performance can be improved greatly although the maximal speedup is not sustained, for 1 workload which does not suffer thrashing, the method introduces negligible runtime overhead. Yan Cui 0002, Weida Zhang, Yu Chen 0004, Yuanchun Shi |
ICPADS | 3 |
| 2010 | Scaling OLTP applications on commodity multi-core platformsabstractMulti-core processor architectures can have significant performance advantage over traditional single core designs, which are limited by power and processor complexity. Predictions based on Moore's Law state that a processor chip may accommodate thousands of cores in 5–10 years. Can software scale with the number of cores and achieve the performance potential? Yan Cui 0002, Yu Chen 0004, Yuanchun Shi |
ISPASS | 2 |
| 2010 | Scalability comparison of commodity operating systems on multi-coresabstractIn this paper, we evaluate and compare the parallel scalability of three commodity operating systems (Linux, Solaris and FreeBSD) on an AMD 32-core platform. Measurements of microbenchmarks and a real-life application reveal that no operating system scales totally better than another for microbenchmarks; for the real-life application, Linux and Solaris are competitive in scalability and perform better than FreeBSD. Related kernel source analysis and performance data suggest that synchronization primitives protecting the shared data structure in kernels are the root cause of the poor scalability on multi-cores. Yan Cui 0002, Yu Chen 0004, Yuanchun Shi, Qingbo Wu 0003 |
ISPASS | 2 |
| 2010 | Reinventing Lock Modeling for Multi-Core SystemsabstractMulti-core architectures have become mainstream. Trends suggest that the number of cores integrated on a single chip will continue to increase. However, lock contention in applications or kernels can degrade the scalability so significantly that the speedup decreases with the increasing number of cores (thrashing). Although the phenomenon can be easily reproduced on real multi-core platforms, existing lock models are not able to do so. To overcome the disadvantage, this paper proposes an analysis model which has the capability of capturing both the sequential execution of critical sections and the overhead of lock implementation. Numerical results indicate that thrashing can be observed by using the proposed model. Furthermore, this model can also be exploited to compare different mechanisms designed for avoiding the lock thrashing. Yan Cui 0002, Weiyi Wu, Yingxin Wang, Xufeng Guo, Yu Chen 0004, Yuanchun Shi |
MASCOTS | 5 |
| 2010 | Aware service base on the assembly of multiple weak location sensorsabstractIndoor location information is important for the awareness computing environment. One single location sensor always is not accurate and applicable enough for providing the right service to the user. In this paper, an OSGi based simulation system is introduced firstly, and then we propose a framework to support the combination of several weak location sensors. With this combination model, multiple weak location sensors can work together to provide the better location aware service to the user. The experiment platform which consists of RFID array location sensor, RF sensor, pressure sensor and other sensors can work together in the OSGi based simulation system, experiment results show that the assembly of multiple weak sensors do better than the solo sensor. It can help the designers to achieve the better cost-performance for the awareness applications between various possible combinations. Pin Tao, Lei Zhang 0060, Yu Chen 0004 |
SMC | 3 |
| 2010 | A Low-Cost Ubiquitous Family Healthcare Framework
Yongqiang Lyu 0001, Lei Zhang 0060, Yu Chen 0004, Yingjie Ren, Weikang Yang, Yuanchun Shi |
UIC | 3 |
| 2009 | CFS Optimizations to KVM Threads on Multi-Core EnvironmentabstractMulti-core architecture provides more on-chip parallelism and powerful computational capability. It helps virtualization achieve scalable performance. KVM (kernel based virtual machine) is different from other virtualization solutions which can make use of the Linux kernel components such as completely fair scheduler (CFS). However, CFS treats the KVM threads as normal tasks without considering about their unique features such as thread allocation mechanism and lock inside guest virtual machine, which may harm the KVM virtualization performance. In this paper, we analyze a phenomenon that some guest multi-threaded applications have very low performance when scheduled by CFS. As a solution to this problem, we introduce two kinds of optimizations in CFS: (1) configuration optimizations (2) lock optimizations. Our contributions are: (1) implement 5 original and 2 newest proposed optimizations in the newest Linux kernel. (2) Classify and compare them, a brief analysis is also given. They are all very simple and general to other virtual machine monitors such as Xen and schedulers as O(1). The performance of our CFS optimizations to KVM threads is measured by running some well-known benchmarks in two guest virtual machines on an 8-core server which models the real world applications. The results indicate our scheduling optimizations can improve the overall system performance. This paper can provide useful advices to KVM developers and virtualization data center administrators. Yisu Zhou, Yan Cui 0002, Yu Chen 0004, Yuanchun Shi, Qingbo Wu 0003 |
ICPADS | 5 |
| 2009 | Fairness and Interactivity of Three CPU Schedulers in LinuxabstractCPU scheduler is a very important subsystem which affects system throughput, interactivity and fairness. The development of Linux kernel is relatively fast-paced. By now, many CPU schedulers have been designed by researchers, hobbyists and kernel hackers. It is necessary to accurately compare and analyze different characteristics among these schedulers, so as to understand and design better CPU schedulers for various applications. However, researchers lack a straight-forward method to compare and analyze these CPU schedulers precisely. In this paper, we systematically analyze and measure fairness, interactivity and multi-processors performance of three schedulers: O(1), RSDL and CFS, by using micro, synthesis and real application benchmarks. They have been ported into one scheduler framework in Linux kernel-2.6.29. Experimental results show that there are notable differences in fairness and interactivity under micro benchmarks, while minor differences in synthesis and real applications. We also analyze the impact of implementations of schedulers on fairness and interactivity of applications, discuss challenges in estimating application resource requirements in different environments, and present some ideas for developing future CPU schedulers. Yu Chen 0004, Yan Cui 0002 |
RTCSA | 2 |
| 2006 | Cicada: A Highly-Precise Easy-Embedded and Omni-Directional Indoor Location Sensing System
Hongliang Gu, Yuanchun Shi, Yu Chen 0004, Bibo Wang, Wenfeng Jiang |
GPC | 3 |
| 2005 | A Component-Based Reflective Middleware Approach to Context-Aware Adaptive Systems
Yanni Wu, Zhenkun Zheng, Xiaoge Wang, Yu Chen 0004 |
ICWE | 5 |
| 2005 | A Soft Real-Time Web News Classification System with Double Control Loops
Huayong Wang, Yu Chen 0004, Yiqi Dai |
WAIM | 2 |