Lulu Yao

dblp:168/9407 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 DMTree: Towards Efficient Tree Indexing on Disaggregated Memory via Compute-side Collaborative Design
Guoli Wei, Yongkun Li 0001, Haoze Song, Lulu Yao, Yinlong Xu 0001, Heming Cui
FAST5
2024 AdaptMD: Balancing Space and Performance in NUMA Architectures With Adaptive Memory Deduplication
abstract
Memory deduplication effectively relieves the memory space bottleneck by removing duplicate pages, especially in virtualized systems in which virtual machines run the same OS and similar applications. However, due to the non-uniform access latencies in NUMA architectures, memory deduplication poses a trade-off between memory savings and access performance: global deduplication across NUMA nodes realizes high memory savings, but leads to frequent cross-node remote access after deduplication and results in performance degradations. In contrast, local deduplication avoids remote access, but limits deduplication effectiveness. We design AdaptMD, an adaptive memory deduplication system that addresses the space-performance trade-off in NUMA architectures. AdaptMD leverages hotness awareness to globally deduplicate only cold pages to reduce remote access. It also migrates similar applications to the same NUMA node to allow local deduplication without remote access. We further make AdaptMD readily configurable to address various deployment scenarios. Experiments show that AdaptMD achieves high memory savings as in global deduplication, while achieving similar access performance as in local deduplication.
Lulu Yao, Yongkun Li 0001, Patrick P. C. Lee, Yinlong Xu 0001
IEEE Trans. Computers1
2023 CARE: A Cost-AwaRe Eviction Strategy for Improving Throughput in Cloud Environments
abstract
To utilize various resources efficiently, cloud data centers usually deploy Latency-Critical (LC) and Best-Effort (BE) applications as containers in their physical machines by assigning higher priorities for LC jobs to use the resources. Due to the increase on the workload, it needs evicting some BE jobs to deprive more resources for LC jobs to meet the Quality of Service (QoS) requirements of LC jobs. Under dynamic workload settings, the frequent BE job eviction results in significant throughput degradation. To this end, we propose a novel Cost-AwaRe Eviction strategy, CARE, which takes into account both the recalculation cost and the remaining time cost of each BE job. CARE fully exploits the two types of different costs of BE jobs and chooses a BE job with the lowest overhead for eviction under different resource demands of LC jobs. Furthermore, CARE improves the throughput of BE jobs without affecting the performance of LC jobs. We prototype and implement CARE atop Kubernetes. Experimental results show that CARE achieves up to 1.70 × throughput gain compared to state-of-the-arts while its negative impact on the performance of LC jobs is negligible.
Xiyuan Liang, Lulu Yao, Si Wu 0003, Yongkun Li 0001, Yinlong Xu 0001
ICPADS2
2023 On Optimizing Traffic Scheduling for Multi-replica Containerized Microservices
abstract
Containerized deployment of microservices has been becoming prevalent, as it provides flexible deployment and elastic resource configuration. For high concurrency and fault tolerance, multiple container replicas are often deployed for each microservice component, but this may induce heavy cross-machine traffic and degrades the performance of microservice applications. Traffic localization tries to put containers with heavy communication traffic on the same machine to reduce cross-machine traffic. However, it is still very common to have the containers with heavy traffic on different machines, especially under multi-replica deployment, due to the insufficient resources of a physical machine. To this end, we develop a network-aware scheduling system OptTraffic, which realizes optimized traffic scheduling for containerized microservices. OptTraffic estimates the traffic between each pair of containers in a lightweight manner by combining a simple math calculation with coarse-grained monitoring, then it proposes an efficient traffic allocation algorithm and leverages dynamic scheduling with multiple optimizations to minimize the cross-machine traffic without sacrificing resource usage balance. Experiments show that under multi-replica deployment, OptTraffic can save up to 47% of the network bandwidth, while reducing the P99 latency by 28%-45%, compared to Kubernetes and existing traffic localization designs for real-world microservice applications.
Xianzhi Zhu, Yongkun Li 0001, Lulu Yao, Zhihao Qi, Yinlong Xu 0001
ICPP3
2023 UniScatter: a Metamaterial Backscatter Tag for Wideband Joint Communication and Radar Sensing
abstract
Millimeter-wave backscatter can simultaneously support high-precision sensing and massive communication and represent one prominent technical evolution in next-generation wireless systems. The backscatter tags should ideally work across a wide mmWave spectrum range with consistent signal strength and angular coverage to accommodate highly diverse application scenarios. However, existing tags made of resonant antennas and RFICs only achieve a few GHz of bandwidth and hardly meet these requirements. In this paper, we present UniScatter, a new backscatter tag structure based on metamaterials. The key design of UniScatter is a graphene-based modulator and a lens-based retroreflector, which have consistent electromagnetic responses across an extensive frequency range and wide angular field-of-view. We have developed a robust fabrication process for UniScatter, and tested it on various mmWave sensing and communication devices. Our field tests show that UniScatter can backscatter signals across a wide frequency band from 24 GHz to 77 GHz with consistently high signal strength and wide angular coverage in 3D space.
Kun Qian 0004, Lulu Yao, Kai Zheng 0003, Xinyu Zhang 0003, Tse Nga Tina Ng
MobiCom2
2023 Towards High Performance and Efficient Memory Deduplication via Mixed Pages
abstract
Large pages are widely supported in modern hardware and OSes to reduce the overhead of TLB misses. However, memory deduplication can be inefficient with large pages, leading to low memory utilization. To simultaneously enjoy the benefits of high performance by accessing memory with large pages (e.g., 2 MB pages) and high deduplication rate by managing memory with base pages (e.g., 4 KB pages), we proposeSmartMemoryDeduplciation (SmartMD), which is an adaptive and efficient memory management scheme via mixed pages. Specifically, we propose lightweight schemes to periodically monitor pages’ access frequency and repetition rate, and present an adaptive conversion scheme to selectively split or reconstruct large pages. We further optimize SmartMD by developing SmartMD$^{+}$, which dynamically adjusts the page scanning cycle by monitoring the TLB miss cost, and reconstructs the split large pages in an on-demand way so as to reduce the CPU overhead of SmartMD. We further implement a prototype system and conduct extensive experiments with various workloads under different system settings. Experiment results show that SmartMD and SmartMD$^{+}$can simultaneously achieve high access performance similar to systems using large pages, and achieve a deduplication rate similar to that applying aggressive deduplication scheme (i.e., KSM) on base pages.
Lulu Yao, Yongkun Li 0001, Fan Guo 0003, Si Wu 0003, Yinlong Xu 0001, John C. S. Lui
IEEE Trans. Computers1
2022 MilliMirror: 3D printed reflecting surface for millimeter-wave coverage expansion
abstract
Next generation wireless networks embrace mmWave technology for its high capacity. Yet, mmWave radios bear a fundamental coverage limitation due to the high directionality and propagation artifacts. In this paper, we explore an economical paradigm based on 3D printing technology for mmWave coverage expansion. We propose MilliMirror, a fully passive metasurface, which can reshape and resteer mmWave beams to anomalous directions to illuminate the coverage blind spots. We develop a closed-form model to efficiently synthesize the MilliMirror design with thousands of unit elements and across a wide frequency band. We further develop an economical process based on 3D printing and metal deposition to fabricate MilliMirror. Our field test results show that MilliMirror can effectively fill the coverage holes and operate transparently to the standard mmWave beam management protocols.
Kun Qian 0004, Lulu Yao, Xinyu Zhang 0003, Tse Nga Tina Ng
MobiCom2
2021 Progressive Memory Adjustment with Performance Guarantee in Virtualized Systems
abstract
As applications are often allocated more memory than they actually used to cope with peak memory demand due to the workload dynamics in virtualized systems running multiple virtual machines, memory adjustment by reclaiming inactive memory in virtual machines is an effective way to enable memory overcommitment so as to reduce the cost. However, existing schemes are designed based on one-shot adjustment, which may reclaim a large size in a single operation, and they are unaware of both memory access dynamics and memory sensitivity of different applications, so they usually result in excessive reclamation and lead to severe performance slowdown. To address this issue, we propose PMA, a progressive memory adjustment scheme that takes into consideration both memory access dynamics and memory sensitivity and leverages virtual machine performance feedback to progressively reclaim inactive memory to avoid performance slowdown. Besides, PMA is designed based on ballooning (i.e., the balloon driver) so it preserves the good isolation between host and virtual machines for full virtualization systems. We also implement a prototype in host user space, and experiments show that PMA effectively limits the performance drop of memory overcommitment (e.g., the performance drop of every virtual machine is limited within 10% for up to 33% memory overcommitment), which is already very close to the optimal case of having enough memory, so PMA is efficient to enable memory overcommitment with performance guarantee in full virtualization systems.
Lulu Yao, Yongkun Li 0001, Weijie Wu, Yinlong Xu 0001
ICPP1
2019 Interactive Multi-camera Soccer Video Analysis System
abstract
Automatic sports video analysis is an active field of research, and accurate player & ball tracking is essential for soccer video analysis and visualization. However, the variations over frames and the scarceness of large-scale well-annotated datasets make it difficult to perform supervised learning using pre-trained models, especially for Multi-Camera Multi-Target Tracking (MCMT). In this paper, we introduce an end-to-end system for multi-camera soccer video analysis that makes heavy use of parallel processing for optimization of the processing workflow. The proposed thread-level parallelism speeds up our system by more than 15 times while maintaining the level of accuracy. The system tracks the trajectories of the ball and the players in a world coordinate system based on soccer videos captured by a set of synchronized cameras. Based on these trajectories, various player-, ball-, and team-related statistics are computed, and the resulting data and visualizations can be interactively explored by the user.
Yunjin Wu, Ziyuan Zhao, Shengqiang Zhang, Lulu Yao, Tom Z. J. Fu
ACM Multimedia4
2015 Hybrid Marte
abstract
MARTE (Modeling and Analysis of Real-Time and Embedded Systems) is a profile of UML (United Modeling Language). MARTE provides support for specification, design and verification of real-time and embedded systems. Even though MARTE time model offers a support to describe multiform clocks, it lacks the ability to model both discrete and continuous behaviors of a hybrid system. To address the problem of hybrid systems modeling, we propose Hybrid MARTE which is an extension to MARTE for hybrid system modeling and analysis. Compare to MARTE, in Hybrid MARTE, we can construct the logical time and chronometric time in a unified way. Besides, a systemic framework for the modeling the requirements and design of a hybrid system is provided. Hybrid MARTE also provides multiple views modeling, something like UML. Hybrid MARTE Class Diagram can be used for description in static view, Hybrid MARTE Sequence Diagram in interactive view and Hybrid MARTE Statechart in dynamic behavioral view. HybridMARTE is successfully used in modeling and analysis of the Train Position Determination of railway control systems.
Lulu Yao, Jing Liu 0012, Yan Zhang 0072, Yuejun Wang
APSEC1
2015 HSD: Hybrid MARTE Sequence Diagram
abstract
Modeling and Analysis of Real-Time and Embedded systems (MARTE) is a profile of United Modeling Language (UML), which provides support for specification, design and verification for Real-Time Embedded Systems (RTES). MARTE sequence diagram can deal with both discrete and dense time in which a clock can be either chronometric or logical. However it lacks the ability to describe the continuous behavior of a hybrid system. We propose a new method named Hybrid MARTE Sequence Diagram (HSD) to describe the communication between participants and the continuous evolution within the execution occurrence of a hybrid system. HSD combines the time model from MARTE and specification of the time-continuous behavior aspects from hybrid automata with MARTE sequence diagram. It improves the MARTE sequence diagram in that: the logical time and the chronometric time are unified. Besides, the description of continuous evolution of a hybrid system is provided. We firstly extend the basic MARTE elements to support both discrete and continuous aspects. Then we define the formal syntax and semantics of HSD based on hybrid transition system. Using this new method, we model an industrial application named Train Position Determination (TPD).
Lulu Yao, Jing Liu 0012, Yan Zhang 0072, Yuejun Wang, Haiying Sun, Qingsheng Wang, Dehui Du, Xiaohong Chen 0007
QRS1