Yao Yue

dblp:55/2531 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 2 first-author · 2 since 2021Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 An Interval-Based Transformer Approach for Carbon Future Prices Forecasting
abstract
We propose a novel interval-based transformer approach called DKformer to forecast quarterly carbon futures prices by treating interval-valued data as inseparable random sets. Our DKformer model leverages the attention mechanism to capture long-term dependencies and complex dynamics in the carbon market. The model integrates the European economic policy uncertainty (EEPU) index and the climate policy uncertainty (CPU) index, utilizing only lags of price values and EPU indices as explanatory variables to effectively capture the impact of policy shifts on carbon futures prices. Empirical tests demonstrate the strong performance of the model in predicting market turning points and responding to various policy change scenarios, including the implementation of the EU ETS Phase 4, the RePowerEU Plan, and the carbon border adjustment mechanism (CBAM). Furthermore, our findings indicate the outstanding performance of our method remains robust to various forecasting period, forecasting horizons, and sets of explanatory variables.
Chuanmiao Yan, Yao Yue, Yuying Sun, Shou-Yang Wang
IEEE Trans. Comput. Soc. Syst.3
2024 SIEVE is Simpler than LRU: an Efficient Turn-Key Eviction Algorithm for Web Caches
Yazhuo Zhang, Juncheng Yang, Yao Yue, Ymir Vigfusson, K. V. Rashmi
NSDI3
2023 LatenSeer: Causal Modeling of End-to-End Latency Distributions by Harnessing Distributed Tracing
abstract
End-to-end latency estimation in web applications is crucial for system operators to foresee the effects of potential changes, helping ensure system stability, optimize cost, and improve user experience. However, estimating latency in microservices-based architectures is challenging due to the complex interactions between hundreds or thousands of loosely coupled microservices. Current approaches either track only latency-critical paths or require laborious bespoke instrumentation, which is unrealistic for end-to-end latency estimation in complex systems.
Yazhuo Zhang, Rebecca Isaacs, Yao Yue, Juncheng Yang, Lei Zhang 0223, Ymir Vigfusson
SoCC3
2023 GL-Cache: Group-level learning for efficient and high-performance caching
Juncheng Yang, Ziming Mao, Yao Yue, K. V. Rashmi
FAST3
2023 FIFO can be Better than LRU: the Power of Lazy Promotion and Quick Demotion
abstract
LRU has been the basis of cache eviction algorithms for decades, with a plethora of innovations on improving LRU's miss ratio and throughput. While it is well-known that FIFO-based eviction algorithms provide significantly better throughput and scalability, they lag behind LRU on miss ratio, thus, cache efficiency.
Juncheng Yang, Ziyue Qiu, Yazhuo Zhang, Yao Yue, K. V. Rashmi
HotOS4
2023 FIFO queues are all you need for cache eviction
abstract
As a cache eviction algorithm, FIFO has a lot of attractive properties, such as simplicity, speed, scalability, and flash-friendliness. The most prominent criticism of FIFO is its low efficiency (high miss ratio).
Juncheng Yang, Yazhuo Zhang, Ziyue Qiu, Yao Yue, K. V. Rashmi
SOSP4
2023 A Deep Learning Method for Motion Artifact Correction in Intravascular Photoacoustic Image Sequence
abstract
In vivo application of intravascular photoacoustic (IVPA) imaging for coronary arteries is hampered by motion artifacts associated with the cardiac cycle. Gating is a common strategy to mitigate motion artifacts. However, a large amount of diagnostically valuable information might be lost due to one frame per cycle. In this work, we present a deep learning-based method for directly correcting motion artifacts in non-gated IVPA pullback sequences. The raw signal frames are classified into dynamic and static frames by clustering. Then, a neural network named Motion Artifact Correction (MAC)-Net is designed to correct motion in dynamic frames. Given the lack of the ground truth information on the underlying dynamics of coronary arteries, we trained and tested the network using a computer-generated dataset. Based on the results, it has been observed that the trained network can directly correct motion in successive frames while preserving the original structures without discarding any frames. The improvement in the visual effect of the longitudinal view has been demonstrated based on quantitative evaluation of the inter-frame dissimilarity. The comparison results validated the motion-suppression ability of our method comparable to gating and image registration-based non-learning methods, while maintaining the integrity of the pullbacks without image preprocessing. Experimental results from in vivo intravascular ultrasound and optical coherence tomography pullbacks validated the feasibility of our method in the in vivo intracoronary imaging scenario.
Du Jiejie, Yao Yue, Sun Huifeng
IEEE Trans. Medical Imaging3
2021 Segcache: a memory-efficient and scalable in-memory key-value cache for small objects
Juncheng Yang, Yao Yue, K. V. Rashmi
NSDI2
2021 A Large-scale Analysis of Hundreds of In-memory Key-value Cache Clusters at Twitter
abstract
Modern web services use in-memory caching extensively to increase throughput and reduce latency. There have been several workload analyses of production systems that have fueled research in improving the effectiveness of in-memory caching systems. However, the coverage is still sparse considering the wide spectrum of industrial cache use cases. In this work, we significantly further the understanding of real-world cache workloads by collecting production traces from 153 in-memory cache clusters at Twitter, sifting through over 80 TB of data, and sometimes interpreting the workloads in the context of the business logic behind them. We perform a comprehensive analysis to characterize cache workloads based on traffic pattern, time-to-live (TTL), popularity distribution, and size distribution. A fine-grained view of different workloads uncover the diversity of use cases: many are far more write-heavy or more skewed than previously shown and some display unique temporal patterns. We also observe that TTL is an important and sometimes defining parameter of cache working sets. Our simulations show that ideal replacement strategy in production caches can be surprising, for example, FIFO works the best for a large number of workloads.
Juncheng Yang, Yao Yue, K. V. Rashmi
ACM Trans. Storage2
2020 A large scale analysis of hundreds of in-memory cache clusters at Twitter
Juncheng Yang, Yao Yue, K. V. Rashmi
OSDI2
2012 Retransmission or redundancy: Transmission reliability study in wireless sensor networks
Hao Wen 0014, Chuang Lin 0002, Fengyuan Ren, Yao Yue, Xiaomeng Huang
Sci. China Inf. Sci.5
2011 Fast checkpoint recovery algorithms for frequently consistent applications
abstract
Advances in hardware have enabled many long-running applications to execute entirely in main memory. As a result, these applications have increasingly turned to database techniques to ensure durability in the event of a crash. However, many of these applications, such as massively multiplayer online games and mainmemory OLTP systems, must sustain extremely high update rates – often hundreds of thousands of updates per second. Providing durability for these applications without introducing excessive overhead or latency spikes remains a challenge for application developers. In this paper, we take advantage of frequent points of consistency in many of these applications to develop novel checkpoint recovery algorithms that trade additional space in main memory for significantly lower overhead and latency. Compared to previous work, our new algorithms do not require any locking or bulk copies of the application state. Our experimental evaluation shows that one of our new algorithms attains nearly constant latency and reduces overhead by more than an order of magnitude for low to medium update rates. Additionally, in a heavily loaded main-memory transaction processing system, it still reduces overhead by more than a factor of two.
Tuan Cao, Marcos Antonio Vaz Salles, Ben Sowell, Yao Yue, Alan J. Demers, Johannes Gehrke, Walker M. White
SIGMOD Conference4
2008 Analyzing the Reliability of Group Transmission in Wireless Sensor Network
abstract
Most previous models about wireless channels are mainly adopted to obtain average packet error rates or to generate artificial network traces. However, when we combine packet group transmission with error correcting mechanisms, the steady packet error rate (PER) is not accurate enough to depict the short-term error event. In this paper, we propose a Markov chain model for group transmission to indicate influences of group length N and initial channel state on transmission reliability. Based on the model, we prove that the difference between steady PER and packet error rate within a group (PERG) is O(1/N), which can be used to estimate group length under a certain error rate constraint. Finally, we apply the model to compare the multipath transmission with the single-path transmission, and investigate an extreme case where the multipath way is less reliable than the single-path way.
Hao Wen 0014, Hongkun Yang, Chuang Lin 0002, Fengyuan Ren, Yao Yue
GLOBECOM5
2007 Retransmission or Redundancy: Transmission Reliability in Wireless Sensor Networks
abstract
As an application-driven network, wireless sensor network generally requires high data reliability to maintain detection and response capabilities. Although two approaches, which are retransmission and redundancy, have been proposed to enhance data reliability, the theoretical work is required to evaluate their impact on transmission reliability and energy efficiency. In this paper, we offer a comprehensive theoretical study on the packet arrival probability and average energy consumption for both approaches. Our analysis indicates that when loss probability remains low or moderate, erasure coding, a scheme based on redundancy, is more reliable and energy efficient than retransmission. However, the performance of erasure coding would largely deteriorate under high packet loss condition. We also demonstrate that its resistance capability against packet loss weakens as hop number increases. Furthermore, with the increase in redundancy, erasure coding has to sacrifice the advantage of energy efficiency for reliability.
Hao Wen 0014, Chuang Lin 0002, Fengyuan Ren, Yao Yue, Xiaomeng Huang
MASS4
2006 Analyzing the Performance and Fairness of BitTorrent-like Networks Using a General Fluid Model
abstract
In this paper, a general fluid model is developed to study the performance and fairness of BitTorrent-like networks. The fluid model incorporates two important features, user settings with multiple groups and inter-group data exchange, in a synthesized way to obtain statistics about the system performance. Our numerical results point out some key parameters of the system, such as the staying time of seeders. Generally, selfish behavior does not receive equal performance degradation, and in some scenarios users have strong incentives of free-riding. We also find content delivery can be greatly deterred when malicious free-riders are overwhelming.
Yao Yue, Chuang Lin 0002, Zhangxi Tan
GLOBECOM1
2006 AntiWorm NPU-based Parallel Bloom filters in Giga-Ethernet LAN
abstract
In this paper, an AntiWorm system based on the Intel IXP Network Processor was implemented using the Parallel Bloom filters technique. The AntiWorm system consists of two components: Bloom filters and Exact Matching engines. The Parallel Bloom filters can identify the suspicious traffic quickly and effectively, and then dispatch them to Exact Matching engines for further investigation. Both the principles and the implementation of the AntiWorm system are introduced in detail. With the consideration of the system performance parameters, two feasible implementation solutions are investigated and the advantages and disadvantages are also compared. The selections of configuration parameters of the AntiWorm system are also discussed. A hash scheme based on MD5's function is proposed for implementing fast hash functions. To test the performance of the AntiWorm system, such as throughput and delay, some experiments are carried out with different simulated traffic condition. The internal statistics of IXP network processor are also collected and analyzed for optimizing the system performance. To demonstrate the operation of the AntiWorm system, assaults by Worm Blaster are used in the test bed, and the experimental results prove the effectiveness of the AntiWorm system. The Software Package WormDetector1.0 is also provided as a software release from the research.
Zhen Chen 0001, Chuang Lin 0002, Jia Ni, Dong-Hua Ruan, Bo Zheng 0007, Zhangxi Tan, Yixin Jiang, Xuehai Peng, An'an Luo, Yao Yue, Yang Wang 0018, Peter D. Ungsunan, Fengyuan Ren
ICC11
2006 Analyzing the performance and fairness of BitTorrent-like networks using a general fluid model
Yao Yue, Chuang Lin 0002, Zhangxi Tan
Comput. Commun.1