Jin Ouyang

dblp:21/6027 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Are Android Developers Following Privacy Guidelines? A Study on Logging Practices of Personal Data
abstract
Logging is a common practice in software development, widely used for debugging, testing, and performance monitoring. However, recording sensitive user data can introduce severe privacy risks. Past incidents involving leaked logs have prompted platforms such as Android to publish strict guidelines discouraging developers from logging personally identifiable information (PII) and other ''linkable'' or ''ambiguous'' data unless strictly required for core functionality. To evaluate real-world compliance, we examined the logging practices of 500 Android applications across six categories. Our findings reveal that 264 apps contain logging violations, from which we identified 864 instances of sensitive data exposure. Notably, 54% of these violations stem from debugging logs that should have been removed before release. The recorded data includes PII such as email addresses and phone numbers, as well as linkable information such as shopping history, health records, and private messages each in direct violation of Android's privacy guidelines. Moreover, removing these logging statements from application source code does not affect app functionality, raising questions about their necessity. Our analysis shows that most violations originate from debugging practices, third-party analytics tracking, and HTTP request logging. Further, by applying a large language model (LLM) to inspect an additional set of 300 applications, we found 240 apps exhibiting sensitive data logging violations. We also discovered that many apps share log data with third-party services, often contradicting their own privacy policies. To mitigate these risks, we provide practical recommendations for both app developers and mobile platforms to enforce responsible and privacy-preserving logging practices.
Jin Ouyang, Tiash Roy, Daqing Hou, Yuzhe Tang, Xueling Zhang
WISEC1
2026 Nappa: NNA-Compatible and Privacy-Preserving DNN Training Framework via Vector Decomposition
abstract
How to preserve the data privacy during the training of deep neural network (DNN) is a key security concern in the artificial intelligence era. However, most existing solutions based on homomorphic encryption and Trusted Execution Environment (TEE) are incompatible with heterogeneous Neural Network Accelerators (NNAs), leading to significant performance loss. We propose a novel method based on vector decomposition to allocate operators across different NNAs, ensuring both throughput and privacy simultaneously. Furthermore, based on this approach, we have designed a compiler that automatically converts front-end model descriptions into backend encrypted computation graphs, which is running securely over trusted and untrusted hardware. This compiler heuristically determines the allocation scheme based on hardware affinity and cross-hardware communication costs, significantly reducing additional overhead. Experimental results demonstrate that our method does not incur extra accuracy costs and achieves a throughput significantly higher than existing methods. Deploying our approach at scale on a platform with a billion users, we have verified its negligible impact on real-world operations while ensuring the privacy protection capability for cross-domain data.
Yan Zhang 0002, Qiushi Li 0002, Ju Ren 0001, Yiqiao Liao, Jin Ouyang, Chengru Song, Honghuan Wu, Kaiqiao Zhan, Ben Wang 0006, Xu Chen 0004, Yaoxue Zhang
IEEE Trans. Dependable Secur. Comput.5
2025 Cauchy: A Cost-Efficient LLM Serving System through Adaptive Heterogeneous Deployment
abstract
Recent advances in large language models (LLMs) have intensified the need for serving LLMs that are cost-efficient and QoS-guaranteed. Existing frameworks often co-locate computationally distinct prefill and decode instances on homogeneous GPUs, overlooking their unique resource demands and under-utilizing heterogeneous GPUs. This leads to suboptimal resource utilization and increased capital expenditure. We present Cauchy, a LLM serving framework that adaptively deploys prefill and decode computation to the most suitable heterogeneous GPUs and dynamically schedules user requests. At the core of Cauchy is choosing proper GPU Combo, a conceptual GPU combination encompassing diverse GPU configurations, for their cost efficiency in running prefill-decode pairs. Cauchy deploys a set of combos to satisfy QoS requirements (e.g., goodput) of LLM inference. Cauchy further employs hierarchical scheduling to handle user requests, using opportunistic scheduling within the allocated GPU Combos and a goodput-weighted round-robin policy across GPU Combos. Dynamic autoscaling is used to stabilize the cost-efficiency in the face of surging requests. Experiments show that Cauchy achieves up to a 38.3% improvement in Tokens/USD efficiency over the state-of-the-art baselines, while maintaining strict Service Level Objectives (SLOs). Our work highlights the importance of leveraging workload and GPU heterogeneity to achieve superior cost-efficient LLM serving.
Renyu Yang, Yuxi Luo, Menghao Zhang 0001, Li Li 0029, Chunming Hu, Tianyu Wo, Chengru Song, Jin Ouyang
SoCC11
2025 Cuckoo: Deadline-Aware Job Packing on Heterogeneous GPUs for DL Model Training
abstract
The growing scale and heterogeneity of GPU clusters pose new challenges to deep learning (DL) job scheduling. While existing schedulers primarily focus on GPU utilization, they often ignore multi-dimensional resource demands of DL workloads and lack precise execution time estimation for co-located jobs. While Muri pioneered the use of interleaving execution to improve resource efficiency, it simplified interference when jobs using one resource simultaneously and is agnostic to the deadline constraints. The job grouping also comes to suboptimal when heterogeneous GPU devices are taken into account. In this paper, we propose Cuckoo, a scheduling system that packs deep learning jobs with stringent deadline requirements over a set of heterogeneous GPU devices where multi-dimensional resources are interleaved and shared by a group of jobs. Specifically, the interleaving execution of simultaneous jobs is characterized and modeled through stage-grained execution time estimation considering the runtime performance interference and the impact of GPU heterogeneity on the job performance. The job packing is formulated as a multi-objective optimization problem which is then solved by the maximum weight matching algorithm. Cuckoo then allocates heterogeneous resources to the packed job groups through a graph-based maximum flow and minimum cut algorithm. Experiments show that Cuckoo improves deadline satisfaction rate by 2.38x and reduces average job completion time (JCT) by 1.81x compared with the state-of-the-art approaches. Cuckoo is implemented based on Kubernetes and has been deployed in Kuaishou to serve thousands of model training jobs that can be interleaved on shared heterogeneous GPU clusters
Yuzheng Zhang, Renyu Yang, Weihan Jiang, Tianyu Ye, Yiqiao Liao, Penghao Zhang, Tiezi Zhang, Tianyu Wo, Chunming Hu, Chengru Song, Jin Ouyang
SoCC13
2025 LLM4Rec-LIGHTNING: High-Throughput Training System for LLM4Rec on Memory-Constrained GPUs
Guowang Zhang, Jin Ouyang, Nong Xiao 0001
ICA3PP (8)3
2025 Hypergraph-Driven Tabular Data Synthesis with Multi-objective Optimization
Jin Ouyang, Sheng Xiang 0001, Ying Zhang 0001, Lu Qin 0001
ICDAR (5)1
2025 KAIOPS: A Platform Solution of End-to-End Multi-Modal AIOps for AI Training at Scale
abstract
The resilience of large-scale AI training platforms are fundamental to enabling contemporary AI innovation and business development. However, with the rapid increase in the scale and complexity of AI model training tasks, anomalies become the norm rather than the exception at scale. Failing to handle them properly may lead to enormous resource waste and prolonged development cycles. Traditional anomaly detection methods struggle to tackle the complex temporal characteristics and extreme class imbalance inherently manifesting in training tasks, and fall short in automated solution to root cause analysis and the follow-up remediation. This paper proposes KAIOPS, an end-to-end automated platform solution for handling anomalies and engineering experience of daily operational maintenance for large-scale AI training clusters at Kuaishou. KAIOPS employs a Temporal Context Encoding mechanism to precisely capture and encode long-term trends and critical temporal context information within fault evolution. The detection model elaborates a dynamic class-weighted loss function for enhancing the detection performance. To deliver a complete end-to-end intelligent processing pipeline, KAIOPS further leverages knowledge graph and LLMs for automated root cause analysis and actionable solution generation. Extensive experiments, on the basis of data collected from Kuaishou’s production-grade training clusters, show the superior performance of our proposed approach. KAIOPS has been deployed in Kuaishou, in both testbed and production grade environments, consisting of with over 10,000 GPUs, and accelerate the reliability assurance for industry-scale model training and serving.
Zeying Wang, Penghao Zhang, Xu Wang 0007, Tianyu Wo, Chunming Hu, Chengru Song, Jin Ouyang, Renyu Yang
ASE9
2025 Kair: A Statistical and Causal Approach to Pinpointing Stragglers in Distributed Model Training
abstract
The distributed deep learning training process within large-scale clusters serves as the foundation of contemporary artificial intelligence. However, its inherent characteristics make it particularly sensitive to stragglers, specifically the presence of slow workers, which can significantly decelerate the entire procedure. Observability tools are essential for identifying stragglers within systems. However, the prevailing system profiling tools are either designed for single-node analysis, lacking visibility across multiple workers, or they recognize stragglers but only deliver high-level symptoms, providing engineers with insufficient insight into the underlying causes.We design Kair, a robust production-standard observability tool. Kair uses an innovative hierarchical approach, transitioning from statistical anomaly detection to causal inference. It employs Kolmogorov-Smirnov statistics for the identification of statistically anomalous workers and implements a causal path tracing algorithm to accurately determine the specific operations, such as computation or communication, that are responsible for the delay. Kair has been evaluated in a production cluster of 2,048 NVIDIA A800 GPUs and demonstrated high effectiveness in detecting latent stragglers at the framework level that are often overlooked by conventional tools. It offers precise suggestions that markedly reduce processing inefficiencies and engineering workload.
Yitang Yang, Jiapeng Chen, Tianyu Wo, Chunming Hu, Chengru Song, Jin Ouyang, Renyu Yang
ASE8
2025 Efficient relation-aware heterogeneous graph neural network for fraud detection
Enxia Li, Jin Ouyang, Sheng Xiang 0001, Lu Qin 0001, Ling Chen 0006
World Wide Web (WWW)2
2024 Kale: Elastic GPU Scheduling for Online DL Model Training
abstract
Large-scale GPU clusters have been widely used for effectively training both online and offline deep learning (DL) jobs. However, elastic scheduling in most cases of resource schedulers is dedicated for offline model training where resource adjustment is planned ahead of time. The native autoscaling policy is on the basis of pre-defined threshold and, if applied directly in online model training, often suffers from belated resource adjustment, leading to diminished model accuracy. In this paper, we present Kale, a novel elastic GPU scheduling system to improve the performance of online DL model training. Through traffic forecasting and resource-throughput modeling, Kale automatically pinpoints the number of required GPUs that best accommodate the on-the-fly data samples before performing stabilized autoscaling. An advanced data shuffling strategy is further employed for balancing uneven samples among different training workers, thereby improving the runtime efficacy. Experiments show that Kale substantially outperforms the state-of-the-art solutions. Compared with the default HPA autoscaling strategy, Kale reduces the accumulated lag and downtime by 69.2% and 33.1%, respectively, whilst lowering the SLO violation rate from 19.57% to just 2.6%. Kale has been deployed at Kuaishou's production-level GPU clusters and successfully underpins real-time video recommendation and advertisement at scale.
Renyu Yang, Jin Ouyang, Weihan Jiang, Tianyu Ye, Menghao Zhang 0001, Sui Huang, Chengru Song, Di Zhang 0026, Tianyu Wo, Chunming Hu
SoCC3
2021 Perph: A Workload Co-location Agent with Online Performance Prediction and Resource Inference
abstract
Striking a balance between improved cluster utilization and guaranteed application QoS is a long-standing research problem in cluster resource management. The majority of current solutions require a large number of sandboxed experimentation for different workload combinations and leverage them to predict possible interference for incoming workloads. This results in non-negligible time complexity that severely restricts its applicability to complex workload co-locations. The nature of pure offline profiling may also lead to model aging problem that drastically degrades the model precision. In this paper, we present Perph, a runtime agent on a per node basis, which decouples ML-based performance prediction and resource inference from centralized scheduler. We exploit the sensitivity of long-running applications to multi-resources for establishing a relationship between resource allocation and consequential performance. We use Online Gradient Boost Regression Tree (OGBRT) to enable the continuous model evolution. Once performance degradation is detected, resource inference is conducted to work out a proper slice of resources that will be reallocated to recover the target performance. The integration with Node Manager (NM) of Apache YARN shows that the throughput of Kafka data-streaming application is 2.0x and 1.82x times that of isolation execution schemes in native YARN and pure cgroup cpu subsystem. In TPC-C benchmarking, the throughput can also be improved by 35% and 23% respectively against YARN native and cgroup cpu subsystem.
Jianyong Zhu, Renyu Yang, Chunming Hu, Tianyu Wo, Shiqing Xue, Jin Ouyang, Jie Xu 0007
CCGRID6
2019 Perphon: a ML-based Agent for Workload Co-location via Performance Prediction and Resource Inference
abstract
Cluster administrators are facing great pressures to improve cluster utilization through workload co-location. Guaranteeing performance of long-running applications (LRAs), however, is far from settled as unpredictable interference across applications is catastrophic to QoS [2]. Current solutions such as [1] usually employ sandboxed and offline profiling for different workload combinations and leverage them to predict incoming interference. However, the time complexity restricts the applicability to complex co-locations. Hence, this issue entails a new framework to harness runtime performance and mitigate the time cost with machine intelligence: i) It is desirable to explore a quantitative relationship between allocated resource and consequent workload performance, not relying on analyzing interference derived from different workload combinations. The majority of works, however, depend on offline profiling and training which may lead to model aging problem. Moreover, multi-resource dimensions (e.g., LLC contention) that are not completely included by existing works but have impact on performance interference need to be considered [3]. ii) Workload co-location also necessitates fine-grained isolation and access control mechanism. Once performance degradation is detected, dynamic resource adjustment will be enforced and application will be assigned an access to specific slices of each resources. Inferring a "just enough" amount of resource adjustment ensures the application performance can be secured whilst improving cluster utilization.
Jianyong Zhu, Renyu Yang, Chunming Hu, Tianyu Wo, Shiqing Xue, Jin Ouyang, Jie Xu 0007
SoCC6
2017 Reliable Computing Service in Massive-Scale Systems through Rapid Low-Cost Failover
abstract
Large-scale distributed systems deployed as Cloud datacenters are capable of provisioning service to consumers with diverse business requirements. Providers face pressure to provision uninterrupted reliable services while reducing operational costs due to significant software and hardware failures. A widely adopted means to achieve such a goal is using redundant system components to implement user-transparent failover, yet its effectiveness must be balanced carefully without incurring heavy overhead when deployed—an important practical consideration for complex large-scale systems. Failover techniques developed for Cloud systems often suffer serious limitations, including mandatory restart leading to poor cost-effectiveness, as well as solely focusing on crash failures, omitting other important types, such as timing failures and simultaneous failures. This paper addresses these limitations by presenting a new approach to user-transparent failover for massive-scale systems. The approach uses soft-state inference to achieve rapid failure recovery and avoid unnecessary restart, with minimal system resource overhead. It also copes with different failures, including correlated and simultaneous events. The proposed approach was implemented, deployed and evaluated within Fuxi system, the underlying resource management system used within Alibaba Cloud. Results demonstrate that our approach tolerates complex failure scenarios while incurring at worst 228.5 microsecond instance overhead with 1.71 percent additional CPU usage.
Renyu Yang, Peter Garraghan, Yihui Feng, Jin Ouyang, Jie Xu 0007, Zhuo Zhang 0015
IEEE Trans. Serv. Comput.5
2016 Hybrid Drowsy SRAM and STT-RAM Buffer Designs for Dark-Silicon-Aware NoC
abstract
The breakdown of Dennard scaling prevents us from powering all transistors simultaneously, leaving a large fraction of dark silicon. This crisis has led to innovative work on power-efficient core and memory architecture designs. However, the research for addressing dark silicon challenges with network-on-chip (NoC), which is a major contributor to the total chip power consumption, is largely unexplored. In this paper, we comprehensively examine the network power consumers and the drawbacks of the conventional power-gating techniques. To overcome the dark silicon issue from the NoC's perspective, we propose DimNoC, a dim silicon scheme, which leverages recent drowsy SRAM design and spin-transfer torque RAM (STT-RAM) technology to replace pure SRAM-based NoC buffers. In particular, we propose two novel hybrid buffer architectures: 1) a hierarchical buffer architecture, which divides the input buffers into a set of levels with different power states and 2) a banked buffer architecture, which organizes the drowsy SRAM and the STT-RAM in different banks, and accesses them in an interleaved fashion to hide the long write latency of STT-RAM. In addition, our hybrid buffer design enables NoC data retention mechanism by storing packets in drowsy SRAM and nonvolatile STT-RAM in a lossless manner. Combined with flow control schemes, the NoC data retention mechanism can improve network performance and power simultaneously. Our experiments over real workloads show that DimNoC can achieve 30.9% network energy saving, 20.3% energy-delay product reduction, and 7.6% router area reduction compared with pure SRAM-based NoC design.
Jia Zhan, Jin Ouyang, Fen Ge, Jishen Zhao, Yuan Xie 0001
IEEE Trans. Very Large Scale Integr. Syst.2
2015 DimNoC: a dim silicon approach towards power-efficient on-chip network
abstract
The diminishing momentum of Dennard scaling leads to the ever increasing power density of integrated circuits, and a decreasing portion of transistors on a chip that can be switched on simultaneously---a problem recently discovered and known as dark silicon. There has been innovative work to address the "dark silicon" problem in the fields of power-efficient core and cache system. However, dark silicon challenges with Network-on-Chip (NoC) are largely unexplored. To address this issue, we propose DimNoC, a "dim silicon" approach, which leverages drowsy SRAM and STT-RAM technologies to replace pure SRAM-based NoC buffers. Specifically, we propose two novel hybrid buffer architectures: 1) a Hierarchical Buffer (HB) architecture, which divides the input buffers into a hierarchy of levels with different memory technologies operating at various power states; 2) a Banked Buffer (BB) architecture, which organizes drowsy SRAM and STT-RAM into separate banks in order to hide the long write-latency of STT-RAM. Our experiments show that the proposed DimNoC can achieve 30.9% network energy saving, 20.3% energy-delay product (EDP) reduction, and 7.6% router area decrease compared with the baseline SRAM-based NoC design.
Jia Zhan, Jin Ouyang, Fen Ge, Jishen Zhao, Yuan Xie 0001
DAC2
2014 Optimizing the NoC Slack Through Voltage and Frequency Scaling in Hard Real-Time Embedded Systems
abstract
Hard real-time embedded systems impose a strict latency requirement on interconnection subsystems. In the case of network-on-chip (NoC), this means each packet of a traffic stream has to be delivered within a time interval. In addition, with the increasing complexity of NoC, it consumes a significant portion of total chip power, which boosts the power footprint of such chips. In this paper, we propose a methodology to minimize the energy consumption of NoC without violating the prespecified latency deadlines of real-time applications. First, we develop a formal approach based on network calculus to obtain the worst-case delay bound of all packets, from which we derive a safe estimate of the number of cycles that a packet can be further delayed in the network without violating its deadline-the worst-case slack. With this information, we then develop an optimization algorithm that trades the slacks for lower NoC energy. Our algorithm recognizes the distribution of slacks for different traffic streams, and assigns different voltages and frequencies to different routers to achieve NoC energy-efficiency, while meeting the deadlines for all packets. Furthermore, we design a feedback-control strategy to enable dynamic frequency and voltage scaling on the network routers in conjunction with the energy optimization algorithm. It can flexibly improve the energy-efficiency of the overall network in response to sporadic traffic patterns at runtime.
Jia Zhan, Nikolay Stoimenov, Jin Ouyang, Lothar Thiele, Narayanan Vijaykrishnan, Yuan Xie 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2013 Designing energy-efficient NoC for real-time embedded systems through slack optimization
abstract
Hard real-time embedded systems impose a strict latency requirement on interconnection subsystems. In the case of network-on-chip (NoC), this means each packet of a traffic stream has to be delivered within a time interval. In addition, with the increasing complexity of NoC, it consumes a significant portion of total chip power, which boosts the power footprint of such chips. In this work, we propose a methodology to minimize the energy consumption of NoC without violating the pre-specified latency deadlines of real-time applications. First, we develop a formal approach based on network calculus to obtain the worst-case delay bound of all packets, from which we derive a safe estimate of the number of cycles that a packet can be further delayed in the network without violating its deadline---the worst-case slack. With this information, we then develop an optimization algorithm that trades the slacks for lower NoC energy. Our algorithm recognizes the distribution of slacks for different traffic streams, and assigns different voltages and frequencies to different routers to achieve NoC energy-efficiency, while meeting the deadlines for all packets.
Jia Zhan, Nikolay Stoimenov, Jin Ouyang, Lothar Thiele, Narayanan Vijaykrishnan, Yuan Xie 0001
DAC3
2011 Enabling quality-of-service in nanophotonic network-on-chip
abstract
With the recent development in silicon photonics, researchers have developed optical network-on-chip (NoC) architectures that achieve both low latency and low power, which are beneficial for future large scale chip-multiprocessors (CMPs). However, none of the existing optical NoC architectures has quality-of-service (QoS) support, which is a desired feature of an efficient interconnection network. QoS support provides contending flows with differentiated bandwidths according to their priorities (or weights), which is crucial to account for application-specific communication patterns and provides bandwidth guarantees for real-time applications. In this paper, we propose a quality-of-service framework for optical network-on-chip based on frame-based arbitration. We show that the proposed approach achieves excellent differentiated bandwidth allocation with only simple hardware additions and low performance overheads. To the best of our knowledge, this is the first work that provides QoS support for optical network-on-chip.
Jin Ouyang, Yuan Xie 0001
ASP-DAC1
2011 A frequent-value based PRAM memory architecture
abstract
Phase Change Random Access Memory (PRAM) has great potential as the replacement of DRAM as main memory, due to its advantages of high density, non-volatility, fast read speed, and excellent scalability. However, poor endurance and high write energy appear to be the challenges to be tackled before PRAM can be adopted as main memory. In order to mitigate these limitations, prior research focuses on reducing write intensity at the bit level. In this work, we study the data pattern of memory write operations, and explore the frequent-value locality in data written back to main memory. Based on the fact that many data are written to memory repeatedly, an architecture of frequent-value storage is proposed for PRAM memory. It can significantly reduce the write intensity to PRAM memory so that the lifetime is improved and the write energy is reduced. The trade-off between endurance and capacity of PRAM memory is explored for different configurations. After using the frequent-value storage, the endurance of PRAM is improved to about 1.6X on average, and the write energy is reduced by 20%.
Guangyu Sun 0003, Dimin Niu, Jin Ouyang, Yuan Xie 0001
ASP-DAC3
2011 F2BFLY: an on-chip free-space optical network with wavelength-switching
abstract
The increasing number of cores in contemporary and future many-core processors will continue to demand high through-put, scalable, and energy efficient on-chip interconnection networks. To overcome the intrinsic inefficiency of electrical interconnects, researchers have leveraged recent developments in chip photonics to design novel optical network-on-chip (NoC). However, existing optical NoCs are mostly based on passively switched, channel-guided optical interconnect in which large amount of power is wasted in heating the micro-rings and maintaining the optical signal integrity.
Jin Ouyang, Dimin Niu, Yuan Xie 0001
ICS1
2010 Evaluation of using inductive/capacitive-coupling vertical interconnects in 3D network-on-chip
abstract
In recent 3DIC studies, through silicon vias (TSV) are usually employed as the vertical interconnects in the 3D stack. Despite its benefit of short latency and low power, forming TSVs adds additional complexities to the fabrication process. Recently, inductive/capactive-coupling links are proposed to replace TSVs in 3D stacking because the fabrication complexities of them are lower. Although state-of-the-art inductive/capacitive-coupling links show comparable bandwidth and power as TSV, the relatively large footprints of those links compromise their area efficiencies. In this work, we study the design of 3D network-on-chip (NoC) using inductive/capacitive-coupling links. We propose three techniques to mitigate the area overhead introduced by using these links: (a) serialization, (b) in-transceiver data compression, and (c) high-speed asynchronous transmission. With the combination of these three techniques, evaluation results show that the overheads of all aspects caused by using inductive/capacitive-coupling vertical links can be bounded under 10%.
Jin Ouyang, Jing Xie 0006, Matthew Poremba, Yuan Xie 0001
ICCAD1
2010 LOFT: A High Performance Network-on-Chip Providing Quality-of-Service Support
abstract
Providing quality-of-service (QoS) for concurrent tasks in many-core architectures is becoming important, especially for real-time applications. QoS support for on-chip shared resources (such as shared cache, bus, and memory controllers)in chip-multiprocessors has been investigated in recent years. Unlike other shared resources, network-on-chip (NoC) does not typically have central arbitration of accesses to the shared resource. Instead, each router shares the responsibility of resource allocation. While such distributed nature benefits the scalable performance of NoC, it also dramatically complicates the problem of providing QoS support for individual flows. Existing approaches to address this problem suffer from various shortcomings such as low network utilization and weak QoS guarantees. In this work, we propose LOFT No architecture which features both high network utilization and strong QoS guarantees. LOFT is based on the combination of two mechanisms: a) locally-synchronized frames (LSF), which is a distributed frame-based scheduling mechanism that provides flexible QoS guarantees to different flows and b) flit-reservation (FRS), which is a flow-control mechanism integrated in LSF that improves network utilization. The experimental results show that LOFT delivers flexible and reliable QoS guarantees while sufficiently utilizes available network capacity to gain high overall throughput.
Jin Ouyang, Yuan Xie 0001
MICRO1
2009 CheckerCore: enhancing an FPGA soft core to capture worst-case execution times
abstract
Embedded processors have become increasingly complex, resulting in variable execution behavior and reduced timing predictability. On such processors, safe timing specifications expressed as bounds on the worst-case execution time (WCET) are generally too loose due to conservative assumptions about complex architectural features, timing anomalies and programmatic complexities. Hence, exploiting the latest architectures may not be an option for embedded systems with hard real-time constraints where deadline misses cannot be tolerated.
Jin Ouyang, Raghuveer Raghavendra, Sibin Mohan, Tao Zhang 0032, Yuan Xie 0001, Frank Mueller 0001
CASES1