Yanghui Ou

dblp:229/0501 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0001-9481-9882ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Neura: A Unified Framework for Hierarchical and Adaptive CGRAs
abstract
Coarse-Grained Reconfigurable Arrays (CGRAs) are a promising solution for energy-efficient acceleration across multiple application domains. Yet, CGRAs face significant scalability challenges that hinder their widespread adoption, stemming from three main concerns: (1) Mapping Scalability — existing mapping algorithms struggle to find feasible and optimal solutions as the design complexity grows; (2) Architectural Limitations — rigid mapping granularity and memory access restrict flexibility and performance; and (3) Dynamic Multi-Kernel Support — dynamic and simultaneous execution of multiple kernels are not thoroughly explored, limiting the applicability of CGRAs in complex multi-kernel scenarios.
Cheng Tan 0002, Miaomiao Jiang, Ruihong Yin, Yanghui Ou, Lei Ju 0001, Jeff Zhang 0001
ASPLOS (2)5
2025 Scaling Co-Packaged Optical Interconnects Using Hybrid 2.5D/3D Integration
abstract
Tightly integrated optical interconnects can provide high-bandwidth, energy-efficient inter-node communication. We describe a novel system which uses hybrid 2.5D/3D integration to compose a state-of-the-art FPGA compute chiplet, three electrical interface chiplets, and three photonic interface chiplets. We use register-transfer-, gate-, transistor-, and device-level simulations to demonstrate the potential for this system to achieve 96Tb/s of bi-directional bandwidth, and we experimentally demonstrate key components including a complete opto-electrical channel. Our results provide a strong case for hybrid 2.5D/3D integration as the key enabler for scaling co-packaged optical interconnects.
Austin Rovinski, Yanghui Ou, Christine Ou, Devesh Khilwani, Yuyang Wang 0003, Songli Wang, Sunwoo Lee 0002, Keren Bergman, Alyosha C. Molnar, Christopher Batten
ISCAS2
2024 ICED: An Integrated CGRA Framework Enabling DVFS-Aware Acceleration
abstract
Coarse-grained reconfigurable arrays (CGRAs) are a promising solution to enable energy-efficient acceleration of applications from different domains. By leveraging reconfiguration at the functional level, they can adapt to significantly different computational patterns. However, the relationships of voltage and frequency with the utilization of CGRA resources and the dynamic management of them are not well explored, leading to inefficient designs. CGRAs have also been successful in accelerating data-dependent streaming applications. However, in these applications, the execution time of each kernel in the pipeline might dynamically vary depending on the characteristics of the input. This also leads to under-utilization of resources for the dynamically changing kernels that do not limit the application throughput. DVFS can also improve energy efficiency for these applications by dynamically changing the voltage and frequency levels of tiles that host non-performance-constraining kernels. This paper proposes ICED - an integrated DVFS-aware framework to map applications on CGRAs that support power islands. ICED proposes a CGRA architecture supporting DVFS islands at varying granularity (from a single tile to a group of tiles) and the related DVFS-aware compilation and mapping toolchain. ICED is the first work that introduces DVFS support for spatio-temporal CGRAs at power-island levels. The experimental evaluation shows that ICED improves average utilization by$\mathbf{2}.\mathbf{3}\times$and energy-efficiency by$\mathbf{1}.\mathbf{32}\times$over a conventional CGRA. With streaming applications, ICED can achieve up to$\mathbf{1}.\mathbf{26}\times$energy-efficiency compared with a state-of-the-art CGRA that introduces partial dynamic reconfiguration to adapt to variations in kernels' throughput.
Cheng Tan 0002, Miaomiao Jiang, Deepak Patil, Yanghui Ou, Zhaoying Li 0004, Lei Ju 0001, Tulika Mitra, Antonino Tumeo, Jeff Zhang 0001
MICRO4
2023 Symbolic Elaboration: Checking Generator Properties in Dynamic Hardware Description Languages
Peitian Pan, Shunning Jiang, Yanghui Ou, Christopher Batten
MEMOCODE3
2022 big.VLITTLE: On-Demand Data-Parallel Acceleration for Mobile Systems on Chip
abstract
Single-ISA heterogeneous multi-core architectures offer a compelling high-performance and high-efficiency solution to executing task-parallel workloads in mobile systems on chip (SoCs). In addition to task-parallel workloads, many data-parallel applications, such as machine learning, computer vision, and data analytics, increasingly run on mobile SoCs to provide real-time user interactions. Next-generation scalable vector architectures, such as the RISC-V Vector Extension and Arm SVE, have recently emerged as unified vector abstractions for both large- and small-scale systems. In this paper, we propose novel area-efficient high-performance architectures called big.VLITTLE that support next-generation vector architectures to efficiently accelerate data-parallel workloads in conventional big.LITTLE systems. big.VLITTLE architectures reconFigure multiple little cores on demand to work as a decoupled vector engine when executing data-parallel workloads. Our results show that a big.VLITTLE system can achieve $1.6\times$ performance speedup over an area-comparable big.LITTLE system equipped with an integrated vector unit across multiple data-parallel applications and $1.7\times$ speedup compared to an aggressive decoupled vector engine for task-parallel workloads.
Tuan Ta, Khalid Al-Hawaj, Nick Cebry, Yanghui Ou, Eric Hall, Courtney Golden, Christopher Batten
MICRO4
2021 UMOC: Unified Modular Ordering Constraints to Unify Cycle- and Register-Transfer-Level Modeling
abstract
We propose unified modular ordering constraints (UMOC), a novel approach that seamlessly unifies method-based cycle-level (CL) modeling and signal-based register-transfer-level (RTL) modeling. Motivated by the challenges in state-of-the-art CL modeling methodologies and existing CL/RTL composition attempts, UMOC successfully breaks the trade-off between model fidelity and scheduling modularity for CL modeling and provides seamless composition of CL and RTL models. Instead of requiring the designer to specify the global intra-cycle ordering of hardware processes, UMOC eliminates this burden using implicit local ordering constraints of RTL signals and explicit local ordering constraints of CL methods. We implement and evaluate UMOC in PyMTL3, a state-of-the-art open-source Python-based hardware modeling framework.
Shunning Jiang, Yanghui Ou, Peitian Pan, Christopher Batten
DAC2
2021 Ultra-Elastic CGRAs for Irregular Loop Specialization
abstract
Reconfigurable accelerator fabrics, including coarse-grain reconfigurable arrays (CGRAs), have experienced a resurgence in interest because they allow fast-paced software algorithm development to continue evolving post-fabrication. CGRAs traditionally target regular workloads with data-level parallelism (e.g., neural networks, image processing), but once integrated into an SoC they remain idle and unused for irregular workloads. An emerging trend towards repurposing these idle resources raises important questions for how to efficiently map and execute general-purpose loops which may have irregular memory accesses, irregular control flow, and inter-iteration loop dependencies. Recent work has increasingly leveraged elasticity in CGRAs to mitigate the first two challenges, but elasticity alone does not address inter-iteration loop dependencies which can easily bottleneck overall performance. In this paper, we address all three challenges for irregular loop specialization and propose ultra-elastic CGRAs (UE-CGRAs), a novel elastic CGRA that accelerates true-dependency bottlenecks and saves energy in irregular loops by overcoming traditional VLSI challenges. UE-CGRAs allow configurable fine-grain dynamic voltage and frequency scaling (DVFS) for each of potentially hundreds of tiny processing elements (PEs) in the CGRA, enabling chains of connected PEs to “rest” at lower voltages and frequencies to save energy, while other chains of connected PEs can “sprint” at higher voltages and frequencies to accelerate through true-dependency bottlenecks. UE-CGRAs rely on a novel ratiochronous clocking scheme carefully overlaid on the inter-PE elastic interconnect to enable low-latency crossings while remaining fully verifiable with commercial static timing analysis tools. We present the UE-CGRA analytical model, compiler, architectural template, and VLSI circuitry, and we demonstrate how UE-CGRAs can specialize for irregular loops and improve performance ($ 1.42-1.50\times$) or energy efficiency $(1.24-2.32\times)$ with reasonable area overhead compared to traditional inelastic and elastic CGRAs, while also improving performance ($ 1.35-3.38\times$) or energy efficiency (up to $1.53\times$) compared to a RISC-V core.
Christopher Torng, Peitian Pan, Yanghui Ou, Cheng Tan 0002, Christopher Batten
HPCA3
2020 Implementing Low-Diameter On-Chip Networks for Manycore Processors Using a Tiled Physical Design Methodology
abstract
Manycore processors are now integrating up to 1000 simple cores into a single die, yet these processors still rely on high-diameter mesh on-chip networks (OCNs) without complex flow-control nor custom circuits due to three reasons: (1) manycores require simple, low-area routers; (2) manycores usually use standard-cell-based design; and (3) manycores use a tiled physical design methodology. In this paper, we explore mesh and torus topologies with internal concentration and/or ruche channels that require low area overhead and can be implemented using a traditional standard-cell-based tiled physical design methodology. We use a combination of analytical and RTL modeling along with layout-level results for both hard macros and a 3×3mm 256-terminal OCN in a 14-nm technology for twelve topologies. Critically, the networks we study use a tiled physical design methodology meaning they: (1) tile a homogeneous hard macro across the chip; (2) implement chip top-level routing between hard macros via short wires to neighboring macros; and (3) use timing closure for the hard macro to quickly close timing at the chip top-level. Our results suggest that a concentration factor of four and a ruche factor of two in a 2D-mesh topology can reduce latency by over 2× at similar area and bisection bandwidth for both small and large messages compared to a 2D-mesh baseline.
Yanghui Ou, Shady O. Agwa, Christopher Batten
NOCS1
2020 Feasibility of Fingertip Oscillometric Blood Pressure Measurement: Model-Based Analysis and Experimental Validation
abstract
The most commonly used oscillometric upper-arm (UA) blood pressure (BP) monitors are not convenient enough for ambulatory BP monitoring, given the large size of the arm cuff and the compression of UA during the measurement. Finger-worn oscillometric BP devices featuring miniaturized finger cuff have been developed and researched as an alternative solution to the UA-based measurement, yet the reliability of the finger-based measurement is still questioned. To investigate the feasibility of oscillometric BP measurements at the finger position, we performed model-based analysis and experimental validation to explore the underlying issues associated with extending the cuff-based oscillometric approach from UA to other alternative sites. The simulation results revealed that a larger bone-to-tissue volume ratio produced a lower pressure transmission efficiency, which can account for the inter-site measurement discrepancies of mean blood pressure (MBP). We also experimentally compared the oscillometric MBP measurements at UA, middle forearm, wrist, finger proximal phalanx, and finger distal phalanx (FD) of 20 young adults, and each position was matched with a cuff of appropriate size and kept at the same height with the heart. The experimental results demonstrated that FD could be a superior alternative position for oscillometric BP measurement, as it requires the smallest cuff size while providing the most consistent MBP with the UA. Our analysis also suggested that further study is demanded to identify the appropriate oscillometric algorithm for reliable systolic blood pressure and diastolic blood pressure measurements at FD.
Jing Liu 0020, Charles G. Sodini, Yanghui Ou, Bryan P. Yan, Yuan-Ting Zhang, Ni Zhao
IEEE J. Biomed. Health Informatics3
2019 PyOCN: A Unified Framework for Modeling, Testing, and Evaluating On-Chip Networks
abstract
There is a growing interest in the open-source hardware movement to amortize non-recurring engineering costs by using plug-and-play system-on-chip (SoC) designs, where the communication among different components is provided by an on-chip interconnection network. Unfortunately, building an on-chip network (OCN) that is suitable for a specific SoC design requires the exploration of a large number of design options and involves diverse research methodologies to evaluate performance, area, energy, and timing. In this paper, we propose PyOCN, a unified framework that vertically integrates multiple research methodologies to enable productively exploring the OCN design space. PyOCN is the first comprehensive framework for modeling (e.g., functional-level, cycle-level, and register-transfer-level), testing (e.g., unit testing, integration testing, and property-based random testing), and evaluating (e.g., simulating, generating, and characterizing) on-chip interconnection networks. We use a case study based on a 64-terminal butterfly network to illustrate the key features of PyOCN and to demonstrate the framework's potential in productively modeling, testing, and evaluating OCNs.
Cheng Tan 0002, Yanghui Ou, Shunning Jiang, Peitian Pan, Christopher Torng, Shady O. Agwa, Christopher Batten
ICCD2