EDBT 2026 Demo / reviewers in the wild / expert
Yuan-Ying Chang
dblp:63/9300
· DBLP profile ↗
6ranked-venue papers
2as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Interconnection networks and networks-on-chip · 82% Performance modeling and evaluation · 18% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Interconnection networks and networks-on-chip
router architecture |
0.2 | 1 | 2013 | TS-Router: On maximizing the Quality-of-Allocation in the On-Chip Network · HPCA 2013 |
Interconnection networks and networks-on-chip › router architecture
router microarchitecture |
0.2 | 1 | 2013 | TS-Router: On maximizing the Quality-of-Allocation in the On-Chip Network · HPCA 2013 |
Interconnection networks and networks-on-chip › network scheduling
switch allocation |
0.2 | 1 | 2013 | TS-Router: On maximizing the Quality-of-Allocation in the On-Chip Network · HPCA 2013 |
Performance modeling and evaluation › simulation › architectural simulation
execution-driven simulation |
0.0 | 1 | 2012 | Attackboard: a novel dependency-aware traffic generator for exploring NoC design space · DAC 2012 |
Performance modeling and evaluation › simulation
simulation-based evaluation |
0.0 | 1 | 2012 | Attackboard: a novel dependency-aware traffic generator for exploring NoC design space · DAC 2012 |
Methods — techniques the papers use, named apart from their topics
verilog implementation · 0.2time series prediction · 0.2full-system simulation · 0.2bloom filter · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Traffic-aware frequency scaling for balanced on-chip networks on GPGPUsabstractGeneral-purpose computing on graphics processing units (GPGPU) can provide orders of magnitude more computing power than general purpose processors (CPU) for highly parallel applications. For such parallel applications, the memory traffic pattern of GPGPUs behaves considerably different from that of CPUs. This gives rise to opportunities for optimizing the on-chip interconnection network (NoC) of GPGPUs. In this work, we first investigate the characteristics of GPGPU memory traffic of typical benchmarks and categorize the memory traffic patterns. Different traffic patterns require different throughput in the request and reply paths of the NoC to match the network load. To meet this requirement, we examine the feasibility of scaling the network frequency dynamically to balance the throughput of the request and reply networks. The decision is guided by monitoring some shader cores to identify the memory traffic pattern. Performance evaluation shows that this dynamic frequency tuning design can achieve up to 27% improvement in terms of execution speedup compared to a baseline setting and 7.4% improvement on average. Chiao-Yun Tu, Yuan-Ying Chang, Chung-Ta King, Chien-Ting Chen, Tai-Yuan Wang |
ICPADS | 2 |
| 2014 | Designing Coalescing Network-on-Chip for Efficient Memory Accesses of GPGPUs
Chien-Ting Chen, Yoshi Shih-Chieh Huang, Yuan-Ying Chang, Chiao-Yun Tu, Chung-Ta King, Tai-Yuan Wang, Janche Sang, Ming-Hua Li |
NPC | 3 |
| 2013 | ShieldUS: A novel design of dynamic shielding for eliminating 3D TSV crosstalk coupling noiseabstract3D IC is a promising technology to meet the demands of high throughput, high scalability, and low power consumption for future generation integrated circuits. One way to implement the 3D IC is to interconnect layers of two-dimensional (2D) IC with Through-Silicon Via (TSV), which shortens the signal lengths. Unfortunately, while TSVs are bundled together as a cluster, the crosstalk coupling noise may lead to transmission errors. As a result, the working frequency of TSVs has to be lowered to avoid the errors, leading to narrower bandwidth that TSVs can provide. In this paper, we first derive the crosstalk noise model from the perspective of 3D chip and then propose ShieldUS, a runtime data-to-TSVs remapping strategy. With ShieldUS, the transition patterns of data over TSVs are observed at runtime, and relatively stable bits will be mapped to the TSVs which act as shields to protect the other bits which have more fluctuations. We evaluate the performance of ShieldUS with address lines from real benchmark traces and data lines of different similarities. The results show that ShieldUS is accurate and flexible. We further study dynamic shielding and our design of Interval Equilibration Unit (IEU) can intelligently select suitable parameters for dynamic shielding, which makes dynamic shielding practical and does not need to predefine parameters. This also improves the practicability of ShieldUS. Yuan-Ying Chang, Yoshi Shih-Chieh Huang, Narayanan Vijaykrishnan, Chung-Ta King |
ASP-DAC | 1 |
| 2013 | TS-Router: On maximizing the Quality-of-Allocation in the On-Chip NetworkabstractSwitch allocation is a critical pipeline stage in the router of an Network-on-Chip (NoC), in which flits in the input ports of the router are assigned to the output ports for forwarding. This allocation is in essence a matching between the input requests and output port resources. Efficient router designs strive to maximize the matching. Previous research considers the allocation decision at each cycle either independently or depending on prior allocations. In this paper, we demonstrate that the matching decisions made in a router along time actually form a time series, and the Quality-of-Allocation (QoA) can be maximized if the matching decision is made across the time series, from past history to future requests. Based on this observation, a novel router design, TS-Router, is proposed. TS-Router predicts future requests to arrive at a router and tries to maximize the matching across cycles. It can be extended easily from most state-of-the-art routers in a lightweight fashion. Our evaluation of TS-Router uses synthetic traffic as well as real benchmark programs in full-system simulator. The results show that TS-Router can have higher number of matchings and lower latency. In addition, a prototype of TS-Router is implemented in Verilog, so that power consumption and area overhead are also evaluated. Yuan-Ying Chang, Yoshi Shih-Chieh Huang, Matthew Poremba, Narayanan Vijaykrishnan, Yuan Xie 0001, Chung-Ta King |
HPCA | 1 |
| 2012 | Attackboard: a novel dependency-aware traffic generator for exploring NoC design spaceabstractNetwork-on-chip (NoC) is very important for many applications, such as many-core architectures and application-specific usages. For exploring the design space, several approaches have been proposed with different considerations. In this paper, inspired by bloom filters, we propose Attackboard, a novel design for exploring the design space of NoC, which satisfies accuracy, space efficiency, and simplicity. To justify the usage of Attackboard, a parallel object detection program is used as the benchmark program to evaluate the performance of a specific NoC. By comparing the results with an execution-based simulator, it shows that Attackboard simultaneously achieves the requirements of fast speed, simplicity, and accuracy. Yoshi Shih-Chieh Huang, Yu-Chi Chang, Tsung-Chan Tsai, Yuan-Ying Chang, Chung-Ta King |
DAC | 4 |
| 2009 | Multiprocessor System-on-Chip Profiling Architecture: Design and ImplementationabstractWith the growing needs for advanced functionalities in modern embedded systems, it is now necessary to integrate multiple processors in the system, preferably on a single chip, to support the required computing complexity. The problem is that such multiprocessor system-on-chip (MPSoC) architecture is very complex and its internal behavior is very difficult to track. An effective tool for profiling the behavior of the MPSoC system is in great need. Such a tool is very usefulduring system designfor exploiting various options and identifying potential bottlenecks. In this paper, we introduce the multiprocessor profiling architecture (MPPA) - a general framework for profiling MPSoC embedded systems. The MPPA framework entails the use of FPGA emulation for the target system, the embedding of performance counters for recording system events, and the development of OS drivers for collecting the profiled data. To demonstrate its use, we show the implementation of an MPSoC emulation system based on Leon3 cores following the MPPA framework. We also show how the MPPA framework and the emulator help the designers to identify performance problems and improve their MPSoC embedded system design. Po-Hui Chen, Chung-Ta King, Yuan-Ying Chang, Shau-Yin Tseng |
ICPADS | 3 |