Qinghao Min

dblp:01/11467 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Performance modeling and evaluation · 76% Processor architecture and microarchitecture · 24%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
simulation
0.322014
DAPs: Dynamic Adjustment and Partial Sampling for Multithreaded/Multicore Simulation · DAC 2014
Transformer: a functional-driven cycle-accurate multicore simulator · DAC 2012
Processor architecture and microarchitecture
multicore design
0.222014
DAPs: Dynamic Adjustment and Partial Sampling for Multithreaded/Multicore Simulation · DAC 2014
Transformer: a functional-driven cycle-accurate multicore simulator · DAC 2012
Performance modeling and evaluation › simulation
architectural simulation
0.212014
DAPs: Dynamic Adjustment and Partial Sampling for Multithreaded/Multicore Simulation · DAC 2014
Performance modeling and evaluation › simulation › architectural simulation
sampled simulation
0.212014
DAPs: Dynamic Adjustment and Partial Sampling for Multithreaded/Multicore Simulation · DAC 2014
Performance modeling and evaluation › simulation › architectural simulation
full-system simulation
0.012012
Transformer: a functional-driven cycle-accurate multicore simulator · DAC 2012

Methods — techniques the papers use, named apart from their topics

partial sampling · 0.2dynamic adjustment · 0.2parallelization · 0.1loosely-coupled functional and timing models · 0.1functional-driven simulation · 0.1
YearPublicationVenuePosition
2014 DAPs: Dynamic Adjustment and Partial Sampling for Multithreaded/Multicore Simulation
abstract
Faced with increasingly large multicore chip designs, architects need fast and accurate simulations for their exploration of design spaces within a limited simulation time budget. In multithreaded applications, threads cannot run simultaneously. Sampling is commonly used to reduce simulation time, but conventional sampling barely detects the instantaneous program variations of synchronization events and the inconsistency between phases of each core. This work proposes a dynamic adjustment and partial sampling technique (DAPs), consisting of aggressive sampling, lazy sampling, and regular sampling, to overcome thread interference in multithreaded applications. Moreover, DAPs partially selects sampling cores to reduce the overhead of sampling inconsistent phases.
Chien-Chih Chen, Yin-Chi Peng, Cheng-Fen Chen, Wei-Shan Wu, Qinghao Min, Pen-Chung Yew, Tien-Fu Chen
DAC5
2014 RPSim: A Rapid Prototyping Full-System Simulator for SoC Software Development
abstract
Nowadays, the release of SoC products has come to a burst. Time-to-market of these products has been shortened to an extreme, nearly 8 to 12 months. To reduce production period, hardware architects generally combine well-tuned IP cores in their designs. To guarantee the process of SoC software development, which will finally decide the release time of products, a fast prototyping simulation platform for SoC software development should be available as soon as possible after hardware design. However, state-of-the-art SoC simulators lack the support for fast integration of IP core and require time-consuming compiler chain modifications for new instructions. In this paper, we present Prism, an extensible and easy-to-use full-system SoC simulation platform for SoC software development. Two mechanisms are designed and implemented to support fast prototyping for new IP core simulation or new instruction extension without compiler tool chain modifications. First, a hardware and software hybrid mechanism is proposed for IP core fast prototyping. A seamless interface is used to eliminate the differences among IP cores. Second, a configurable library mechanism is designed for new instruction extension. Register dependence can be maintained for detailed timing simulation without compiler tool chain modification. In such a design, the major effort for extension is to specify the elaborate common customization interface. Experimental results show these mechanisms only involve about 0.36% runtime overhead. Based on RPSim, a graduate student only needs write about 40 lines of code and takes less than half an hour to extend a new IP core simulation in RPSim.
Haojun Wang, Qinghao Min
NAS2
2012 Transformer: a functional-driven cycle-accurate multicore simulator
abstract
Full-system simulators are extremely useful in evaluating design alternatives for multicore. However, state-of-the-art multicore simulators either lack good extensibility due to their tightly-coupled design between functional model (FM) and timing model (TM), or cannot guarantee cycle-accuracy. This paper conducts a comprehensive study on factors affecting cycle-accuracy and uncovers several contributing factors ignored before. Based on the study, we propose a loosely-coupled functional-driven full-system simulator for multicore, namely Transformer. To ensure extensibility and cycle-accuracy, Transformer leverages an architecture-independent interface between FM and TM and uses a lightweight scheme to detect and recover from execution divergence between FM and TM. Based on Transformer, a graduate student only needs to write about 180 lines of code and takes about two months to extend an X86 functional model (QEMU) in Transformer. Moreover, the loosely-coupled design also removes the complex interaction between FM and TM and opens the opportunity to parallelize FM and TM to improve performance. Experimental results show that Transformer achieves an average of 8.4% speedup over GEMS while guaranteeing the cycle-accuracy. A further parallelization between FM and TM leads to 35.3% speedup.
Zhenman Fang, Qinghao Min, Keyong Zhou, Yibin Hu, Haibo Chen 0001, Jian Li 0059, Binyu Zang
DAC2