Xianqing Yao

dblp:179/0131 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Reconfigurable computing and FPGAs · 41% Electronic design automation · 32% Memory systems · 27%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › memory architecture
multi-bank memory
0.622017
Conflict-Free Loop Mapping for Coarse-Grained Reconfigurable Architecture with Multi-Bank Memory · IEEE Trans. Parallel Distributed Syst. 2017
Joint Modulo Scheduling and Memory Partitioning with Multi-Bank Memory for High-Level Synthesis (Abstract Only) · FPGA 2017
Electronic design automation › high-level synthesis › memory synthesis
memory partitioning
0.422017
Joint Modulo Scheduling and Memory Partitioning with Multi-Bank Memory for High-Level Synthesis (Abstract Only) · FPGA 2017
Conflict-Free Loop Mapping for Coarse-Grained Reconfigurable Architecture with Multi-Bank Memory · IEEE Trans. Parallel Distributed Syst. 2017
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture
0.312017
Conflict-Free Loop Mapping for Coarse-Grained Reconfigurable Architecture with Multi-Bank Memory · IEEE Trans. Parallel Distributed Syst. 2017
Electronic design automation
high-level synthesis
0.312017
Joint Modulo Scheduling and Memory Partitioning with Multi-Bank Memory for High-Level Synthesis (Abstract Only) · FPGA 2017
Reconfigurable computing and FPGAs › coarse-grained reconfigurable architecture
loop mapping
0.312017
Conflict-Free Loop Mapping for Coarse-Grained Reconfigurable Architecture with Multi-Bank Memory · IEEE Trans. Parallel Distributed Syst. 2017
Reconfigurable computing and FPGAs
modulo scheduling
0.312017
Joint Modulo Scheduling and Memory Partitioning with Multi-Bank Memory for High-Level Synthesis (Abstract Only) · FPGA 2017

Methods — techniques the papers use, named apart from their topics

modulo scheduling · 0.6operator scheduling · 0.3memory partitioning · 0.3joint memory partitioning · 0.3
YearPublicationVenuePosition
2017 Joint Modulo Scheduling and Memory Partitioning with Multi-Bank Memory for High-Level Synthesis (Abstract Only)
Shouyi Yin, Xianqing Yao, Zhicong Xie, Leibo Liu, Shaojun Wei
FPGA3
2017 Memory fartitioning-based modulo scheduling for high-level synthesis
abstract
High-Level Synthesis (HLS) has been widely recognized as an efficient compilation process targeting FPGAs for algorithm evaluation and product prototyping. However, the massively parallel memory access demands and the extremely expensive cost of single-bank memory with multi-port have impeded loop pipelining performance. Thus, based on an alternative multi-bank memory architecture, a joint approach that employs memory-aware force directed scheduling and multi-cycle memory partitioning is formally proposed to achieve legitimate pipelining kernel and valid bank mapping with less resource consumption and optimal pipelining performance. The experimental results over a variety of benchmarks show that our approach can achieve the optimal pipelining performance and meanwhile reduce the number of multiple independent memory banks by 55.1% on average, compared with the state-of-the-art approaches.
Shouyi Yin, Xianqing Yao, Zhicong Xie, Leibo Liu, Shaojun Wei
ISCAS3
2017 Conflict-Free Loop Mapping for Coarse-Grained Reconfigurable Architecture with Multi-Bank Memory
abstract
Coarse-grained reconfigurable architecture (CGRA) is a promising architecture with high performance, high power-efficiency and attraction of flexibility. The computation-intensive parts of an application (e.g., loops) are often mapped on CGRA for acceleration. Due to the high parallel data access demands, the architecture with multi-bank memory is proposed to improve parallelism. For CGRA with multi-bank memory, a joint solution, which simultaneously considers the memory partitioning and modulo scheduling, is proposed to achieve a valid mapping with better performance. In this solution, the modulo scheduling and operator scheduling are used to achieve a valid loop mapping and a valid data placement without any memory access conflicts. By avoiding the pipelining stalls caused by conflicts, the performance of loop mapping is greatly improved. The experimental results on benchmarks of the Livermore, Polybench and Mediabench show that our approach can improve the performance of loops on CGRA to 1.89×, 1.49× and 1.37× compared with REGIMap, HTDM and REGIMap with memory partitioning, at cost of an acceptable increase in compilation time.
Shouyi Yin, Xianqing Yao, Dajiang Liu, Jiangyuan Gu, Leibo Liu, Shaojun Wei
IEEE Trans. Parallel Distributed Syst.2
2016 Joint loop mapping and data placement for coarse-grained reconfigurable architecture with multi-bank memory
abstract
Coarse-Grained Reconfigurable Architecture (CGRA) is a promising architecture with high performance, high power-efficiency and attraction of flexibility. The compute-intensive parts of an application (e.g. loops) are often mapped onto CGRA for acceleration. Since the high-parallel demands of PEs and the extremely expensive cost of single-bank memory with multi-port, the architecture with multi-bank memory is favored increasingly. Based on this purpose, a joint solution, which simultaneously considers modulo scheduling and data placement, is proposed to achieve a valid mapping with better performance. The experimental results on loops from Livermore, Polybench and Mediabench show that our approach can significantly improve the performance of the kernels on CGRA compared with REGIMap, HTDM and REGIMap+MP, with an acceptable increase in compilation time.
Shouyi Yin, Xianqing Yao, Leibo Liu, Shaojun Wei
ICCAD2
2016 Memory-Aware Loop Mapping on Coarse-Grained Reconfigurable Architectures
abstract
The coarse-grained reconfigurable architectures (CGRAs) are a promising class of architectures with the advantages of high performance and high power efficiency. The compute-intensive parts of an application (e.g., loops) are often mapped onto the CGRA for acceleration. Due to the extra overhead of memory access and the limited communication bandwidth between the processing element (PE) array and local memory, previous works trying to solve the routing problem are mainly confined in the internal resources of PE arrays (e.g., PEs and registers). Inevitably, routing with PEs or registers will consume a lot of computational resources and cause the increase of the initiation interval. To solve this problem, this paper makes two contributions: 1) establishing a precise formulation for the CGRA mapping problem while using shared local data memory as a routing resource and 2) extracting an effective approach for mapping loops to CGRAs. The experimental results on loops of the SPEC2006, Livermore, and MiBench show that our approach (called MEMMap) can improve the performance of the kernels on CGRA up to 1.62×, 1.58×, 1.28×, and 1.23× compared with the edge-centric modulo scheduling, EPIMap, REGIMap, and force-directed map, respectively, with an acceptable increase in compilation time.
Shouyi Yin, Xianqing Yao, Dajiang Liu, Leibo Liu, Shaojun Wei
IEEE Trans. Very Large Scale Integr. Syst.2