Colin Yu Lin

dblp:16/6334 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 5 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Reconfigurable computing and FPGAs · 36% Integrated circuit design · 22% Hardware accelerators and domain-specific architectures · 22%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Reconfigurable computing and FPGAs › FPGA-based heterogeneous computing
CPU-FPGA platform
0.512021
Sparse Tucker Tensor Decomposition on a Hybrid FPGA-CPU Platform · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
Hardware accelerators and domain-specific architectures › tensor accelerator
tensor decomposition accelerator
0.512021
Sparse Tucker Tensor Decomposition on a Hybrid FPGA-CPU Platform · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
Integrated circuit design
digital signal processing circuits
0.212016
A Computationally Efficient Reconfigurable FIR Filter Architecture Based on Coefficient Occurrence Probability · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Integrated circuit design
low-power circuit design
0.212016
A Computationally Efficient Reconfigurable FIR Filter Architecture Based on Coefficient Occurrence Probability · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2016
Data mining
high-dimensional data analysis
0.112021
Sparse Tucker Tensor Decomposition on a Hybrid FPGA-CPU Platform · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
Energy-efficient computing › power-performance tradeoff
energy-delay product optimization
0.112012
Operation scheduling and architecture co-synthesis for energy-efficient dataflow computations on FPGAs (abstract only) · FPGA 2012
Reconfigurable computing and FPGAs
FPGA architecture
0.112012
Operation scheduling and architecture co-synthesis for energy-efficient dataflow computations on FPGAs (abstract only) · FPGA 2012
Reconfigurable computing and FPGAs
FPGA power reduction
0.112012
Operation scheduling and architecture co-synthesis for energy-efficient dataflow computations on FPGAs (abstract only) · FPGA 2012
Electronic design automation
high-level synthesis
0.112012
Operation scheduling and architecture co-synthesis for energy-efficient dataflow computations on FPGAs (abstract only) · FPGA 2012
Electronic design automation › high-level synthesis › scheduling
operation scheduling
0.112012
Operation scheduling and architecture co-synthesis for energy-efficient dataflow computations on FPGAs (abstract only) · FPGA 2012

Methods — techniques the papers use, named apart from their topics

tucker decomposition · 1.0kronecker product · 1.0QR decomposition with column pivoting · 1.0common subexpression sharing · 0.2canonic signed digit · 0.2genetic algorithm · 0.1dataflow graph folding · 0.1
YearPublicationVenuePosition
2021 Sparse Tucker Tensor Decomposition on a Hybrid FPGA-CPU Platform
abstract
Recommendation systems, social network analysis, medical imaging, and data mining often involve processing sparse high-dimensional data. Such high-dimensional data are naturally represented as tensors, and they cannot be efficiently processed by conventional matrix or vector computations. Sparse Tucker decomposition is an important algorithm for compressing and analyzing these sparse high-dimensional datasets. When energy efficiency and data privacy are major concerns, hardware accelerators on resource-constraint platforms become crucial for the deployment of tensor algorithms. In this work, we propose a hybrid computing framework containing CPU and FPGA to accelerate sparse Tucker factorization. This algorithm has three main modules: 1) tensor-times-matrix (TTM); 2) Kronecker products; and 3) QR decomposition with column pivoting (QRP). In addition, we accelerate the former two modules on a Xilinx FPGA and the latter one on a CPU. Our hybrid platform achieves$23.6 \times \sim 1091\times $speedup and over 93.519% ~ 99.514% energy savings compared with CPU on the synthetic and real-world datasets.
Weiyun Jiang, Kaiqi Zhang 0002, Colin Yu Lin, Feng Xing, Zheng Zhang 0005
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2016 A Computationally Efficient Reconfigurable FIR Filter Architecture Based on Coefficient Occurrence Probability
abstract
Reconfigurable digital filter is being widely used in applications such as communication and signal processing. Its performance, power consumption, and logic resource utilization are the major factors to be taken into consideration when designing the filters. This paper proposes a concise canonic signed digit coefficient grouping method aiming at reducing the number of common subexpressions (CSs). Further, we statistically analyze every CS occurance for numerous sorts of the finite-impulse response (FIR) filters and obtain characterization of the distribution behavior for all the possible CS patterns in a 16-bit coefficient. Thus, a novel processing element structure is proposed to form a medium-grain array for computationally efficient realization of reconfigurable FIR filter. The experiment results suggest such design implementations typically achieve 21% reduction in silicon area, 20% decrease in power consumption, and 14% improvement in operation speed in comparison to other conventional FIR architectures.
Rui Jia, Haigang Yang, Colin Yu Lin, Rui Chen 0014, Zhenhong Guo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2015 A technology mapper for depth-constrained FPGA logic cells
abstract
In the last decade, progress in logic synthesis has brought about new advantageous circuit representations. These representations, such as And-Inverter Graphs in the ubiquitous open-source synthesizer ABC, have inspired new designs of Field Programmable Gate Arrays (FPGAs), which, instead of using Look-Up Tables (LUTs), mimic the topology of the circuit representation in the basic logic cells. More recent examples are Majority-Inverter Graphs, another uniform representation which has triggered considerable interest in synthesis and which naturally suggests new logic cells. Yet, in this paper we observe how naïvely adapting technology mapping solutions for classic LUT-based FPGAs to these new architectures incurs severe shortcomings. The key issue is that LUTs are inherently input-constrained (the logic function they implement is irrelevant) and have generally a single output; on the other hand, logic cells made of uniform networks of some fundamental logic function (e.g., And-Invert) are constrained in terms of logic depth and multiple outputs are an integral feature. We introduce novel and effective solutions to address these differences; the result is a highly versatile mapper—thus enabling further research in these new architectures—with a significantly better performance than what is described in literature for one such architecture. Specifically, when we compare with the state of the art on one sample architecture, we obtain a significant decrease in area (on average 18% over several benchmarks) while also improving slightly the critical path (a reduction of 3%).
Zhenghong Jiang, Grace Zgheib, Colin Yu Lin, David Novo, Liqun Yang, Haigang Yang, Paolo Ienne
FPL3
2014 A survey of open source processors for FPGAs
abstract
FPGA-based SoCs have been gaining great favors among traditional applications and expanding the application domains to some new areas. Reusing and sharing components is an attractive and pratical methodology for system designers to reduce the design complexity of the SoC architectures. Open source hardware has become an effective method of improving the design productivity. As the core functional component, processors affects the performance of SoC systems. This paper investigates existing open source processors, and gives an overview. From the points of usability and stability, the main features of open source processors are summarized. Following these features, some open source processors with high usability and stability are selected. These open source processors and existing vendor-provided soft processors are implemented on Stratix V and Virtex-7 FPGAs using corresponding EDA tools, Quart us n and ISE. The implementation results are compared and discussed.
Rui Jia, Colin Yu Lin, Zhenhong Guo, Rui Chen 0014, Tongqiang Gao, Haigang Yang
FPL2
2014 Exploring architecture parameters for dual-output LUT based FPGAs
abstract
Dual-output lookup tables (LUTs) are mainstream in the design of commercial FPGA products. A detailed exploration of architectural parameters of FPGAs based on dualoutput LUTs is presented. Different from traditional single-output LUT based architecture, “shared inputs” between the sub-LUTs is a new parameter specific to dual-output architecture. In this paper, we focus on the effect of ratio of shared inputs on the performance and area-efficiency. First, we study the required cluster inputs and derive a relationship between cluster inputs, LUT size and cluster size under different ratios of shared inputs. Secondly, our evaluation results show that a FPGA with 4-LUTs and a shared input ratio of two thirds is preferred for area-efficiency, while a large LUT size of 9 with no shared inputs achieves best performance. Finally, we determine that a LUT size of 4, a cluster size from 3 to 8, and a shared input ratio between 1/3 and 2/3, provide the best area-delay product for dual-output LUT based FPGAs.
Zhenghong Jiang, Colin Yu Lin, Liqun Yang, Haigang Yang
FPL2
2014 A semi-supervised modeling approach for performance characterization of FPGA architectures
abstract
An approach to estimate the performance of FPGA architectures is proposed based on semi-supervised model tree algorithm. The proposed approach avoids synthesizing, mapping, packing, placing and routing, which are essential steps in a traditional flow to obtain the performance of FPGA. Thus it is time efficient while the performance predicted maintains quite close to the result obtained through the traditional method (a tool flow called VTR). This can be utilized effectively during the early FPGA design stage to choose an optimal architecture under a certain metric. Comparisons are made between the performance obtained by the proposed approach and by VTR on a commercial 40nm technology. Results show that the proposed approach has MRE below 7.62% compared to VTR, and improves the time cost by thousands of times when utilized in architecture design space exploration.
Liqun Yang, Haigang Yang, Wei Li 0008, Colin Yu Lin
FPL6
2014 Size aware placement for island style FPGAs
abstract
In this paper we first examine the impact of FPGA size on overall performance and run-time of placement and routing in the context of cluster-based island-style FPGAs. Based on the observations, an FPGA placement algorithm, Min-Size, is introduced to alleviate the deterioration of performance and run-time of placement and routing when using a large FPGA to implement a circuit. We achieve this by allowing Min-Size to generate a more compact placement of logic, I/O and hard blocks. Our experimental results have shown a 3X and AX speedup in placement and routing run-time, a 38% and 41% reduction in wire length, and a 8% and 5% improvement in critical path delay when FPGA size increases 10 times.
Junying Huang, Colin Yu Lin, Haigang Yang
FPT2
2013 A Soft Coarse-Grained Reconfigurable Array Based High-level Synthesis Methodology: Promoting Design Productivity and Exploring Extreme FPGA Frequency
abstract
Compared to the use of a typical software development flow, the productivity of developing FPGA-based compute applications remains much lower. Although the use of high-level synthesis (HLS) tools may partly alleviate this shortcoming, the lengthy low-level FPGA implementation process remains a major obstacle to high productivity computing, limiting the number of compile-debug-edit cycles per day. Furthermore, high-level application developers often lack the intimate hardware engineering experience that is needed to achieve high performance on FPGAs, therefore undermining their usefulness as accelerators. To address the productivity and performance problems, a HLS methodology that utilizes soft coarse-grained reconfigurable arrays (SCGRAs) as an intermediate compilation step is presented. Instead of compiling high-level applications directly to circuits, the compilation process is reduced to an operation scheduling task targeting the SCGRA.
Cheng Liu 0008, Colin Yu Lin, Hayden Kwok-Hay So
FCCM2
2013 Timing-constrained minimum area/power FPGA memory mapping
abstract
Physical block memory is one of the earliest hardened blocks in modern FPGAs. FPGA memory mapping utilizes memory blocks to construct user's logic memory designs. Previous mapping methods optimized for circuit area or power consumption. However, timing performance becomes more important for large and critical logic memory designs. In this work, a critical path delay model will be presented for the first time to estimate implementation performance during memory mapping. Experiment results showed that over 90% estimation errors by the model were within 10% and the average was only 3.5%. Based on the delay model, we proposed the first timing-constrained FPGA memory mapping algorithm. The algorithm generates memory configurations that meet the user's performance requirement. The algorithm also achieves optimal area/power subject to the user's timing constraint. The area or power optimum was validated with a commercial FPGA memory mapper.
Fangqing Du, Colin Yu Lin, Xiuhai Cui, Jiabin Sun, Fei Liu 0011, Haigang Yang
FPL2
2012 Operation scheduling and architecture co-synthesis for energy-efficient dataflow computations on FPGAs (abstract only)
abstract
Compiling high-level user applications for execution on FPGAs often involves synthesizing dataflow graphs beyond the size of the available on-chip computational resources. One way to address this is by folding the execution of the given dataflow graphs onto an array of directly connected simple configurable processing elements (CPEs). Under this scenario, the performance and energy-efficiency of the resulting system depends not only on the mapping schedule of the compute operations on the CPEs, but also on the topology of the interconnect array that connects the CPEs. This paper presents a framework in which the operation scheduler and the underlying CPE interconnect network topology are co-optimized on a per-application basis for energy-efficient FPGA computation. Given the same application, more than 2.5x difference in energy-efficiency was achievable by the use of different common regular array topologies to connect the CPEs. Moreover, by using irregular application-specific interconnect topologies derived from a genetic algorithm, up to 50% improvement in energy-delay-product was achievable when compared to the use of even the best regular topology. The use of such framework is anticipated to serve as part of a rapid high-level FPGA application compiler since minimum hardware place-and-route is needed to generate the optimal schedule and topology.
Colin Yu Lin, Ngai Wong 0001, Hayden Kwok-Hay So
FPGA1
2011 A Model for Peak Matrix Performance on FPGAs
Colin Yu Lin, Hayden Kwok-Hay So, Philip H. W. Leong
FCCM1
2011 A Model for Matrix Multiplication Performance on FPGAs
abstract
Computations involving matrices form the kernel of a large spectrum of computationally demanding applications for which FPGAs have been utilized as accelerators. Their performance is related to their underlying architectural and system parameters such as computational resources, memory and I/O bandwidth. A simple analytic model that gives an estimate of the performance of FPGA-based sparse matrix-vector and matrix-matrix multiplication is presented, dense matrix multiplication being a special case. The efficiency of existing implementations are compared to the model and performance trends for future technologies examined.
Colin Yu Lin, Hayden Kwok-Hay So, Philip H. W. Leong
FPL1
2010 Design space exploration for sparse matrix-matrix multiplication on FPGAs
abstract
The design and implementation of a sparse matrix-matrix multiplication architecture on FPGAs is presented. Performance of the design, in terms of computational latency, as well as the associated power-delay and energy-delay tradeoff are studied. Taking advantage of the sparsity of the input matrices, the proposed design allows user-tunable power-delay and energy-delay tradeoffs by employing different number of processing elements (PEs) in the architecture design and different block size in the blocking decomposition. Such ability allows designers to employ different on-chip computational architecture for different system power-delay and energy-delay requirements. It is in contrast to conventional dense matrix-matrix multiplication architectures that always favor the maximum number of PEs and largest block size. In our implementation, the better energy consumption and power-delay product favors less PEs and smaller block size for the 90%-sparsity matrix-matrix multiplications. While in order to achieve better energy-delay product, more PEs and larger block size are preferred.
Colin Yu Lin, Zheng Zhang 0005, Ngai Wong 0001, Hayden Kwok-Hay So
FPT1
2009 Operation scheduling for FPGA-based reconfigurable computers
abstract
Many high-performance applications involve large data sets that are impossible to fit entirely within on-chip memories of even the largest FPGAs. As a result, they must be stored in off-chip SDRAMs and loaded onto the FPGAs as computations progress. Because of the high latency and energy consumption associated with off-chip memory accesses, it is important to develop efficient operation schedules that not only minimize latency of computations, but also the amount of data I/Os. We formulate this problem as a modified resource-constrained job scheduling problem. The problem is then solved using a list scheduling algorithm that takes advantage of the fast burst-mode access of SDRAMs. Results have shown that for large problem sizes, the performance of our algorithm is within 1% of a hand-optimized matrix-matrix multiplication implementation, with no memory overhead, and is within 0.03% of the theoretical minimum latency of an 8-by-8 cofactor matrix computation.
Colin Yu Lin, Ngai Wong 0001, Hayden Kwok-Hay So
FPL1