Wei Jiang 0035

dblp:21/3839-35 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
0since 2021 · last 2011
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-authorSoftware engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Electronic design automation · 79% Reconfigurable computing and FPGAs · 15% Processor architecture and microarchitecture · 6%
Databases, data mining, and information retrieval
1 paper
Data mining · 50% Graph data management · 50%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
high-level synthesis
0.332011
Pattern-Mining for Behavioral Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Pattern-based behavior synthesis for FPGA resource reduction · FPGA 2008
Behavior and communication co-optimization for systems with sequential communication media · DAC 2006
Electronic design automation › high-level synthesis › behavioral transformation
behavioral synthesis
0.222011
Pattern-Mining for Behavioral Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Pattern-based behavior synthesis for FPGA resource reduction · FPGA 2008
Reconfigurable computing and FPGAs
FPGA accelerator
0.112010
Accelerating Monte Carlo based SSTA using FPGA · FPGA 2010
Electronic design automation › timing analysis
static timing analysis
0.112010
Accelerating Monte Carlo based SSTA using FPGA · FPGA 2010
Electronic design automation › timing analysis › statistical timing analysis
statistical static timing analysis
0.112010
Accelerating Monte Carlo based SSTA using FPGA · FPGA 2010
Reconfigurable computing and FPGAs
FPGA resource optimization
0.112008
Pattern-based behavior synthesis for FPGA resource reduction · FPGA 2008
Electronic design automation › high-level synthesis
resource binding
0.112008
Pattern-based behavior synthesis for FPGA resource reduction · FPGA 2008
Processor architecture and microarchitecture
chip multiprocessor
0.112007
Synthesis of an application-specific soft multiprocessor system · FPGA 2007
Electronic design automation › system-level design
system synthesis
0.112007
Synthesis of an application-specific soft multiprocessor system · FPGA 2007
Electronic design automation › system-level design
communication synthesis
0.112006
Behavior and communication co-optimization for systems with sequential communication media · DAC 2006
Graph data management › graph similarity
graph edit distance
0.012011
Pattern-Mining for Behavioral Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Data mining
pattern mining
0.012011
Pattern-Mining for Behavioral Synthesis · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2011
Electronic design automation
hardware verification and test
0.012010
Accelerating Monte Carlo based SSTA using FPGA · FPGA 2010
Electronic design automation › high-level synthesis
datapath generation
0.012008
Pattern-based behavior synthesis for FPGA resource reduction · FPGA 2008
Electronic design automation
logic synthesis
0.012007
Synthesis of an application-specific soft multiprocessor system · FPGA 2007
Electronic design automation › logic synthesis
technology mapping
0.012007
Synthesis of an application-specific soft multiprocessor system · FPGA 2007

Methods — techniques the papers use, named apart from their topics

graph edit distance · 0.3feature-based filtering · 0.2pattern matching · 0.1mathematical programming · 0.1locality-sensitive hashing · 0.1packing · 0.1labeling · 0.1clustering · 0.1SCOOP algorithm · 0.1
YearPublicationVenuePosition
2011 Pattern-Mining for Behavioral Synthesis
abstract
Pattern-based synthesis has drawn wide interest from researchers who tried to utilize the regularity in applications for design optimizations. In this letter, we present a general pattern-based behavior synthesis framework which can efficiently extract similar structures in programs. Our approach is very scalable in benefit of advanced pruning techniques. The similarity of structures is captured by a mismatch-tolerant metric: the graph edit distance. The graph edit distance can naturally capture different program variations such as bit-width, structure, and port variations. In addition, we further our approach to handle control-intensive applications, and this leads to more opportunities for optimization. Our algorithm uses a feature-based filtering approach for fast pruning, and a graph similarity metric called the generalized edit distance for measuring variations in control-data flow graphs. Furthermore, we apply our pattern-based synthesis system to the resource optimization problem in behavioral synthesis. Considering knowledge of discovered patterns, the resource binding step can intelligently generate the data-path to reduce interconnect costs. Experiments show that our approach can, on average, reduce the total area by about 20% with 7% latency overhead with our pattern techniques on the Xilinx Virtex-4 field-programmable gate arrays, compared to the traditional behavioral synthesis flow.
Jason Cong, Hui Huang 0001, Wei Jiang 0035
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2011 Automatic memory partitioning and scheduling for throughput and power optimization
abstract
Memory bottleneck has become a limiting factor in satisfying the explosive demands on performance and cost in modern embedded system design. Selected computation kernels for acceleration are usually captured by nest loops, which are optimized by state-of-the-art techniques like loop tiling and loop pipelining. However, memory bandwidth bottlenecks prevent designs from reaching optimal throughput with respect to available parallelism. In this paper we present an automatic memory partitioning technique which can efficiently improve throughput and reduce energy consumption of pipelined loop kernels for given throughput constraints and platform requirements. Also, our proposed algorithm can handle general array access beyond affine array references. Our partition scheme consists of two steps. The first step considers cycle accurate scheduling information to meet the hard constraints on memory bandwidth requirements specifically for synchronized hardware designs. An ILP formulation is proposed to solve the memory partitioning and scheduling problem optimally for small designs, followed by a heuristic algorithm which is more scalable and equally effective for solving large scale problems. Experimental results show an average 6× throughput improvement on a set of real-world designs with moderate area increase (about 45% on average), given that less resource sharing opportunities exist with higher throughput in optimized designs. The second step further partitions the memory banks for reducing the dynamic power consumption of the final design. In contrast to previous approaches, our technique can statically compute memory access frequencies in polynomial time with little or no profiling. Experimental results show about 30% power reduction on the same set of benchmarks.
Jason Cong, Wei Jiang 0035, Bin Liu 0006, Yi Zou 0001
ACM Trans. Design Autom. Electr. Syst.2
2010 A generalized control-flow-aware pattern recognition algorithm for behavioral synthesis
abstract
Pattern recognition has many applications in design automation. A generalized pattern recognition algorithm is presented in this paper which can efficiently extract similar patterns in programs. Compared to previous pattern-based techniques, our approach overcomes their limitation in handling control-flow-aware patterns, and leads to more opportunities for optimization. Our algorithm uses a feature-based filtering approach for fast pruning, and an elegant graph similarity metric called the generalized edit distance for measuring variations in CDFGs. Furthermore, our pattern recognition algorithm is applied to solve the area optimization problem in behavioral synthesis. Our experimental results show up to a 40% area reduction on a set of real-world benchmarks with a moderate 9% latency overhead, compared to synthesis results without pattern extractions; and up to a 30% area reduction, compared to the results using only data-flow patterns.
Jason Cong, Hui Huang 0001, Wei Jiang 0035
DATE3
2010 Accelerating Monte Carlo based SSTA using FPGA
abstract
Monte Carlo based SSTA serves as the golden standard against alternative SSTA algorithms, but it is seldom used in practice due to its high computation time. In this paper, we accelerate Monte Carlo based SSTA using the FPGA platform. A simple dataflow pipeline technique will not work well due to the excessive usage of FPGA logic slices. We leverage the recently proposed pattern matching method to identify common circuit structures, and further use a mathematical programming based formulation to explore the trade-off between performance and logic slices consumption. The proposed design provides two orders of magnitude speedup compared to the CPU-based implementation.
Jason Cong, Karthik Gururaj, Wei Jiang 0035, Bin Liu 0006, Kirill Minkovich, Yi Zou 0001
FPGA3
2009 Automatic memory partitioning and scheduling for throughput and power optimization
abstract
Hardware acceleration is crucial in modern embedded system design to meet the explosive demands on performance and cost. Selected computation kernels for acceleration are usually captured by nest loops, which are optimized by state-of-the-art techniques like loop tiling and loop pipelining. However, memory bandwidth bottlenecks prevent designs to reach optimal throughput with respect to available parallelism. In this paper we present an automatic memory partitioning technique which can efficiently improve throughput and reduce energy consumption of pipelined loop kernels for given throughput constraints and platform requirement. Our partition scheme consists of two steps, the first step considers cycle accurate scheduling information to meet the hard constraints on memory bandwidth requirements specifically for synchronized hardware designs. Experimental results show an average 6X throughput improvement on a set of real world designs with moderate area increase (about 45% on average), given that less resource sharing opportunities exist with higher throughput in optimized designs. The second step further partitions the memory banks for reducing the dynamic power consumption of the final design. In contrast with previous approaches, our technique can statically compute memory access frequencies in polynomial time with little to none profiling. Experimental results show about 30% power reduction on the same set of benchmarks.
Jason Cong, Wei Jiang 0035, Bin Liu 0006, Yi Zou 0001
ICCAD2
2009 Synthesis Algorithm for Application-Specific Homogeneous Processor Networks
abstract
The application specific multiprocessor system-on-a-chip is a promising design alternative because of its high degree of flexibility, short development time, and potentially high performance attributed to application specific optimizations. However, designing an optimal application specific multiprocessor system is still challenging because there are a number of important metrics, such as throughput, latency, and resource usage, which need to be explored and optimized. This paper addresses the problem of synthesizing an application-specific multiprocessor system for stream-oriented embedded applications to minimize system latency under the throughput constraint. We employ a novel framework for this problem, similar to that of technology mapping in the logic synthesis domain, and develop a set of efficient algorithms, including labeling and clustering for efficient generation of the multiprocessor architecture with application specific optimized latency. Specifically, the result of our algorithm is latency optimal for directed acyclic task graphs. Application of our approach to the Motion JPEG example on Xilinx's Virtex II Pro platform FPGA shows interesting design tradeoffs.
Jason Cong, Karthik Gururaj, Guoling Han, Wei Jiang 0035
IEEE Trans. Very Large Scale Integr. Syst.4
2008 Scheduling with integer time budgeting for low-power optimization
abstract
In this paper we present a mathematical programming formulation of the integer time budgeting problem for directed acyclic graphs. In particular, we formally prove that our constraint matrix has a special property that enables a polynomial-time algorithm to solve the problem optimally with a guaranteed integral solution. Our theory can be directly applied to solving a scheduling problem in behavioral synthesis with the objective of minimizing the system power consumption. Given a set of scheduling constraints and a collection of convex power-delay tradeoff curves for each type of operation, our scheduler can intelligently schedule the operations to appropriate clock cycles and simultaneously select the module implementations that lead to low-power solutions. Experiments demonstrate that our proposed technique can produce near-optimal results (within 6% of the optimum by the ILP formulation), with 40x+ speedup.
Wei Jiang 0035, Zhiru Zhang, Miodrag Potkonjak, Jason Cong
ASP-DAC1
2008 Pattern-based behavior synthesis for FPGA resource reduction
abstract
Pattern-based synthesis has drawn wide interest from researchers who tried to utilize the regularity in applications for design optimizations. In this paper we present a general pattern-based behavior synthesis framework which can efficiently extract similar structures in programs. Our approach is very scalable in benefit of advanced pruning techniques that include locality sensitive hashing and characteristic vectors. The similarity of structures is captured by a mismatch-tolerant metric: graph edit distance. The edit distance between two graphs is the minimum number of vertex/edge insertion, deletion, substitution operations to transform one graph into the other. Graph edit distance can naturally handle various program variations such as bit-width variations, structure variations and port variations. In addition, we apply our pattern-based synthesis system to FPGA resource optimization with the observation that multiplexors are particularly expensive on FPGA platforms. Considering knowledge of discovered patterns, the resource binding step can intelligently generate the data-path to reduce interconnect costs. Experiments show our approach can, on average, reduce the total area by about 20% with 7% latency overhead on the Xilinx Virtex-4 FPGAs, compared to the traditional behavior synthesis flow
Jason Cong, Wei Jiang 0035
FPGA2
2007 Synthesis of an application-specific soft multiprocessor system
abstract
The application-specific multiprocessor System-on-a-Chip is a promising design alternative because of its high degree of flexibility, short development time, and potentially high performance attributed to application-specific optimizations. However, designing an optimal application-specific multiprocessor system is still challenging because there are a number of important metrics, such as throughput, latency, and resource usage, that need to be explored and optimized. This paper addresses the problem of synthesizing the application-specific multiprocessor system to minimize latency and resource usage under the throughput constraint. We employ a novel framework for this problem, similar to that of technology mapping in the logic synthesis domain, and develop a set of efficient algorithms, including labeling, clustering and packing, for efficient generation of the multiprocessor architecture with application-specific optimized latency and resources. Specifically, the result of our algorithm is latency-optimal for directed acyclic task graphs. Application of our approach to the Motion JPEG example on Xilinx’s Virtex II Pro platform FPGA shows interesting design tradeoffs.
Jason Cong, Guoling Han, Wei Jiang 0035
FPGA3
2006 Behavior and communication co-optimization for systems with sequential communication media
abstract
In this paper we propose a new communication synthesis approach targeting systems with sequential communication media (SCM). Since SCMs require that the reading sequence and writing sequence must have the same order, different transmission orders may have a dramatic impact on the final performance. However, the problem of determining the best possible communication order for SCMs is not adequately addressed by prior work. The goal of our work is to consider behaviors in communication synthesis for SCM, detect appropriate transmission order to optimize latency, automatically transform the behavior descriptions, and automatically generate driver routines and glue logics to access physical channels. Our algorithm, named SCOOP, successfully achieves these goals by behavior and communication co-optimization. Compared to the results without optimization, we can achieve an average 20% improvement in total latency on a set of real-life benchmarks.
Jason Cong, Yiping Fan, Guoling Han, Wei Jiang 0035, Zhiru Zhang
DAC4
2006 Platform-based resource binding using a distributed register-file microarchitecture
abstract
Behavior synthesis and optimization beyond the register transfer level require an efficient utilization of the underlying platform features. This paper presents a platform-based resource-binding approach using a distributed register-file microarchitecture (DRFM) that makes efficient use of distributed embedded memory blocks as register files in modern FPGAs. A DRFM contains multiple islands, each having a local register file, a functional unit pool and data-routing logic. Compared with the traditional discrete-register counterpart, a DRFM allows use of the platform-featured on-chip memory or register-file IP blocks to implement its local register files, and this results in substantial saving of multiplexing logic and global interconnects. DRFM provides a useful architectural template and a direct optimization objective for minimizing inter-island connections for synthesis algorithms. Based on DRFM, we propose a novel binding algorithm focusing on the minimization of the inter-island connections. By applying our approach, significant reductions on multiplexors and global-interconnections are observed. On the Xilinx Virtex II FPGA platform, our experimental results show a 2X logic area reduction and a 7.8% performance improvement, compared with the traditional discrete-register-based approach.
Jason Cong, Yiping Fan, Wei Jiang 0035
ICCAD3