EDBT 2026 Demo / reviewers in the wild / expert
Yuan Wen
dblp:77/3361
· DBLP profile ↗
26ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 3 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Threads to Tiles: T2T, a Compiler for CUDA-to-NPU Translation via 2D VectorizationabstractCUDA’s programming model, exposing massive parallelism via fine-grained scalar threads, has become the de facto standard for GPU computing. Concurrently, NPUs are emerging as highly efficient accelerators, but their architecture is fundamentally different, relying on coarse-grained, explicit 2-D tile-based instructions. This creates a critical challenge: bridging the semantic gap "From Threads to Tiles". A direct translation is infeasible, as it requires lifting the implicit parallelism of CUDA’s scalar model into the explicit, multi-dimensional vector space of NPUs, a problem we formalize as a lifting challenge.This paper introduces T2T, a compiler framework that automates this "Threads to Tiles" translation via the 2-D Vectorization technique. T2T first transforms a CUDA kernel’s implicit SIMT parallelism into a structured, explicit loop nest via our Unified Parallelism Abstraction (UPA), making the parallelism analyzable. From this representation, T2T’s core vectorization engine systematically selects optimal pairs of loops and maps them onto the NPU’s 2-D tile instructions to maximize hardware utilization. To ensure correctness and handle performance-critical CUDA features, a final set of semantics-preserving optimizations is applied, including efficient control-flow management and vectorization of warp-level intrinsics.We implement T2T based on Polygeist and evaluate representative NPU architectures. On a diverse set of benchmarks, kernels translated by T2T achieve up to 73% of native CUDA performance on an A100 GPU and outperform baseline translation approaches by up to 6.9×. Our work demonstrates that a systematic, compiler-driven approach to 2-D vectorization is a principled and high-performance path for porting the rich CUDA ecosystem to the evolving landscape of NPU accelerators. Shuaijiang Li, Ying Liu 0055, Shuoming Zhang, Yijin Li, Yangyu Zhang, Runyu Zhou, Xiyu Shi, Chunwei Xia, Yuan Wen, Xiaobing Feng 0002, Huimin Cui |
CGO | 12 |
| 2026 | Symbiotic MLLM Serving: Dynamically Balancing Parallelism Across GPUs and Resources Within GPUs
Yangyu Zhang, Zhaolin Duan, Shuoming Zhang, Shuaijiang Li, Donglin Yu, Yuan Wen, Chunwei Xia, Xiyu Shi, Huimin Cui |
ISCA | 10 |
| 2026 | ParDiff: Efficiently Parallelizing Reverse-Mode Automatic Differentiation with Direct IndexingabstractAutomatic Differentiation (AD) is a technique that computes the derivatives of numerical programs by systematically applying the chain rule, playing a critical role in domains such as machine learning, simulation, and control systems. However, parallelizing differentiated programs remains a significant challenge due to the conflict between tapes (a data structure for intermediate variable storage) and summations: the differentiation process inherently introduces inter-thread summation patterns, which require prohibitively expensive atomic operations; and traditional tape designs tightly couple data retrieval with the program’s control flow, preventing code restructuring needed to eliminate these costly dependencies. Shuhong Huang, Shizhi Tang, Yuan Wen, Huanqi Cao, Ruibai Tang, Yidong Chen 0003, Jiping Yu, Jidong Zhai |
PPoPP | 3 |
| 2025 | SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMsabstractRecent multimodal large language models (MLLMs) marry modality-specific
vision or audio encoders with a shared text decoder. While the encoder is compute-
intensive but memory-light, the decoder is the opposite, yet state-of-the-art serving
stacks still time-multiplex these complementary kernels, idling SMs or HBM in
turn. We introduce SpaceServe, a serving system that space-multiplexes MLLMs:
it decouples all modality encoders from the decoder, and co-locates them on the
same GPU using fine-grained SM partitioning available in modern runtimes. A
cost-model-guided Space-Inference Scheduler (SIS) dynamically assigns SM slices,
while a Time-Windowed Shortest-Remaining-First (TWSRFT) policy batches en-
coder requests to minimise completion latency and smooth decoder arrivals.
Evaluation shows that SpaceServe reduces time-per-output-token by 4.81×
on average and up to 28.9× on Nvidia A100 GPUs. SpaceServe is available at
https://github.com/gofreelee/SpaceServe Shuoming Zhang, Xiyu Shi, Yangyu Zhang, Shuaijiang Li, Donglin Yu, Zheming Yang, Yuan Wen, Huimin Cui |
NeurIPS | 10 |
| 2025 | O.O: Optimized one-die placement for face-to-face bonded 3D ICsabstractAs the miniaturization of integrated circuits (ICs) reaches its physical limits, the industry is entering a “more-than-Moore” era, demanding new Electronic Design Automation (EDA) tools. Existing TSV-based 3D placers focus on minimizing cuts while burgeoning F2F-bonded ICs feature dense interconnection between two planar die. Towards this novel structure, we proposed an integrated adaptation methodology upon mature one-die-based placement strategies. First, we instructively utilized a one-die placer to provide a statistical looking-ahead net diagnosis. The netlist henceforth shall be coarsened topologically and geometrically using a multi-level framework. Our multi-objective gain formulation guides a level-by-level refinement of the partition. This formulation considers factors like cut expectation, heterogeneous row heights, and balanced cell distribution, enabling efficient incremental calculations at each level. Given the partition, we synchronized the behavior of analytical planar placers by balancing the density and wirelength objective function among asymmetric layers. Finally, the result will be further improved by heuristic detail placement of bonding terminals and a post-place partition adjustment. Experimental results demonstrate that our fine-grained fusion of partitioning and placement techniques are competitive compared with the top three winners of the 2022 ICCAD CAD Contest, achieving the best normalized average wirelength with competitive runtime under various 3D architectural constraints . Xingyu Tong 0001, Yuhao Ren, Zhijie Cai, Yuan Wen, Zhifeng Lin, Jianli Chen |
Integr. | 6 |
| 2025 | Dynamic Power Management Through Multi-agent Deep Reinforcement Learning for Heterogeneous SystemsabstractPower management and optimization play a significant role in modern computer systems, from battery-powered devices to servers running in data centers. Existing approaches for power capping fail to meet the requirements presented by dynamic workloads, and the situation becomes even more severe, given the divergent energy efficiency of workloads on heterogeneous hardware platforms. Adaptively optimizing energy consumption for dynamic workloads presents a great challenge to heterogeneous systems. To tackle this challenge, we present a machine learning based method to improve system-level power efficiency. We employ multi-agent deep reinforcement learning (MADRL) to automatically explore the relationship between long-term performance and the power budget for workloads of different types on classic CPU-GPU heterogeneous platforms. Our framework equips each device with an agent, enabling decentralized control over its power budget while maintaining centralized coordination to maximize the running time of applications within a power cap. We evaluate our approach against state-of-the-art methods on CPU-GPU platforms. Experimental results show that our method improves performance by an average of 8.5%. Additionally, our method is significantly more stable compared to the state-of-the-art heuristic approach. Yiming Wang 0010, Weizhe Zhang, Meng Hao 0002, Weizhi Kong, Yuan Wen |
ACM Trans. Archit. Code Optim. | 5 |
| 2024 | O.O: Optimized One-die Placement for Face-to-face Bonded 3D ICsabstractThe expansion of the IC dimension is ushering in a more-than-Moore era, necessitating corresponding EDA tools. Existing TSV-based 3D placers focus on minimizing cuts, while burgeoning F2F-bonded ICs features dense interconnection between two planar die. Towards this novel structure, we proposed an integrated adaptation methodology upon mature one-die-based placement strategies. First, we instructively utilized a one-die placer to provide a statistical looking-ahead net diagnosis. The netlist henceforth shall be coarsened topologically and geometrically with a multi-level framework. Level by level, the partition will be refined according to a multi-objective gain formulation, including cut expectation, heterogeneous row height, and balanced cell distribution. Given the partition, we synchronized the behavior of analytical planar placers by balancing the density and wirelength objective function among asymmetric layers. Finally, the result will be further improved by heuristic bonding terminals’ detail placement and a post-place partition adjustment. Compared to the top three winners of the 2022 CAD Contest at ICCAD, experiment results show that our fine-grained fusion upon partitioning and placement gets the best normalized average wirelength with a fairly reasonable runtime under all 3D architectural constraints. Xingyu Tong 0001, Zhijie Cai, Yuan Wen, Zhifeng Lin, Jianli Chen |
ASPDAC | 5 |
| 2024 | Effective Analytical Placement for Advanced Hybrid-Row-Height Circuit DesignsabstractRecently, hybrid-row-height designs have been introduced to achieve performance and area co-optimization in advanced nodes. Hybrid-row-height designs incur challenging issues to layout due to the heterogeneous cell and row structures. In this paper, we present an effective algorithm to address the hybrid-row-height placement problem in two major stages: (1) global placement, and (2) legalization. Inspired by the multi-channel processing method in convolutional neural networks (CNN), we use the feature extraction technique to equivalently transform the hybrid-row-height global placement problem into two sub-problems that can be solved effectively. We propose a multi-layer nonlinear framework with alignment guidance and a self-adaptive parameter adjustment scheme, which can obtain a high-quality solution to the hybrid-row-height global placement problem. In the legalization stage, we formulate the hybrid-row-height legalization problem into a convex quadratic programming (QP) problem, then apply the robust modulus-based matrix splitting iteration method (RMMSIM) to solve the QP efficiently. After RMMSIM-based global legalization, Tetris-like allocation is used to resolve remaining physical violations. Compared with the state-of-the-art work, experiments on the 2015 ISPD Contest benchmarks show that our algorithm can achieve 7%; shorter final total wirelength and $2.23 \times $ speedup. Yuan Wen, Benchao Zhu, Zhifeng Lin, Jianli Chen |
ASPDAC | 1 |
| 2024 | Optimizing Deep Learning Inference via Global Analysis and Tensor ExpressionsabstractOptimizing deep neural network (DNN) execution is important but becomes increasingly difficult as DNN complexity grows. Existing DNN compilers cannot effectively exploit optimization opportunities across operator boundaries, leaving room for improvement. To address this challenge, we present Souffle, an open-source compiler that optimizes DNN inference across operator boundaries. Souffle creates a global tensor dependency graph using tensor expressions, traces data flow and tensor information, and partitions the computation graph into subprograms based on dataflow analysis and resource constraints. Within a subprogram, Souffle performs local optimization via semantic-preserving transformations, finds an optimized program schedule, and improves instruction-level parallelism and data reuse. We evaluated Souffle using six representative DNN models on an NVIDIA A100 GPU. Experimental results show that Souffle consistently outperforms six state-of-the-art DNN optimizers by delivering a geometric mean speedup of up to 3.7× over TensorRT and 7.8× over Tensorflow XLA. Chunwei Xia, Qianqi Sun, Zheng Wang 0001, Yuan Wen, Xiaobing Feng 0002, Huimin Cui |
ASPLOS (1) | 5 |
| 2023 | Sgap: towards efficient sparse tensor algebra compilation for GPU
Genghan Zhang, Yuetong Zhao, Yanting Tao, Zhongming Yu, Guohao Dai 0001, Sitao Huang, Yuan Wen, Pavlos Petoumenos, Yu Wang 0002 |
CCF Trans. High Perform. Comput. | 7 |
| 2023 | Federated clustering for recognizing driving styles from private trajectories
Lin Lu 0002, Yuan Wen, Jinxiong Zhu, Shengwu Xiong 0001 |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | Performance of the Semi-Empirical Precipitable Water Vapor Retrieval Algorithm Developed for Polarized Scanning Atmospheric Corrector (PSAC) in the Presence of Sensor DecayabstractPolarized Scanning Atmospheric Corrector (PSAC) is an optical sensor onboard HuanjingJianzai-2 (HJ-2) A/B satellites. One of its missions is to monitor precipitable water vapor (PWV) by using its near-infrared (NIR) channels. Since the accuracy of the commonly used NIR PWV retrieval algorithm developed based on radiative transfer model (RTM) would be significantly affected by radiometric decay of sensors, and the recalibration of decayed sensors is a complex process, it is interesting and necessary to find a robust PWV retrieval algorithm that is not affected by sensor decay. At present, a semi-empirical algorithm constructed based on the matching results between ground-based PWV data and the actual PSAC observations has been used for the PWV retrieval of PSAC. Since the systematic calibration error of PSAC is considered in constructing the algorithm, it should be able to remove the negative effects of sensor decay on PWV retrieval results. Because the above inference has not been confirmed quantitatively, it is necessary to evaluate the accuracy of the algorithm in the presence of sensor decay. The evaluation results based on simulated data show that the accuracy of the semi-empirical algorithm does not change regardless of the presence or absence of radiometric decay in PSAC. Moreover, the algorithm is used for PWV retrieval of MODIS to test its effectiveness. Compared with the official PWV data developed based on RTM, the MODIS PWV data developed by using the semi-empirical algorithm are reduced by more than 50% in both absolute and relative errors. Yanqing Xie, Yuan Wen, Yunduan Li, Weizhen Hou, Zhenhai Liu, Xuefeng Lei, Zhongzheng Hu, Zhengqiang Li |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Analytical Placement with 3D Poisson's Equation and ADMM-based Optimization for Large-scale 2.5D Heterogeneous FPGAsabstractAs design complexity keeps increasing, the 2.5D field-programmable gate array (FPGA) with large logic capacity has become popular in modern circuit applications. A 2.5D FPGA consists of multiple dies connected through super long lines (SLLs) on an interposer. Each die contains heterogeneous logic blocks and ASIC-like clocking architectures to achieve better skew and timing. Existing works consider these problems separately and thus may lead to serious timing issues or routing failure. This article presents an analytical placement algorithm for the 2.5D FPGA to simultaneously minimize the number of inter-die SLL signals and intra-die clocking violations. Using a lifting dimension technique, we first formulate the 2.5D global placement problem as a three-dimensional continuous and differential minimization problem, where the SLL-aware block distribution is modeled by 3D Poisson’s equation and directly solved to obtain an analytical solution. Then, we further reformulate the minimization problem as a separable optimization problem with linear constraints. Based on the proximal alternating direction method of multipliers optimization method, we efficiently optimize the separable subproblems one by one in an alternating fashion. Finally, clock-aware legalization and detailed placement are applied to legalize and improve our placement results. Compared with the state-of-the-art works, experimental results show that our algorithm can resolve all clocking constraints and reduce the number of SLL crossing signals by 36.9% with similar wirelength in a comparable running time. Xingyu Tong 0001, Yuan Wen, Jianli Chen, Jun Yu 0010, Wenxing Zhu, Yao-Wen Chang |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2020 | An Improved Heterogeneous Dynamic List Schedule Algorithm
Wei Hu 0001, Yu Gan 0004, Yuan Wen, Xiangyu Lv, Yonghao Wang, Meikang Qiu |
ICA3PP (1) | 3 |
| 2020 | Design of a Convolutional Neural Network Instruction Set Based on RISC-V and Its Microarchitecture Implementation
Qiang Jiao, Wei Hu 0001, Yuan Wen, Yong Dong, Zhenhao Li 0004, Yu Gan 0004 |
ICA3PP (2) | 3 |
| 2020 | TASO: Time and Space Optimization for Memory-Constrained DNN InferenceabstractConvolutional neural networks (CNNs) are used in many embedded applications, from industrial robotics and automation systems to biometric identification on mobile devices. State-of-the-art classification is typically achieved by large networks, which are prohibitively expensive to run on mobile and embedded devices with tightly constrained memory and energy budgets. We propose an approach for ahead-of-time domain specific optimization of CNN models, based on an integer linear programming (ILP) for selecting primitive operations to implement convolutional layers. We optimize the trade-off between execution time and memory consumption by: 1) attempting to minimize execution time across the whole network by selecting data layouts and primitive operations to implement each layer; and 2) allocating an appropriate work space that reflects the upper bound of memory footprint per layer. These two optimization strategies can be used to run any CNN on any platform with a C compiler. Our evaluation with a range of popular ImageNet neural architectures (GoogleNet, AlexNet, VGG, ResNetand SqueezeNet) on the ARM Cortex-A15 yields speedups of 8× compared to a greedy algorithm based primitive selection, reduces memory requirement by 2.2× while sacrificing only 15% of inference time compared to a solver that considers inference time only. In addition, our optimization approach exposes a range of optimal points for different configurations across the Pareto frontier of memory and latency trade-off, which can be used under arbitrary system constraints. Yuan Wen, Andrew Anderson 0001, Valentin Radu, Michael F. P. O'Boyle, David Gregg |
SBAC-PAD | 1 |
| 2020 | Research on Plant Disease Recognition Based on Deep Complementary Feature Classification NetworkabstractTraditional convolutional neural network classification models often only focus on the most distinguishing feature regions of the image and ignore the weaker feature regions. However, the image position distribution of plant diseases is very uneven. If we use convolutional neural network for plant disease recognition, there will be insufficient feature response, which will cause recognition errors. Aiming at such problems, we have designed a deep complementary feature classification network. First, the network uses DeepLabv3+ and Conditional Random Field (CRF) to generate disease part detection frames in a weakly supervised manner and combines semantic segmentation to extract disease object instances. Then we designed Complementary Feature Part Generation Models. Finally, it uses a bidirectional Gated Recurrent Unit (Bi-GRU) to perform the classification and recognition of the complementary features described above. We performed experiments on the PlantVillage dataset. The experimental results show that the proposed network recognition accuracy is 99.21%, which is 4.2% higher than the baseline model xception-65 used. We also performed experiments on the grape disease data set that we created. The accuracy of the proposed network recognition is 93.46%, which is 7.2% higher than the baseline model xception-65. In addition, compared with the better algorithms for plant disease identification in recent years, the accuracy performance has also been improved. JiaYou Chen, Hong Guo 0005, Wei Hu 0001, Juanjuan He, Yonghao Wang, Yuan Wen |
SMC | 6 |
| 2020 | A Improved List Heuristic Scheduling Algorithm for Heterogeneous Computing SystemsabstractWhen the traditional heterogeneous multi-core scheduling algorithm performs tasks with high resource density, a large amount of idle time often occurs on the processor core. Therefore, based on the environment of heterogeneous multi-core processors, this paper studies the static heuristic table scheduling algorithm, and proposes an optimization approach for the problem of single priority assignment and too simple task assignment. We design optimization in the static heuristic scheduling algorithm list generation phase and task allocation phase, and propose a hybrid task allocation method with three strategies to improve the standby time utilization of processor core. Then, DVFS technology is used to optimize the scheduling results, so that the task can run with lower energy consumption without increasing makespan. Finally, the new algorithm is compared with three traditional scheduling algorithms through design experiments, and it is proved that the new algorithm has better performance when executing more tasks. Wei Hu 0001, Yu Gan 0004, Xiangyu Lv, Yonghao Wang, Yuan Wen |
SMC | 5 |
| 2020 | Generative Adversarial Training for Weakly Supervised Nuclei Instance SegmentationabstractNuclei segmentation occupies an important position in medical image analysis, which helps to predict and diagnose diseases. With the further research of deep learning, the task of nuclei segmentation has been automated. However, most existing methods require a great deal of manually marked full masks for training, which is time-consuming and labor-intensive, and can only be done by professional personnel. For the purpose of reducing the cost of labeling, we propose a weakly supervised method using generative adversarial training for segmentation of nucleus. In the case of no boundary, but only the centroid of the nucleus, the proposed method segmented the nucleus region with blurred boundaries. We first use the generative adversarial network(GAN) to generate the likelihood map of the nuclear centroid, then use Guided Backpropagation to visualize the pixels that contributes to the detection of the centroid of each nucleus, and finally obtain the segmentation mask of the nucleus by graph-cut. In addition, for the purpose of training the network better, we performed stain normalization on each pathological image. We have verified the proposed method on a multi-organ nuclei dataset. The final experiment results show that our advanced method achieves better segmentation performance than other weakly supervised methods, and can even reach the level of full supervision. Wei Hu 0001, Huanhuan Sheng, Jing Wu 0019, Yonghao Wang, Yuan Wen |
SMC | 7 |
| 2019 | POSTER: Space and Time Optimal DNN Primitive Selection with Integer Linear ProgrammingabstractConvolutional neural networks (CNNs) are used in many applications, from industrial robotics to biometric identification on mobile devices. But they can be too resource-hungry for mobile and embedded devices with tightly constrained memory and energy budgets. We propose an ahead-of-time primitive selection for CNNs, based on integer linear programming (ILP). Under a tight memory budget, our ILP solver selects the optimal primitive for each layer such that the entire network is optimized for execution time subject to a memory budget, or vice versa. Our method yields significant speedup and memory reduction compared to existing methods. Yuan Wen, Andrew Anderson 0001, Valentin Radu, Michael F. P. O'Boyle, David Gregg |
PACT | 1 |
| 2014 | Smart multi-task scheduling for OpenCL programs on CPU/GPU heterogeneous platformsabstractHeterogeneous systems consisting of multiple CPUs and GPUs are increasingly attractive as platforms for high performance computing. Such platforms are usually programmed using OpenCL which provides program portability by allowing the same program to execute on different types of device. As such systems become more mainstream, they will move from application dedicated devices to platforms that need to support multiple concurrent user applications. Here there is a need to determine when and where to map different applications so as to best utilize the available heterogeneous hardware resources. In this paper, we present an efficient OpenCL task scheduling scheme which schedules multiple kernels from multiple programs on CPU/GPU heterogeneous platforms. It does this by determining at runtime which kernels are likely to best utilize a device. We show that speedup is a good scheduling priority function and develop a novel model that predicts a kernel's speedup based on its static code structure. Our scheduler uses this prediction and runtime input data size to prioritize and schedule tasks. This technique is applied to a large set of concurrent OpenCL kernels. We evaluated our approach for system throughput and average turn-around time against competitive techniques on two different platforms: a Core i7/Nvidia GTX590 and a Core i7/AMD Tahiti 7970 platforms. For system throughput, we achieve, on average, a 1.21x and 1.25x improvement over the best competitors on the NVIDIA and AMD platforms respectively. Our approach reduces the turnaround time, on average, by at least 1.5x and 1.2x on the NVIDIA and AMD platforms respectively, when compared to alternative approaches. Yuan Wen, Zheng Wang 0001, Michael F. P. O'Boyle |
HiPC | 1 |
| 2014 | QoE-based bandwidth allocation with SDN in FTTH networksabstractIn the High Speed Internet (HSI) service of the Fiber-To-The-Home (FTTH) networks, there are increasingly various applications, such as browsing, video streaming, large downloads and online games. They are competing for the fixed bandwidth on a best-effort basis, and finally resulting in the network congestion and poor quality of experience (QoE). Users want to improve the quality of certain applications. However, today's network service controller (e.g. Broadband Remote Access Server, BRAS) lacks mechanisms to meet the users' desire to enhance the QoE of specific applications. Moreover, BRAS still lacks mechanisms to allocate the bandwidth resources properly for users' different applications according to their “sweet points”. “Sweet points” is a specific bandwidth value. The QoE gets worse quickly when the bandwidth is smaller than the “sweet point”, and keeps the same approximately when the bandwidth is larger than the “sweet point”. In this paper, we proposed a novel BRAS architecture using Software-Defined Networking (SDN) technology, which can improve the user's QoE by adjusting the bandwidth of a specific application to its “sweet point” according to their requirements. To demonstrate the feasibility of our proposed novel BRAS, we built a prototype using SDN to help the user to adjust the bandwidth for the specific application and improve the users' QoE. The experimental results show that users could enhance the QoE of specific applications according to users' preference. Wei Guo 0003, Yuan Wen, Chengjun Li, Weisheng Hu |
NOMS | 4 |
| 2009 | HMMer acceleration using systolic array based reconfigurable architectureabstractHMMer is a widely-used bioinformatics software package that uses profile Hidden Markov Models (HMMs) to model the primary structure consensus of a family of protein or nucleic acid sequences. However, with the rapid growth of both sequence and model databases, it is more and more time-consuming to run HMMer on traditional computer architecture. With the development of modern field programmable gate array (FPGA) technology, applications can be accelerated using CPU-FPGA cooperative system by mapping computational-intensive work onto FPGA. In this paper, the computation kernel of HMMer, P7Viterbi, is selected to be accelerated by FPGA. After carefully data dependency analysis, we proposed a systolic array based reconfigurable architecture to exploit both inter-module and intra-module parallelism. There is an infrequent feedback loop in P7Viterbi to update the value of beginning state (B state), which limits further parallelization. Previous work either ignored the feedback loop or serialized the process, leading to loss of either precision or efficiency. Our proposed architecture can exploit maximum parallelism without loss of precision. The proposed architecture speculatively runs with fully parallelism assuming that the feedback loop does not take place. If the rare feedback case actually occurs, a rollback mechanism is used to ensure correctness. Results show that by using Xilinx Virtex-5 110T FPGA, the proposed architecture can achieve about a 56.8 times speedup compared with that of Intel Core2 Duo 2.33GHz CPU. Yanteng Sun, Peng Li 0031, Guochang Gu, Yuan Wen |
FPGA | 4 |
| 2009 | Accelerating HMMer on FPGAs using systolic array based architectureabstractHMMer is a widely-used bioinformatics software package that uses profile HMMs (Hidden Markov Models) to model the primary structure consensus of a family of protein or nucleic acid sequences. However, with the rapid growth of both sequence and model databases, it is more and more time-consuming to run HMMer on traditional computer architecture. In this paper, the computation kernel of HMMer, P7Viterbi, is selected to be accelerated by FPGA. There is an infrequent feedback loop in P7Viterbi to update the value of beginning state (B state), which limits further parallelization. Previous work either ignored the feedback loop or serialized the process, leading to loss of either precision or efficiency. Our proposed syslolic array based architecture with a parallel data providing unit can exploit maximum parallelism of the full version of P7Viterbi. The proposed architecture speculatively runs with fully parallelism assuming that the feedback loop does not take place. If the rare feedback case actually occurs, a rollback mechanism is used to ensure correctness. Results show that by using Xilinx Virtex-5 110T FPGA, the proposed architecture with 20 PEs can achieve about a 56.8 times speedup compared with that of Intel Core2 Duo 2.33 GHz CPU. Yanteng Sun, Peng Li 0031, Guochang Gu, Yuan Wen |
IPDPS | 4 |
| 2009 | Multi-dimensional Data Visualization using Concentric Coordinates
Jiawan Zhang, Yuan Wen, Quang Vinh Nguyen 0002, Mao Lin Huang, Jiadong Yang |
VINCI | 2 |
| 2005 | Nonlinear least-square solution to flat-top pattern synthesis using arbitrary linear array
Yuan Wen, Woon-Seng Gan, Jun Yang 0004 |
Signal Process. | 1 |