VLDB 2026 Research / reviewers in the wild / expert
Dong Wang 0040
dblp:40/3934-40
· DBLP profile ↗
27ranked-venue papers
6as first author
9since 2021 · last 2027
0000-0002-0068-8824ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 6 first-authorArtificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
8 papers |
Hardware accelerators and domain-specific architectures · 45% Reconfigurable computing and FPGAs · 30% Integrated circuit design · 10% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 22 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Reconfigurable computing and FPGAs
FPGA accelerator |
0.8 | 2 | 2020 | DSP-Efficient Hardware Acceleration of Convolutional Neural Network Inference on FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 ABM-SpConv: A Novel Approach to FPGA-Based Acceleration of Convolutional Neural Network Inference · DAC 2019 |
Reconfigurable computing and FPGAs
coarse-grained reconfigurable architecture |
0.6 | 3 | 2015 | An Energy-Efficient Coarse-Grained Reconfigurable Processing Unit for Multiple-Standard Video Decoding · IEEE Trans. Multim. 2015 A Mixed-Grained Reconfigurable Computing Platform for Multiple-Standard Video Decoding (Abstract Only) · FPGA 2015 Implementation of multi-standard video decoder on a heterogeneous coarse-grained reconfigurable processor · Sci. China Inf. Sci. 2014 |
Machine learning › Efficient and distributed learning › model compression › quantization
mixed-precision quantization |
0.6 | 1 | 2022 | MultiQuant: Training Once for Multi-bit Quantization of Neural Networks · IJCAI 2022 |
Machine learning › Efficient and distributed learning
model compression |
0.6 | 1 | 2022 | MultiQuant: Training Once for Multi-bit Quantization of Neural Networks · IJCAI 2022 |
Machine learning › Efficient and distributed learning › model compression › quantization
multi-bit quantization |
0.6 | 1 | 2022 | MultiQuant: Training Once for Multi-bit Quantization of Neural Networks · IJCAI 2022 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.6 | 1 | 2022 | MultiQuant: Training Once for Multi-bit Quantization of Neural Networks · IJCAI 2022 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
CNN inference accelerator |
0.4 | 1 | 2020 | DSP-Efficient Hardware Acceleration of Convolutional Neural Network Inference on FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
convolution acceleration |
0.4 | 1 | 2020 | DSP-Efficient Hardware Acceleration of Convolutional Neural Network Inference on FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.4 | 1 | 2020 | DSP-Efficient Hardware Acceleration of Convolutional Neural Network Inference on FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
CNN accelerator |
0.4 | 1 | 2019 | ABM-SpConv: A Novel Approach to FPGA-Based Acceleration of Convolutional Neural Network Inference · DAC 2019 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.4 | 1 | 2019 | ABM-SpConv: A Novel Approach to FPGA-Based Acceleration of Convolutional Neural Network Inference · DAC 2019 |
Hardware accelerators and domain-specific architectures › video coding accelerator
video decoding accelerator |
0.3 | 3 | 2015 | A Mixed-Grained Reconfigurable Computing Platform for Multiple-Standard Video Decoding (Abstract Only) · FPGA 2015 An Energy-Efficient Coarse-Grained Reconfigurable Processing Unit for Multiple-Standard Video Decoding · IEEE Trans. Multim. 2015 Implementation of multi-standard video decoder on a heterogeneous coarse-grained reconfigurable processor · Sci. China Inf. Sci. 2014 |
Integrated circuit design › digital circuit design
arithmetic circuit design |
0.3 | 2 | 2014 | Complex Function Approximation Using Two-Dimensional Interpolation · IEEE Trans. Computers 2014 A Radix-16 Combined Complex Division/Square Root Unit with Operand Prescaling · IEEE Trans. Computers 2012 |
Energy-efficient computing › low-power design
low-power processor design |
0.2 | 1 | 2015 | An Energy-Efficient Coarse-Grained Reconfigurable Processing Unit for Multiple-Standard Video Decoding · IEEE Trans. Multim. 2015 |
Embedded and real-time systems
multimedia processing |
0.2 | 1 | 2014 | Implementation of multi-standard video decoder on a heterogeneous coarse-grained reconfigurable processor · Sci. China Inf. Sci. 2014 |
Performance modeling and evaluation
simulation |
0.2 | 1 | 2013 | ReSSIM: a mixed-level simulator for dynamic coarse-grained reconfigurable processor · Sci. China Inf. Sci. 2013 |
Processor architecture and microarchitecture › computer arithmetic
digit-recurrence algorithm |
0.1 | 1 | 2012 | A Radix-16 Combined Complex Division/Square Root Unit with Operand Prescaling · IEEE Trans. Computers 2012 |
Reconfigurable computing and FPGAs › FPGA resource optimization
FPGA resource utilization |
0.1 | 1 | 2020 | DSP-Efficient Hardware Acceleration of Convolutional Neural Network Inference on FPGAs · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020 |
Image and video coding › video coding standards
H.264/AVC |
0.1 | 1 | 2015 | A Mixed-Grained Reconfigurable Computing Platform for Multiple-Standard Video Decoding (Abstract Only) · FPGA 2015 |
Memory systems
lookup table |
0.1 | 1 | 2014 | Complex Function Approximation Using Two-Dimensional Interpolation · IEEE Trans. Computers 2014 |
Electronic design automation › hardware simulation
multilevel simulation |
0.0 | 1 | 2013 | ReSSIM: a mixed-level simulator for dynamic coarse-grained reconfigurable processor · Sci. China Inf. Sci. 2013 |
Reconfigurable computing and FPGAs
FPGA implementation |
0.0 | 1 | 2012 | A Radix-16 Combined Complex Division/Square Root Unit with Operand Prescaling · IEEE Trans. Computers 2012 |
Methods — techniques the papers use, named apart from their topics
convolution transformation · 0.8quantization-aware training · 0.6monte carlo sampling · 0.6genetic algorithm · 0.6accuracy predictor · 0.6static reconfiguration · 0.4dynamic reconfiguration · 0.4reduced precision arithmetic · 0.4line-switched mesh connect · 0.2hierarchical configuration context · 0.2two-dimensional convolution · 0.2lagrange interpolation · 0.2bipartite table · 0.2HW/SW partitioning · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | DisQ: Distribution-aware quantization for ultra-low-bit large language models
Dong Wang 0040 |
Expert Syst. Appl. | 2 |
| 2025 | Dual-Branch Cross-Layer Information Flow Network for Camouflaged Object Detection in Complex ScenesabstractCamouflaged Object Detection (COD) aims to accurately detect objects that blend in with their backgrounds. While deep learning-based methods have significantly improved detection accuracy, the persistent similarity in chromatic and textural characteristics between camouflaged objects and their environments continues to pose substantial challenges, particularly in complex scenarios. Existing COD methods still encounter two primary challenges: (1) the difficulty in effectively fusing detailed features with semantic representations, leading to suboptimal balance between local and global features in the fused features; (2) progressive information dilution in conventional top-down decoder architectures, resulting in imprecise boundary localization. To address these dual challenges of multi-scale feature enhancement and semantic information dilution, we propose a D ual-Branch C ross-Layer I nformation F low Network (DCIFNet) for COD in complex scenes. Specifically, our framework integrates a hybrid CNN-ViT architecture for parallel extraction of localized patterns and global contextual features. Subsequently, a Global Local Fusion (GLF) module is employed to iteratively and adaptively fuse the global and local features at each stage using gating weights. Additionally, a novel Dense Connection Decoder (DCD) module is designed to mitigate the information dilution problem associated with traditional top-down decoders. Extensive experiments show that DCIFNet outperforms other mainstream COD models in most performance metrics on four widely used benchmark datasets. On the COD10K dataset, compared with the latest SOTA work of PRNet, DCIFNet improves the metrics of \({S_{\alpha}}\) , \({E_{\phi}}\) , and \(F_{\beta}^{w}\) by 0.11%, 0.11%, and 1.38%, respectively, demonstrating a significant performance enhancement. The code is available https://github.com/WateverOk/DCIFNet . Dong Wang 0040 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | BNN-SAM: Improving generalization of binary object detector by Seeking Flat Minima
Han Pu, Dezheng Zhang 0004, Ke Xu 0011, Ruchan Mo, Zhihong Yan, Dong Wang 0040 |
Appl. Intell. | 6 |
| 2024 | End-to-end acceleration of the YOLO object detection framework on FPGA-only devices
Dezheng Zhang 0004, Aibin Wang, Ruchan Mo, Dong Wang 0040 |
Neural Comput. Appl. | 4 |
| 2024 | Ulit-BiDet: An Ultralightweight Object Detector for SAR Images Based on Binary Neural NetworksabstractSynthetic aperture radar (SAR) target detection has extensively utilized convolutional neural networks (CNNs). Nonetheless, CNN-based methods often achieve favorable detection accuracy at the cost of high model complexity, hindering deployment of the algorithm in real-time application scenarios, such as maritime rescue and military decision making. To deal with this problem, we are the first to propose an ultra-lightweight object detector named Ulit-BiDet for SAR images, which incorporates binary neural network with very low storage and computation costs. First, we creatively design a performance and cost-scalable binary backbone that adapts to the diverse resources and computational capacities on practical devices. Then, the backbone structure is optimized with a new Non-Local Module to enhance semantic contextual information, thereby alleviating false detections caused by the interference from land clutter and sea clutter. Third, considering the SAR imaging mechanism, the interference near the ship boundary with similar scattering power probably affects the localization accuracy due to the interfered object-related contour information. To tackle the localization issue, we uniquely propose to utilize valuable and extra object-related contour semantics to guide representation learning of ship targets. The scheme compels the model to generate features that highlight object contour, thereby promoting accurate boundary localization in ship target detection. We validated the robustness of the proposed network in three mostly-cited publicly available datasets. Experimental results demonstrate that our model achieves 97.2%, 95.2%, and 77.3% detection accuracy with only 1.27M parameters and 0.22G OPs on SSDD, SAR-Ship, and AIR–SARShip2.0 ship detection datasets, respectively. Han Pu, Zhengwen Zhu, Dong Wang 0040 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | MultiQuant: Training Once for Multi-bit Quantization of Neural NetworksabstractQuantization has become a popular technique to compress deep neural networks (DNNs) and reduce computational costs, but most prior work focuses on training DNNs at each individual fixed bit-width and accuracy trade-off point. How to produce a model with flexible precision is largely unexplored. This work proposes a multi-bit quantization framework (MultiQuant) to make the learned DNNs robust for different precision configuration during inference by adopting Lowest-Random-Highest bit-width co-training method. Meanwhile, we propose an online adaptive label generation strategy to alleviate the problem of vicious competition under different precision caused by one-hot labels in the supernet training. The trained supernet model can be flexibly set to different bit widths to support dynamic speed and accuracy trade-off. Furthermore, we adopt the Monte Carlo sampling-based genetic algorithm search strategy with quantization-aware accuracy predictor as evaluation criterion to incorporate the mixed precision technology in our framework. Experiment results on ImageNet datasets demonstrate MultiQuant method can attain the quantization results under different bit-widths comparable with quantization-aware training without retraining. Ke Xu 0011, Qiantai Feng, Xingyi Zhang 0001, Dong Wang 0040 |
IJCAI | 4 |
| 2022 | TA-BiDet: Task-aligned binary object detector
Han Pu, Ke Xu 0011, Dezheng Zhang 0004, Dong Wang 0040 |
Neurocomputing | 6 |
| 2021 | Edge-Wise One-Level Global Pruning on NAS Generated Networks
Qiantai Feng, Ke Xu 0011, Yuhai Li, Dong Wang 0040 |
PRCV (4) | 5 |
| 2021 | GenExp: Multi-objective pruning for deep neural network based on genetic algorithm
Ke Xu 0011, Dezheng Zhang 0004, Jianjing An, Dong Wang 0040 |
Neurocomputing | 6 |
| 2020 | DSP-Efficient Hardware Acceleration of Convolutional Neural Network Inference on FPGAsabstractField-programmable gate array (FPGA)-based accelerators for convolutional neural network (CNN) inference have received significant attention in recent years. The reported designs tend to adopt a similar underlying approach based on multiplier-accumulator (MAC) arrays, which yields strong demand for the available on-chip DSP blocks, while leaving FPGA logic and memory resources underutilized. The practical outcome is that the computational roof of the accelerator is bound by the number of DSP blocks offered by the target FPGA. In addition, integrating the CNN accelerator with other functional units that may also need DSP blocks would degrade the inference performance. Leveraging the robustness of inference accuracy to limited arithmetic precision, we propose a transformation to the convolution computation, which leads to transformation of the accelerator design space and relaxes the pressure on the required DSP resources. Through analytical and empirical evaluations, we demonstrate that our approach enables us to strike a favorable balance between utilization of the FPGA on-chip memory, logic, and DSP resources, due to which, our accelerator considerably outperforms state of the art. We report the effectiveness of our approach on a variety of FPGA devices, including Cyclone-V, Stratix-V, and Arria-10, which are used in large number of applications, ranging from embedded settings to high performance computing. Our proposed technique yields 1.5x throughput improvement and 4x DSP resource reduction compared to the best frequency domain convolution-based accelerator, and 2.5x boost in raw arithmetic performance and 8.4x saving in DSPs compared to a state-of-the-art sparse convolution-based accelerator. Dong Wang 0040, Ke Xu 0011, Jingning Guo, Soheil Ghiasi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2019 | ABM-SpConv: A Novel Approach to FPGA-Based Acceleration of Convolutional Neural Network InferenceabstractHardware accelerators for convolutional neural network (CNN) inference have been extensively studied in recent years. The reported designs tend to utilize a similar underlying architecture based on multiplier-accumulator (MAC) arrays, which has the practical consequence of limiting the FPGA-based accelerator performance by the number of available on-chip DSP blocks, while leaving other resource under-utilized. To address this problem, we consider a transformation to the convolution computation, which leads to transformation of the accelerator design space and relaxes the pressure on the required DSP resources. We demonstrate that our approach enables us to strike a judicious balance between utilization of the on-chip memory, logic, and DSP resources, due to which, our accelerator considerably outperforms state of the art. We report the effectiveness of our approach on a Stratix-V GXA7 FPGA, which shows 55% throughput improvement, while using 6.25% less DSP blocks, compared to the best reported CNN accelerator on the same device. Dong Wang 0040, Ke Xu 0011, Qun Jia, Soheil Ghiasi |
DAC | 1 |
| 2019 | A Scalable OpenCL-Based FPGA Accelerator for YOLOv2abstractThis paper implements an OpenCL-based FPGA accelerator for YOLOv2 on Arria-10 GX1150 FPGA board. The hardware architecture adopts a scalable pipeline design to support multi-resolution input image, and improves resource utilization by full 8-bit fixed-point computation and CONV+BN+Leaky-ReLU layer fusion technology. The proposed design achieves a peak throughput of 566 GOPs under 190 MHz working frequency. The accelerator could run YOLOv2 inference with 288×288 input resolution and tiny YOLOv2 with 416×416 input resolution at the speed of 35 and 71 FPS, respectively. Ke Xu 0011, Xiaoyun Wang 0001, Dong Wang 0040 |
FCCM | 3 |
| 2019 | Training Low Bitwidth Model with Weight Normalization for Convolutional Neural Networks
Haoxin Fan, Jianjing An, Dong Wang 0040 |
PRCV (1) | 3 |
| 2017 | PipeCNN: An OpenCL-based open-source FPGA accelerator for convolution neural networksabstractConvolutional neural networks (CNNs) have been employed in many applications, such as image classification, video analysis and speech recognition. Being compute-intensive, CNNs are widely accelerated by GPUs with high power dissipations. Recently, studies were carried out exploiting FPGA as CNN accelerator because of its reconfigurability and advantage on energy efficiency over GPU, especially when OpenCL-based high-level synthesis tools are now available providing fast verification and implementation flows. In this paper, we demonstrate PipeCNN - an efficient FPGA accelerator that can be implemented on a variety of FPGA platforms with reconfigurable performance and cost. The PipeCNN project is openly accessible, and thus can be used either by researchers as a generic framework to explore new hardware architectures or by teachers as a off-the-self design example for any academic courses related to FPGAs. Dong Wang 0040, Ke Xu 0011, Diankun Jiang |
FPT | 1 |
| 2015 | A Mixed-Grained Reconfigurable Computing Platform for Multiple-Standard Video Decoding (Abstract Only)abstractA mixed-grained reconfigurable computing platform targeting multiple-standard video decoding is proposed in this paper. The platform integrates eight coarse-grained Reconfigurable Processing Units (RPUs), each of which consists of 16×16 multi-functional Processing Elements (PEs) and are implemented in TSMC 65 nm technology and two Altera Stratix IV EP4SE820 FPGAs. By exploiting dynamic reconfiguration of the RPUs and static reconfiguration of the FPGAs, the proposed platform achieves scalable performances and cost trade-offs to support a variety of video coding standards, including H.264, MPEG-2, AVS and HEVC. Two types of platform configuration are tested in this work. One configuration utilizes two RPUs and targets multiple-standard high-definition (HD) video decoding, while the other utilizes only one RPU, which works under a lower frequency and targets at standard resolution (SD) decoding. The HD configuration can decode 1920×1080 H.264 video streams at 30 frames per second (fps) under 200 MHz and 1920×1080 HEVC video streams at 30 fps under 236 MHz. It achieves a 25% performance gain over an industrial coarse-grained reconfigurable processor for H.264 decoding, and a 3.85× performance boosts over the Intel i5 general-purpose CPU for HEVC decoding. Leibo Liu, Victor Y. Chen, Dong Wang 0040, Min Zhu 0001, Shouyi Yin, Shaojun Wei |
FPGA | 3 |
| 2015 | An Energy-Efficient Coarse-Grained Reconfigurable Processing Unit for Multiple-Standard Video DecodingabstractA coarse-grained reconfigurable processing unit (RPU) consisting of 16 ×16 multi-functional processing elements (PEs) interconnected by an area-efficient line-switched mesh connect (LSMC) routing is implemented on a 5.4 mm ×3.1 mm die in TSMC 65 nm LP1P8M CMOS technology. A hierarchical configuration context (HCC) organization scheme is proposed to reduce the implementation overhead and the energy dissipation spent on fast reconfiguration. The proposed RPU is integrated into two system-on-a-chips (SoCs), targeting multiple-standard video decoding. The high-performance chip, comprising two RPU processors (named REMUS_HPP), can decode 1920 ×1080 H.264 video streams at 30 frames per second (fps) under 200 MHz. REMUS_HPP achieves a 25% performance gain over the XPP-III reconfigurable processor with only 280 mW power consumption, resulting in a 14.3 × improvement on energy efficiency. The other chip (named REMUS_LPP), targeting low power applications, integrates only one RPU processor. REMUS_LPP can decode 720 ×480 H.264 video streams at 35fps with 24.5 mW under 75 MHz, achieving a 76% reduction in power dissipation and a 3.96 × improvement on energy efficiency compared with the ADRES reconfigurable processor. Leibo Liu, Dong Wang 0040, Min Zhu 0001, Yansheng Wang, Shouyi Yin, Peng Cao 0002, Jun Yang 0006, Shaojun Wei |
IEEE Trans. Multim. | 2 |
| 2015 | Correction to "An Energy-Efficient Coarse-Grained Reconfigurable Processing Unit for Multiple-Standard Video Decoding"
Leibo Liu, Dong Wang 0040, Min Zhu 0001, Yansheng Wang, Shouyi Yin, Peng Cao 0002, Jun Yang 0006, Shaojun Wei |
IEEE Trans. Multim. | 2 |
| 2014 | A novel secure MIMO cognitive networkabstractThe existing cognitive radio networks (CRNs) lack the ability to deal with the electronic attack of malicious users, which can destroy the normal communication of the cognitive users. To solve the problem, we propose a novel secure CRN by combining MIMO technique and low density parity check (LDPC) codes at the base-station (BS) and mobile stations (MSs), the electronic interference of the malicious users can be effectively weakened. The capacity and the outage probability of the proposed MIMO CRN are analyzed. The simulation results in quasi-static Rayleigh flat-fading channel show that the proposed secure MIMO CRN can effectively cancel electronic interference from the malicious users. Yang Xiao 0004, Pengpeng Lan, Dong Wang 0040 |
ISCAS | 3 |
| 2014 | The diffserv cognitive network node with Controlled-UDPabstractThe existing cognitive network (CN) nodes do not support Differentiate-Serve (Diffserv) Quality of Service (QoS), and the existing Diffserv QoS scheme has not considered the link capacities between UDP subscribers and TCP subscribers of CN nodes. To solve these problems, a Diffserv CN node and Controlled-UDP (C-UDP) based fluid model are proposed for hybrid traffics' CN. The proposed C-UDP and CN node limit the data rates of UDP subscribers according to their priorities, and allocate the link capacities to different classes of TCP traffics and C-UDP traffics. The dynamic simulation results demonstrate the proposed CN node and C-UDP for the TCP/UDP hybrid traffics for CN to be valid. Yang Xiao 0004, Jinfeng Zou, Dong Wang 0040 |
ISCAS | 3 |
| 2014 | Implementation of multi-standard video decoder on a heterogeneous coarse-grained reconfigurable processor
Leibo Liu, Victor Y. Chen, Dong Wang 0040, Shouyi Yin, Peng Cao 0002, Shaojun Wei |
Sci. China Inf. Sci. | 3 |
| 2014 | Complex Function Approximation Using Two-Dimensional InterpolationabstractThis paper presents a new scheme for evaluating complex reciprocal and exponential functions in hardware. The proposed method utilizes a two-dimensional convolution algorithm to interpolate bivariate functions from tabulated function values in the complex domain. To reduce the memory requirements for lookup tables, the interpolation is decomposed into independent row and column computations, such that the same coefficient table can be shared. Three different interpolation kernels from degree-1 (linear) to degree-2 (quadratic Lagrange) and degree-3 (cubic Lagrange) are explored to find the optimal design parameters and the most acceptable trade-offs between performance and hardware resources. Moreover, a generic hardware architecture is designed to provide scalable implementation capabilities for computation precision and interpolation degree. To verify the proposed architecture, eight complex reciprocal and eight complex exponential design instances are implemented. The ASIC- and FPGA-based experimental results show that the proposed scheme can efficiently approximate the complex reciprocal and exponential functions with up to 16-bit precision, as well as achieve a considerable reduction of memory requirements compared with traditional bipartite and multipartite schemes. The proposed method is also applicable to other complex functions. Dong Wang 0040, Milos D. Ercegovac, Yang Xiao 0004 |
IEEE Trans. Computers | 1 |
| 2014 | SimRPU: A Simulation Environment for Reconfigurable Architecture ExplorationabstractTo assist the system architects with fast exploration and performance evaluation of the reconfigurable software/hardware architectures, this paper presents a system-level simulator, named after SimRPU, for the reconfigurable processing unit (RPU), which is the major computing engine in reconfigurable processor. The proposed simulator consists of a simulation kernel, a software compiler, a system profiler providing performance, area and power information for the desired architectures, and a system debugger supporting inspecting and modification of the internal state of the RPU. Object-oriented hierarchical and parameterized architecture modeling techniques are proposed to satisfy the requirements for a fast and comprehensive evaluation. Cycle-accurate simulation mechanisms are developed to improve the accuracy of the profiled performance data. Compared with the traditional register transfer level (RTL) based simulation scheme, the proposed simulator could achieve an average speedup of 18.5× with only 3.5% reduction on performance estimation accuracy. One reconfigurable processor targeted at high-definition multimedia decoding applications (such as H.264, MPEG2, AVS, etc.) is implemented with Taiwan Semiconductor Manufacturing Company 65-nm process using the proposed exploration and design flow. The measured results show that the implemented architecture has obvious advantages in terms of both performance and power consumption than the reference designs in multimedia decoding applications. Leibo Liu, Dong Wang 0040, Shouyi Yin, Victor Y. Chen, Min Zhu 0001, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2013 | Implementation of multi-standard video decoding algorithms on a coarse-grained reconfigurable multimedia processorabstractThis paper proposed a THPHP (Task-based Hybrid Parallels and Hybrid Pipelines) scheme to implement multistandard video decoding algorithms, i.e. MPEG-2, H.264 and AVS (Audio Video coding Standard), on a heterogeneous coarsegrained reconfigurable multimedia processor called REMUS (REconfigurable MUltimedia System). Multiple level parallelism and multiple level pipeline techniques are proposed in this scheme. Simulation results show that the video decoder can support H.264 HP (High Profile) 1920×1080@30fps (frame per second) streams, AVS JP (Jizhun Profile) 1920×1080@39fps streams, and MPEG-2 MP (Main Profile) 1920×1080@41fps streams when exploiting a 200MHz working frequency. Leibo Liu, Victor Y. Chen, Shouyi Yin, Dong Wang 0040, Shaojun Wei, Li Zhou 0015, Peng Cao 0002 |
ISCAS | 4 |
| 2013 | Battery-Aware MAC Analytical Modeling for Extending Lifetime of Low Duty-Cycled Wireless Sensor NetworkabstractEmerging techniques and systems for Wireless Sensor Network (WSN) are developed in the last decade for various application fields. In WSN, the sensor nodes are usually distributed over a large area and are powered by batteries with limited energy, maintaining a long service lifetime for the entire network becomes a challenging task. In this paper, a novel battery aware MAC analytical model is proposed for low duty-cycled WSN. The proposed analytical model takes the characteristics of actual battery into account and targets the optimal sleep interval with a reasonable trade-offs between the energy dissipation on sending the preamble and idle listening. The simulation results demonstrate that the proposed approach can improve the energy efficiency as well as guarantee low latency and high reliability. Shouyi Yin, Leibo Liu, Shaojun Wei, Dong Wang 0040 |
NAS | 5 |
| 2013 | ReSSIM: a mixed-level simulator for dynamic coarse-grained reconfigurable processor
Leibo Liu, Wen Jia, Shouyi Yin, Dong Wang 0040, Guanyi Sun, Eugene Tang, Shaojun Wei |
Sci. China Inf. Sci. | 4 |
| 2012 | A Radix-16 Combined Complex Division/Square Root Unit with Operand PrescalingabstractWe present a novel design of a radix-16 combined unit for complex division and square root in fixed-point format. A new digit-recurrence algorithm with two-step operand prescaling is developed for complex square root to avoid postscaling of the result. A combined recurrence algorithm is generalized and a scalable hardware architecture is proposed. Designs with different operand precision are implemented in Altera Stratix-II FPGA and cost and performance are evaluated and compared with reference designs of complex division or square root implemented combined or separately. The results show advantages of the proposed combined design in cost and performance. Dong Wang 0040, Milos D. Ercegovac |
IEEE Trans. Computers | 1 |
| 2009 | A radix-8 complex divider for FPGA implementationabstractWe present a design of a radix-8 complex division for fixed-point operands suitable for FPGA implementation. The design, consisting of operands' prescaling and digit recurrence, shares logic resources and optimizes the use of 6-input LUTs of FPGA devices for efficient design. An optimized single table for prescaling factors is developed. The design is implemented in Altera Stratix-II FPGA for several operands precisions and compared in cost, latency and power with a design using non-shared resources and with an IP-based design. The results show advantages of the proposed design in cost, delay, and power. Dong Wang 0040, Milos D. Ercegovac, Nanning Zheng 0001 |
FPL | 1 |