VLDB 2026 Research / reviewers in the wild / expert
Dake Liu
dblp:98/4851
· DBLP profile ↗
46ranked-venue papers
1as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 28 · 7 since 2021Computer networks · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Instruction Fusion for ASIPs: A Low-Cost Approach to Structural Hazard ControlabstractMoore’s law is reaching saturation, while the demand for computing capacity continues to rise. Application-Specific Instruction-Set Processors (ASIPs) are therefore an essential technology for embedded processing in areas such as communication, media, gaming, and control systems, due to their high area and energy efficiency. The most effective method for ASIPs acceleration is known as instruction fusion. Instruction fusion enhances performance but also introduces challenges related to the complexity of structural hazard control. The difficulty in this research field is achieving fine-grained structural hazard control using low-cost hardware. The state-of-the-art technologies have not been able to effectively address this challenge. Firstly, this paper systematically analyzes the rationale behind instruction fusion and explores methods to optimize its performance. Secondly, based on the design of a 5G micro-base station baseband processor, we present an instruction fusion pipeline design example. Furthermore, to address the structural hazards that arise from instruction fusion, we propose a lightweight Hardware Resource Table (HRT) to address the issue and outline its benefits. These benefits include utilizing low silicon costs to achieve performance improvements and reducing on-chip program memory. According to benchmark evaluations, area efficiency improves by 23%, accompanied by a 108% increase in energy efficiency. Xinbing Zhou, Tianlang Liu, Tiancheng Tang, Yi Man, Wei Chen 0100, Dake Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2026 | Low-Latency and Low-Overhead Fault-Tolerant Network-on-Chip for Embedded Real-Time Heterogeneous Parallel Digital Signal ProcessorsabstractAt present, most network on chips (NoCs) are designed for symmetric multiprocessing architectures for general purpose computing, it is accompanied by substantial routing delays and a relatively high cost associated with reorder buffer memory. How to design NoC with high bandwidth transmission, ultra-low latency, and low silicon overhead for embedded real-time heterogeneous parallel digital signal processor, which faces major challenge. In order to solve this challenge, we hence propose an innovative NoC architecture. It is designed to meet the high throughput, low latency and silicon overhead requirements of embedded real-time heterogeneous parallel digital signal processor systems. The significant difference in data transmission length is a characteristic of heterogeneous multi-core systems. Our designed NoC has the advantages of flexibility and low latency. It can not only significantly improve the transmission performance of long data packets, but also improve the transmission efficiency of short data packets. In addition, we design efficient routing methods and fault-tolerant data transmission mechanisms to improve reliability and performance of NoC, which can mitigate the impact of router failures, reduce the need of reorder buffer memory and power consumption, and improve the robustness of NoC. The experiments showed the performance advantages of our design in terms of latency, throughput and silicon overhead in comparison with the state-of-the-art NoC, our design respectively reduced the latency, area and power consumption by 93%, 97.5% and 94.2% respectively, and improved the throughput and MTTF (Mean Time To Failure) by 22% and 9.4 times respectively. This verifies its effectiveness in high-end applications. Wei Chen 0100, Dake Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | Efficient and Flexible Deep Learning Processor for Embedded Devices With 99.8% PE Utilization Under Low Silicon CostabstractThe flexibility and efficiency of accelerator hinder the realization of DNN (Deep Neural Network) on embedded devices. How to design efficient and flexible deep learning processors to meet low-power requirements while supporting various iteratively updated deep neural networks, which is a major challenge in the current field. Moreover, the significant power consumption and latency generated by data access between on-chip and off-chip memory have been the bottleneck of DNN accelerator development. We hence design the flexible and efficient deep learning processor (DLP) based on our designed specific neural network instruction set for DNN inference, which can be used to flexibly accelerate the inference of iteratively updated neural networks. Our design includes the design of hardware architecture, neural network acceleration instruction set, efficient SIMD control path and data path, which can be used to accelerate the execution of the DNN programs, and support hiding the data transfer time between on-chip and off-chip memory. In addition, we propose a parallel and conflict-free data access scheme to reduce data access overhead in memory system. Moreover, we also propose a scheduling framework to improve performance and minimize power consumption under the hardware resource constraints. The experimental results indicate that our design’s power consumption efficiency is 75.3 TOPS/W at TSMC 65 nm process. Our power efficiency is 7.6, 5.5 and 4 times higher than the state-of-art DNN accelerators respectively. Wei Chen 0100, Dake Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2026 | Sayram: A Hardware-software Co-design to Accelerate Wireless Baseband ProcessingabstractMicro base stations, with limited antennas and extensive deployment, require scaled-down hardware. Software-defined radio solutions (e.g., CPU, many-core systems, GPU) offer flexibility but incur high area and power costs, while traditional DSP lacks efficient acceleration for smaller configurations. The key challenge for micro base stations is achieving minimal area and power overhead while meeting 5G requirements. This article presents a hardware-software co-designed architecture, Sayram, which minimizes overhead for 5G physical layer processing. Sayram integrates an instruction fusion mechanism, along with the compiler for simplified programming, a Vector Indirect Addressing Memory (VIAM) to minimize memory access cycles, and an improved vector register design to accelerate small-scale matrix computation, thereby improving overall processor efficiency. Operating at 1 GHz, Sayram achieves 158 GOPS with a 1.18 mm \(^2\) area, supporting 2T2R and 4T4R Physical Uplink Shared Channel (PUSCH) processing in single-core and dual-core modes, respectively. Evaluations show that Sayram’s area efficiency is 3× and 9× higher than traditional DSP and CGRA architectures, respectively, with power efficiency improvements of 44× and 6×. Sayram’s energy and area efficiency surpass CPU solutions by orders of magnitude. Xinbing Zhou, Shaobo Shi, Shao-Han Liu, Yunxiang Tang, Tiancheng Tang, Yi Man, Dake Liu |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2026 | Efficient NoC for Embedded Heterogeneous Multi-Core DSPsabstractAt present, most networks on chips (NoCs) are designed for symmetric multiprocessing architectures for general-purpose computing. While currently designed NoCs provide considerable flexibility, they are accompanied by substantial routing delays and a relatively high cost associated with reorder buffer memory. The significant difference in data transmission length is a characteristic of heterogeneous multi-core systems. We propose an innovative NoC architecture for embedded heterogeneous multi-core digital signal processor (DSP) systems, which includes the design of top-level architecture, router, network interface (NI), and NoC pipeline. Our designed NoC architecture has the advantages of flexibility and low latency. It can not only significantly improve the transmission performance of long data packets but also improve the transmission efficiency of short data packets. Moreover, we propose efficient routing methods, which include an efficient routing algorithm, a hybrid transmission method, a packet-connected circuit transmission method, a short packet transmission method, and multicast and broadcast transmission methods. We also propose fault-tolerant data transmission mechanisms, which include a fault-tolerant data transmission algorithm and a deadlock avoidance method. Our proposed method and mechanism can improve the performance and reliability of NoC, mitigate the impact of router failures, eliminate the need for reorder buffer memory, reduce power consumption, and improve the robustness of NoC. The experiments show the performance advantages of our design in terms of latency, throughput, and silicon overhead in comparison with state-of-the-art NoCs. Our design reduced the latency, area, and power consumption by 47%, 92.6%, and 78.9%, respectively, and improved the throughput by 27%. This verifies its effectiveness in high-end applications. Wei Chen 0100, Dake Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | Task Scheduling for Heterogeneous Multi-Core Processors Based on Deep Reinforcement LearningabstractHeterogeneous multicore processor systems are commonly used for scheduling tasks of DAG applications. Deep reinforcement learning, with its superior ability to perceive decisions directly and handle high‐dimensional state actions, has become a prevalent solution for scheduling these systems. However, the incomplete environment models and large action spaces of deep reinforcement learning present significant challenges to scheduling. This paper investigates a scheduling problem in a heterogeneous multicore processor environment. Initially, system environment information is extracted and encoded using a graph convolutional neural network based on integrating adapter and AdapterFusion into the transformer architecture. Then, by separating task selection and processor allocation, the decision space is reduced: the former uses a deep neural network to learn to select nodes, and the latter allocates processors using a heuristic scheduling algorithm combining earliest completion time‐based node replication and rolling technology. The entire scheduling process is a Markov decision problem. Therefore, the PPO algorithm with dynamic adjustment of the clipping factor, combined with an advantage actor‐critic network, is employed for training, optimizing, and evaluating the algorithm to find the optimal scheduling strategy. The training process adopts a reward function for the time and power consumption required for completed task scheduling to ensure that multiple DAG application task scheduling can achieve optimal performance. Experiments conducted in various environments with different parameters show that, compared to other algorithms, this algorithm reduces the overall execution time and power consumption cost of heterogeneous multicore processor tasks by 11.09%. Qiguang Tan, Wei Chen 0100, Dake Liu |
Int. J. Intell. Syst. | 3 |
| 2025 | Hardware Sharing Design Method of Rate Matching and Interleaving for Wireless Terminal in Industrial Internet of ThingsabstractReducing the power consumption of baseband chips in wireless terminals of the Industrial Internet of Things (IIoT) is a great challenge. Among them, the logic function of the rate matching and interleaving hardware module is very complex, which occupies a considerable portion of power consumption in the baseband chip, so it is of great significance to design a low-power rate matching and interleaving hardware. Due to differences in interleaving algorithms and throughput, the interleavers used in 4G long term evolution (LTE) and 5G NR/6G have not been merged into a single architecture. Switching between different standards provides new options for interleaver design, and by configuring the architecture of different interleavers, it is possible to use the same hardware for different standards to reduce hardware resources and power consumption. This article studies the block interleaving and rate matching of turbo codes, convolutional codes, polar codes, and low-density parity check (LDPC) codes used in 4G LTE and 5G NR/6G communication links. With regard to the different algorithms for these four types of encoding, shared design is carried out on the hardware structure. In this experiment, according to the proposed memory and interleaving sharing scheme for hardware design and hardware simulation, the area overhead of 0.1$\mu $m2 and power consumption of 4.31 mW are obtained by Synopsys synthesis at the SMIC 28-nm process and the frequency of 50 MHz. This achieves the maximum hardware reuse of four encoding schemes in the downlink communication link of 4G LTE and 5G NR/6G, and reduce power consumption. Wei Chen 0100, Kejia Huo, Dake Liu |
IEEE Internet Things J. | 3 |
| 2025 | AIKII: An AI-Enhanced Knowledge Interactive Interface for Knowledge Representation in Educational GamesabstractABSTRACT The use of generative AI to create responsive and adaptive game content has attracted considerable interest within the educational game design community, highlighting its potential as a tool for enhancing players' understanding of in‐game knowledge. However, designing effective player‐AI interaction to support knowledge representation remains unexplored. This paper presents AIKII, an AI‐enhanced Knowledge Interaction Interface designed to facilitate knowledge representation in educational games. AIKII employs various interaction channels to represent in‐game knowledge and support player engagement. To investigate its effectiveness and user learning experience, we implemented AIKII into The Journey of Poetry, an educational game centered on learning Chinese poetry, and conducted interviews with university students. The results demonstrated that our method fosters contextual and reflective connections between players and in‐game knowledge, enhancing player autonomy and immersion. Dake Liu, Huiwen Zhao, Wen Tang 0004 |
Comput. Animat. Virtual Worlds | 1 |
| 2025 | High-Throughput LDPC Decoder for Multiple Wireless StandardsabstractIt is a great challenge to design an LDPC decoder with multi-standard compatibility, flexibility and low silicon overhead. This paper presents the efficient and low-overhead design of an LDPC decoder tailored for multi-standard, which include WLAN, 5G NR and WiMAX. We follow the design principles of Application-Specific Instruction-set Processor (ASIP). In order to enhance throughput, we double the computational speed by reducing memory speed from double logic speed to logic speed. By proposing the optimized hybrid scheduling algorithm based on matrix reordering, we further solve scheduling problems and eliminate pipeline conflicts. Through performing logic synthesis utilizing the 28 nm SMIC CMOS cell library, synthesis results show that the core area of our designed decoder is 0.86 mm2, the logic gate count is 1716 K, and our design achieves impressive throughput rates, that is up to 9.96 Gbps for WLAN, 7.69 Gbps for WiMAX, and 33 Gbps for 5G NR. Compared with other state-of-the-art LDPC decoders, the experimental results show that our proposed decoder has up to$4.5\times $higher throughput,$3.9\times $better area efficiency and$5.8\times $better energy efficiency than these state-of-the-art implementations. Wei Chen 0100, Dake Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | B5G/6G URLLC Latency Reduction Method for Multisensor Industrial Internet of ThingsabstractHow to enable B5G/6G ultra reliable low-latency communications (URLLC) to meet the requirements of low latency, ultrahigh-speed and real-time industrial automatic control for multisensor Industrial Internet of Things (IIoT), it is one of the most important challenges of multisensor IIoT. The round-trip time (RTT) is significantly reduced by semi-persistent scheduling (SPS) or short transmission time interval (TTI) method, yet these methods may cost extra frequency and time resources. To reduce latency without the cost of extra spectrum and time resources, such research is almost blank in both academia and industry for B5G/6G URLLC scenario of multisensor IIoT. We hence propose the latency reduction method for the B5G/6G URLLC scenario of multisensor IIoT. Our research fills this research gap. The proposed reduction method includes a storage planning method for global data and local data, a dynamic data configuration method and an execution time minimization method. Based on these methods, we can obtain the minimum data space of the program and the data storage arrangement approaching to the minimum worst-case execution time (WCET). Under the main frame of SPS and short TTI method, our proposed method can further reduce latency without reducing spectrum and time resources utilization by minimizing data storage and access delay. The experimental results show that our method did not reduce spectrum and time resources utilization, and when our method was applied in SPS and short TTI method, our method further reduced the physical layer computing latency by about 21% to 28%. Wei Chen 0100, Dake Liu, Yong Bai 0002 |
IEEE Internet Things J. | 2 |
| 2024 | Conflict-Free Parallel Data Access Technology for Matrix Calculation in Memory System of ASIP of 5G/6G Macro Base StationsabstractAmong the physical layer baseband algorithms in macro base stations, the matrix processing has the dominant computing cost, large data access overhead, and complicated addressing mode. The existing data access methods are not the best solution for memory system of application-specific instruction set processor (ASIP) in 5G/6G macro stations. We hence proposed a parallel conflict-free data access method for matrix. Moreover, we proposed a nonredundant access method for positive-definite matrix that stores only trigonometric part, and the corresponding parallel addressing method with low storage overhead. Our method solves the problem of data conflict and minimizes the memory space of ASIP, and supports ASIP to approach the performance limit of architecture to the maximum extent under the constraint of architecture and data parallelism. We implemented the proposed method as a static memory optimizer, which provides a tool for the design and optimization of memory system of ASIP in 5G/6G macro base stations with low overhead. Experimental results show that for the${64}\mathbf {\times }{64}$positive-definite matrix inversion algorithm based on Cholesky decomposition, our addressing method can save 32% of the execution time, and the cost of memory space is reduced by half. Wei Chen 0100, Dake Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Hardware Sharing for Channel Interleavers in 5G NR StandardabstractInterleaver module is an important part of modern mobile communication system. It plays an important role in reducing bit error rate and improving transmission efficiency over fading channels. In 5G NR (5th Generation New Radio) standards, LDPC (low-density parity-check) and polar channel codes are employed for data channels and control channels, respectively. If multiple interleavers are implemented separately for them, the cost increases significantly. To address this issue, a hardware multiplexing scheme for channel interleavers based on LDPC and polar codes is proposed in this paper. Firstly, the formulas for the processes of the control channel interleaving and data channel interleaving are derived with respect to 5G NR standard. Then, the hardware implementation structures of the two interleavers are given. Subsequently, hardware reuse is proposed by sharing the similar or identical parts between the two hardware structures. Simulation results verify the correctness of our proposed scheme and demonstrate that it can realize the hardware sharing of the two kinds of channel interleavers to reduce the cost of silicon. Xiaokang Xiong, Yuhang Dai, Zhuhua Hu, Kejia Huo, Yong Bai 0002, Hui Li 0039, Dake Liu |
Secur. Commun. Networks | 7 |
| 2019 | A High-Flexible Low-Latency Memory-Based FFT Processor for 4G, WLAN, and Future 5GabstractA high-throughput programmable fast Fourier transform (FFT) processor is designed supporting 16- to 4096-point FFTs and 12- to 2400-point discrete Fourier transforms (DFTs) for 4G, wireless local area network, and future 5G. A 16-path data parallel memory-based architecture is selected as a tradeoff between throughput and cost. To implement a hardware-efficient high-speed processor, several improvements are provided. To maximally reuse the hardware resource, a reconfigurable butterfly unit is proposed to support computing including eight radix-2 in parallel, four radix-3/4 in parallel, two radix-5/8 in parallel, and a radix-16 in one clock cycle. Twiddle factor multipliers using different schemes are optimized and compared, wherein modified coordinate rotation digital computer scheme is finally implemented to minimize the hardware cost while supporting both FFTs and DFTs. An optimized conflict-free data access scheme is also proposed to support multiple butterflies at any radices. The processor is designed as a general IP and can be implemented using a processor synthesizer (application-specific instruction-set processor designer). The electronic design automation synthesis result based on a 65-nm technology shows that the processor area is 1.46 mm2. The processor supports 972 MS/s 4096-point FFT at 250 MHz with a power consumption of 68.64 mW and a signal-to-quantization-noise ratio of 66.1 dB. The proposed processor has better-normalized throughput per area unit than the state-of-the-art available designs. Shao-Han Liu, Dake Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | High-throughput area-efficient processor for 3GPP LTE cryptographic core algorithmsabstractThere are three sets of cryptographic algorithms working on LTE technology and each set based on one core algorithm. In high-end embedded systems, it is necessary to implement the three core algorithms: block cipher AES-128 and stream ciphers SNOW 3G and ZUC, with high performance and low silicon cost. This paper proposes a high throughput ASIP (application-specific instruction-set processor) design (CP-LTE) for the three core algorithms. Yuanhong Huo, Dake Liu |
ASAP | 2 |
| 2017 | A 3-coil simultaneous power and uplink data transmission inductive link for battery-less implantable devicesabstractTraditional 2-coil inductive link, for battery-less medical implantable devices, could hardly be used for the simultaneous wireless power and data transmission when the coupling is ultra-weak, since the strong power carrier could overwhelm the uplink signal and saturate the receiver. We thus proposed a 3-coil inductive link to transmit power and downlink data, as well as uplink data from the in-body device to the out-body part simultaneously. That is to add one extra out-body coil for the uplink receiving to minimize the power interference to the uplink signal by balanced canceling interference magnetic field. The inbody coil is still shared by the wireless power transfer (WPT) and the uplink communication. Two carriers (i.e. 2 MHz power carrier and 500 kHz uplink data carrier) are used for the WPT and the communication separately. Through optimizing the relative position of the two out-body coils and the circuit parameters, we achieved satisfied signal quality and sufficient efficiency of WPT. The prototype of this link sent -10 dBm power carrying 10 kbps uplink data rate under 10 mW power carrier interference with much relaxed -23 dB signal-to-interference ratio (SIR). This SIR was 67 dB higher than that of the 2-coil solution in our previous work. At the same time, the uplink caused less than 14% relative drop of the WPT efficiency compared with the pure WPT link without the uplink circuits. Min Li 0006, Dake Liu, Chen Gong 0003, Wan Qiao |
ISCAS | 2 |
| 2015 | High performance table-based architecture for parallel CRC calculationabstractA high performance table-based architecture implementation for CRC (cyclic redundancy check) algorithms is proposed. The architecture is designed based on a highly parallel CRC algorithm. The algorithm first divides a given message with any length into bytes. Then it performs CRC computation using lookup tables among the divided bytes in parallel. At last, the results are XORed to obtain the CRC value of the given message. The algorithm is table-based and can accelerate different CRC algorithms. Based on the algorithm, the architecture is designed to accelerate CRC algorithms with high parallelism and flexibility. The architecture is configurable and can support CRC algorithms such as CRC32, CRC24, CRC-CCITT, CRC16, CRC8. CRC value of 128-bit input data can be generated in one cycle. Our method also allows calculation over data that is less than 128-bit wide without increasing hardware cost. With 128-bit input each clock cycle, the throughput of the proposed architecture reaches up to 100 Gbps by utilizing 16 KB SRAM (Static Random Access Memory) with about 12% area reduction compared with previous work. Yuanhong Huo, Wei Wang 0128, Dake Liu |
LANMAN | 4 |
| 2015 | High-Throughput Trellis Processor for Multistandard FEC DecodingabstractTrellis codes, including Low-Density Parity-Check (LDPC), turbo, and convolutional code (CC), are widely adopted in advanced wireless standards to offer high-throughput forward error correction (FEC). Designing a multistandard FEC decoder is of great challenge. In this paper, a trellis application specified instruction-set processor (TASIP) is presented for multistandard trellis decoding. A unified forward-backward recursion kernel with an eight-state parallel trellis structure is proposed. Based on the kernel, a datapath for multialgorithm and a shared memory subsystem are introduced. The flexibility and the compatibility are guaranteed by a programmable decoding flow and the trellis decoding instruction set. Synthesis results show that the area consumption is 2.12 mm2(65 nm). TASIP provides trimode FEC decoding ability with the throughput of 533, 186, and 225 Mb/s for LDPC, turbo, and 64 states CC under the clock frequency of 200 MHz, which outperforms other trimode proposals both in area efficiency and recursion efficiency. TASIP provides high-throughput decoding for current standards, including 3rd Generation Partnership Project-Long Term Evolution, 802.16e, and 802.11n, with unified architecture and high compatibility. Zhenzhi Wu, Dake Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | Flexible multistandard FEC processor design with ASIP methodologyabstractDesigning decoder for forward error correction (FEC) is more and more challenging because of the requirements on simultaneous supporting of various wireless standards within one IC module. The flexibility, silicon cost and throughput efficiency are all necessary to be traded off. In this paper, by using ASIP methodology, software-hardware co-design is introduced to offer sufficient flexibility of FEC decoding. The decoding procedure can be programmable for decoding QC-LDPC, Turbo and Convolutional Codes. Firstly, the common features from all mentioned algorithms and their corresponding datapaths are analyzed and a unified multi-standard datapath is introduced. Based on it, an application specific instruction-set is proposed and an ASIP (Application Specific Instruction-set Processor) for the FEC algorithms is designed. The firmware FEC codes are developed to adapt to standards. Synthesis results show that the proposed FEC processor is 1.54mm2under 65nm CMOS process. It offers QC-LDPC decoding for WiMAX, Turbo decoding for 3GPP-LTE, and 64 states Convolutional code (CC) decoding at the throughput of 193 Mbps, 62 Mbps and 60 Mbps respectively under clock frequency of 200 MHz. The proposed ASIP provides programmable high throughput compared to other tri-mode hardware modules. Zhenzhi Wu, Dake Liu |
ASAP | 2 |
| 2014 | FPGA implementation of a multi-algorithm parallel FEC for SDR platformsabstractForward Error Correction (FEC) consumes excessive computation in a Software Defined Radio (SDR) system. In this work, a high-throughput flexible FEC processor is proposed for the decoding acceleration. The FEC processor enables Turbo/QC-LDPC/Convolutional Code decoding with software-hardware co-reconfigurability. A multi-algorithm unified trellis processing unit is introduced for resource sharing. A parallel architecture is proposed for high-throughput decoding. The Software Defined FEC (SD-FEC) with Application Specific Instruction-set Processor architecture is introduced for improving flexibility and enabling fast reconfiguration. The proposed SD-FEC can be applied to both low-cost low power applications and high performance applications. Results show that the proposed tri-mode FEC processor achieves high decoding efficiency and enough flexibility, which suits for the flexible SDR platforms. Zhenzhi Wu, Dake Liu, Qingying Wang |
FPL | 2 |
| 2014 | A Vision of IoT: Applications, Challenges, and Opportunities With China PerspectiveabstractInternet of Things (IoT), which will create a huge network of billions or trillions of “Things” communicating with one another, are facing many technical and application challenges. This paper introduces the status of IoT development in China, including policies, R&D plans, applications, and standardization. With China's perspective, this paper depicts such challenges on technologies, applications, and standardization, and also proposes an open and general IoT architecture consisting of three platforms to meet the architecture challenge. Finally, this paper discusses the opportunity and prospect of IoT. Shanzhi Chen, Dake Liu, Bo Hu 0003, Hucheng Wang |
IEEE Internet Things J. | 3 |
| 2013 | Conflict-free data access for multi-bank memory architectures using paddingabstractFor high performance computation memory access is a major issue. Whether it is a supercomputer, a GPGPU device, or an Application Specific Instruction set Processor (ASIP) for Digital Signal Processing (DSP) parallel execution is a necessity. A high rate of computation puts pressure on the memory access, and it is often non-trivial to maximize the data rate to the execution units. Many algorithms that from a computational point of view can be implemented efficiently on parallel architectures fail to achieve significant speed-ups. The reason is very often that the speed-up possible with the available execution units are poorly utilized due to inefficient data access. This paper shows a method for improving the access time for sequences of data that are completely static at the cost of extra memory. This is done by resolving memory conflicts by using padding. The method can be automatically applied and it is shown to significantly reduce the data access time for sorting and FFTs. The execution time for the FFT is improved with up to a factor of 3.4 and for sorting by a factor of up to 8. Joar Sohl, Jian Wang 0035, Andreas Rosenblad, Dake Liu |
HiPC | 4 |
| 2012 | Convolutional Decoding on Deep-pipelined SIMD Processor with Flexible Parallel MemoryabstractSingle Instruction Multiple Data (SIMD) architecture has been proved to be a suitable parallel processor architecture for media and communication signal processing. But the computing overhead such as memory access latency and vector data permutation limit the performance of conventional SIMD processor. Solutions such as combined VLIW and SIMD architecture are designed with an increased complexity for compiler design and assembly programming. This paper introduces the SIMD processor in the ePUMA1 platform which uses deep execution pipeline and flexible parallel memory to achieve high computing performance. Its deep pipeline can execute combined operations in one cycle. And the parallel memory architecture supports conflict free parallel data access. It solves the problem of large vector permutation in a short vector SIMD machine in a more efficient way than conventional vector permutation instruction. We evaluate the architecture by implementing the soft decision Viterbi algorithm for convolutional decoding. The result is compared with other architectures, including TI C54x, CEVA TeakLike III, and PowerPC AltiVec, to show ePUMA's computing efficiency advantage. Jian Wang 0035, Andreas Rosenblad, Joar Sohl, Dake Liu |
DSD | 4 |
| 2012 | Automatic Permutation for Arbitrary Static Access PatternsabstractA significant portion of the execution time on current SIMD and VLIW processors is spent on data access rather than instructions that perform actual computations. The ePUMA architecture provides features that allow arbitrary data elements to be accessed in parallel as long as the elements reside in different memory banks. Using permutation to move data elements that are accessed in parallel, the overhead from memory access can be greatly reduced; and, in many cases completely removed. This paper presents a practical method for automatic permutation based on Integer Linear Programming (ILP). No assumptions are made about the structure of the access patterns other than their static nature. Methods for speeding up the solution time for periodic access patterns and reusing existing solutions are also presented. Benchmarks for e.g. FFTs show speedups of up to 3.4 when using permutation compared to regular implementations. Joar Sohl, Jian Wang 0035, Andreas Rosenblad, Dake Liu |
ISPA | 4 |
| 2011 | Case Study of Efficient Parallel Memory Access Programming for the Embedded Heterogeneous Multicore DSP Architecture ePUMAabstractWe consider the challenges in writing efficient code for ePUMA, a novel domain-specific heterogeneous multicore architecture with SIMD DSP slave cores, multi-banked on-chip vector register files for parallel access and configurable permutation hardware that decouples memory access from computation. Suitable data layout in memory and in vector registers, combined with using ePUMA's powerful addressing modes, is key to exploiting SIMD units efficiently and achieving the throughput required for prospective applications in 4G mobile telecommunication and multimedia. Erik Hansson, Joar Sohl, Christoph W. Kessler, Dake Liu |
CISIS | 4 |
| 2010 | NoGapCL: A flexible common language for processor hardware descriptionabstractFlexible Application Specific Instruction set Processors (ASIP) are starting to replace monolithic ASICs in a wide variety of fields. However the construction of an ASIP is today associated with a substantial design effort. NoGap (Novel Generator of Micro Architecture and Processor) is a tool for ASIP designs, utilizing hardware multiplexed data paths. One of the main advantages of NoGap compared to other EDA tools for processor design, is that NoGap impose few limits on the architecture and thus design freedom. NoGap does not assume a fixed processor template and is not a data flow synthesizer. To reach this flexibility NoGap makes heavy use of the compositional design principle. This paper describe NoGapCL, a flexible common language for processor hardware description. A RISC processor using NoGapCLhas been constructed with NoGap in less than a working day and synthesized to an FPGA. With no FPGA specific optimizations this processor met timing closure at 178 MHz in a Virtex-4 LX80 speedgrade 12. Wenbiao Zhou, Per Karlström, Dake Liu |
DDECS | 3 |
| 2010 | Software Programmable Data Allocation in Multi-bank Memory of SIMD ProcessorsabstractThe host-SIMD style heterogeneous multi-processor architecture offers high computing performance and user friendly programmability. It explores both task level parallelism and data level parallelism by the on-chip multiple SIMD coprocessors. For embedded DSP applications with predictable computing feature, this architecture can be further optimized for performance, implementation cost and power consumption. The optimization could be done by improving the SIMD processing efficiency and reducing redundant memory accesses and data shuffle operations. This paper introduces one effective approach by designing a software programmable multi-bank memory system for SIMD processors. Both the hardware architecture and software programming model are described in this paper, with an implementation example of the BLAS syrk routine. The proposed memory system offers high SIMD data access flexibility by using lookup table based address generators, and applying data permutations on both DMA controller interface and SIMD data access. The evaluation results show that the SIMD processor with this memory system can achieve high execution efficiency, with only 10% to 30% overhead. The proposed memory system also saves the implementation cost on SIMD local registers, in our system, each SIMD core has only 8 128-bit vector registers. Jian Wang 0035, Joar Sohl, Olof Kraigher, Dake Liu |
DSD | 4 |
| 2010 | Implementation of a Floating Point Adder and Subtracter in NoGAP, A Comparative Case StudyabstractFlexible Application Specific Instruction-set Processors (ASIPs) are starting to replace monolithic Application Specific Integrated Circuits (ASICs) in a wide variety of fields. However the construction of an ASIP is today associated with a substantial design effort. Novel Generator of Accelerators And Processors (NoGap) is a tool for ASIP design utilizing hardware multiplexed data paths. One of the main advantages of NoGap compared to other EDA tools for processor design, is that NoGap imposes few limits on the architecture and thus design freedom. To prove that NoGap can be used to design complex data paths a reimplementation of a floating point adder/subtracter previously implemented using Verilog with FPGA specific optimizations was reimplemented using the NoGap Common Language (NoGapCL). The adder/subtracter implemented in Verilog can operate at a frequency of 377 MHz in a Virtex-4SX35 (speed grade -12) as compared with the NoGap implementation which had a maximum operation frequency of 276 Mhz, using the hand optimized mantissa adder from the original Verilog code, the NoGap implementation reached timing closure at 326 Mhz. Per Karlström, Wenbiao Zhou, Dake Liu |
EUC | 3 |
| 2010 | Architectural Support for Reducing Parallel Processing Overhead in an Embedded MultiprocessorabstractThe host-multi-SIMD chip multiprocessor (CMP) architecture has been proved to be an efficient architecture for high performance signal processing which explores both task level parallelism by multi-core processing and data level parallelism by SIMD processors. Different from the cache-based memory subsystem in most general purpose processors, this architecture uses on-chip scratchpad memory (SPM) as processor local data buffer and allows software to explicitly control the data movements in the memory hierarchy. This SPM-based solution is more efficient for predictable signal processing in embedded systems where data access patterns are known at design time. The predictable performance is especially important for real time signal processing. According to Amdahl's law, the nonparallelizable part of an algorithm has critical impact on the overall performance. Implementing an algorithm in a parallel platform usually produces control and communication overhead which is not parallelizable. This paper presents the architectural support in an embedded multiprocessor platform to maximally reduce the parallel processing overhead. The effectiveness of these architecture designs in boosting parallel performance is evaluated by an implementation example of 64×64 complex matrix multiplication. The result shows that the parallel processing overhead is reduced from 369% to 28%. Jian Wang 0035, Joar Sohl, Dake Liu |
EUC | 3 |
| 2009 | Large Matrix Multiplication on a Novel Heterogeneous Parallel DSP Architecture
Joar Sohl, Jian Wang 0035, Dake Liu |
APPT | 3 |
| 2009 | Memory Conflict Analysis and Interleaver Design for Parallel Turbo Decoding Supporting HSPA EvolutionabstractHSPA evolution has raised the requirements for WCDMA based systems where turbo code has been adapted to perform the error correction. Many parallel turbo decoding architectures have recently been proposed to enhance the channel throughput but the interleaving algorithm used in WCDMA based systems does not freely allows to use them due to high percentage of memory conflicts. This paper provides a comprehensive analysis for reduction of interleaver memory conflicts while generating more than one address in a single clock cycle. It also provides trade-off analysis in terms of area and power efficiency for multiple architectures for different functions involved in the interleaver design. The final architecture supports processing of two parallel SISO blocks and manages the conflicts by applying different approaches like stream misalignment, memory division and small FIFO buffer. The proposed architecture is low cost and consumes 4.3 K gates at a frequency of 150 MHz. This work also focuses on reduction of preprocessing overheads by introducing the segment based modulo computation, thus providing further relaxation to SISO decoding process. Rizwan Asghar, Di Wu 0003, Johan Eilert, Dake Liu |
DSD | 4 |
| 2009 | An ASIC perspective on FPGA optimizationsabstractIn this paper we discuss how various design components perform in both FPGAs and standard cell based ASICs. We also investigate how various common FPGA optimizations will effect the performance and area of an ASIC port. We find that most techniques that are used to optimize a design for an FPGA will not have a negative impact on the area in an ASIC. The intended audience for this paper are engineers charged with creating designs or IP cores that are optimized for both FPGAs and ASICs. Andreas Ehliar, Dake Liu |
FPL | 2 |
| 2009 | Low Complexity Hardware Interleaver for MIMO-OFDM based Wireless LANabstractA low complexity hardware interleaver architecture is presented for MIMO-OFDM based wireless LAN e.g. 802.11n. Novelty of the presented architecture is twofold; 1) Flexibility to choose interleaver implementation with different modulation scheme and different size for different spatial streams in a multi antenna system, 2) Complexity to compute on the fly interleaver address is reduce by using recursion and is supported by mathematical formulation. The proposed interleaver architecture is implemented on 65 nm CMOS process and it consumes 0.035 mm2area. The proposed architecture supports high speed communication with maximum throughput of 900 Mbps at a clock rate of 225 MHz. Rizwan Asghar, Dake Liu |
ISCAS | 2 |
| 2009 | Implementation Aspects of Fixed-Complexity Soft-Output MIMO DetectionabstractThis paper discusses implementation aspects of a recently proposed fixed-complexity soft-output (FCSO) symbol detector for MIMO systems (Larsson and Jalden, 2008). A further approximation to the FCSO detector is proposed which substantially reduces the complexity at the cost of a minor performance loss. With the resulting method, it is possible to carry out close-to ML detection for MIMO systems with a large number antennas (e.g. 4times4) using higher-order modulation schemes (e.g. 64-QAM) at low silicon cost in real-time. Furthermore, the parallelism inherited by the FCSO algorithm allows massive parallel processing which makes the method suitable for implementation in multi-core baseband signal processing hardware architectures. Di Wu 0003, Erik G. Larsson, Dake Liu |
VTC Spring | 3 |
| 2008 | A high performance microprocessor with DSP extensions optimized for the Virtex-4 FPGAabstractAs the use of FPGAs increases, the importance of highly optimized processors for FPGAs will increase. In this paper we present the microarchitecture of a soft microprocessor core optimized for the Virtex-4 architecture. The core can operate at 357 MHz, which is significantly faster than Xilinxpsila Microblaze architecture on the same FPGA. At this frequency it is necessary to keep the logic complexity down and this paper shows how this can be done while retaining sufficient functionality for a high performance processor. Andreas Ehliar, Per Karlström, Dake Liu |
FPL | 3 |
| 2008 | Implementation of a programmable linear MMSE detector for MIMO-OFDMabstractThis paper presents a linear minimum mean square error (LMMSE) symbol detector for MIMO-OFDM enabled mobile terminals. The detector is implemented using a programmable baseband processor aimed for software-defined radio (SDR). Owing to the dynamic range supplied by the floating-point SIMD datapath, special algorithms can be adopted to reduce the computational latency of detection. The programmable solution not only supports different transmit/receive antenna configurations, but also allows hardware multiplexing to obtain silicon and power efficiency. Compared to several existing fixed-functional solutions, the one proposed in this paper is smaller, more flexible and faster. Johan Eilert, Di Wu 0003, Dake Liu |
ICASSP | 3 |
| 2008 | Cost Analysis of Channel Estimation in MIMO-OFDM for Software Defined RadioabstractChannel state information (CSI) is critical for the overall performance of wireless systems. Meanwhile, the estimation of CSI forms one of the most intensive tasks in radio baseband signal processing. This paper investigates the real-time implementation of channel estimation for MIMO-OFDM systems using programmable hardware aimed for software defined radio. Based on the programmable hardware architecture proposed by us, several prevalent channel estimation methods such as Least Square (LS), Minimum Mean Square Error (MMSE) and Pilot-Symbol-Aided (PSA) are evaluated from both the performance and computational latency perspectives. By utilizing the symmetric feature of the covariance matrix, a simplified two-sided Jacobi rotation method is adopted to speed up the complex-valued singular value decomposition involved in the MMSE channel estimation. Di Wu 0003, Johan Eilert, Dake Liu |
WCNC | 4 |
| 2007 | An fpga based open source network-on-chip architectureabstractNetworks on Chip (NoC) has long been seen as a potential solution to the problems encountered when implementing large digital hardware designs. In this paper we describe an open source FPGA based NoC architecture with low area overhead, high throughput and low latency compared to other published works. The architecture has been optimized for Xilinx FPGAs and the NoC is capable of operating at a frequency of 260 MHz in a Virtex-4 FPGA. We have also developed a bridge so that generic Wishbone bus compatible IP blocks can be connected to the NoC. Andreas Ehliar, Dake Liu |
FPL | 2 |
| 2007 | Efficient Complex Matrix Inversion for MIMO Software Defined RadioabstractComplex matrix inversion is a very computationally demanding operation in advanced multi-antenna wireless communications. Traditionally, systolic array-based QR decomposition (QRD) is used to invert large matrices. However, the matrices involved in MIMO baseband processing in mobile handsets are generally small which means QRD is not necessarily efficient. In this paper, a new method is proposed using programmable hardware units which not only achieves higher performance but also consumes less silicon area. Furthermore, the hardware can be reused for many other operations such as complex matrix multiplication, filtering, correlation and FFT/IFFT. Johan Eilert, Di Wu 0003, Dake Liu |
ISCAS | 3 |
| 2007 | Lattice-Reduction Aided Multi-User STBC Decoding with Resource ConstraintsabstractLattice-reduction aided decoders have been proposed in MIMO system to achieve near maximum likelihood decoder performance while maintaining reasonable complexity. This paper studies the implementation of lattice-reduction aided linear decoders on a programmable device for multi-user space-time block coding (MU-STBC). By reloading software, the device can be configured to use different decoding schemes according to the amount of resources available, which is an important feature of cognitive radio. In this paper, two different lattice-reduction aided linear decoding methods namely SQRD-LR and AQRD-LR for MU-STBC are evaluated based on their BER performance and computational complexity. Furthermore, the effect of deadline constraint on LR is evaluated and based on the evaluation, a new method namely adaptive decoding is proposed by us to allow mode-switching of the decoder according to the environment parameters, so that the best decoder performance can always be achieved while fulfiling the resource constraints. Di Wu 0003, Johan Eilert, Dake Liu |
PIMRC | 3 |
| 2006 | Accelerating CABAC encoding for multi-standard media with configurabilityabstractThis paper presents the study of how to accelerate CABAC encoding for emerging heterogeneous multimedia applications. The latest image and video compression standards such as JPEG2000 and H.264 both have adopted context adaptive binary arithmetic coding to achieve performance enhancement. However, CABAC requires high computing power. After investigating computational complexity of CABAC coding, firstly, instruction level acceleration is elaborated. Secondly, a configurable accelerator for CABAC encoding in multiple standards is proposed. Benchmarking performance and implementation cost is also addressed. Oskar Flordal, Di Wu 0003, Dake Liu |
IPDPS | 3 |
| 2006 | MIPS cost estimation for OFDM-VBLAST systemsabstractThe focus of this paper is to investigate the feasibility of using programmable DSP processors for MIMO based radio systems. Several detection algorithms were evaluated and MIPS costs for low complexity detection and channel estimation algorithms for OFDM-VBLAST MIMO systems where calculated. Based on the MIPS cost estimation a feasible hardware architecture was derived. The result shows the feasibility of implementing MIMO radio systems in a programmable architecture Haiyan Jiao, Anders Nilsson 0001, Eric Tell, Dake Liu |
WCNC | 4 |
| 2004 | Reduced floating point for MPEG1/2 layer III decodingabstractA new approach to decode MPEG 1/2-layer III, mp3, is presented. Instead of converting the algorithm to fixed point, we propose a 16-bit floating point implementation. These 16 bits include 1 sign bit and 15 bits of both mantissa and exponent. The dynamic range is increased by using this 16-bit floating point as compared to both 24 and 32-bit fixed point. The 16-bit floating point is also suitable for fast prototyping. Usually, new algorithms are developed in 64-bit floating point. Instead of using scaling and double precision as in fixed point implementations we can use this 16-bit floating point easily. In addition, this format works well even for memory compiling. The intention of this approach is a fast, simple, low power, and low silicon area implementation for consumer products like cellular phones and PDAs. Both listening tests and tests versus the psychoacoustic model have been completed. Mikael Olausson, Andreas Ehliar, Johan Eilert, Dake Liu |
ICASSP (5) | 4 |
| 2004 | Using low precision floating point numbers to reduce memory cost for MP3 decodingabstractThe purpose of our work has been to evaluate the practicality of using a 16-bit floating point representation to store the intermediate sample values and other data in memory during the decoding of MP3 bit streams. A floating point number representation offers a better trade-off between dynamic range and precision than a fixed point representation for a given word length. Using a floating point representation means that smaller memories can be used which leads to smaller chip area and lower power consumption without reducing sound quality. We have designed and implemented a DSP processor based on 16-bit floating point intermediate storage. The DSP processor is capable of decoding all MP3 bit streams at 20 MHz and this has been demonstrated on an FPGA prototype. Johan Eilert, Andreas Ehliar, Dake Liu |
MMSP | 3 |
| 2003 | Implementation of fast CRC calculationabstractAbstract- CRC is important for error detection in communication systems. With transmission speeds of several Gb/s the highspeed implementation is a bottleneck. A circuit with two parallel calculation units has been implemented in a 0.35 micron process. They use 32 bits and 64 bits parallel input respectively. Chip measurements prove throughput higher than 5.76 Gb/s, which indicates that 10 Gb/s throughput is possible in more modern processes. I. Tomas Henriksson, Dake Liu |
ASP-DAC | 2 |
| 2003 | A general DSP processor at the cost of 23K gates and 1/2 a man-year design timeabstractThis paper describes the design and implementation of a 16-bit fixed point DSP processor. The processor is intended as a platform for hardware accelerators and allows additional computational units and assembler instructions to be added. The I/O facilities can also be customized to the needs of a specific application. Benchmarking has shown that the processor, without any hardware accelerators, has a performance comparable to single MAC commercial DSP processors. The architecture has been successfully synthesized in a 0.13 /spl mu/m process, resulting in a net-list of about 23000 gates, and a clock frequency of 195 MHz, making the performance/gate count ratio very competitive. It is also small enough to integrate 100 heterogeneous processors on a chip for example for communication infrastructure applications. The complete design time, including architecture and instruction set planning, assembler, debugger, instruction set simulator, RTL code and complete verification was about half a person-year. Eric Tell, Mikael Olausson, Dake Liu |
ICASSP (2) | 3 |
| 2002 | Embedded Protocol Processor for Fast and Efficient Packet ReceptionabstractComputer network equipment presents a bottleneck for further increasing the capacity in the networks. Terminals have problems keeping up with network speed when using general purpose processors for protocol processing. We present a novel processor architecture, that works in-line with the data flow and does not use a traditional von Neuman architecture. The program is contained in three lookup tables within the processor core, which allows for one cycle if-then-else and switch-case-case... execution. The processor is estimated to be able to handle a 10 Gb/s Ethernet connection when implemented in 0.18 micron technology. Tomas Henriksson, Ulf Nordqvist, Dake Liu |
ICCD | 3 |