Chunhua Xiao

dblp:49/4409 · DBLP profile ↗
← Back
26ranked-venue papers
11as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 8 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 MC-CGRA: A Memory-Computation Coordinated CGRA Framework for Stream Processing
abstract
Coarse-Grained Reconfigurable Arrays (CGRAs) have emerged as a promising platform for domain-specific accelerators. However, traditional designs face significant limitations in large-scale stream processing. Current decoupled software-hardware partitioning approaches often result in underutilized hardware parallelism and suboptimal memory organization, which severely constrains the scalability of kernel implementations. Consequently, large-scale kernels fail to fully exploit their inherent data locality, resulting in frequent off-chip memory accesses and overhead from repeated invocations, thereby substantially degrading overall performance. To address these challenges, this paper introduces MC-CGRA, a memory-computation coordinated CGRA framework. By leveraging its novel Chain-of-Computation (CoC) model, which uniformly represents operations within kernels as stream nodes, MC-CGRA achieves seamless coordination between memory access and pipelined computation through a software-defined approach. The framework incorporates a stream-centric CGRA microarchitecture to minimize frequent data exchanges between large-scale stream computing kernels and off-chip memory. An MC-CGRA prototype with an 8×10 PE array has been implemented on the AMD/Xilinx VCU118 platform. Experimental results show that the prototype combines fast compilation with sustained high throughput for stream processing kernels of varying scales, underscoring its efficiency in real-time scenarios. The prototype attains an average performance of 29.73 GOPS, outperforming state-of-the-art solutions by 1.55× and 1.62× in FFT and FIR workloads, respectively.
Chunhua Xiao, Han Diao, Weijie Yuan 0008
DATE2
2024 Adaptive Hybrid FFT: A Novel Pipeline and Memory-Based Architecture for Radix-2k FFT in Large Size Processing
abstract
In the field of digital signal processing, the fast Fourier transform (FFT) is a fundamental algorithm, with its processors being implemented using either the pipelined architecture, well-known for high-throughput applications but weak in hardware utilization, or the memory-based architecture, designed for area-constrained scenarios but failing to meet stringent throughput requirements. Therefore, we propose an adaptive hybrid FFT, which leverages the strengths of both pipelined and memory-based architectures. In this paper, we propose an adaptive hybrid FFT processor that combines the advantages of both architectures, and it has the following features. First, a set of radix-2kmulti-path delay commutators (MDC) units are developed to support high-performance large-size processing. Second, a conflict-free memory access scheme is formulated to ensure a continuous data flow without data contention. Third, We demonstrate the existence of a series of bit-dimension permutations for reordering input data, satisfying the generalized constraints of variable-length, high-radix, and any level of parallelism for wide adaptivity. Furthermore, the proposed FFT processor has been implemented on a field-programmable gate array (FPGA). As a result, the proposed work outperforms conventional memory-based FFT processors by requiring fewer computation cycles. It achieves higher hardware utilization than pipelined FFT architectures, making it suitable for highly demanding applications.
Fangyu Zhao, Chunhua Xiao, Xiaohua Du
ISPA2
2024 Optimizing Batched Small Matrix Multiplication on Multi-core DSP Architecture
abstract
General Matrix Multiplication (GEMM) is a critical computational operation in scientific computing and machine learning domains. While traditional GEMM performs well on large matrices, it is inefficient in terms of data transfer and computation for small matrices. Many High-Performance Computing (HPC) tasks can be decomposed into large batches of small matrix multiplication operations. Multi-core Digital Signal Processors (DSPs) are commonly used to accelerate high-performance computing. We present a design for batched fusion small matrix multiplication (BFMM) tailored for multi-core DSP architecture. To address the inefficiencies and redundancy in storage and computational operations associated with batch small matrix multiplications, we designed several strategies. We design a matrix fusion concatenation strategy, an access coordination mechanism, and a mechanism for fragment aggregation. BFMM supports an efficient K-dimension multi-core parallelization strategy. The parameter constraint model makes BFMM highly portable. BFMM also includes a performance evaluation model that facilitates assessment and verification. Experimental results demonstrate that, compared to traditional GEMM (TGEMM) on multi-core DSP and traditional GEMM with concatenated data access (TGEMM Op), BFMM exhibits superior performance. For large batches of small matrices, our design achieves 1.21x to 18x higher performance than TGEMM Op on single-core DSP, while on multi-core DSP, it outperforms TGEMM Op by 1.14x to 18.1x.
Xiaohan Zuo, Chunhua Xiao
ISPA2
2023 PBFL: Communication-Efficient Federated Learning via Parameter Predicting
abstract
Abstract Federated learning (FL) is an emerging privacy-preserving technology for machine learning, which enables end devices to cooperatively train a global model without uploading their local sensitive data. Because of limited network bandwidth and considerable communication overhead, communication efficiency has become an essential bottleneck for FL. Existing solutions attempt to improve this situation by reducing communication rounds while usually come with more computation resource consumption or model accuracy deterioration. In this paper, we propose a parameter Prediction-Based DL (PBFL). In which an extended Kalman filter-based prediction algorithm, a practical prediction error threshold setting mechanism and an effective global model updating strategy are included. Instead of collecting all updates from participants, PBFL takes advantage of predicting values to aggregate the model, which substantially reduces required communication rounds while guaranteeing model accuracy. Inspired by the idea of prediction, each participant checks whether its prediction value is out of the tolerance threshold limits and only uploads local updates that have an inaccurate prediction value. In this way, no additional local computational resources are required. Experimental results on both multilayer perceptrons and convolutional neural networks show that PBFL outperforms the state-of-the-art methods and improves the communication efficiency by >66% with 1% higher model accuracy.
Kaiju Li, Chunhua Xiao
Comput. J.2
2022 Cop-Flash: Utilizing hybrid storage to construct a large, efficient, and durable computational storage for DNN training
abstract
Traditional computing architectures that separate computing from storage face severe limitations when processing the data that is continuously produced in the cloud and at the edge. Recently, the computational storage device (CSD) is becoming one of the critical cloud infrastructures which can overcome these limitations. Many studies utilize CSD for DNN training to extract useful information and knowledge from the data quickly and efficiently. However, all previous work has used homogeneous storage, which is not fully considered the requirements of DNN training on CSD. Thus, we exploit the leverage of hybrid NAND flash memory to optimize this problem. Nevertheless, typical hybrid storage architectures have limitations when used for DNN training. Moreover, their management strategies can not fully exploit the heterogeneity of hybrid flash memory. To address this issue, we propose a novel SLC-TLC flash memory called Co-Partitioning Flash (Cop-Flash), which utilizes two different hybrid flash memory partitioning methods to divide storage into three different properties of flash memory. Meanwhile, two key technologies are included in Cop-Flash: 1) lifetime-based I/O identifier is proposed to identify data hotness according to data lifetime to maximize the benefits of heterogeneity and minimize the impact of garbage collection. 2) Erase-aware Adaptive Dual-zone Management is proposed to increase bandwidth utilization and guarantee system reliability. We compared Cop-Flash with two related state-of-the-art hybrid storage using hard partitioning and soft partitioning as well as TLC-only flash memory under real DNN training workloads. Experimental results show that Cop-Flash improves the performance by 29.1%, 38.8%, 56.6% and outperforms them by 2.3x, 1.29x, and 8.3x in terms of lifespan.
Chunhua Xiao, Dandan Xu
CLOUD1
2022 SDST-Accelerating GEMM-based Convolution through Smart Data Stream Transformation
abstract
The development of flexible Convolutional Neural Network (CNN) accelerators is critical for large-scale inference and training. Accelerators based on the General Matrix Multiplication (GEMM) kernel have gained popularity due to their ability to accelerate the most prevalent convolutional and fully connected layers in CNNs. However, the convolution inputs must be reshaped and packed into redundant matrices, which is performed by the im2col (image to column) algorithm. As the performance of the GEMM kernel improves, it increases latency and gradually becomes a bottleneck. To address this issue, we propose Smart Data Stream Transformation (SDST), a technique that eliminates explicit data transformation through data stream manipulation. SDST divides the input data into conflict-free streams based on the locality of data redundancy. Additionally, we design the continuity-friendly data layout to unify the transformations across data streams. Our design is evaluated by running the YoloV3-tiny model on an FPGA-based prototype system. Experimental results show that SDST improves the performance of convolutional acceleration by a factor of 1.12 to 5.69 compared to explicit im2col performed on the CPU.
Chunhua Xiao, Dandan Xu, Fangzhu Lin, Kun Ning
CCGRID1
2022 PASM: Parallelism Aware Space Management strategy for hybrid SSD towards in-storage DNN training acceleration
Chunhua Xiao, Shi Qiu 0012, Dandan Xu
J. Syst. Archit.1
2022 Contention Minimization in Emerging SMART NoC via Direct and Indirect Routes
abstract
SMART (Single-cycle Multi-hop Asynchronous Repeated Traversal) Network-on-Chip (NoC), a recently proposed dynamically reconfigurable NoC, enables single-cycle long-distance communication by building single-bypass paths directly between distant communication pairs. However, such a single-cycle single-bypass path will be readily broken when contention occurs. Thus, packets will be buffered at intermediate routers with blocking latency from other contending packets, and extra router-stage latency to rebuild the remaining path when available. In this article, we propose an effective contention-minimized routing algorithm to achieve maximal bypassing. Specifically, we identify two potential routes: direct route, with which packets can reach the destination in a single bypass; and indirect route, with which packets can reach the destination in multiple bypasses via an(multiple) intermediate router(s). The novel feature is that, contrary to an intuitive approach, not the routes with minimal distance but the indirect routes via the arbitrary intermediate routers (even if they may be non-minimal) that avoid contentions yield the minimized end-to-end latency. Evaluation on realistic benchmarks demonstrates the effectiveness of the proposed routing strategy, which achieves average performance improvement by 35.48 percent in communication latency, 28.31 percent in application schedule length, and 37.59 percent in network throughput, compared with the current routing in SMART NoCs.
Peng Chen 0027, Hui Chen 0016, Mengquan Li, Weichen Liu 0001, Chunhua Xiao, Yiyuan Xie, Nan Guan
IEEE Trans. Computers6
2022 Fast and Low Overhead Metadata Operations for NVM-Based File System Using Slotted Paging
abstract
Existing nonvolatile memory (NVM)-based file systems can fully leverage the characteristics of NVM to obtain better performance than traditional disk-based file systems. It has the potential capacity to efficiently manage metadata and perform fast metadata operations. However, most NVM-based file systems mainly focus on managing file metadata (inode), while pay little attention to directory metadata (dentry), which also has a noticeable impact on the file system performance. Besides, the traditional journaling technique that guarantees metadata consistency may not yield satisfactory performance on NVM-based file systems. To solve these problems, in this article we propose a fast and low overhead metadata operation mechanism, called FLOMO. It first adopts a novel slotted-paging structure in NVM to reorganize dentry for efficiently performing dentry operations, and utilizes the red–black tree in DRAM to accelerate dentry lookup and the search process of dentry deletion. Moreover, FLOMO presents a selective journaling scheme for metadata updates, which partially logs the changes related to dentry in the proposed slotted page, thereby, mitigating the redundant journaling overhead. To verify FLOMO, we implement it in a typical NVM-based file system, the persistent memory file system (PMFS). Experimental results show that FLOMO accelerates the metadata operations in PMFS by 34.4%$\sim 59$%, and notably reduces the journaling overhead for metadata, shortening the latency by 59% on average. For real-world applications, FLOMO has higher throughput compared with PMFS, PMFS without journal, and NOVA, achieving up to$2.1\times $,$1.1\times $, and$1.3\times $performance improvement, respectively.
Fangzhu Lin, Chunhua Xiao, Weichen Liu 0001, Kun Ning
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 A Light-Weight Deployment Methodology for DNN Re-training in Resource-Constraint Scenarios
Songtao Guo, Chunhua Xiao, Zilan Liao
WASA (1)3
2021 Contention-Aware Routing for Thermal-Reliable Optical Networks-on-Chip
abstract
Optical network-on-chip (ONoC) architecture offers ultrahigh bandwidth, low latency, and low power dissipation for new-generation manycore systems. However, the benefits in communication performance and energy efficiency will be diminished by communication contention. The intrinsic thermal susceptibility is another challenge for ONoC designs. Under on-chip temperature variations, core functional devices suffer from significant thermal-induced optical power loss, which seriously threatens ONoCs' reliability. In this article, we develop novel routing techniques to resolve both issues for ONoCs. By analyzing the thermal effect in ONoCs, we first present a routing criterion at the network level. Combined with device-level thermal tuning, it can implement thermal-reliable ONoCs. Two routing approaches, including a mixed-integer linear programming (MILP) model and a heuristic algorithm (called CAR), are further proposed to minimize communication conflicts based on guaranteed thermal reliability, and meanwhile, maximize the communication energy efficiency in the presence of on-chip thermal variations. By applying the criterion, our approaches achieve excellent performance with largely reduced complexity of design space exploration. The evaluation results based on both synthetic traffic patterns and realistic benchmarks validate the effectiveness of our approaches with an average of 126.95% improvement in communication performance and 16.12% reduction in energy overhead compared to state-of-the-art techniques. CAR only introduces 7.20% performance difference compared to the MILP model and is more scalable to large-size ONoCs.
Mengquan Li, Weichen Liu 0001, Luan H. K. Duong, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2020 Mobi-PMFS: An Efficient and Durable In-Memory File System for Mobile Devices
abstract
Emerging byte-addressable non-volatile memory (NVM) has the advantages of fast, cheap and persistent, and is considered as the next generation of persistent memory. However, existing NVM-based filesystems cannot adapt well for mobile devices, and not to mention the consideration for mobile application characteristics. In this paper, we propose an efficient and durable in-memory file system named as Mobi-PMFS for mobile devices. Proposed Mobi-PMFS is not only adaptive to ARM architecture, but also customized according to mobile application features. A wear-aware three-list space management scheme including a switching allocation algorithm is proposed to provide optimum performance for mobile systems while keeping the durability of NVM. Experimental results show that Mobi-PMFS is 8x times and 1.2x faster than EXT4-SSD and EXT4-DAX, and provides 11x wear-leveling improvement compared with the original PMFS.
Chunhua Xiao, Fangzhu Lin, Xiaoxiang Fu, Ting Wu 0012, Yuanjun Zhu, Weichen Liu 0001
COMPSAC1
2020 COSMA: An Efficient Concurrency-Oriented Space Management Scheme for In-memory File Systems
abstract
Emerging file systems have been designed for fully exploring NVM's advanced features. However, with the development of big data, these file systems suffer from the performance degradation in highly concurrent environment, which are caused by serious access conflicts in space management. To solve this problem, we propose an efficient concurrency-oriented space management scheme named as COSMA. Along with novel data structure design, COSMA is able to greatly reduce request congestion among multiple threads through hierarchical space allocation scheme. Furthermore, COSMA provides 3 reclamation strategies to improve space utilization, and can also adapt to different systems which varied in NVM capacities. To ensure the system reliability, COSMA is capable of keeping wear leveling among multiple NVMs slots. We implement COSMA in a representative persistent file system, PMFS. Experimental results show that COSMA can improve the IOPS of PMFS by 15%, the write throughput of PMFS by 7.6% and the concurrent processing performance of PMFS by 50 %. Besides, it can also achieve wear-leveling among multiple NVMs.
Chunhua Xiao, Zipei Feng, Ting Wu 0012, Xiaoxiang Fu, Weichen Liu 0001
ICCD1
2020 Separable Binary Convolutional Neural Network on Embedded Systems
abstract
We have witnessed the tremendous success of deep neural networks. However, this success comes with the considerable memory and computational costs which make it difficult to deploy these networks directly on resource-constrained embedded systems. To address this problem, we propose TaijiNet, a separable binary network, to reduce the storage and computational overhead while maintaining a comparable accuracy. Furthermore, we also introduce a strategy called partial binarized convolution which binarizes only unimportant kernels to efficiently balance network performance and accuracy. Our approach is evaluated on the CIFAR-10 and ImageNet datasets. The experimental results show that with the proposed TaijiNet, the separable binary versions of AlexNet and ResNet-18 can achieve 26× and 6.4× compression rates with comparable accuracy when comparing with the full-precision versions respectively. In addition, by adjusting the PCA threshold, the xnor version of Taiji-AlexNet improves accuracy by 4-8 percent comparing with other state-of-the-art methods.
Renping Liu 0002, Xianzhang Chen, Duo Liu 0002, Yingjian Ling, Weilue Wang, Yujuan Tan, Chunhua Xiao, Chaoshu Yang, Runyu Zhang 0002, Liang Liang 0002
IEEE Trans. Computers7
2019 Thermal Sensing Using Micro-ring Resonators in Optical Network-on-Chip
abstract
In this paper, we for the first time utilize the micro-ring resonators (MRs) in optical networks-on-chip (ONoCs) to implement thermal sensing without requiring additional hardware or chip area. The challenges in accuracy and reliability that arise from fabrication-induced process variations (PVs) and device-level wavelength tuning mechanism are resolved. We quantitatively model the intrinsic thermal sensitivity of MRs with finegrained consideration of wavelength tuning mechanism. Based on it, a novel PV-tolerant thermal sensor design is proposed. By exploiting the hidden ‘redundancy’ in wavelength division multiplexing (WDM) technique, our sensor achieves accurate and efficient temperature measurement with the capability of PV tolerance. Evaluation results based on professional photonic component and circuit simulations show an average of 86.49% improvement in measurement accuracy compared to the state-of-the-art on-chip thermal sensing approach using MRs. Our thermal sensor achieves stable performance in the ONoCs employing dense WDM with an inaccuracy of only 0.8650 K.
Weichen Liu 0001, Mengquan Li, Wanli Chang 0001, Chunhua Xiao, Yiyuan Xie, Nan Guan, Lei Jiang 0001
DATE4
2019 Wear-aware Memory Management Scheme for Balancing Lifetime and Performance of Multiple NVM Slots
abstract
Emerging Non-Volatile Memory (NVM) has many advantages, such as near-DRAM speed, byte-addressability, and persistence. Modern computer systems contain many memory slots, which are exposed as a unified storage interface by shared address space. Since NVM has limited write endurance, many wear-leveling techniques are implemented in hardware. However, existing hardware techniques can only effective in a single NVM slot, which cannot ensure wear-leveling among multiple NVM slots. This paper explores how to optimize a storage system with multiple NVM slots in terms of performance and lifetime. We show that simple integration of multiple NVMs in traditional memory policies results in poor reliability. We also reveal that existing hardware wear-leveling technologies are ineffective for a system with multiple NVM slots. In this paper, we propose a common wear-aware memory management scheme for in-memory file system. The proposed memory scheme enables wear-aware control of NVM slot use which minimizes the cost of performance and lifetime. We implemented the proposed memory management scheme and evaluated their effectiveness. The experiments show that the proposed wear-aware memory management scheme can outperform wear-leveling effect by more than 2600x, and the lifetime of NVM can be prolonged by 2.5x, the write performance can be improved by up to 15%.
Chunhua Xiao, Linfeng Cheng, Lei Zhang 0072, Duo Liu 0002, Weichen Liu 0001
MSST1
2019 Energy-efficient crypto acceleration with HW/SW co-design for HTTPS
Chunhua Xiao, Lei Zhang 0072, Weichen Liu 0001, Neil W. Bergmann, Yuhua Xie
Future Gener. Comput. Syst.1
2019 NV-eCryptfs: Accelerating Enterprise-Level Cryptographic File System with Non-Volatile Memory
abstract
The development of cloud computing and big data results in a large amount of data transmitting and storing. In order to protect sensitive data from leakage and unauthorized access, many cryptographic file systems are proposed to transparently encrypt file contents before storing them on storage devices, such as eCryptfs. However, the time-consuming encryption operations cause serious performance degradation. We found that compared with non-crypto file system EXT4, the performance slowdown could be up to 58.53 and 86.89 percent respectively for read and write with eCryptfs. Although prior work has proposed techniques to improve the efficiency of cryptographic file system through computation acceleration, no solution focused on the inefficiency working flow, which is demonstrated to be a major factor affecting system performance. To address this open problem, we present NV-eCryptfs, an asynchronous software stack for eCryptfs, which utilizes NVM as a fast storage tier on top of slower block devices to fully parallelize encryption and data I/O. We design an efficient NVM management scheme to support the fast parallel cryptographic operations. Besides providing an address space that can be directly accessed by the hardware accelerators, our designed mechanism is able to record the memory allocation states, and supplies a backup plan to deal with the situation of NVM shortage. The additional index structure is built to accelerate lookup operations to determine if a given data block resides in NVM. Moreover, we integrate an adaptive scheduling in NV-eCryptfs to process I/O requests dynamically according to access pattern and request size, which is able to take full utilization of both software and hardware acceleration to boost crypto performance. Our evaluation shows the proposed NV-eCryptfs outperforms the original eCryptfs with software routine 23.41× and 5.82× respectively for read and write.
Chunhua Xiao, Lei Zhang 0072, Weichen Liu 0001, Linfeng Cheng, Pengda Li, Yanyue Pan, Neil W. Bergmann
IEEE Trans. Computers1
2018 AEAS - Towards High Energy-efficiency Design for OpenSSL Encryption Acceleration through HW/SW Co-design
abstract
Entering the Big Data Era leads to the rapid development of web applications which provide high performance sensitive access on large data centers. OpenSSL has been widely deployed as a freely available implementation of SSL/TLS protocol that secures transactions over the Internet. In order to accelerate the speed of OpenSSL, many alternative encryption approaches are designed. However, energy consumption has been ignored in the rush for performance. Energy efficiency becomes a challenge with the increasing demands for performance and energy saving in data centers. In this paper, we present the Adaptive Encryption Acceleration System (AEAS), an OpenSSL encryption acceleration scheme. It provides high energy-efficiency encryption through HW/SW co-design. The essential idea is exerting the superiorities of energy efficiency for different encryption approaches and making full use of system resource through the Dynamic Management Mechanism including RequestAllocation algorithm and DynamicScheduler algorithm. Specifically, this scheme supports instruction set and hardware to process the computation compatibly by the Adaptive Control Crypto (ac_crypto) engine. Experimental results show that AEAS can improve energy efficiency by up to 933.5%, 68.8%, and 483.7% comparing with software, AES-NI and QAT, respectively.
Chunhua Xiao, Yuhua Xie, Lei Zhang 0072
ACM Great Lakes Symposium on VLSI1
2018 User Experience-Enhanced and Energy-Efficient Task Scheduling on Heterogeneous Multi-Core Mobile Systems
abstract
Heterogeneous Multi-Core Mobile Systems has been widely used to improve performance. However, it faces with the challenge of tradeoff between energy saving and user experience. ARM big. LITTLE architecture, a heterogeneous computing architecture, is a power-optimization technology. In most big. LITTLE devices, however, it still cannot achieve excellent user experience and higher energy saving. In this paper, we propose an improved task scheduling (UCES-GTS) by introducing the concept of user-centric task on big. LITTLE mobile device. In order to enhance user experience, the response time of user-centric tasks is shortened with reducing slack time of them properly. We then present a detailed algorithm to compute appropriate frequency and allocate the CPU resources to each task. The experimental evaluation results show that our improved global task scheduling model can achieve 17 % and 8 % energy saving average compared with the clustered switching scheduling and the original global task scheduling respectively. And the response time of user-centric tasks can decrease 27 % average, which means excellent user experience.
Weichen Liu 0001, Mengquan Li, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao, Yaoyao Ye
ICPADS6
2018 Fine-Grained Task-Level Parallel and Low Power H.264 Decoding in Multi-Core Systems
abstract
In the past few years, the extinction of Moore's Law makes people reconsider the solutions for dealing with the low computing resource utilization of applications on multicore processor systems. However, making good use of computing resources in multi-core processors systems is not easy due to the differences between single-core and multi-core architecture. Nowadays short video apps like Instagram and Tik Tok have successfully caught people's eyes by fascinating short videos, typically just 10 to 30 seconds long, uploaded by the users of apps. And almost all of these videos are recorded by their mobile devices, which are typically HD (High Definition) or FHD (Full High Definition) videos, which prefer to be encoded/decoded by H.264/AVC rather then HEVC (High Efficiency Video Coding) on mobile devices in view of the energy consumption and decoding speed. How to dive the huge potential of the computing resource on multi-core mobile devices to speed up decoding these videos while consuming low energy, is a big challenge. In our previous work [1], a relatively simple parallel framework was proposed to implement a parallel H.264/ AV C decoder. This work further proposes a more detailed systematic task-level parallel framework, together with an energy saving strategy based on this framework, to research a new H.264/AVC decoder on multi-core processor systems. The proposed parallel method is composed of a set of rules to guide parallel software programming (PSPR) and a software parallelization framework (SPF). The PSPR is applied in pre-processing steps to address the potential issues limiting the inherent parallelism, and the SPF is applied to parallelize the original serial programs. After the parallelization is successfully deployed, DVFS technique would be applied to decrease the power dissipation based on the SPF. Results show that proposed solutions make a significant improvement in decoding speed of 32% at 720p, 27% at 1080p and 29% at 2160p, and in energy savings of 25% at 720p, 25% at 1080p and 23% at 2160p on a four-core workstation running Linux, compared to the original serial H.264/ AV C decoder. The results demonstrate our methods are effective and scalable, served as a reference for future parallel software development.
Wenyang Liu, Weichen Liu 0001, Mengquan Li, Peng Chen 0027, Lei Yang 0018, Chunhua Xiao, Yaoyao Ye
ICPADS6
2018 Hardware/Software Adaptive Cryptographic Acceleration for Big Data Processing
abstract
Along with the explosive growth of network data, security is becoming increasingly important for web transactions. The SSL/TLS protocol has been widely adopted as one of the effective solutions for sensitive access. Although OpenSSL could provide a freely available implementation of the SSL/TLS protocol, the crypto functions, such as symmetric key ciphers, are extremely compute-intensive operations. These expensive computations through software implementations may not be able to compete with the increasing need for speed and secure connection. Although there are lots of excellent works with the objective of SSL/TLS hardware acceleration, they focus on the dedicated hardware design of accelerators. Hardly of them presented how to utilize them efficiently. Actually, for some application scenarios, the performance improvement may not be comparable with AES-NI, due to the induced invocation cost for hardware engines. Therefore, we proposed the research to take full advantages of both accelerators and CPUs for security HTTP accesses in big data. We not only proposed optimal strategies such as data aggregation to advance the contribution with hardware crypto engines, but also presented an Adaptive Crypto System based on Accelerators (ACSA) with software and hardware codesign. ACSA is able to adopt crypto mode adaptively and dynamically according to the request character and system load. Through the establishment of 40 Gbps networking on TAISHAN Web Server, we evaluated the system performance in real applications with a high workload. For the encryption algorithm 3DES, which is not supported in AES-NI, we could get about 12 times acceleration with accelerators. For typical encryption AES supported by instruction acceleration, we could get 52.39% bandwidth improvement compared with only hardware encryption and 20.07% improvement compared with AES-NI. Furthermore, the user could adjust the trade-off between CPU occupation and encryption performance through MM strategy, to free CPUs according to the working requirements.
Chunhua Xiao, Lei Zhang 0072, Yuhua Xie, Weichen Liu 0001, Duo Liu 0002
Secur. Commun. Networks1
2016 An Efficient Technique of Application Mapping and Scheduling on Real-Time Multiprocessor Systems for Throughput Optimization
abstract
Multiprocessor systems are becoming ubiquitous in today’s embedded systems design. In this article, we address the problem of mapping an application represented by a Homogeneous Synchronous Dataflow (HSDF) graph onto a real-time multiprocessor platform with the objective of maximizing total throughput. We propose that the optimal solution to the problem is composed of three components: actor-to-processor mapping, retiming, and actor ordering on each processor. The entire problem is systematically modeled into a Boolean Satisfiability (SAT) problem such that the optimal solution can be guaranteed theoretically. In order to explore the vast solution space more efficiently, we develop a specific HSDF theory solver based on the special characteristics of the timed HSDF, and integrate it into the general search framework of the SAT solver. Two alternative integration methods based on branch-and-bound are presented to achieve early branch pruning in the search space; thus, the scalability is greatly improved. Extensive performance evaluation on synthetic examples and a case study on the realistic H.264 Video Decoder show that our approach provides as much as 76.9% throughput improvement, and is scalable to industry-sized applications.
Weichen Liu 0001, Chunhua Xiao
ACM Trans. Embed. Comput. Syst.2
2013 Stream arbitration: Towards efficient bandwidth utilization for emerging on-chip interconnects
abstract
Alternative interconnects are attractive for scaling on-chip communication bandwidth in a power-efficient manner. However, efficient utilization of the bandwidth provided by these emerging interconnects still remains an open problem due to the spatial and temporal communication heterogeneity. In this article, a Stream Arbitration scheme is proposed, where at runtime any source can compete for any communication channel of the interconnect to talk to any destination. We apply stream arbitration to radio frequency interconnect (RF-I). Experimental results show that compared to the representative token arbitration scheme, stream arbitration can provide an average 20% performance improvement and 12% power reduction.
Chunhua Xiao, Mau-Chung Frank Chang, Jason Cong, Michael Gill, Zhangqin Huang, Chunyue Liu, Glenn Reinman, Hao Wu 0026
ACM Trans. Archit. Code Optim.1
2011 A framework of multi-characteristics fuzzy dynamic scheduling for parallel video processing on MPSoC architecture
abstract
This paper addresses the inherent unreliability and instability of the multiple uncertain characteristics of complex embedded multiprocessor systems, such as MPSoC (multi-processor system on chip) systems. In this work, we propose a fuzzy sets description for the multiple uncertain characteristics of system, and using fuzzy set membership calculation to determine the scheduling priorities of tasks and resources, in order to improve the capability of concurrent executions of tasks. And also we present a method estimation of comprehensive analysis on the multiple performances in order to increase utilization factor and balancing loads on processors. Through simulation, we demonstrate that the fuzzy dynamic scheduling algorithm can handle wide variety requirements of multiple performances, and we also show an improved approach to solve its demerit of partial adjustment. We implemented a prototype MPSoC system, which utilizes the proposed fuzzy dynamic scheduling algorithm to allocating tasks onto the multiprocessors. And in our case studies, we design a parallel architecture h.264 encoder, running on a multicore MPSoC system on FPGA. The speedup ratio of this prototype application system is up to 12.69.
Da Li 0004, Yibin Hou, Zhangqin Huang, Chunhua Xiao
FUZZ-IEEE4
2005 Estimating wheat grain protein content from ground-based hyperspectral data using a improved detecting method
Yanli Lu, Shaokun Li, Ruizhi Xie, Shiju Gao, Keru Wang, Chunhua Xiao
IGARSS7