Wenping Zhu

dblp:144/0494 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-3276-4019ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 High-throughput Side-Channel-Protected Stream Cipher Hardware for 6G Systems
Yuluan Cao, Cankun Zhao, Bohan Yang 0001, Wenping Zhu, Hanning Wang, Min Zhu 0001, Leibo Liu
ACNS (3)4
2025 Chameleon-SAT: An Adaptive Boolean Satisfiability Accelerator Using Mixed-Signal In-Memory Computing for Versatile SAT Problems
abstract
Boolean satisfiability (SAT), the first proven nondeterministic polynominal-complete problem, is crucial in dataintensive applications. Different applications have a wide spectrum of SAT problem sets (scale, complexity) and also various solution requirements (algorithm completeness, speed). Current SAT solvers are insufficient for providing ideal solutions under different scenarios. This work presents the Chameleon-SAT, the first ASIC-based SAT accelerator that can support local search, Davis-Putnam- Logemann-Lovel, Conflict-Driven Clause Learning algorithms, while leveraging the efficient mixed-signal inmemory computing architecture to achieve orders-of-magnitude improvements in speed compared to the prior SAT solvers. By judiciously selecting the reconfiguration mode, Chameleon-SAT is able to solve a wide range of the SAT problems to achieve smallscale, high-complexity cases ($\geq 90 \times$ for 20 variables/ 86 clauses, satisfiable problems), medium-scale, structured cases ($\geq 19 \times$ for 50 variables/ 215 clauses, unsatisfiable problems), and largescale, high-complexity cases ($\geq 7 \times$ for 100 variables/ 430 clauses, satisfiable problems).
Iris Ying Chou, Hao Kong 0003, Yi Huang 0036, Jianfeng Zhu 0001, Wenping Zhu, Shaojun Wei, Aoyang Zhang, Leibo Liu
DAC5
2025 EFFACT: A Highly Efficient Full-Stack FHE Acceleration Platform
abstract
Fully Homomorphic Encryption (FHE) is a set of powerful cryptographic schemes that allows computation to be performed directly on encrypted data with an unlimited depth. Despite FHE’s promising in privacy-preserving computing, yet in most FHE schemes, ciphertext generally blows up thousands of times compared to the original message, and the massive amount of data load from off-chip memory for bootstrapping and privacy-preserving machine learning applications (such as HELR, ResNet-20), both degrade the performance of FHE-based computation. Several hardware designs have been proposed to address this issue, however, most of them require enormous resources and power. An acceleration platform with easy programmability, high efficiency, and low overhead is a prerequisite for practical application. This paper proposes EFFACT, a highly efficient full-stack FHE acceleration platform with a compiler that provides comprehensive optimizations and vector-friendly hardware. We start by examining the computational overhead across different real-world benchmarks to highlight the potential benefits of reallocating computing resources for efficiency enhancement. Then we make a design space exploration to find an optimal SRAM size with high utilization and low cost. On the other hand, EFFACT features a novel optimization named streaming memory access which is proposed to enable high throughput with limited SRAMs. Regarding the software-side optimization, we also propose a circuit-level function unit reuse scheme, to substantially reduce the computing resources without performance degradation. Moreover, we design novel NTT and automorphism units that are suitable for a cost-sensitive and highly efficient architecture, leading to low area. For generality, EFFACT is also equipped with an ISA and a compiler backend that can support several FHE schemes like CKKS, BGV, and BFV. We provide both FPGA and ASIC versions of EFFACT. On account of our full stack design, FPGA-EFFACT outperforms the SOTA FPGA accelerators in gmean by $1.22 \times$. Meanwhile, ASIC-EFFACT shows increased improvements in terms of the performance per chip area and the performance per Watt compared with the SOTA ASIC works.
Yi Huang 0036, Xinsheng Gong, Dibei Chen, Jianfeng Zhu 0001, Wenping Zhu, Liangwei Li, Mingyu Gao 0001, Shaojun Wei, Aoyang Zhang, Leibo Liu
HPCA6
2023 Mckeycutter: A High-throughput Key Generator of Classic McEliece on Hardware
abstract
Classic McEliece is a code-based quantum-resistant public-key scheme characterized with relative high encapsulation/decapsulation speed and small ciphertexts, with an in-depth analysis on its security. However, slow key generation with large public key size make it hard for wider applications. Based on this observation, Mckeycutter, a high-throughput key generator in hardware, is proposed to accelerate the key generation in Classic McEliece based on algorithm-hardware co-design. Meanwhile the storage overhead caused by large-size keys is also minimized. First, compact large-size GF(2) Gauss elimination method is presented by adopting naive processing array and memory-friendly scheduling strategy. Second, an optimized constant-time hardware sorter is proposed to support regular memory accesses with less comparators and storage. Third, algorithmlevel pipeline is enabled for high-throughput processing, allowing for concurrent key generations. Our FPGA implementation results achieve around 4× improvements in throughput with 9~14× less memory-time product compared with the existing FPGA solutions.
Yihong Zhu, Wenping Zhu, Chen Chen 0083, Min Zhu 0001, Zhengdong Li, Shaojun Wei, Leibo Liu
DAC2
2022 BitCluster: Fine-Grained Weight Quantization for Load-Balanced Bit-Serial Neural Network Accelerators
abstract
Convolutional neural network (CNN) has demonstrated great success in pattern recognition scenarios at the cost of nearly billions of parameters and consequent convolution operations. Various dedicated hardware designs are proposed to accelerate the CNN computation in more energy-efficient manners. Especially, the bit-serial accelerator (BSA) is one of the most effective approaches on resource-limited platforms by eliminating zero-bit computations. However, the irregular distribution and varying number of effectual (nonzero) bits in weights significantly cause hardware underutilization, impeding further performance improvement of state-of-the-art BSAs. To address this issue, BitCluster, a hardware-friendly quantization method, is proposed to make each weight with the identical number of effectual bits for load-balanced computation. Considering distinct sensitivities to weight precision in different neural layers, layer-level BitCluster is proposed to design further for fine-grained weight quantization. It systematically determines the layerwise quantization configurations, which significantly improve the overall performance with$1.6\times $higher hardware utilization and$3.4\times $speedup on average than state-of-the-art BSAs, with$5\times $better energy efficiency on average.
Ang Li 0033, Huiyu Mo, Wenping Zhu, Eric Q. Li, Shouyi Yin, Shaojun Wei, Leibo Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2021 LWRpro: An Energy-Efficient Configurable Crypto-Processor for Module-LWR
abstract
Saber, the only module-learning with rounding-based algorithm in NIST's third round of post-quantum cryptography (PQC) standardization process, is characterized by simplicity and flexibility. However, energy-efficient implementation of Saber is still under investigation since the commonly used number theoretic transform can not be utilized directly. In this manuscript, an energy-efficient configurable crypto-processor supporting multi-security-level key encapsulation mechanism of Saber, is proposed. First, an 8-level hierarchical Karatsuba framework is utilized to reduce degree-256 polynomial multiplication to the coefficient-wise multiplication. Second, a hardware-efficient Karatsuba scheduling strategy and an optimized pre-/post-processing structure is designed to reduce the area overheads of scheduling strategy. Third, a task-rescheduling-based pipeline strategy and truncated multipliers are proposed to enable fine-grained processing. Moreover, multiple parameter sets are supported in LWRpro to enable configurability among various security scenarios. Enabled by these optimizations, LWRpro requires 1066, 1456 and 1701 clock cycles for key generation, encapsulation, and decapsulation of Saber768. The post-layout version of LWRpro is implemented with TSMC 40 nm CMOS process within 0.38 mm2. The throughput for Saber768 is up to 275k encapsulation operations per second and the energy efficiency is 0.15 uJ/encapsulation while operating at 400 MHz, achieving nearly 50× improvement and 31× improvement, respectively compared with current PQC hardware solutions.
Yihong Zhu, Min Zhu 0001, Bohan Yang 0001, Wenping Zhu, Chenchen Deng, Chen Chen 0083, Shaojun Wei, Leibo Liu
IEEE Trans. Circuits Syst. I Regul. Pap.4
2021 A 460 GOPS/W Improved Mnemonic Descent Method-Based Hardwired Accelerator for Face Alignment
abstract
The mnemonic descent method (MDM) algorithm is the first end-to-end recurrent convolutional system for high-accuracy face alignment. However, the heavy computational complexity and high memory access demands make it difficult to satisfy the requirements of real-time applications. To address this problem, an improved MDM (I-MDM) algorithm is proposed for efficient hardware implementation based on several hardware-oriented optimizations. First, a patch merging mechanism is introduced to dynamically cluster and eliminate redundant landmarks, which significantly reduces computational complexity with minimal accuracy loss. Second, a dedicated convolutional layer is inserted to halve the number of computations and memory access of the subsequent fully connected layer, yielding a 4.42% decrease in the failure rate. Third, a lightweight preprocessing method named dual regressors is proposed to reinitialize face images, which can greatly improve the overall accuracy. Moreover, compared with a similar method, the DR method can reduce computations and memory storage by nearly 99.9%. Overall and compared with the MDM algorithm, I-MDM not only reduces the number of computations by 23.5% but also decreases the failure rate by 17.9% on the 300 W test set. Based on the proposed I-MDM algorithm, an I-MDM-based hardwired accelerator is presented using the TSMC 65 nm CMOS process. First, compared with similar solutions, the gradient calculation operation is rearranged and loaded pixels are reused in the HoG feature extraction to eliminate all division operations and 25% off-chip memory access. Second, patch-independent central activations are used to enable patch-level pipelined operations, yielding a 2× acceleration in the overall process. This accelerator achieves 460 GOPS/W energy efficiency at 330 MHz, which is 38× higher than the most recent face alignment accelerator with the same process.
Huiyu Mo, Leibo Liu, Wenping Zhu, Eric Q. Li, Shouyi Yin, Shaojun Wei
IEEE Trans. Multim.3
2020 TFE: Energy-efficient Transferred Filter-based Engine to Compress and Accelerate Convolutional Neural Networks
abstract
Although convolutional neural network (CNN) models have greatly enhanced the development of many fields, the untenable number of parameters and computations in these models yield significant performance and energy challenges in hardware implementations. Transferred filter-based methods, as very promising techniques that have not yet been explored in the architecture domain, can substantially compress CNN models. However, their straightforward hardware implementation inherently incurs massive redundant computations, causing significant energy and time consumption. In this work, a highly efficient transferred filter-based engine (TFE) is developed to alleviate this deficiency, with CNN models compressed and accelerated. First, the filters of CNN models are flexibly transferred according to specific tasks to reduce the model size. Then, two hardware-friendly mechanisms are proposed in the TFE to remove duplicate computations caused by transferred filters, which can further accelerate transferred CNN models. The first mechanism exploits the shared weights hidden in each row of transferred filters and reuses the corresponding same partial sums, reducing at least 25% of repetitive computations in each row. The second mechanism can intelligently schedule and access the memory system to reuse the repetitive partial sums among different rows of the transferred filters with at least 25% of computations eliminated. Furthermore, an efficient hardware architecture is proposed in the TFE to fully reap the benefits of the two proposed mechanisms such that different types of networks are flexibly supported. To achieve high energy efficiency, the sub-array-based filter mapping method (SAFM) is proposed, where the process element (PE) subarray is used as the elementary computational unit to support various filters. Therein, input data can be efficiently broadcast in each PE sub-array and the load can be stripped from each PE and intensively alleviated, which can dramatically reduce the area and power consumption. Excluding MobileNet-like networks that adopt depth-wise convolution, most mainstream networks can be compressed and accelerated by the proposed TFE. Two state-of-the-art transferred filter-based methods, i.e., doubly CNN and symmetry CNN are implemented by exploiting the TFE. Compared with Eyeriss, average speedup improvements of 2.93× and 3.17× are achieved in the convolutional layers of various modern CNNs. The overall energy efficiency can be improved by 12.66× and 13.31× on average. Compared with other state-of-the-art related works, the TFE can maximally achieve a parameter reduction of 4.0×, a speedup of 2.72× and an energy efficiency improvement of 10.74× on VGGNet.
Huiyu Mo, Leibo Liu, Wenjing Hu, Wenping Zhu, Eric Q. Li, Ang Li 0033, Shouyi Yin, Xiaowei Jiang, Shaojun Wei
MICRO4
2020 Deterministic conversion rule for CNNs to efficient spiking convolutional neural networks
Xu Yang 0017, Wenping Zhu, Shuangming Yu, Nanjian Wu
Sci. China Inf. Sci.3
2020 A Multi-Task Hardwired Accelerator for Face Detection and Alignment
abstract
Face detection and alignment are two fundamental tasks for facial applications and the corresponding accelerators have been designed to enable energy-efficient acceleration. However, these dedicated accelerators are always designed separately, thereby ignoring the inherent correlation between face detection and alignment and causing additional communication and area overhead. Based on this motivation, a multi-task cascaded convolutional networks (MTCNN) algorithm-based accelerator is presented in this work to support both face detection and alignment for multiple faces. First, multiply-accumulate (MAC) operations and memory access of the magnification process in the resize module are reduced by 22.8% and 24.8% on average, respectively, when compared with those of similar methods. Second, clustering non-maximum suppression (C-NMS) is proposed to significantly reduce the intersection over union computation and eliminate the hardware-inference sorting process in NMS, yielding a 16.0% speedup in the overall process. Third, an efficient pipeline architecture is proposed to implement a complexity- and memory-intensive proposal network of MTCNN in a more computationally efficient manner, with 38.3% less memory capacity than a similar solution. Meanwhile, only approximately half of the multipliers are needed to achieve the same throughput with high pipeline utilization. Fourth, considering the variable number of faces in each input, a batch schedule mechanism is proposed to improve the fully-connected layer hardware utilization by 16.7% on average in the batch process. Based on a simulation with the TSMC 28 nm CMOS process, this accelerator consumes only 10.9ms at 400 MHz to simultaneously process 5 faces. The power efficiency reaches 4.80 TOPS/W, which is$227.4\times $higher than that of the state-of-the-art solution.
Huiyu Mo, Leibo Liu, Wenping Zhu, Eric Q. Li, Shouyi Yin, Shaojun Wei
IEEE Trans. Circuits Syst. Video Technol.3
2019 L-MPC: A LUT based Multi-Level Prediction-Correction Architecture for Accelerating Binary-Weight Hourglass Network
abstract
A binary-weight hourglass network (B-HG) accelerator for landmark detection, built on the proposed look-up-table (LUT) based multi-level prediction-correction approach, is enabled for high-speed and energy-efficient processing on IoT edge devices. First, LUT with a unified mode is adopted to support convolutional neural network with fully variable weight bit precision to minimize operations of B-HG, which achieves 1.33×-1.50× speedup on multi-bit weight CNN relative to the similar solution. Second, multi-level prediction-correction model is proposed to achieve computational-efficient convolution with adaptive precision. The operations saved can be increase by about 30% than the two-stage model. Besides, nearly 77.4% of the operations in B-HG can be saved by using the combination of these two methods, yielding a 2.3× inference speedup. Third, block computing based pipeline is designed to improve the residual block deficiency in B-HG. It can not only reduce about 66.2% off-chip memory access than the baseline, but also save 60% and 31% on-chip memory space and access compared to the similar fused-layer accelerator. The proposed B-HG accelerator achieves 450 fps at 500MHz based on the simulation in TSMC 28 nm process. Meanwhile, the power efficiency is up to 8.5 TOPS/W, which is two orders of magnitude higher than the dedicated face landmark detection accelerator.
Leibo Liu, Wenping Zhu, Eric Q. Li, Huiyu Mo, Shaojun Wei
DAC3
2019 A 1.17 TOPS/W, 150fps Accelerator for Multi-Face Detection and Alignment
abstract
Face detection and alignment are highly-correlated, computation-intensive tasks, without being flexibly supported by any facial-oriented accelerator yet. This work proposes the first unified accelerator for multi-face detection and alignment, along with the optimizations on multi-task cascaded convolutional networks algorithm, to implement both multi-face detection and alignment. First, the clustering non-maximum suppression is proposed to significantly reduce intersection over union computation and eliminate the hardware-interfer-ence sorting process, bringing 16.0% speed-up without any loss. Second, a new pipeline architecture is presented to implement the proposal network in more computation-efficient manner, with 41.7% less multiplier usage and 38.3% decrease in memory capacity compared with the similar method. Third, a batch schedule mechanism is proposed to improve hardware utilization of fully-connected layer by 16.7% on average with variable input number in batch process. Based on the TSMC 28 nm CMOS process, this accelerator only consumes 6.7ms at 400 MHz to simultaneously process 5 faces for each image and achieves 1.17 TOPS/W power efficiency, which is 54.8× higher than the state-of-the-art solution.
Huiyu Mo, Leibo Liu, Wenping Zhu, Eric Q. Li, Wenjing Hu, Shaojun Wei
DAC3
2019 A Binary-Feature-Based Object Recognition Accelerator With 22 M-Vector/s Throughput and 0.68 G-Vector/J Energy-Efficiency for Full-HD Resolution
abstract
Considering that the binary-feature-based approximate nearest neighbor (ANN) search technique has not been fully exploited to date, a multisegment binary feature-based hierarchical clustering tree model is proposed to achieve fast binary feature matching (FM). In addition, the multisegment vocabulary forest, is developed for the ease of hardware-oriented implementation. During the ANN searching process, the corresponding leaf nodes of each segment of the query feature are returned simultaneously to improve processing speed and accuracy. Furthermore, a hierarchical decomposition based on the term frequency-inverse document frequency is used to reduce the run-time search space and total memory footprint for object database storage. Finally, a fine-grained feature-level fully pipelined object recognition accelerator is implemented based on a dedicated design between FM and object scoring. The performance of the proposed object recognition accelerator is evaluated based on TSMC 65 nm CMOS technology. The accelerator achieves 22 M-vec/s and 6.8 × 108vec/J in throughput and energy efficiency for full-HD resolution, respectively; these results represent a 10.6× and 9× improvement, respectively, relative to current state-of-the-art solutions. The average power consumption is 32.6 mW when operating at 200 MHz.
Leibo Liu, Wenping Zhu, Shouyi Yin, Shaojun Wei
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 A Face Alignment Accelerator Based on Optimized Coarse-to-Fine Shape Searching
abstract
The coarse-to-fine shape searching (CFSS) framework is a recently developed algorithm that achieves relatively high accuracy in face alignment by alleviating the poor initialization problem facing traditional cascaded regression approaches. However, its high computational complexity and memory access demands make it difficult for CFSS to satisfy the requirements of real-time processing. To address this issue, a fast shape searching face alignment (F-SSFA) accelerator is presented based on the optimization of the CFSS algorithm and an efficient hardware implementation. First, the learning-based low-dimensional speeded-up robust features method, based on the correlations between the SURF features and the regression targets, is introduced to distill the feature set down to the only most distinct features to reduce the computing load. Second, the partial keypoints Euclidean distance and shape affine transformation are introduced to replace feature extraction and support vector machine classification, thereby accelerating the shape searching process. Compared with CFSS, F-SSFA achieves a $5.8\times $ speedup while achieving similar accuracy. Moreover, a VLSI architecture is proposed to realize the fixed-point F-SSFA algorithm. Multiple descriptors located in adjacent regions are simultaneously generated in a single access to the corresponding image data. Therefore, repeated memory access operations are avoided. The optimal parameter configuration for hardware implementation is also exploited based on a tradeoff between accuracy and hardware performance. Simulated with TSMC 65-nm 1P8M technology within a 3.6 mm2area, a post-layout simulation shows that 700 fps can be achieved while consuming 300 mW at 200 MHz.
Leibo Liu, Wenping Zhu, Huiyu Mo, Shouyi Yin, Yiyu Shi 0001, Shaojun Wei
IEEE Trans. Circuits Syst. Video Technol.3
2019 Face Alignment With Expression- and Pose-Based Adaptive Initialization
abstract
Face alignment is a critical task in many multimedia and vision applications that use face-based algorithms. Recent research has focused on achieving efficient initialization to improve performance; however, the use of facial attributes and the extent of their correlation with initialization have not been fully exploited. This paper presents a lightweight method called expression- and pose-based adaptive initialization (EXPAI), in which facial attributes, that is, expression and pose information, are used as priors. This approach can significantly improve the face alignment performance. In addition, reliable expression and head pose information can be derived simultaneously in the same framework. First, an expression- and pose-based template dictionary is formed by augmenting the mean shape across three degrees of freedom, thereby substantially improving the robustness of the initial shape with respect to large head pose variations. Second, each the template corresponds to an image of interest, which is jointly determined using a shape-constrained multiclass classifier and binary classifiers, and is assigned a pretrained confidence coefficient. The initial shape that is thus generated for subsequent cascaded regression is more adaptive and enables higher accuracy. Furthermore, EXPAI enables initialization with significantly increased computational efficiency because of its independence from the original dataset. The experimental results obtained on the widely used 300-W dataset show that our method achieves very competitive performance compared with that of state-of-the-art methods. In particular, for the challenging subset of 300-W, EXPAI reduces errors by more than 14% compared with coarse-to-fine shape searching (CFSS), which currently exhibits the best performance among regression-based approaches. Furthermore, a speed increase of more than 10 times compared with CFSS is achieved.
Huiyu Mo, Leibo Liu, Wenping Zhu, Shouyi Yin, Shaojun Wei
IEEE Trans. Multim.3
2017 A 700fps Optimized Coarse-to-Fine Shape Searching Based Hardware Accelerator for Face Alignment
abstract
In this work, a fast shape searching face alignment (F-SSFA) algorithm based accelerator is proposed to achieve real-time processing. Firstly, a learning based low-dimensional SURF feature is introduced to reduce the computation cost in the cascaded regression. Then the Euclidean distance and shape affine transformation are utilized to accelerate the shape searching procedure. F-SSFA therefore greatly reduces the computational complexity while keeping the same accuracy. Also, a fixed-point F-SSFA based VLSI architecture is designed with approximately 80% decrease in the data transmission traffic. The throughput of this accelerator achieves 700 fps, which is especially suitable for high-speed facial-related applications.
Leibo Liu, Wenping Zhu, Huiyu Mo, Chenchen Deng, Shaojun Wei
DAC3
2016 A 135-frames/s 1080p 87.5-mW Binary-Descriptor-Based Image Feature Extraction Accelerator
abstract
Binary image descriptors, which derive image feature description from the local image patches directly, are widely adopted in the mobile and embedded applications due to lower computational complexity and memory requirement. With the aim of improving the computation efficiency without degrading recognition performance, a lightweight binary robust descriptor is proposed based on the analysis of the state-of-the-art binary descriptors in this paper. A directional edge detection and optimized keypoint score function are developed to refine the keypoints. In addition, rotation invariance is achieved by executing circular symmetric-based descriptor generation and a coarse-grained orientation calculation method concurrently. The experimental results demonstrate that the proposed keypoint detector and binary descriptor achieve more than two times speedup and at least 23.6% improvement in processing speed with comparable performance, respectively. Furthermore, a very large scale integration architecture is also designed based on in-depth exploration of bit-level and task-level parallelism. Based on the postlayout simulation in a TSMC 65-nm CMOS process, the accelerator can achieve 135 frames/s on 1080p image while only consuming 87.5 mW at a 200-MHz operating frequency.
Wenping Zhu, Leibo Liu, Guangli Jiang, Shouyi Yin, Shaojun Wei
IEEE Trans. Circuits Syst. Video Technol.1
2015 A 127 fps in full hd accelerator based on optimized AKAZE with efficiency and effectiveness for image feature extraction
abstract
Visual feature extraction is a fundamental technique in vision-based application. This paper proposes an effective and efficient VLSI architecture based on optimized accelerated KAZE (AKAZE) for real-time feature extraction. AKAZE is a new feature detection algorithm with strong robustness for object recognition. To extract feature more robustly and reduce hardware resource, a two-dimensional pipeline array named Loop-Snake Architecture is presented. It takes advantage of computational similarity in different octaves and provides flexibility in precision-speed tradeoff on the fly. Furthermore, Polar Local Difference Binary descriptor and the corresponding structure are proposed to greatly reduce the memory bandwidth requirement and improve the speed. The experimental results indicate the optimized algorithm keeps the same accuracy compared with the original algorithm. The whole hardware system achieves 127fps in 1080p resolution at 200 MHz frequency. The throughput is twice faster than the state-of-the-art solutions.
Guangli Jiang, Leibo Liu, Wenping Zhu, Shouyi Yin, Shaojun Wei
DAC3
2014 A 65 nm uneven-dual-core SoC based platform for multi-device collaborative computing
abstract
Multiple mobile device-based collaborative computing emerges with the rapid proliferation of various smart mobile devices such as smartphones and tablets, which provide always-on connectivity, information and communication. However, due to severe resource poverty and poor network connectivity, lots of traditional embedded electronic devices with attracting features cannot be incorporated into this computing paradigm conveniently. In this paper, an uneven-dual-core SoC, which integrates a CPU core and a MCU core on a single chip with multiple operating system support, is proposed to realize loosely-coupled multiple heterogeneous device collaboration. A network file system, MRFS (Multi-client Raindrop File System), and FAT-X (File Allocation Table eXtension) are also proposed to provide client-centric cross-device data consistency and virtual file access respectively. Comprehensive mobile services are enabled by offloading appropriate tasks from existing smart mobile devices to involved traditional embedded devices. The SoC is implemented onto a 16.65 mm2silicon with 65 nm CMOS technology. This paper also presents three typical applications to illustrate the universality and huge potential for innovative usage model of the proposed system.
Wenping Zhu, Leibo Liu, Shouyi Yin, Shaojun Wei, Eugene Tang, Jiqiang Song, Jinzhan Peng
ISCAS1