Yifan He 0002

dblp:70/6707-2 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
13since 2021 · last 2024
0009-0001-8664-037XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 2Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Diversity-Authenticity Co-constrained Stylization for Federated Domain Generalization in Person Re-identification
abstract
This paper tackles the problem of federated domain generalization in person re-identification (FedDG re-ID), aiming to learn a model generalizable to unseen domains with decentralized source domains. Previous methods mainly focus on preventing local overfitting. However, the direction of diversifying local data through stylization for model training is largely overlooked. This direction is popular in domain generalization but will encounter two issues under federated scenario: (1) Most stylization methods require the centralization of multiple domains to generate novel styles but this is not applicable under decentralized constraint. (2) The authenticity of generated data cannot be ensured especially given limited local data, which may impair the model optimization. To solve these two problems, we propose the Diversity-Authenticity Co-constrained Stylization (DACS), which can generate diverse and authentic data for learning robust local model. Specifically, we deploy a style transformation model on each domain to generate novel data with two constraints: (1) A diversity constraint is designed to increase data diversity, which enlarges the Wasserstein distance between the original and transformed data; (2) An authenticity constraint is proposed to ensure data authenticity, which enforces the transformed data to be easily/hardly recognized by the local-side global/local model. Extensive experiments demonstrate the effectiveness of the proposed DACS and show that DACS achieves state-of-the-art performance for FedDG re-ID.
Fengxiang Yang, Zhun Zhong, Zhiming Luo, Yifan He 0002, Shaozi Li, Nicu Sebe
AAAI4
2024 Selective Domain-Invariant Feature for Generalizable Deepfake Detection
abstract
With diverse presentation forgery methods emerging continually, detecting the authenticity of images has drawn growing attention. Although existing methods have achieved impressive accuracy in training dataset detection, they still perform poorly in the unseen domain and suffer from forgery of irrelevant information such as background and identity, affecting generalizability. To solve this problem, we proposed a novel framework Selective Domain-Invariant Feature (SDIF), which reduces the sensitivity to face forgery by fusing content features and styles. Specifically, we first use a Farthest-Point Sampling (FPS) training strategy to construct a task-relevant style sample representation space for fusing with content features. Then, we propose a dynamic feature extraction module to generate features with diverse styles to improve the performance and effectiveness of the feature extractor. Finally, a domain separation strategy is used to retain domain-related features to help distinguish between real and fake faces. Both qualitative and quantitative results in existing benchmarks and proposals demonstrate the effectiveness of our approach.
Yingxin Lai, Yifan He 0002, Zhiming Luo, Shaozi Li
ICASSP3
2024 A Collaborative Framework Using Multimodal Data and Adaptive Noise for Human Behavior Anomaly Detection
abstract
Human behavior anomaly detection in video aims to identify unusual behaviors that are crucial for public safety. Recently, there has been an increase in reconstruction or prediction-based methods that integrate diverse modal features to enhance anomaly detection. However, they use methods that independently or directly fusion multimodal features without fully considering the collaborative potential between multimodal features, which are susceptible to interference from semantic differences, thereby impacting detection performance. In contrast, we design a collaborative framework using multimodal data and adaptive noise for behavior anomaly detection. Our framework detects anomalies by analyzing the contrastive differences between two modalities alongside single-frame reconstruction errors. Specifically, we first learn the correlation between RGB and skeletal modalities for normal behavior through contrastive learning and use inter-modal contrast difference to detect motion anomalies. Additionally, we propose a single-frame reconstruction network that adaptively adds noise based on the importance of foreground features to detect appearance anomalies. Anomalies often occur in the motion foreground, and increasing noise in this area can make it more difficult to reconstruct anomalies. Extensive experiments validate the state-of-the-art performance of our method on three public datasets.
Jianzhe Gao, Kejia Zhang 0003, Yifan He 0002, Zhiming Luo, Shaozi Li
IJCNN4
2024 A Multilevel Guidance-Exploration Network and Behavior-Scene Matching Method for Human Behavior Anomaly Detection
abstract
Human behavior anomaly detection aims to identify unusual human actions, playing a crucial role in intelligent surveillance and other areas. The current mainstream methods still adopt reconstruction or future frame prediction techniques. However, reconstructing or predicting low-level pixel features easily enables the network to achieve overly strong generalization ability, allowing anomalies to be reconstructed or predicted as effectively as normal data. Different from their methods, inspired by the Student-Teacher Network, we propose a novel framework called the Multilevel Guidance-Exploration Network (MGENet), which detects anomalies through the difference in high-level representation between the Guidance and Exploration network. Specifically, we first utilize the Normalizing Flow that takes skeletal keypoints as input to guide an RGB encoder, which takes unmasked RGB frames as input, to explore latent motion features. Then, the RGB encoder guides the mask encoder, which takes masked RGB frames as input, to explore the latent appearance feature. Additionally, we design a Behavior-Scene Matching Module to detect scene-related behavioral anomalies. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on ShanghaiTech and UBnormal datasets, with AUC of 86.9% and 74.3%, respectively. The code is available at https://github.com/molu-ggg/GENet.
Zhiming Luo, Jianzhe Gao, Yingxin Lai, Yifan He 0002, Shaozi Li
ACM Multimedia6
2024 SoftAct: A High-Precision Softmax Architecture for Transformers Supporting Nonlinear Functions
abstract
Transformer-based deep learning networks are revolutionizing our society. The convolution and attention co-designed (CAC) Transformers have demonstrated superior performance compared to the conventional Transformer-based networks. However, CAC Transformer networks contain various nonlinear functions, such as softmax and complex activation functions, which require high precision hardware design yet typically with significant cost in area and power consumption. To address these challenges,SoftAct, a compact and high-precision algorithm-hardware co-designed architecture, is proposed to implement both softmax and nonlinear activation functions in CAC Transformer accelerators. An improved softmax algorithm with penalties is proposed to maintain precision in hardware. A stage-wise full zero detection method is developed to skip redundant computation in softmax. A compact and reconfigurable architecture with a symmetrically designed linear fitting module is proposed to achieve nonlinear functions. TheSoftActarchitecture is designed in an industrial 28-nm CMOS technology with the MobileViT-xxs network classifying the ImageNet-1k dataset as the benchmark. Compared with the state of the art,SoftActimproves up to 5.87% network accuracy under 8-bit quantization, 153.2× area efficiency, and 1435× overall efficiency.
Yuzhe Fu, Changchun Zhou 0001, Tianling Huang, Eryi Han, Yifan He 0002, Hailong Jiao
IEEE Trans. Circuits Syst. Video Technol.5
2024 Toward Effective Traffic Sign Detection via Two-Stage Fusion Neural Networks
abstract
Automatic detection of traffic signs is crucial for Advanced Driving Assistance Systems (ADAS). Current two-stage approaches consist of a preliminary object detection step, where the traffic signs are categorized within broader families (e.g., speed limits), and then sub-classes (e.g., speed limit 40). However, these cascading methods fail to achieve satisfying performance, especially in more realistic driving scenarios where images are acquired under more challenging conditions. Under such conditions, the first-stage detection step is likely to provide inaccurate predictions, making the subsequent classification step useless. In this paper, we propose a simple yet effective two-stage fusion framework for traffic sign detection. Different from the previous cascading method, our framework directly predicts categories in the first-stage detection and fuse the two-stage category predictions to improves overall robustness. Besides, in order to filter the false detection boxes under low-resolution inputs, we also propose an effective post-processing method called Surrounding-Aware Non-Maximum Suppression (SA-NMS) as an alternative technique for the first-stage detection. After combining the above proposed methods, our framework obtains good detection performance. Experimental results on the widely used Tsinghua-Tencent 100K (TT100K) traffic sign dataset, which contains images of traffic signs collected under a variety of challenging conditions, show that the proposed framework outperforms current approaches in both accuracy and inference speed, achieving 89.7 mAP and 65 FPS for${608\times608}$low resolution images.
Zhishan Li, Battista Biggio, Yifan He 0002, Haoran Cai, Fabio Roli, Lei Xie 0007
IEEE Trans. Intell. Transp. Syst.4
2023 An Energy-Efficient 3D Point Cloud Neural Network Accelerator With Efficient Filter Pruning, MLP Fusion, and Dual-Stream Sampling
abstract
Three-dimensional (3D) point cloud has been employed in a wide range of applications recently. As a powerful weapon for point cloud analysis, point-based point cloud neural networks (PNNs) have demonstrated superior performance with less computation complexity and parameters, compared to sparse 3D convolution-based networks and graph-based convolutional neural networks. However, point-based PNNs still suffer from high computational redundancy, large off-chip memory access, and low parallelism in hardware implementation, thereby hindering the applications on edge devices. In this paper, to address these challenges, an energy-efficient 3D point cloud neural network accelerator is proposed for on-chip edge computing. An efficient filter pruning scheme is used to skip the redundant convolution of pruned filters and zero-value feature channels. A block-wise multi-layer perceptron (MLP) fusion method is proposed to increase the on-chip reuse of features, thereby reducing off-chip memory access. A dual-stream blocking technique is proposed for higher parallelism while maintaining inference accuracy. Implemented in an industrial 28-nm CMOS technology, the proposed accelerator achieves an effective energy efficiency of 12.65 TOPS/W and 0.13 mJ/frame energy consumption for PointNeXt-S at 100 MHz, 0.9 V supply voltage, and 8-bit data width. Compared to the state-of-the-art point cloud neural network accelerators, the proposed accelerator enhances the energy efficiency by up to 66.6× and reduces the energy consumption per frame by up to 70.2×.
Changchun Zhou 0001, Yuzhe Fu, Siyuan Qiu, Ge Li 0002, Yifan He 0002, Hailong Jiao
ICCAD6
2023 Sagitta: An Energy-Efficient Sparse 3D-CNN Accelerator for Real-Time 3-D Understanding
abstract
Three-dimensional (3-D) understanding or inference has received increasing attention, where 3-D convolutional neural networks (3D-CNNs) have demonstrated superior performance compared to 2D-CNNs, since 3D-CNNs learn features from all three dimensions. However, 3D-CNNs suffer from intensive computation and data movement. In this article, Sagitta, an energy-efficient low-latency on-chip 3D-CNN accelerator, is proposed for edge devices. Locality and small differential value dropout are leveraged to increase the sparsity of activations. A full-zero-skipping convolutional microarchitecture is proposed to fully utilize the sparsity of weights and activations. A hierarchical load-balancing scheme is also introduced to increase the hardware utilization. Specialized architecture and computation flow are proposed to enhance the effectiveness of the proposed techniques. Fabricated in a 55-nm CMOS technology, Sagitta achieves 3.8 TOPS/W for C3D at a latency of 0.1 s and 4.5 TOPS/W for 3D U-Net at a latency of 0.9 s at 100 MHz and 0.91-V supply voltage. Compared to the state-of-the-art 3D-CNN and 2D-CNN accelerators, Sagitta enhances the energy efficiency by up to$379.6\times $and$11\times $, respectively.
Changchun Zhou 0001, Siyuan Qiu, Xugang Cao, Yuzhe Fu, Yifan He 0002, Hailong Jiao
IEEE Internet Things J.6
2023 CNN Accelerator at the Edge With Adaptive Zero Skipping and Sparsity-Driven Data Flow
abstract
An energy-efficient convolutional neural network (CNN) accelerator is proposed for low-power inference on edge devices. An adaptive zero skipping technique is proposed to dynamically skip the zeros in either activations or weights, depending on which has the higher sparsity. The characteristic of non-zero data aggregation is explored to enhance the effectiveness of adaptive zero skipping in performance boosting. To mitigate the load imbalance issue after zero skipping, a sparsity-driven data flow and low-complexity dynamic task allocation are employed for different convolution layers. Facilitated further by a two-stage distiller, the proposed accelerator achieves$5.42\times $,$3.41\times $, and$3.42\times $performance boosting for VGG16, AlexNet, and Mobilenet-v1, respectively, compared to the baseline. Implemented in a 55-nm low power CMOS technology, the proposed accelerator achieves an effective energy efficiency of 2.41 TOPS/W, 2.35 TOPS/W, and 0.64 TOPS/W for VGG16, AlexNet, and Mobilenet-v1, respectively, at 100 MHz and 1.08 V supply voltage.
Ming Liu 0022, Changchun Zhou 0001, Siyuan Qiu, Yifan He 0002, Hailong Jiao
IEEE Trans. Circuits Syst. Video Technol.4
2022 An Efficient Framework for Detection and Recognition of Numerical Traffic Signs
abstract
Due to the variety of categories and uneven distribution of available samples, automatic traffic sign detection and recognition is still a challenging task. For those categories with less training data, existing deep learning methods cannot achieve desirable performance, and the overall detection effect is not satisfactory as well. In this letter, we fully explore the relationship between different traffic signs with digital characters and transform the category objects into multi-level classes to alleviate the uneven distribution of samples. We design a lightweight two-stage object detection framework with high real-time performance. The first stage network is proposed to obtain the category groups of traffic signs, and then we construct another object detection network to identify the digital characters of the detected traffic signs. To make the prediction in the first stage more accurate, we put forward a boxes fusion algorithm in the post-processing process and a refine module to improve the recognition performance. Experimental results show that our approach possesses significantly improved performance compared with the latest object detection networks and other traffic sign detectors. Even some traffic signs that only exist in testset can also be recognized accurately by our method.
Zhishan Li, Mingmu Chen, Yifan He 0002, Lei Xie 0007
ICASSP3
2022 PBDE: an effective post-processing method based on box density for object detection
Zhishan Li, Baozhi Jia, Yifan He 0002, Lei Xie 0007
Appl. Intell.3
2021 An Energy-Efficient Low-Latency 3D-CNN Accelerator Leveraging Temporal Locality, Full Zero-Skipping, and Hierarchical Load Balance
abstract
Three-dimensional convolutional neural network (3D-CNN) has demonstrated outstanding classification performance in video recognition compared to two-dimensional CNN (2D-CNN), since 3D-CNN not only learns the spatial features of each frame, but also learns the temporal features across all frames. However, 3D-CNN suffers from intensive computation and data movement. To solve these issues, an energy-efficient low-latency 3D-CNN accelerator is proposed. Temporal locality and small differential value dropout are used to increase the sparsity of activation. Furthermore, to fully utilize the sparsity of weight and activation, a full zero-skipping convolutional microarchitecture is proposed. A hierarchical load-balancing scheme is also introduced to improve resource utilization. With the proposed techniques, a 3D-CNN accelerator is designed in a 55-nm low-power CMOS technology, bringing in up to 9.89× speedup compared to the baseline implementation. Benchmarked with C3D, the proposed accelerator achieves an energy efficiency of 4.66 TOPS/W at 100 MHz and 1.08 V supply voltage.
Changchun Zhou 0001, Siyuan Qiu, Yifan He 0002, Hailong Jiao
DAC4
2021 An Effective Face Anti-Spoofing Method via Stereo Matching
abstract
Various algorithms based on Convolutional Neural Network (CNN) have achieved great performance in the task of face anti-spoofing (FAS). However, the issue of most approaches is that the robustness in unknown scenes is not strong due to the different quality of attack images and environmental factors. In this letter, we propose a real-time face anti-spoofing method based on stereo matching. We input the left and right views of a pair of infrared face images into our proposed lightweight stereo matching network to get a disparity map. Then, we use the disparity map as input for the classification network to obtain face living information. Experimental results show that the proposed approach has significantly improved performance compared with the latest state-of-the-art face anti-spoofing methods. Considering the real-time requirement, our method is superior to most single-image-based models in inference time and far less than those based on other large scale stereo matching networks in computational complexity. The code is available athttps://github.com/lizhishan1997/StereoMatching_FAS.
Zhishan Li, Jiayan Yuan, Baozhi Jia, Yifan He 0002, Lei Xie 0007
IEEE Signal Process. Lett.4
2018 Datawidth-Aware Energy-Efficient Multipliers: A Case for Going Sign Magnitude
abstract
Algorithms from many application domains, such as linear algebra and image/signal processing, heavily use the multiplication operator. Despite hardware support which is present in most modern cores, multiplication remains one of the most energy hungry arithmetic operations. This work explores how the energy efficiency of hardware multipliers can be improved by taking into account that the operands of a multiplication typically do not utilize the full width of the datapath. Seven datawidth-aware multiplier designs are implemented and evaluated. Post-layout energy analysis is performed to obtain the energy efficiency of each design for a number of representative benchmarks targeting the consumer market. The results show a significant improvement in energy efficiency compared to a 32-bit Baugh-Wooley baseline multiplier. A 32-bit sign-magnitude based design, integrated in a two's complement datapath, is shown to have a 1.38 times better energy efficiency than a baseline two's complement multiplier. In the best case (JPEG encoding), the energy efficiency is increased by a factor 2.25, demonstrating that a sign-magnitude multiplier, and datawidth-aware multipliers in general, are an attractive option for ultra low-energy designs.
Luc Waeijen, Hailong Jiao, Henk Corporaal, Yifan He 0002
DSD4
2016 MacSim: A MAC-Enabled High-Performance Low-Power SIMD Architecture
abstract
Single-Instruction-Multiple-Data (SIMD) architectures, which exploit data-level parallelism (DLP), are widely used to achieve high-performance and low-power computing. In most of streaming applications, such as CNN-based detection and recognition, color space conversion and various kinds of filters, multiply-accumulate is one of the most important and expensive operations to be executed. In this paper, we propose a high-performance low-power SIMD architecture with advanced multiply accumulator (MAC) support (MacSim) to improve the computational efficiency. In addition, a smart loop tiling scheme is proposed. To support this tiling even further, the MAC unit is equipped with multiple accumulator registers. According to the Design Space Exploration (DSE) of the proposed MAC unit, a MAC instance with four accumulator registers (MAC4reg) is selected as a good choice for target kernels. In this paper, a 64-PE 16-bit (processing element) SIMD instance without MAC support is taken as the baseline. For a head-to-head comparison, a 64-PE 16-bit SIMD with MAC4reg (MacSim4) and the baseline SIMD are all implemented in HDL and synthesized with a TSMC 40nm low-power library. Five streaming application kernels are mapped to both architectures. Our experimental results show with MAC4reg the runtime and energy consumption are reduced up to 38% and 42% respectively. Besides, a 4-layer CNN-based detection application is also fully mapped onto the proposed MacSim4. Working at 950MHz, MacSim4 reaches a throughput of 62.4 GOPS, which meets the requirement of real-time (720P HD, 30fps) detection. The energy consumption per PE per operation is very low, 4.7pJ/Op excluding SRAM (Static Random Access Memory) and 4.8pJ/Op including a 2k-entry SRAM bank. As a prototype, the proposed SIMD is mapped into an FPGA and can run all the kernels.
Tong Geng, Luc Waeijen, Maurice Peemen, Henk Corporaal, Yifan He 0002
DSD5
2014 Reduction Operator for Wide-SIMDs Reconsidered
abstract
It has been shown that wide Single Instruction Multiple Data architectures (wide-SIMDs) can achieve high energy efficiency, especially in domains such as image and vision processing. In these and various other application domains, reduction is a frequently encountered operation, where multiple input elements need to be combined into a single element by an associative operation, e.g. addition or multiplication. There are many applications that require reduction such as: partial histogram merging, matrix multiplication and min/max-finding. Wide-SIMDs contain a large number of processing elements (PEs), which in general are connected by a minimal form of interconnect for scalability reasons. To efficiently support reduction operations on wide-SIMDs with such a minimal interconnect, we introduce two novel reduction algorithms which do not rely on complex communication networks or any dedicated hardware. The proposed approaches are compared with both dedicated hardware and other software solutions in terms of performance, area, and energy consumption. A practical case study demonstrates that the proposed software approach has much better generality, flexibility and no additional hardware cost. Compared to a dedicated hardware adder tree, the proposed software approach saves 6.8% area with a performance penalty of only 6.5%.
Luc Waeijen, Dongrui She, Henk Corporaal, Yifan He 0002
DAC4
2013 MAMPSX: A demonstration of rapid, predictable HMPSOC synthesis
abstract
Heterogeneous Multiprocessor systems-on-chip (HMPSoC) are becoming popular as a means of meeting energy efficiency requirements of modern embedded systems. However, as these HMPSoCs run multimedia applications as well, they also need to meet realtime requirements. Designing HMPSoCs with predictable timing behavior is a key challenge, as the current design methods for these platforms are semi-automated, non-predictable, or support limited heterogeneity. In this demonstration, we present a design framework to rapidly generate and implement predictable HMPSoC designs. It takes the application specifications and the architecture model as input and generates the entire HMPSoC, for FPGA prototyping, that meets the throughput constraints of the application. We also present results of a case study that computes the performance-power tradeoffs of an industrial vision application. A tool-chain targeting the Xilinx Zynq FPGA is also presented.
Shakith Fernando, Mark Wijtvliet, Firew Siyoum, Yifan He 0002, Sander Stuijk, Akash Kumar 0001, Henk Corporaal
FPL4
2013 MAMPSx: A design framework for rapid synthesis of predictable heterogeneous MPSoCs
abstract
Heterogeneous Multiprocessor System-on-Chips (HMPSoC) are becoming popular as a means of meeting energy efficiency requirements of modern embedded systems. However, as these HMPSoCs run multimedia applications as well, they also need to meet real-time requirements. Designing these predictable HMPSoCs is a key challenge, as the current design methods for these platforms are either semi-automated, non-predictable, or have limited heterogeneity. In this paper, we propose a design framework to generate and program HMPSoC designs in a rapid and predictable manner. It takes the application specifications and the architecture model as input and generates the entire HMPSoC, for FPGA prototyping, that meets the throughput constraints. The experimental results show that our framework can provide a conservative bound on the worst-case throughput of the FPGA implementation. We also present results of a case study that computes the area-power trade-offs of an industrial vision application. The entire design space exploration of all configurations was completed in 8 hours. A tool-chain targeting the Xilinx Zynq FPGA is also presented.
Shakith Fernando, Firew Siyoum, Yifan He 0002, Akash Kumar 0001, Henk Corporaal
RSP3
2013 Efficient communication support in predictable heterogeneous MPSoC designs for streaming applications
Yifan He 0002, Dongrui She, Sander Stuijk, Henk Corporaal
J. Syst. Archit.1
2013 An energy-efficient method of supporting flexible special instructions in an embedded processor with compact ISA
abstract
In application-specific processor design, a common approach to improve performance and efficiency is to use special instructions that execute complex operation patterns. However, in a generic embedded processor with compact Instruction Set Architecture (ISA), these special instructions may lead to large overhead such as: ( i ) more bits are needed to encode the extra opcodes and operands, resulting in wider instructions; ( ii ) more Register File (RF) ports are required to provide the extra operands to the function units. Such overhead may increase energy consumption considerably. In this article, we propose to support flexible operation pair patterns in a processor with a compact 24-bit RISC-like ISA using: ( i ) a partially reconfigurable decoder that exploits the pattern locality to reduce opcode space requirement; ( ii ) a software-controlled bypass network to reduce operand encoding bit and RF port requirement. An energy-aware compiler backend is designed for the proposed architecture that performs pattern selection and bypass-aware scheduling to generate energy-efficient codes. Though the proposed design imposes extra constraints on the operation patterns, the experimental results show that for benchmark applications from different domains, the average dynamic instruction count is reduced by over 25%, which is only about 2% less than the architecture without such constraints. The proposed architecture reduces total energy by an average of 15.8% compared to the RISC baseline, while the one without constraints achieves almost no improvement due to its high overhead. When high performance is required, the proposed architecture is able to achieve a speedup of 13.8% with 13.1% energy reduction compared to the baseline by introducing multicycle SFU operations.
Dongrui She, Yifan He 0002, Henk Corporaal
ACM Trans. Archit. Code Optim.2
2012 Energy efficient special instruction support in an embedded processor with compact isa
abstract
The use of special instructions that execute complex operation patterns is a common approach in application specific processor design to improve performance and efficiency. However, in an embedded generic processor with compact instruction set architecture (ISA), such instructions may lead to large overhead as: i) more bits are needed to encode the extra opcodes and operands, resulting in wider instructions; ii) more register file (RF) ports are required to provide the extra operands to the function units. Such overhead may increase energy consumption considerably.
Dongrui She, Yifan He 0002, Henk Corporaal
CASES2
2012 Scheduling for register file energy minimization in explicit datapath architectures
abstract
In modern processor architectures, the register file (RF) consumes considerable amount of the processor power. It is well known that by allowing software to have explicit fine-grained control over the datapath, the transport-triggered architectures (TTAs) can substantially reduce the RF traffic, thereby minimizing the RF energy. However, it is important to make sure that the gain in RF is not cancelled out by the overhead due to the fine-grained datapath control, in particular, the deterioration of code density in conventional TTAs. In this paper, we analyze the potential of minimizing RF energy in MOVE-Pro, a TTA-based processor framework. We present a flexible compiler backend, which performs energy-aware instruction scheduling to push the limit of RF energy reduction. The experimental results show that with the proposed energy-aware compiler backend, MOVE-Pro is able to significantly reduce RF energy compared to its RISC/VLIW counterparts, by up to 80%. Meanwhile the code density of MOVE-Pro remains at the same level as its RISC/VLIW counterparts, allowing the energy saving in RF to be successfully transferred to total energy saving.
Dongrui She, Yifan He 0002, Bart Mesman, Henk Corporaal
DATE2
2011 Feasibility Analysis of Ultra High Frame Rate Visual Servoing on FPGA and SIMD Processor
Yifan He 0002, Zhenyu Ye, Dongrui She, Bart Mesman, Henk Corporaal
ACIVS1
2011 From Xetal-II to Xetal-Pro: On the Road Toward an Ultralow-Energy and High-Throughput SIMD Processor
abstract
Looking forward to the next generation of mobile streaming computing, the demanded energy efficiency of end-user terminals will become ever stringent. The Xetal-Pro processor, which is the latest member of the Xetal low-power single-instruction multiple data (SIMD) processor family from Philips, is presented in this paper. The predecessor of Xetal-Pro, known as Xetal-II, already ranks as one of the most computational-efficient [in terms of giga operations per second (GOPS)/Watt] processors available today, however, it cannot yet achieve the demanded energy efficiency (less than 1 pJ per operation). Unlike Xetal-II, Xetal-Pro supports ultrawide supply voltage (Vdd) scaling from the nominal supply to the subthreshold region. Although aggressiveVddscaling causes severe throughput degradation, this can be partly compensated for by the massive parallelism in the Xetal family. Xetal-II includes a large on-chip frame memory (FM), which cannot be scaled well to an ultralowVddhence creating a big obstacle to increase energy efficiency. Therefore, we investigate both different FM realizations and memory organization alternatives. A hybrid memory system (HMS), which reduces the non-local memory traffic and enables furtherVddscaling, is proposed. For design space exploration of the right number of the scratchpad memory (SM) entries, the corresponding data locality analysis is provided, too. Moreover, some unique circuit implementation issues of Xetal-Pro such as the customized level-shifter are also discussed. Compared to Xetal-II operating at the nominal voltage, Xetal-Pro provides up to two times energy efficiency improvement even withoutVddscaling (essentially a consequence of data localization in the SM) when delivering the same amount of ultrahigh throughput. WithVddscaling into the sub/near threshold region, Xetal-Pro could gain more than ten times energy reduction while still delivering a high throughput of 0.69 GOPS (counting multiply and add operations only). The new insight of Xetal-Pro sheds light on the direction of future ultralow-energy SIMD processors.
Yu Pu, Yifan He 0002, Zhenyu Ye, Sebastian M. Londono, Anteneh A. Abbo, Richard P. Kleihorst, Henk Corporaal
IEEE Trans. Circuits Syst. Video Technol.2
2010 Xetal-Pro: an ultra-low energy and high throughput SIMD processor
abstract
This paper presents Xetal-Pro SIMD processor, which is based on Xetal-II, one of the most computational-efficient (in terms of GOPS/Watt) processors available today. Xetal-Pro supports ultra wide V DD scaling from nominal supply to the sub-threshold region. Although aggressive V DD scaling causes severe throughput degradation, this can be compensated by the nature of massive parallelism in the Xetal family. The predecessor of Xetal-Pro, Xetal-II, includes a large on-chip frame memory (FM), which cannot operate reliably at ultra low voltage. Therefore we investigate both different FM realizations and memory organization alternatives. We propose a hybrid memory architecture which reduces the non-local memory traffic and enables further V DD scaling. Compared to Xetal-II operating at nominal voltage, we could gain more than 10× energy reduction while still delivering a sufficiently high throughput of 0.69 GOPS (counting multiply and add operations only). This work gives a new insight to the design of ultra-low energy SIMD processors, which are suitable for portable streaming applications.
Yifan He 0002, Yu Pu, Richard P. Kleihorst, Zhenyu Ye, Anteneh A. Abbo, Sebastian M. Londono, Henk Corporaal
DAC1
2008 Real-Time Hough Transform on 1-D SIMD Processors: Implementation and Architecture Exploration
Yifan He 0002, Zoran Zivkovic, Richard P. Kleihorst, Alexander Danilin, Henk Corporaal, Bart Mesman
ACIVS1