Fang Su

dblp:09/567 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 Look Before You Leap : Precision Instruction Supply via SmartScout
abstract
Modern high-performance processors extensively employ Fetch-Directed Instruction Prefetching (FDIP) to mitigate instruction supply bottlenecks. However, the efficacy of FDIP is fundamentally constrained by the accuracy of the Branch Prediction Unit (BPU). As the critical component within the BPU, the Branch Target Buffer (BTB) faces severe capacity bottlenecks. While pre-decoding-based prefetching offers a remedy, existing approaches suffer from two critical impediments: (1) The Noise Dilemma: Suboptimal trade-off between coverage and accuracy. (2) Inefficient Miss Resolution: Current designs rely on reactive recovery or stalls, failing to leverage available front-end slack for proactive correction.
Peng Qu 0001, Tingji Zhang, Fang Su, Zhe Pan 0001, Youhui Zhang
ICS4
2025 Hierarchical Prefetching: A Software-Hardware Instruction Prefetcher for Server Applications
abstract
The large working set of instructions in server-side applications causes a significant bottleneck in the front-end, even for high-performance processors equipped with fetch-directed instruction prefetching (FDIP). Prefetchers specifically designed for server scenarios typically rely on a record-and-replay mechanism that exploits the repetitiveness of instruction sequences. However, the efficacy of these techniques is compromised by discrepancies between actual and predicted control flows, resulting in loss of coverage and timeliness. This paper proposes Hierarchical Prefetching, a novel approach that tackles the limitations of existing prefetchers. It identifies common coarse-grained functionality blocks (called Bundles) within the server code and prefetches them as a whole. Bundles are significantly larger than typical prefetch targets, encompassing tens to hundreds of kilobytes of code. The approach combines simple software analysis of code for bundle formation and light-weight hardware for record-and-replay prefetching. The prefetcher requires under 2KB of on-chip storage by keeping most of the metadata in main memory. Experiments with 11 popular server workloads reveal that Hierarchical Prefetching significantly improves miss coverage and timeliness over prior techniques, achieving a 6.6% average performance gain over FDIP.
Tingji Zhang, Boris Grot, Wenjian He, Yashuai Lv, Peng Qu 0001, Fang Su, Guowei Zhang 0002, Youhui Zhang
ASPLOS (2)6
2025 A TRRIP Down Memory Lane: Temperature-Based Re-Reference Interval Prediction For Instruction Caching
abstract
Modern mobile CPU software pose challenges for conventional instruction cache replacement policies due to their complex runtime behavior causing high reuse distance between executions of the same instruction.Mobile code commonly suffers from large amounts of stalls in the CPU frontend and thus starvation of the rest of the CPU resources.Complexity of these applications and their code footprint are projected to grow at a rate faster than available on-chip memory due to power and area constraints, making conventional hardware-centric methods for managing instruction caches to be inadequate.We present a novel software-hardware co-design approach called TRRIP (Temperature-based Re-Reference Interval Prediction) that enables the compiler to analyze, classify, and transform code based on "temperature" (hot/cold), and to provide the hardware with a summary of code temperature information through a well-defined OS interface based on using code page attributes.TRRIP's lightweight hardware extension employs code temperature attributes to optimize the instruction cache replacement policy resulting in the eviction rate reduction of hot code.TRRIP is designed to be practical and adoptable in real mobile systems that have strict feature requirements on both the software and hardware components.TRRIP can reduce the L2 MPKI for instructions by 26.5% resulting in geomean speedup of 3.9%, on top of RRIP cache replacement running mobile code already optimized using PGO.
Henry Kao, Nikhil Sreekumar, Prabhdeep Singh Soni, Ali Sedaghati, Fang Su, Maziar Goudarzi
MICRO5
2025 Effective Dimension Extraction Mechanism: A novel mechanism for meta-heuristic algorithms in solving complex high-dimensional problems
Fang Su
Expert Syst. Appl.1
2025 How do IT service vendors build organizational resilience in response to major exogenous shocks?
Fang Su, Ji-Ye Mao
Inf. Manag.1
2024 A Manifold-Guided Gravitational Search Algorithm for High-Dimensional Global Optimization Problems
abstract
Gravitational Search Algorithm (GSA) is a well‐known physics‐based meta‐heuristic algorithm inspired by Newton’s law of universal gravitation and performs well in solving optimization problems. However, when solving high‐dimensional optimization problems, the performance of GSA may deteriorate dramatically due to severe interference of redundant dimensional information in the high‐dimensional space. To solve this problem, this paper proposes a Manifold‐Guided Gravitation Search Algorithm, called MGGSA. First, based on the Isomap, an effective dimension extraction method is designed. In this mechanism, the effective dimension is extracted by comparing the dimension differences of the particles located in the same sorting position both in the original space and the corresponding low‐dimensional manifold space. Then, the gravitational adjustment coefficient is designed, so that the particles can be guided to move in a more appropriate direction by increasing the effect of effective dimension, reducing the interference of redundant dimension on particle motion. The performance of the proposed algorithm is tested on 35 high‐dimensional (dimension is 1000) benchmark functions from CEC2010 and CEC2013, and compared with eleven state‐of‐art meta‐heuristic algorithms, the original GSA and four latest GSA’s variants, as well as three well‐known large‐scale global optimization algorithms. The experimental results demonstrate that MGGSA not only has a fast convergence rate but also has high solution accuracy. Besides, MGGSA is applied to three real‐world application problems, which verifies the effectiveness of MGGSA on practical applications.
Fang Su, Yance Wang, Yuxing Yao
Int. J. Intell. Syst.1
2020 An Energy-Efficient Flexible Capacitive Pressure Sensing System
abstract
Flexible capacitive pressure sensing system (FCPSS) is promising in the area of healthcare, robotics, and Internet of Things (IoT). As the size of the sensing array increases, designing energy-efficient FCPSS is getting challenging. This work provides a comprehensive solution for low-power FCPSS design, where major contributions are as follows. 1) Crosstalk-induced measurement error in a crossbar structure FCPSS is first studied and an accurate and low-power linear iterative algorithm is proposed for on-chip sensing array calibration (SAC). 2) Binary Neural Network (BNN)-based spatial-temporal adaptive sensing scheme for FCPSS is first proposed to utilize the sparsity of sampling and to further improve energy efficiency. Combined with the clock-gating-friendly low-power sensor interface, the system consumes 31.39 μJ energy and gains 95.04% capacitor measurement accuracy for each sensing operation on a 10×10 array, achieving 116× energy reduction compared with the state-of-the-art technology.
Qinghang Zhao, Xiyuan Tang, Fang Su, Nan Sun 0001, Huazhong Yang, Yongpan Liu
ISCAS4
2020 Multi-channel precision-sparsity-adapted inter-frame differential data codec for video neural network processor
abstract
Activation I/O traffic is a critical bottleneck of video neural network processor. Recent works adopted an inter-frame difference method to reduce activation size. However, current methods can't fully adapt to the various precision and sparsity in differential data. In this paper, we propose the multi-channel precision-sparsity-adapted codec, which will separate the differential activation and encode activation in multiple channels. We analyze the most adapted encoding of each channel, and select the optimal channel number with the best performance. A two-channel codec hardware has been implemented in the ASIC accelerator, which can encode/decode activations in parallel. Experiment results show that our coding achieves 2.2x-18.2x compression rate in three scenarios with no accuracy loss, and the hardware has 42x/174x improvement on speed and energy-efficiency compared with the software codec.
Yixiong Yang, Fang Su, Fanyang Cheng, Zhuqing Yuan, Huazhong Yang, Yongpan Liu
ISLPED3
2019 AERIS: area/energy-efficient 1T2R ReRAM based processing-in-memory neural network system-on-a-chip
abstract
ReRAM-based processing-in-memory (PIM) architecture is a promising solution for deep neural networks (NN), due to its high energy efficiency and small footprint. However, traditional PIM architecture has to use a separate crossbar array to store either positive or negative (P/N) weights, which limits both energy efficiency and area efficiency. Even worse, imbalance running time of different layers and idle ADCs/DACs even lower down the whole system efficiency. This paper proposes AERIS, an Area/Energy-efficient 1T2R ReRAM based processing-In-memory NN System-on-a-chip to enhance both energy and area efficiency. We propose an area-efficient 1T2R ReRAM structure to represent both P/N weights in a single array, and a reference current cancelling scheme (RCS) is also presented for better accuracy. Moreover, a layer-balance scheduling strategy, as well as the power gating technique for interface circuits, such as ADCs/DACs, is adopted for higher energy efficiency. Experiment results show that compared with state-of-the-art ReRAM-based architectures, AERIS achieves 8.5x/1.3x peak energy/area efficiency improvements in total, due to layer-balance scheduling for different layers, power gating of interface circuits, and 1T2R ReRAM circuits. Furthermore, we demonstrate that the proposed RCS compensates the non-ideal factors of ReRAM and improves NN accuracy by 5.2% in the XNOR net on CIFAR-10 dataset.
Jinshan Yue, Yongpan Liu, Fang Su, Shuangchen Li, Zhibo Wang 0004, Wenyu Sun, Xueqing Li 0002, Huazhong Yang
ASP-DAC3
2019 A Task Failure Rate Aware Dual-Channel Solar Power System for Nonvolatile Sensor Nodes
abstract
In line with the rapid development of the Internet of Things (IoT), the maintenance of on-board batteries for a trillion sensor nodes has become prohibitive both in time and costs. Energy harvesting is a promising solution to this problem. However, conventional energy-harvesting systems with storage suffer from low efficiency because of conversion loss and storage leakage. Direct supply systems without energy buffer provide higher efficiency, but fail to satisfy quality of service (QoS) due to mismatches between input power and workloads. Recently, a novel dual-channel photovoltaic power system has paved the way to achieve both high energy efficiency and QoS guarantee. This article focuses on the design-time and run-time co-optimization of the dual-channel solar power system. At the design stage, we develop a task failure rate estimation framework to balance design costs and failure rate. At run-time, we propose a task failure rate aware QoS tuning algorithm to further enhance energy efficiency. Through the experiments on both a simulation platform and a prototype board, this study demonstrates a 27% task failure rate reduction compared with conventional architectures with identical design costs. And the proposed online QoS tuning algorithm brings up to 30% improvement in energy efficiency with nearly zero failure rate penalty.
Fang Su, Yongpan Liu, Xiao Sheng, Hyung Gyu Lee, Naehyuck Chang, Huazhong Yang
ACM Trans. Embed. Comput. Syst.1
2018 Cross-Domain Attribute Representation Based on Convolutional Neural Network
Gaoyuan Liang, Fang Su, Fanxin Qu
ICIC (3)3
2018 Domain transfer convolutional attribute embedding
abstract
In this paper, we study the problem of transfer learning with the attribute data. In the transfer learning problem, we want to leverage the data of the auxiliary and the target domains to build an effective model for the classification problem in the target domain. Meanwhile, the attributes are naturally stable cross different domains. This strongly motives us to learn effective domain transfer attribute representations. To this end, we proposed to embed the attributes of the data to a common space using the powerful convolutional neural network (CNN) model. The convolutional representations of the data points are mapped to the corresponding attributes so that they can be effective embedding of the attributes. We also represent the data of different domains by a domain-independent CNN, ant a domain-specific CNN and combine their outputs with the attribute embedding to build the classification model. An joint learning framework is constructed to minimise the classification errors, the attribute mapping error, the mismatching of the domain-independent representations cross different domains, and to encourage the neighbourhood smoothness of representations in the target domain. The minimisation problem is solved by an iterative algorithm based on gradient descent. Experiments over benchmark data-sets of person re-identification, bankruptcy prediction and spam email detection show the effectiveness of the proposed method.
Fang Su
J. Exp. Theor. Artif. Intell.1
2017 Nonvolatile processors: Why is it trending?
abstract
Energy harvesting has become a promising solution to power up Internet-of-Things (IoT) devices. In this scenario, the constrained power budget and frequent absence of ambient energy cause severe reliability issues and performance degradation on conventional CMOS computing circuits. Fortunately, the advent of nonvolatile processor (NVP) opens the possibility to compute continuously using an intermittent power supply. It is considered as a key component of the next generation IoT edge devices. In this work, we provide insights to the evolution of the NVP and its application in real world scenarios. Efforts on improving the performance of NVP and future research prospects are also discussed in this paper.
Fang Su, Kaisheng Ma, Xueqing Li 0002, Tongda Wu, Yongpan Liu, Narayanan Vijaykrishnan
DATE1
2017 CNN-based pattern recognition on nonvolatile IoT platform for smart ultraviolet monitoring: (Invited paper)
abstract
Intelligent computing and maintenance-free powering are two desirable characteristics of wearable IoT devices. Energy harvesting nonvolatile intelligent processor (NIP) with neural network computation capability has the potential to advance these goals. Individual ultraviolet (UV) exposure monitoring progressively becomes one conspicuous application of wearable devices. In resource constrained wearable sensor nodes, we can alleviate the data transmission burden via convolutional neural networks (CNNs) based pattern recognition. Nevertheless, in spite of the substantially improved computing capability of NIP, typically computational and memory intensive CNNs are still too bulky for on-node implementation. We develop an CNN-based pattern recognition system for nonvolatile IoT platform for smart UV monitoring, and propose a optimization method to achieve extremely tiny and efficient CNNs. Experimental results show that the offline-trained CNN can recognize individual UV exposure patterns with accuracy of 85%, and the simplified on-node CNN can achieve 93.2% parameters reduction with only 5% accuracy loss.
Jinyang Li 0002, Qingwei Guo, Fang Su, Jinshan Yue, Jingtong Hu, Huazhong Yang, Yongpan Liu
ICCAD3
2016 Design of nonvolatile processors and applications
abstract
Energy harvesting is under intense investigation as a promising substitute for batteries. However, given the erratic nature of ambient energy sources, temporary status in conventional CMOS circuits can be lost upon a sudden power outage. Taking advantage of emerging nonvolatile memories (NVMs), nonvolatile processor (NVP) backs up system contexts when power failure occurs, and recalls pre-stored data on resumption. It has become a hot topic for the capability to survive power variations and to guarantee forward progress on computation tasks. This paper acts as a brief guide introducing the concepts, the current status, the challenges and opportunities, as well as emerging applications of NVPs. Through this paper, we expect to help researchers who are new in this area better understand the development trends and future research prospects. We also hope to attract researchers to join and explore more innovative applications of NVP.
Fang Su, Zhibo Wang 0004, Jinyang Li 0002, Meng-Fan Chang, Yongpan Liu
VLSI-SoC1
2011 Incremental manifold learning by spectral embedding methods
Housen Li, Hao Jiang 0001, Roberto Barrio, Xiangke Liao, Lizhi Cheng, Fang Su
Pattern Recognit. Lett.6