Jianfeng An

dblp:73/2472 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-6559-1196ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RWIP: Region-Level Write-Intensity Prediction for GPU L2 Cache Based on Hybrid-Retention STT-MRAM
Yujie Pu, Chen Zhao 0009, Jianfeng An, Wenzhe Zhao 0001, Pengju Ren
Euro-Par (1)3
2025 PRDSE: A Prior-Driven Design Space Exploration Method
abstract
Design space exploration is essential for optimizing deep neural network accelerators, which face increasing computational and energy demands as model complexity grows. Previous approaches rely heavily on local insights, often neglecting the need for extensive exploration to improve the global perspective. This leads to challenges such as blind exploration and a higher likelihood of getting trapped in local optima. Without dynamic adjustments or adaptive strategies, these methods struggle to navigate large, complex design spaces effectively. In this paper, we propose PRDSE, a design space exploration framework based on reinforcement learning that integrates both intrinsic and extrinsic metrics, guided by prior knowledge. The proposed method incorporates an adaptive adjustment mechanism that dynamically balances intrinsic and extrinsic rewards based on the progress of exploration, improving both search efficiency and optimization performance. Compared to state-of-the-art methods, PRDSE achieves substantial improvements, with latency speedups of up to$3.49 \mathrm{x}$in the cloud environment and$3.43 \mathrm{x}$in the edge environment, respectively. This work demonstrates that PRDSE effectively balances exploration and optimization objectives, providing a more efficient and scalable approach to design space exploration in accelerator design.
Junda Zhu 0007, Xiaoya Fan, Jianfeng An, Kaijie Feng
ASAP3
2025 CSDSE: An efficient design space exploration framework for deep neural network accelerator based on cooperative search
Kaijie Feng, Xiaoya Fan, Jianfeng An, Haoyang Wang 0014, Chuxi Li
Neurocomputing3
2023 CSDSE: Apply Cooperative Search to Solve the Exploration-Exploitation Dilemma of Design Space Exploration
Kaijie Feng, Xiaoya Fan, Jianfeng An, Haoyang Wang 0014, Chuxi Li
ICA3PP (4)3
2023 A gated multi-hierarchical feature fusion network for recognizing steel plate surface defects
Huanjie Tao, Minghao Lu, Zhenwu Hu, Jianfeng An
Multim. Syst.4
2023 An improved interaction-and-aggregation network for person re-identification
Huanjie Tao, Wenjie Bao, Qianyue Duan, Zhenwu Hu, Jianfeng An
Multim. Tools Appl.5
2023 An Adaptive Interference Removal Framework for Video Person Re-Identification
abstract
Video person re-identification (V-ReID) can leverage rich spatial-temporal information embedded in sequence data to achieve better accuracy. However, it is vulnerable to the interference of inaccurate and redundant noisy frames in each sequence as well as the background clutter and person-irrelevant pixels in each frame. To solve the above issues, this paper presents an adaptive interference removal framework (IRF) to learn discriminative feature representations by removing various interference. Our IRF mainly consists of two modules including an attention-guided adaptive interference frame removal module (IFRM) and an attention-guided adaptive interference pixel removal module (IPRM). IFRM and IPRM are designed to locate task-relevant keyframes and key pixels, respectively. IFRM adopts the attention mechanism to predict frame-wise scores to characterize the contribution of each frame to the final identification task. IPRM collaboratively utilizes camera identity classification loss, person identity classification loss, target attention loss, and person mask adversarial loss for extracting pure pedestrian representations. A progressive mask augmentation strategy is designed to restrain the data distribution of the generated person masks to further guide model training. Extensive experiments demonstrate that our models outperform state-of-the-art accuracy on seven person ReID datasets.
Huanjie Tao, Qianyue Duan, Jianfeng An
IEEE Trans. Circuits Syst. Video Technol.3
2023 ACDSE: A Design Space Exploration Method for CNN Accelerator based on Adaptive Compression Mechanism
abstract
Customized accelerators for Convolutional Neural Network (CNN) can achieve better energy efficiency than general computing platforms. However, the design of a high-performance accelerator should take into account a variety of parameters and physical constraints. The increasing parameters and tighter constraints gradually complicate the design space, which poses new challenges to the capacity and efficiency of design space exploration methods. In this paper, we provide a novel design space exploration method named ACDSE for optimizing the design process of CNN accelerators. ACDSE implements the adaptive compression mechanism to dynamically adjust the search range and prune low-value design points according to the exploration states. As a result, it can focus on valuable subspace while also improving exploration capacity and efficiency. Additionally, we implement ACDSE to address the problem of CNN accelerator latency optimization. The experiment indicates that, compared to former DSE methods, ACDSE can reduce latency and increase efficiency by 1.39x-5.07x and 2.07x-43.87x, respectively, under the most stringent constraint conditions, demonstrating its superior adaptability to the complicated design space.
Kaijie Feng, Xiaoya Fan, Jianfeng An, Chuxi Li, Kaiyue Di, Jiangfei Li
ACM Trans. Embed. Comput. Syst.3
2022 An Automatic-Addressing Architecture With Fully Serialized Access in Racetrack Memory for Energy-Efficient CNNs
abstract
Racetrack memory, an emerging low-power magnetic memory, promises a competitive replacement for traditional memory in the accelerators. However, random access in racetrack memory is time and energy expenditure for CNN accelerators because of its large amount of invalid-shifts. In this article, we propose an automatic-addressing architecture that builds a novel data layout to guarantee that the next round of memory access can be always satisfied at the in-situ or rigorously adjacent cells of current round, producing a fully serialized access footprint that can drive instant port-alignment without any invalid-shifts in racetrack memory. By this way, original address-based access degrades to the selections repeated among the three candidates, i.e., onein-situcell and two neighbor cells. Based on this simplification, a lightweight access management can generate the sequence of one-out-three selections according to the deterministic access behaviors defined by CNN hyper-parameters. The evaluation shows that, when deploying the five popular CNN applications to our architecture, the physical shifts of racetrack is curtailed by 74.64 percent over legacy layout, which achieves 54.2 and 42.1 percent energy reduction on read and write, respectively. A case study of YOLOv2 indicates that our architecture performs 6.503 GOp/J that achieves$18.5 \times$improvement to server-level GPUs.
Danghui Wang, Jianfeng An, Xiaoya Fan
IEEE Trans. Computers4