EDBT 2026 Demo / reviewers in the wild / expert
Hongjia Li 0003
dblp:14/10237-3
· DBLP profile ↗
15ranked-venue papers
4as first author
5since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | TAAS: a timing-aware analytical strategy for AQFP-capable placement automationabstractAdiabatic Quantum-Flux-Parametron (AQFP) is a superconducting logic with extremely high energy efficiency. AQFP circuits adopt the deep pipeline structure, where the four-phase AC-power serves as both the energy supply and the clock signal and transfers the data from one clock phase to the next. However, the deep pipeline structure causes the stage delay of the data propagation is comparable to the delay of the zigzag clocking, which triggers timing violations easily. In this paper, we propose a timing-aware analytical strategy for the AQFP placement, TAAS, that immensely reduces timing violations under specific spacing constraints and wirelength constraints of AQFP. TAAS includes two main characteristics: 1) a timing-aware objective function that incorporates a four-phase timing model for the analytical global placement. 2) a unique detailed placement including the timing-aware dynamic programming technique and the time-space cell regularization. To validate the effectiveness of TAAS, various representative circuits are adopted as benchmarks. As shown in the experimental results, our strategy can increase the maximum operating frequency by up to 30% ~ 40% with a negligible wirelength increase -3.41%~1%. Peiyan Dong, Yanyue Xie, Hongjia Li 0003, Mengshu Sun, Olivia Chen, Nobuyuki Yoshikawa, Yanzhi Wang 0001 |
DAC | 3 |
| 2021 | YOLObile: Real-Time Object Detection on Mobile Devices via Compression-Compilation Co-DesignabstractThe rapid development and wide utilization of object detection techniques have aroused attention on both accuracy and speed of object detectors. However, the current state-of-the-art object detection works are either accuracy-oriented using a large model but leading to high latency or speed-oriented using a lightweight model but sacrificing accuracy. In this work, we propose YOLObile framework, a real-time object detection on mobile devices via compression-compilation co-design. A novel block-punched pruning scheme is proposed for any kernel size. To improve computational efficiency on mobile devices, a GPU-CPU collaborative scheme is adopted along with advanced compiler-assisted optimizations. Experimental results indicate that our pruning scheme achieves 14x compression rate of YOLOv4 with 49.0 mAP. Under our YOLObile framework, we achieve 17 FPS inference speed using GPU on Samsung Galaxy S20. By incorporating our proposed GPU-CPU collaborative scheme, the inference speed is increased to 19.1 FPS, and outperforms the original YOLOv4 by 5x speedup. Source code is at: https://github.com/nightsnack/YOLObile. Yuxuan Cai 0001, Hongjia Li 0003, Geng Yuan, Wei Niu 0002, Yanyu Li, Xulong Tang, Bin Ren 0002, Yanzhi Wang 0001 |
AAAI | 2 |
| 2021 | A Compression-Compilation Co-Design Framework Towards Real-Time Object Detection on Mobile DevicesabstractThe rapid development and wide utilization of object detection techniques have aroused requirements for both accuracy and speed of object detectors. In this work, we propose a compression-compilation co-design framework to achieve real-time YOLOv4 inference on mobile devices. We propose a novel fine-grained structured pruning, which maintain high accuracy while achieving high hardware parallelism. Our pruned YOLOv4 achieves 48.9 mAP and 17 FPS inference speed on an off-the-shelf Samsung Galaxy S20 smartphone, which is 5.5x faster than the original state-of-the-art detector YOLOv4. Yuxuan Cai 0001, Geng Yuan, Hongjia Li 0003, Wei Niu 0002, Yanyu Li, Xulong Tang, Bin Ren 0002, Yanzhi Wang 0001 |
AAAI | 3 |
| 2021 | Real-Time Mobile Acceleration of DNNs: From Computer Vision to Medical ApplicationsabstractWith the growth of mobile vision applications, there is a growing need to break through the current performance limitation of mobile platforms, especially for computationally intensive applications, such as object detection, action recognition, and medical diagnosis. To achieve this goal, we present our unified real-time mobile DNN inference acceleration framework, seamlessly integrating hardware-friendly, structured model compression with mobile-targeted compiler optimizations. We aim at an unprecedented, realtime performance of such large-scale neural network inference on mobile devices. A fine-grained block-based pruning scheme is proposed to be universally applicable to all types of DNN layers, such as convolutional layers with different kernel sizes and fully connected layers. Moreover, it is also successfully extended to 3D convolutions. With the assist of our compiler optimizations, the fine-grained block-based sparsity is fully utilized to achieve high model accuracy and high hardware acceleration simultaneously. To validate our framework, three representative fields of applications are implemented and demonstrated, object detection, activity detection, and medical diagnosis. All applications achieve real-time inference using an off-the-shelf smartphone, outperforming the representative mobile DNN inference acceleration frameworks by up to 6.7x in speed. The demonstrations of these applications can be found in the following link: https://bit.ly/39lWpYu. Hongjia Li 0003, Geng Yuan, Wei Niu 0002, Yuxuan Cai 0001, Mengshu Sun, Zhengang Li 0001, Bin Ren 0002, Xue Lin 0001, Yanzhi Wang 0001 |
ASP-DAC | 1 |
| 2021 | Towards AQFP-Capable Physical Design AutomationabstractAdiabatic Quantum-Flux-Parametron (AQFP) superconducting technology exhibits a high energy efficiency among superconducting electronics, however lacks effective design automation tools. In this work, we develop the first, efficient placement and routing framework for AQFP circuits considering the unique features and constraints, using MIT-LL technology as an example. Our proposed placement framework iteratively executes a fixed-order, row-wise placement algorithm, where the row-wise algorithm derives optimal solution with polynomial-time complexity. To address the maximum wirelength constraint issue in AQFP circuits, a whole row of buffers (or even more rows) is inserted. A* routing algorithm is adopted as the backbone algorithm, incorporating dynamic step size and net negotiation process to reduce the computational complexity accounting for AQFP characteristics, improving overall routability. Extensive experimental results demonstrate the effectiveness of our proposed framework. Hongjia Li 0003, Mengshu Sun, Tianyun Zhang, Olivia Chen, Nobuyuki Yoshikawa, Bei Yu 0001, Yanzhi Wang 0001, Yibo Lin |
DATE | 1 |
| 2020 | Database and Benchmark for Early-stage Malicious Activity Detection in 3D PrintingabstractIncreasing malicious users have sought practices to leverage 3D printing technology to produce unlawful tools in criminal activities. It is of vital importance to enable 3D printers to identify the objects to be printed and terminate at early stage if illegal objects are identified. Deep learning yields significant rises in performance in the object recognition tasks. However, the lack of large-scale databases in 3D printing domain stalls the advancement of automatic illegal weapon recognition. This paper presents a new 3D printing image database, namely C3PO, which compromises two subsets for the different system working scenarios. We extract images from the numerical control programming code files of 22 3D models, and then categorize the images into 10 distinct labels. These two sets are designed for identifying: (i). printing knowledge source (G-code) at beginning of manufacturing, (ii). printing procedure during manufacturing. Importantly, we demonstrate that the weapons can be recognized in either scenario using deep learning based approaches using our proposed database. The quantitative results are promising, and the future exploration of the database and the crime prevention in 3D printing are demanding tasks. Zhe Li 0001, Hongjia Li 0003, Qiyuan An, Qinru Qiu, Wenyao Xu, Yanzhi Wang 0001 |
ASP-DAC | 3 |
| 2020 | An Image Enhancing Pattern-Based Sparsity for Real-Time Inference on Mobile Devices
Wei Niu 0002, Tianyun Zhang, Sijia Liu 0001, Sheng Lin 0001, Hongjia Li 0003, Wujie Wen, Xiang Chen 0010, Jian Tang 0008, Kaisheng Ma, Bin Ren 0002, Yanzhi Wang 0001 |
ECCV (13) | 6 |
| 2020 | ASAP: An Analytical Strategy for AQFP PlacementabstractAdiabatic Quantum-Flux-Parametron (AQFP) is a superconducting logic with very low energy dissipation. Each AQFP cell is driven by AC-power to serve as both power supply and clock signal. The clock signals trigger the data flow from one clock phase to the next clock phase, and the delay for each output in the same phase has to be equal. At the same time, the signal current attenuates as the wire becomes longer. When a wire exceeds a maximum length, the weak current causes incorrect data. Thus, rows of buffers have to be inserted as repeaters to satisfy both delay synchronization and wirelength constraint. These inserted buffers significantly increase the power consumption and also the total delay of AQFP circuits. In this paper, we propose an analytical strategy for AQFP placement (ASAP) to provide effective placement results that greatly reduce the number of additional inserted buffers. ASAP includes two main characteristics: 1) a new wire-length function for analytical global placement and 2) detailed placement including fixed-order Lagrangian relaxation and cell balancing algorithm. Experimental results show the efficiency of ASAP framework and a 53% reduction of buffers over the state-of-the-art method. Yi-Chen Chang, Hongjia Li 0003, Olivia Chen, Yanzhi Wang 0001, Nobuyuki Yoshikawa, Tsung-Yi Ho |
ICCAD | 2 |
| 2020 | New Passive and Active Attacks on Deep Neural Networks in Medical ApplicationsabstractSecurity of deep neural network (DNN) inference engines, i.e., trained DNN models on various platforms, has become one of the biggest challenges in deploying artificial intelligence in domains where privacy, safety, and reliability are of paramount importance, such as in medical applications. In addition to classic software attacks such as model inversion and evasion attacks, recently a new attack surface---implementation attacks which include both passive side-channel attacks and active fault injection and adversarial attacks---is arising, targeting implementation peculiarities of DNN to breach their confidentiality and integrity. This paper presents several novel passive and active attacks on DNN we have developed and tested over medical datasets. Our new attacks reveal a largely under-explored attack surface of DNN inference engines. Insights gained during attack exploration will provide valuable guidance for effectively protecting DNN execution against reverse-engineering and integrity violations. Cheng Gongye, Hongjia Li 0003, Majid Sabbagh, Geng Yuan, Xue Lin 0001, Thomas Wahl, Yunsi Fei |
ICCAD | 2 |
| 2019 | ADMM-based Weight Pruning for Real-Time Deep Learning Acceleration on Mobile DevicesabstractDeep learning solutions are being increasingly deployed in mobile applications, at least for the inference phase. Due to the large model size and computational requirements, model compression for deep neural networks (DNNs) becomes necessary, especially considering the real-time requirement in embedded systems. In this paper, we extend the prior work on systematic DNN weight pruning using ADMM (Alternating Direction Method of Multipliers). We integrate ADMM regularization with masked mapping/retraining, thereby guaranteeing solution feasibility and providing high solution quality. Besides superior performance on representative DNN benchmarks (e.g., AlexNet, ResNet), we focus on two new applications facial emotion detection and eye tracking, and develop a top-down framework of DNN training, model compression, and acceleration in mobile devices. Experimental results show that with negligible accuracy degradation, the proposed method can achieve significant storage/memory reduction and speedup in mobile devices. Hongjia Li 0003, Ning Liu 0007, Sheng Lin 0001, Shaokai Ye, Tianyun Zhang, Xue Lin 0001, Wenyao Xu, Yanzhi Wang 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | Fast and Accurate Trajectory Tracking for Unmanned Aerial Vehicles based on Deep Reinforcement LearningabstractContinuous trajectory control of fixed-wing unmanned aerial vehicles (UAVs) is complicated when considering hidden dynamics. Due to UAV multi degrees of freedom, tracking methodologies based on conventional control theory, such as Proportional-Integral-Derivative (PID) has limitations in response time and adjustment robustness, while a model based approach that calculates the force and torques based on UAV's current status is complicated and rigid. We present an actor-critic reinforcement learning framework that controls UAV trajectory through a set of desired waypoints. A deep neural network is constructed to learn the optimal tracking policy and reinforcement learning is developed to optimize the resulting tracking scheme. The experimental results show that our proposed approach can achieve 58.14% less position error, 21.77% less system power consumption and 9.23% faster attainment than the baseline. The actor network consists of only linear operations, hence Field Programmable Gate Arrays (FPGA) based hardware acceleration can easily be designed for energy efficient real-time control. Yilan Li, Hongjia Li 0003, Zhe Li 0001, Haowen Fang, Amit K. Sanyal, Yanzhi Wang 0001, Qinru Qiu |
RTCSA | 2 |
| 2018 | FFT-based deep learning deployment in embedded systemsabstractDeep learning has delivered its powerfulness in many application domains, especially in image and speech recognition. As the backbone of deep learning, deep neural networks (DNNs) consist of multiple layers of various types with hundreds to thousands of neurons. Embedded platforms are now becoming essential for deep learning deployment due to their portability, versatility, and energy efficiency. The large model size of DNNs, while providing excellent accuracy, also burdens the embedded platforms with intensive computation and storage. Researchers have investigated on reducing DNN model size with negligible accuracy loss. This work proposes a Fast Fourier Transform (FFT)-based DNN training and inference model suitable for embedded platforms with reduced asymptotic complexity of both computation and storage, making our approach distinguished from existing approaches. We develop the training and inference algorithms based on FFT as the computing kernel and deploy the FFT-based inference model on embedded platforms achieving extraordinary processing speed. Sheng Lin 0001, Ning Liu 0007, Mahdi Nazemi, Hongjia Li 0003, Caiwen Ding, Yanzhi Wang 0001, Massoud Pedram |
DATE | 4 |
| 2017 | Deep reinforcement learning: Framework, applications, and embedded implementations: Invited paperabstractThe recent breakthroughs of deep reinforcement learning (DRL) technique in Alpha Go and playing Atari have set a good example in handling large state and actions spaces of complicated control problems. The DRL technique is comprised of (i) an offline deep neural network (DNN) construction phase, which derives the correlation between each state-action pair of the system and its value function, and (ii) an online deep Q-learning phase, which adaptively derives the optimal action and updates value estimates. In this paper, we first present the general DRL framework, which can be widely utilized in many applications with different optimization objectives. This is followed by the introduction of three specific applications: the cloud computing resource allocation problem, the residential smart grid task scheduling problem, and building HVAC system optimal control problem. The effectiveness of the DRL technique in these three cyber-physical applications have been validated. Finally, this paper investigates the stochastic computing-based hardware implementations of the DRL framework, which consumes a significant improvement in area efficiency and power consumption compared with binary-based implementation counterparts. Hongjia Li 0003, Tianshu Wei, Ao Ren, Qi Zhu 0002, Yanzhi Wang 0001 |
ICCAD | 1 |
| 2016 | Dynamic converter reconfiguration for near-threshold non-volatile processors using in-door energy harvestingabstractEnergy harvesting is becoming a preferred choice for future wearable embedded systems compared to batteries because of size, longevity, and maintenance convenience. However, harvested energy is intrinsically unstable. In order to overcome this drawback, non-volatile processors (NVPs) have been proposed to bridge intermittent program execution. However, the harvested power is limited even with multiple energy harvesters when they are in-door. Therefore, a near-threshold processor is ideal to maintain low power consumption. One of the biggest challenges in realizing near-threshold non-volatile processor is to provide a required high write voltage to non-volatile memories when there is a power failure and checkpoint is needed. In order to address this challenge, in this paper, we propose a dynamic converter reconfiguration for ambient energy harvesting-based NVPs to support near-threshold computing. We further investigate thorough optimization techniques to achieve high robustness in reconfiguration and checkpointing, high conversion efficiency, and low ripple magnitude. Experimental results demonstrate that the proposed techniques can significantly reduce the power consumption and improve the performance of energy harvesters and NVPs. Caiwen Ding, Hongjia Li 0003, Jingtong Hu, Yongpan Liu, Yanzhi Wang 0001 |
ICCD | 2 |
| 2016 | Luminescent solar concentrator-based photovoltaic reconfiguration for hybrid and plug-in electric vehiclesabstractAlong with growing public concerns over the energy crisis, hybrid and plug-in electric vehicles (HPEVs) are becoming increasingly popular. However, the total carbon footprint cannot be significantly reduced yet due to the relatively high carbon footprint of batteries in HPEVs. On-board PV systems, which mount PV cells on hood, roof, trunk, and door panels of an HPEV, can assist propelling the vehicle and enable battery charging whenever there is sunlight, and therefore, better mileage can be achieved for HPEVs. A reconfigurable on-board PV system has been proposed to tackle the output power degradation under a non-uniform distribution of solar irradiance levels on different vehicle panels. However, there are still some limitations for mounting PV cells on HPEVs even with the reconfiguration technique such as low efficiency, high cost, and appearance. To address these limitations, we propose to use semiconductor nanomaterials-based luminescent solar concentrators (LSC)-enhanced PV cells for the reconfigurable on-board PV systems. We properly optimize the size of the LSC-enhanced PV cell, the size of macrocells, and the reconfiguration period to achieve a balance between system performance and computation complexity, energy overhead, and capital cost. Furthermore, due to the transparency and flexibility of LSC polymer, we consider employing LSC-enhanced PV cells on vehicle windows. Experiments demonstrate up to 2.49× performance improvement of the proposed LSC-based PV system comparing with the baseline PV system. Caiwen Ding, Hongjia Li 0003, Yanzhi Wang 0001, Naehyuck Chang, Xue Lin 0001 |
ICCD | 2 |