Kai Huang 0002

dblp:86/489-2 · DBLP profile ↗
← Back
37ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0003-2295-5433ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 8 first-author · 14 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorSecurity and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Adaptive All-Digital Clock Data Calibration Circuit for Chiplet Interfaces
Zhenming Li, Pengkang Luo, Guowei Xing, Xiaowen Jiang 0001, Wei Xi 0001, Kai Huang 0002
ISCAS6
2026 An Efficient BIST and Fault Diagnosis Circuit for TSVs
Guowei Xing, Xiaowen Jiang 0001, Pengkang Luo, Zhenming Li, Wei Xi 0001, Kai Huang 0002
ISCAS6
2026 A Piezoelectric Energy-Harvesting Sensor Interface IC With High-Efficient Power Management and Readout Circuitry for Structural Health Monitoring
Dehong Wang, Siyao Cao, Jiankao Pan, Kai Huang 0002, Sijun Du, Zhichao Tan, Menglian Zhao, Shuang Song 0003
IEEE Trans. Circuits Syst. I Regul. Pap.6
2025 A Structure-Aware Irregular Blocking Method for Sparse LU Factorization
Dongliang Xiong, Kai Huang 0002, Changjun Wu, Xiaowen Jiang 0001
ICA3PP (3)3
2025 Quantitative Cost Model and Cost Optimization Methods for Multi-Technology-Node Architecture
abstract
In multi-chiplet design research, cost reduction is a key advantage of heterogeneous integration. Accurate system cost modeling is essential for making informed decisions early in the design process. However, previous cost models overlooked the impact of mixing different technology nodes on overall system cost. In this paper, we propose a quantitative cost model for estimating the cost of the multi-chiplet systems across multiple technology nodes and various chiplet partition granularities. In addition, three optimization strategies are introduced, focusing on partition granularity, technology node, and wafer metal layer, which work synergistically to improve the cost efficiency of 2.5D design. Case studies demonstrate that applying optimization strategies results in 18% to 31% cost savings compared to non-optimized 2.5D designs.
Qimin Yuan, Kai Huang 0002, Xiaowen Jiang 0001, Dongliang Xiong
ICCAD2
2025 An Adaptive Input Voltage Current-Balanced Analog Frontend System for Multiple Cell Li-ion Battery Electrochemical Impedance Monitoring
abstract
This paper proposes a dual-mode AFE system that supports both voltage and electrochemical impedance spectroscopy (EIS) monitoring for multiple cell Li-ion batteries. The proposed current-balanced IA adapts the common mode voltage of the selected cell by taking the same voltage from the same cell, enabling multiple cell monitoring with minimum current. The system also includes a DC-servo loop cancelling the DC component from the AC voltage excited by a current generator. To the best of the author’s knowledge, it is the first BMS AFE system supporting multiple cell EIS monitoring, providing better SoC/SoH estimation and safety for Li-ion batteries.
Yutong Zhang 0015, Dehong Wang, Jiankao Pan, Kai Huang 0002, Menglian Zhao, Shuang Song 0003
ISCAS7
2025 ATEP: An Asynchronous Timing Error Prediction Circuit With Adaptive Voltage and Frequency Scaling
abstract
Timing error prediction circuits have demonstrated greater efficiency in reducing the worst case timing margins of conventional circuits. However, prior works of timing error prediction circuits have complicated the clock tree or introduced a substantial number of delay cells along data paths, resulting in a considerable increase in area overhead. This work introduces asynchronous timing error prediction circuit (ATEP), an ATEP that integrates timing error prediction technology with bundle-data asynchronous templates. The proposed circuit leverages delay lines in the request wire to generate the warning detection window (WDW) independent of clock signals, thereby reducing area overhead and streamlining the clock tree. In addition, we present an adaptive voltage and frequency scaling (AVFS) controller, which evaluates the likelihood of warnings or the quantity of warning paths based on path activation rates to determine when to cease adjustments. This strategy helps to identify the frequency closer to the point of first failure (PoFF). Furthermore, we propose a dynamic warning detector gating strategy to gate warning detectors based on the current environment, further decreasing power consumption. Implementing this circuit on an RISC-V processor, targeting 28-nm CMOS technology, yields up to a 56.8% performance improvement with only a 3.79% area cost and up to a 28.0% reduction in power consumption.
Dongliang Xiong, Huibo Gao, Kai Huang 0002
IEEE Trans. Very Large Scale Integr. Syst.6
2024 Blade-DA: Click-Based Asynchronous Resilient Circuit With Dynamic Delay Adjustment
abstract
Asynchronous resilient circuits have demonstrated higher efficiency and robustness due to their flexible local control mechanism and ability to eliminate metastability, compared to the synchronous resilient circuits. This paper proposes the Blade-DA template, which is an asynchronous resilient circuit that incorporates a delay adjustment scheme. The Blade-DA template utilizes the Click template to simplify circuit design with a larger timing resilient window (TRW). Additionally, the Blade-DA template integrates a dynamic delay adjustment controller, which adaptively tunes the TRW based on real-time timing error rates. In this way, this controller reduces the timing error rate to improve performance under unfavorable Process, Voltage, and Temperature (PVT) corners. Moreover, we propose an optimization method of adding extra constraints to change the delay distribution, which further improves the performance of Blade-DA template. Compared to the previous controller with a large TRW, the proposed asynchronous controller offers a significant 81.65% improvement in maximum frequency and a 61.93% reduction in area while operating at the same frequency. Furthermore, the Blade-DA template is applied to a three-stage RISC-V processor, resulting in a 7.92%-14.78% performance improvement under unfavorable PVT conditions compared to the fixed TRW asynchronous template, while incurring only a 1.63% additional area cost.
Dongliang Xiong, Kai Huang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 A 20.3μW 1.9GΩ Input Impedance Capacitively-Coupled Chopper-Stabilized Amplifier for Bio-Potential Readout
abstract
This paper presents a low-power chopper-stabilized capacitively-coupled frontend amplifier with auxiliary-path-based input impedance ( Z$_{\mathbf{in}}$) boosting. In order to achieve a high Z$_{\mathbf{in}}$, a low noise and a small chip area, techniques on both system level and circuit level are implemented. On the system level, small capacitors ( C$_{\mathbf{in}}$$=$0.5 pF, C$_{\mathbf{fb}}$$=$25 fF) are used with a biased pseudo-resistor ( R$=$2.5 G$\Omega $) fed back to an amplifier internal node. As a result, a high achievable Z$_{\mathbf{in}}$and low high-pass corner frequency are achieved. On the circuit level, an input capacitance shielded current feedback (CSCF) topology achieving effective 10 fF C$_{\mathbf{amp}}$is proposed as the core of the capacitive feedback amplifier in order not to increase the input noise. Moreover, the design space of auxiliary-path-based boosting is explored to obtain the optimal value of buffer bandwidth and auxiliary capacitor size to save power. The amplifier and its Z$_{\mathbf{in}}$boosting circuit are implemented in a standard 55 nm CMOS process and characterized experimentally. Measurement results show that the proposed amplifier provides an input noise density of 50 nV/$\surd$Hz, and an integrated noise of 0.85$\mu$V$_{\mathbf{rms}}$in 200 Hz band. The Z$_{\mathbf{in}}$is boosted to 1.92 G$\Omega $at DC and 1.02 G$\Omega $at 50 Hz with only 1.0$\mu$A in each auxiliary path buffer. The amplifier also archives 77 dB CMRR and 76 dB PSRR while consuming 20.3$\mu$W in total.
Yizhao Zhou, Shuang Song 0003, Yipeng Cao, Feijun Zheng, Kai Huang 0002, Zhichao Tan, Menglian Zhao
IEEE Trans. Circuits Syst. I Regul. Pap.8
2023 Structured Term Pruning for Computational Efficient Neural Networks Inference
abstract
The state-of-the-art convolutional neural network accelerators are showing a growing interest in exploiting the bit-level sparsity and eliminating the ineffectual computations of zero bits. However, the excessive redundancy and the irregular distribution of nonzero bits limit the real speedup in the accelerators. To address this, we propose an algorithm-architecture codesign, named structured term pruning (STP), to boost the computation efficiency of neural networks inference. Specifically, we enhance the bit sparsity by guiding the weights toward the value with fewer power-of-two terms. Then, we structure the terms with layer-wise group budgets. Retraining is adopted to recover the accuracy drop. We also design the hardware of the group processing element and the fast signed-digital encoder for efficient implementation of STP networks. The system design of STP is realized with some easy alterations on an input stationary systolic array design. Extensive evaluation results demonstrate that STP can reduce significant inference computation costs, and achieve$2.35\times $computational energy saving for the ResNet18 network on the ImageNet dataset.
Kai Huang 0002, Bowen Li 0017, Siang Chen, Luc Claesen, Wei Xi 0001, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong, Xiaolang Yan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Efficient Halftoning via Deep Reinforcement Learning
abstract
Halftoning aims to reproduce a continuous-tone image with pixels whose intensities are constrained to two discrete levels. This technique has been deployed on every printer, and the majority of them adopt fast methods (e.g., ordered dithering, error diffusion) that fail to render structural details, which determine halftone's quality. Other prior methods of pursuing visual pleasure by searching for the optimal halftone solution, on the contrary, suffer from their high computational cost. In this paper, we propose a fast and structure-aware halftoning method via a data-driven approach. Specifically, we formulate halftoning as a reinforcement learning problem, in which each binary pixel's value is regarded as an action chosen by a virtual agent with a shared fully convolutional neural network (CNN) policy. In the offline phase, an effective gradient estimator is utilized to train the agents in producing high-quality halftones in one action step. Then, halftones can be generated online by one fast CNN inference. Besides, we propose a novel anisotropy suppressing loss function, which brings the desirable blue-noise property. Finally, we find that optimizing SSIM could result in holes in flat areas, which can be avoided by weighting the metric with the contone's contrast map. Experiments show that our framework can effectively train a light-weight CNN, which is 15x faster than previous structure-aware methods, to generate blue-noise halftones with satisfactory visual quality. We also present a prototype of deep multitoning to demonstrate the extensibility of our method.
Haitian Jiang, Dongliang Xiong, Xiaowen Jiang 0001, Kai Huang 0002
IEEE Trans. Image Process.6
2023 Structured Dynamic Precision for Deep Neural Networks Quantization
abstract
Deep Neural Networks (DNNs) have achieved remarkable success in various Artificial Intelligence applications. Quantization is a critical step in DNNs compression and acceleration for deployment. To further boost DNN execution efficiency, many works explore to leverage the input-dependent redundancy with dynamic quantization for different regions. However, the sensitive regions in the feature map are irregularly distributed, which restricts the real speed up for existing accelerators. To this end, we propose an algorithm-architecture co-design, named Structured Dynamic Precision (SDP). Specifically, we propose a quantization scheme in which the high-order bit part and the low-order bit part of data can be masked independently. And a fixed number of term parts are dynamically selected for computation based on the importance of each term in the group. We also present a hardware design to enable the algorithm efficiently with small overheads, whose inference time mainly scales with the precision proportionally. Evaluation experiments on extensive networks demonstrate that compared to the state-of-the-art dynamic quantization accelerator DRQ, our SDP can achieve 29% performance gain and 51% energy reduction for the same level of model accuracy.
Kai Huang 0002, Bowen Li 0017, Dongliang Xiong, Haitian Jiang, Xiaowen Jiang 0001, Xiaolang Yan, Luc Claesen, Dehong Liu, Junjian Chen, Zhili Liu
ACM Trans. Design Autom. Electr. Syst.1
2022 Halftoning with Multi-Agent Deep Reinforcement Learning
abstract
Deep neural networks have recently succeeded in digital halftoning using vanilla convolutional layers with high parallelism. However, existing deep methods fail to generate halftones with a satisfying blue-noise property and require complex training schemes. In this paper, we propose a halftoning method based on multi-agent deep reinforcement learning, called HALFTONERS, which learns a shared policy to generate high-quality halftone images. Specifically, we view the decision of each binary pixel value as an action of a virtual agent, whose policy is trained by a low-variance policy gradient. Moreover, the blue-noise property is achieved by a novel anisotropy suppressing loss function. Experiments show that our halftoning method produces high-quality halftones while staying relatively fast.
Haitian Jiang, Dongliang Xiong, Xiaowen Jiang 0001, Aiguo Yin, Kai Huang 0002
ICIP6
2022 Structured precision skipping: Accelerating convolutional neural networks with budget-aware dynamic precision selection
Kai Huang 0002, Siang Chen, Bowen Li 0017, Luc Claesen, Hao Yao, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong
J. Syst. Archit.1
2022 Acceleration-Aware Fine-Grained Channel Pruning for Deep Neural Networks via Residual Gating
abstract
Deep neural networks have achieved remarkable advancement in various intelligence tasks. However, the massive computation and storage consumption limit applications on resource-constrained devices. While channel pruning has been widely applied to compress models, it is challenging to reach very deep compressions for such a coarse-grained pruning structure without significant performance degradation. In this article, we propose an acceleration-aware fine-grained channel pruning (AFCP) framework for accelerating neural networks, which optimizes trainable gate parameters by estimating residual errors between pruned and original channels with hardware characteristics. Our fine-grained concept consists of both algorithm and structure levels. Different from existing methods that leverage a predefined pruning criterion, AFCP explicitly considers both zero-out and similar criteria for each channel, and adaptively selects the suitable one via residual gate parameters. For structure level, AFCP adopts a fine-grained channel pruning strategy for residual neural networks and a decomposition-based structure, which further extends the pruning optimization space. Moreover, instead of using theoretical computation costs, such as floating-point operations, we propose the hardware predictor that bridges the gap between realistic acceleration and pruning procedure to guide the learning of pruning, which improves the efficiency of model pruning when deployed on accelerators. Extensive evaluation results demonstrate that AFCP outperforms state-of-the-art methods, and achieves a favorable balance between model performance and computation cost.
Kai Huang 0002, Siang Chen, Bowen Li 0017, Luc Claesen, Hao Yao, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2021 DPOQ: Dynamic Precision Onion Quantization
abstract
With the development of deployment platforms and application scenarios for deep neural networks, traditional fixed network architectures cannot meet the requirements. Meanwhile the dynamic network inference becomes a new research trend. Many slimmable and scalable networks have been proposed to satisfy different resource constraints (e.g., storage, latency and energy). And a single network may support versatile architectural configurations including: depth, width, kernel size, and resolution. In this paper, we propose a novel network architecture reuse strategy enabling dynamic precision in parameters. Since our low-precision networks are wrapped in the high-precision networks like an onion, we name it dynamic precision onion quantization (DPOQ). We train the network by using the joint loss with scaled gradients. To further improve the performance and make different precision network compatible with each other, we propose the precision shift batch normalization (PSBN). And we also propose a scalable input-specific inference mechanism based on this architecture and make the network more adaptable. Experiments on the CIFAR and ImageNet dataset have shown that our DPOQ achieves not only better flexibility but also higher accuracy than the individual quantization.
Bowen Li 0017, Kai Huang 0002, Siang Chen, Dongliang Xiong, Luc Claesen
ACML2
2021 Expected Energy Optimization for Real-Time Multiprocessor SoCs Running Periodic Tasks with Uncertain Execution Time
abstract
Energy optimization plays an increasingly critical role in designing an embedded real-time multiprocessor System on Chip (MPSoC). Dynamic Voltage Frequency Scaling (DVFS) and Dynamic Power Management (DPM) are preferable techniques to optimize energy consumption. However, previous DVFS and DPM algorithms were mostly designed for inter-task scheduling, without sufficient exploration on intra-task scheduling for further energy reduction. This paper presents a new intra-task scheduling approach considering the probabilistic distribution of task execution time, and it optimizes the mathematical expectation of power consumption (expected power consumption) for periodic dependent tasks with uncertain execution time running on MPSoCs using DVFS and DPM. The energy-efficient scheduling problem can be formulated by means of mixed integer linear programming (MILP) with the proposed technique. Moreover, we also propose a technique to compress the exploration space by reorganizing the probabilistic profiling information of all tasks. Our experimental results on synthetic and realistic benchmarks show that the proposed approach achieves up to 30 percent energy savings compared with other existing methods.
Kai Huang 0002, Ke Wang 0034, Dandan Zheng 0001, Xiaowen Jiang 0001, Rongjie Yan, Xiaolang Yan
IEEE Trans. Sustain. Comput.1
2020 DFQF: Data Free Quantization-aware Fine-tuning
abstract
Data free deep neural network quantization is a practical challenge, since the original training data is often unavailable due to some privacy, proprietary or transmission issues. The existing methods implicitly equate data-free with training-free and quantize model manually through analyzing the weights’ distribution. It leads to a significant accuracy drop in lower than 6-bit quantization. In this work, we propose the data free quantization-aware fine-tuning (DFQF), wherein no real training data is required, and the quantized network is fine-tuned with generated images. Specifically, we start with training a generator from the pre-trained full-precision network with inception score loss, batch-normalization statistics loss and adversarial loss to synthesize a fake image set. Then we fine-tune the quantized student network with the full-precision teacher network and the generated images by utilizing knowledge distillation (KD). The proposed DFQF outperforms state-of-the-art post-train quantization methods, and achieve W4A4 quantization of ResNet20 on the CIFAR10 dataset within 1% accuracy drop.
Bowen Li 0017, Kai Huang 0002, Siang Chen, Dongliang Xiong, Haitian Jiang, Luc Claesen
ACML2
2020 Fine-Grained Channel Pruning for Deep Residual Neural Networks
Siang Chen, Kai Huang 0002, Dongliang Xiong, Bowen Li 0017, Luc Claesen
ICANN (2)2
2020 Trigger Identification Using Difference-Amplified Controllability and Dynamic Transition Probability for Hardware Trojan Detection
abstract
To remain dormant in the validation and manufacturing test, Trojans tend to have at least one trigger signal at the gate-level netlist with a very low transition probability. Our paper exploits this stealthy nature of trigger signals to detect Trojans using static and dynamic transition probabilities. The proposed trigger identification is a reference-free scheme, and no prior knowledge of a Trojan-free design is required. First, we reveal the relation between combinational 0/1-controllability and 0/1-probability and propose a static transition probability analysis based on our proposed difference-amplified controllability, which can be easily obtained by the Sandia Controllability/Observability Analysis Program. The k-means clustering method is adopted for potential trigger classification to extend the scalability and adaptability to different circuit sizes. Second, we propose to utilize the transition probability of a dynamic simulation for correction of the results. Experiments show that the proposed detection scheme can obtain a 0% false negative rate and a maximum 11.7% false positive rate on Trust-HUB benchmarks.
Kai Huang 0002
IEEE Trans. Inf. Forensics Secur.1
2019 A Scalable and Adaptable ILP-Based Approach for Task Mapping on MPSoC Considering Load Balance and Communication Optimization
abstract
Task mapping has been a hot topic in multiprocessor system-on-chip software design for decades. During the mapping process, load balance (LB) and communication optimization have been two important performance optimization factors. This paper studies the relations between LB, interprocessor communications, and communication pipeline technique during the mapping process, and proposes an integer linear programming (ILP)-based static task mapping approach, which considers both LB and communication optimization. The approach consists of an optimized ILP model for task mapping with fewer variables compared to previous ILP mapping works. Moreover, to enhance the scalability of the ILP task mapping, the task-processor-cluster algorithm is proposed to reduce the scale of the task graph and the number of processors and then solve the coarse-grained input by the ILP mapping. To increase the adaptability of the ILP task mapping, the improved augmented E-constraint method is further integrated with the ILP formulations to select the best mapping for different applications. Experimental results on a 2/4/8/16/24-CPU platform of both synthetic and real-life benchmarks demonstrate the efficiency of the proposed approach.
Kai Huang 0002, Dandan Zheng 0001, Min Yu 0006, Xiaowen Jiang 0001, Xiaolang Yan, Lisane B. de Brisolara, Ahmed Amine Jerraya
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2019 A Real-Time High-Quality Complete System for Depth Image-Based Rendering on FPGA
abstract
Depth image-based rendering (DIBR) techniques have recently drawn more attention in various 3D applications. In this paper, a real-time high-quality DIBR system that consists of disparity estimation and view synthesis is proposed. For disparity estimation, a local approach that focuses on depth discontinuities and disparity smoothness is presented to improve the disparity accuracy. For view synthesis, a method that contains view interpolation and extrapolation is proposed to render high-quality virtual views. Moreover, the system is designed with an optimized parallelism scheme to achieve a high throughput, and can be scaled up easily. It is implemented on an Altera Stratix IV FPGA at a processing speed of 45 frames per second for 1080p resolution. Evaluated on selected image sets of the Middlebury benchmark, the average error rate of the disparity maps is 6.02%; the average peak signal to noise ratio and structural similarity values of the virtual views are 30.07 dB and 0.9303, respectively. The experimental results indicate that the proposed DIBR system has the top-performing processing speed and its accuracy performance is among the best of state-of-the-art hardware implementations.
Luc Claesen, Kai Huang 0002, Menglian Zhao
IEEE Trans. Circuits Syst. Video Technol.3
2017 High-quality view interpolation based on depth maps and its hardware implementation
abstract
Three dimensional (3D) vision applications have drawn more attention nowadays and many products are entering the mass market. View interpolation is a crucial step to generate intermediate viewpoints from reference images. However, it is still challenging to achieve good performance in both processing speed and image quality for various 3D applications. In this paper, a hardware-compatible view interpolation algorithm is proposed, which can produce high-quality intermediate images by disparity warping and color blending. Moreover, a fully pipelined hardware architecture is designed based on the algorithm. A prototype of the proposed architecture has been implemented on an Altera Stratix-IV FPGA board, achieving 65 frames per second (fps) with a full HD (1920 × 1080) resolution. It is evaluated on the Middlebury benchmark quantitatively, and visual results of real-world images are also provided.
Kai Huang 0002, Luc Claesen
FPL2
2017 User Perceived Value-Aware Cloud Pricing for Profit Maximization of Multiserver Systems
abstract
With the rapid deployment of cloud computing infrastructures, understanding the economics of cloud computing has becoming a pressing issue for cloud service providers. However, existing pricing models rarely consider the dynamic interaction between user requests and the cloud service provider, thus can not accurately reflect the law of supply and demand in marketing. In this paper, we propose a pricing model based on the concept of user perceived value in the domain of economics that accurately capture the real supply and demand situation in the cloud service market. We then design a profit maximization scheme based on the presented dynamic pricing model that optimizes profit of the cloud service provider without violating user service-level agreement. Extensive experiments using data extracted from real-world applications validate the effectiveness of the proposed user perceived value-based pricing model. The proposed profit maximization scheme achieves 24.44% more profit as compared to the state of the art benchmarking methods.
Peijin Cong, Liying Li 0002, Gaoyuan Shao, Junlong Zhou, Mingsong Chen 0001, Kai Huang 0002, Tongquan Wei
ICPADS6
2017 A Hybrid Multi-objective Evolutionary Algorithm for Energy-Aware Allocation and Scheduling Optimization of MPSoCs
abstract
MPSoCs are increasingly being adopted in the design of emerging complex embedded systems. Resource limitations require designers to find optimizations among various design considerations. Task mapping and scheduling become one of the key issues in designing such systems. To meet the requirements of makespan minimization and workload balance for energy-aware MPSoCs, the paper presents a unified formulation to find satisfied task mapping and scheduling solutions. The model considers both computation and communication cost, and enables applying dynamic power management (DPM) for energy optimization. To efficiently approximate the Pareto front of the optimization problem, we propose a multi-objective hybrid algorithm (MOHA) by integrating a Pareto local search into an evolutionary process, with a problem-specific initialization. Experimental results from realistic benchmarks demonstrate that the proposed techniques are able to generate high-quality solutions of realistic applications on the target architecture, compared with state-of-the-art methods.
Rongjie Yan, Yupeng Zhou, Yige Yan, Minghao Yin, Min Yu 0006, Feifei Ma, Kai Huang 0002
ICTAI7
2017 Providing Predictable Performance via a Slowdown Estimation Model
abstract
Interapplication interference at shared main memory slows down different applications differently. A few slowdown estimation models have been proposed to provide predictable performance by quantifying memory interference, but they have relatively low accuracy. Thus, we propose a more accurate slowdown estimation model called SEM at main memory. First, SEM unifies the slowdown estimation model by measuring IPC directly. Second, SEM uses the per-bank structure to monitor memory interference and improves estimation accuracy by considering write interference, row-buffer interference, and data bus interference. The evaluation results show that SEM has significantly lower slowdown estimation error (4.06%) compared to STFM (30.15%) and MISE (10.1%).
Dongliang Xiong, Kai Huang 0002, Xiaowen Jiang 0001, Xiaolang Yan
ACM Trans. Archit. Code Optim.2
2016 SoC and FPGA oriented high-quality stereo vision system
abstract
Stereo matching is a crucial step for acquiring depth information from stereo images. However, it is still challenging to achieve good performance in both speed and accuracy for various stereo vision applications. In this paper, a hardware-compatible stereo matching algorithm is proposed; its associated hardware implementation is also presented. The proposed algorithm can produce high-quality disparity maps with the use of mini-census transform, segmentation-based adaptive support weight and effective refinement. Moreover, the proposed implementation is optimized as a fully pipelined and scalable hardware system. The proposed design is evaluated based on the Middlebury benchmarks and the average overall error rate is 6.10%. The experimental results indicate that the accuracy is competitive with some state-of-art software implementations.
Kai Huang 0002, Luc Claesen
FPL2
2016 SoC oriented real-time high-quality stereo vision system
abstract
Stereo matching is a crucial step to extract depth information from stereo images. However, it is still challenging to achieve good performance in both speed and accuracy for various stereo vision applications. In this paper, a hardware-compatible stereo matching algorithm is proposed and its associated hardware implementation is also presented. The proposed algorithm can produce high-quality disparity maps with the combined use of the mini-census transform, segmentation-based adaptive support weight and effective refinement. Moreover, the proposed architecture is optimized as a fully pipelined and scalable hardware system. Implemented on an Altera Stratix-IV FPGA board, it can achieve 65 frames per second (fps) for 1024 × 768 stereo images and a 64 pixel disparity range. The proposed architecture is evaluated based on the Middlebury benchmarks and the average error rate is 6.56%. The experimental results indicate that the accuracy is competitive with some state-of-the-art software implementations.
Kai Huang 0002, Luc Claesen
VLSI-SoC2
2016 Memory Access Scheduling Based on Dynamic Multilevel Priority in Shared DRAM Systems
abstract
Interapplication interference at shared main memory severely degrades performance and increasing DRAM frequency calls for simple memory schedulers. Previous memory schedulers employ a per-application ranking scheme for high system performance or a per-group ranking scheme for low hardware cost, but few provide a balance. We propose DMPS, a memory scheduler based on dynamic multilevel priority. First, DMPS uses “memory occupancy” to measure interference quantitatively. Second, DMPS groups applications, favors latency-sensitive groups, and dynamically prioritizes applications by employing a per-level ranking scheme. The simulation results show that DMPS has 7.2% better system performance and 22% better fairness over FRFCFS at low hardware complexity and cost.
Dongliang Xiong, Kai Huang 0002, Xiaowen Jiang 0001, Xiaolang Yan
ACM Trans. Archit. Code Optim.2
2015 Profiling and annotation combined method for multimedia application specific MPSoC performance estimation
abstract
Accurate and fast performance estimation is necessary to drive design space exploration and thus support important design decisions. Current techniques are either time consuming or not accurate enough. In this paper, we solve these problems by presenting a hybrid method for multimedia multiprocessor system-on-chip (MPSoC) performance estimation. A general coverage analysis tool GNU gcov is employed to profile the execution statistics during the native simulation. To tackle the complexity and keep the analysis and simulation manageable, the orthogonalization of communication and computation parts is adopted. The estimation result of the computation part is annotated to a transaction accurate model for further analysis, by which a gradual refinement of MPSoC performance estimation is supported. The implementation and its experimental results prove the feasibility and efficiency of the proposed method.
Kai Huang 0002, Siwen Xiu, Dandan Zheng 0001, Min Yu 0006, De Ma, Kai Huang 0001, Gang Chen 0023, Xiaolang Yan
Frontiers Inf. Technol. Electron. Eng.1
2015 Communication Optimizations for Multithreaded Code Generation from Simulink Models
abstract
Communication frequency is increasing with the growing complexity of emerging embedded applications and the number of processors in the implemented multiprocessor SoC architectures. In this article, we consider the issue of communication cost reduction during multithreaded code generation from partitioned Simulink models to help designers in code optimization to improve system performance. We first propose a technique combining message aggregation and communication pipeline methods, which groups communications with the same destinations and sources and parallelizes communication and computation tasks. We also present a method to apply static analysis and dynamic emulation for efficient communication buffer allocation to further reduce synchronization cost and increase processor utilization. The existing cyclic dependency in the mapped model may hinder the effectiveness of the two techniques. We further propose a set of optimizations involving repartition with strongly connected threads to maximize the degree of communication reduction and preprocessing strategies with available delays in the model to reduce the number of communication channels that cannot be optimized. Experimental results demonstrate the advantages of the proposed optimizations with 11--143% throughput improvement.
Kai Huang 0002, Min Yu 0006, Rongjie Yan, Xiaolang Yan, Lisane B. de Brisolara, Ahmed Amine Jerraya, Jiong Feng
ACM Trans. Embed. Comput. Syst.1
2014 Annotation and analysis combined cache modeling for native simulation
abstract
To accelerate the speed of performance estimation and raise its accuracy for MPSoC, we propose a static analysis and dynamic annotation combined method to efficiently model cache mechanism in native simulation. We use a new cache model to statically analyze segmental profiling results to speed up simulation, and utilize a dynamic annotation technique to exactly trace the addresses of local variables. Experimental results show the efficiency of the proposed techniques for more accurate system performance estimation.
Rongjie Yan, De Ma, Kai Huang 0002, Siwen Xiu
ASP-DAC3
2013 High throughput VLSI architecture for H.264/AVC context-based adaptive binary arithmetic coding (CABAC) decoding
abstract
Context-based adaptive binary arithmetic coding (CABAC) is the major entropy-coding algorithm employed in H.264/AVC. In this paper, we present a new VLSI architecture design for an H.264/AVC CABAC decoder, which optimizes both decode decision and decode bypass engines for high throughput, and improves context model allocation for efficient external memory access. Based on the fact that the most possible symbol (MPS) branch is much simpler than the least possible symbol (LPS) branch, a newly organized decode decision engine consisting of two serially concatenated MPS branches and one LPS branch is proposed to achieve better parallelism at lower timing path cost. A look-ahead context index (ctxIdx) calculation mechanism is designed to provide the context model for the second MPS branch. A head-zero detector is proposed to improve the performance of the decode bypass engine according to UEG k encoding features. In addition, to lower the frequency of memory access, we reorganize the context models in external memory and use three circular buffers to cache the context models, neighboring information, and bit stream, respectively. A pre-fetching mechanism with a prediction scheme is adopted to load the corresponding content to a circular buffer to hide external memory latency. Experimental results show that our design can operate at 250 MHz with a 20.71k gate count in SMIC18 silicon technology, and that it achieves an average data decoding rate of 1.5 bins/cycle.
Kai Huang 0002, De Ma, Rongjie Yan, Haitong Ge, Xiaolang Yan
J. Zhejiang Univ. Sci. C1
2013 Performance Estimation Techniques With MPSoC Transaction-Accurate Models
abstract
Efficient design of multiprocessor system-on-chip (MPSoC) requires early, fast, and accurate performance estimation techniques. In this paper, we present new techniques based on fine-grained code analysis to estimate accurate performance during simulation of MPSoC transaction accurate models. First, a GCC profiling tool is applied in the native simulation process. Based on the profiling result, an instruction analyzer of the target CPU architecture is proposed to analyze the cycle cost of C code under estimation. In addition, a memory analyzer is used to further estimate memory access latency including both instruction/data cache time cost and global memory access cycles. Both data and instruction cache models are proposed to estimate cache miss penalty, and a segment-based strategy is adopted to update the cache models more efficiently. Furthermore, an equalized access model is presented to imitate the memory access behavior of processors for estimating global memory access latency caused by bus contention and memory bandwidth. We have applied these techniques on an H.264 decoder application with different hardware architectures. The experimental results show that applying these techniques can obviously improve estimation accuracy of transaction accurate models close to that of the virtual prototype models, with a tolerable overhead on simulation speed.
De Ma, Rongjie Yan, Kai Huang 0002, Min Yu 0006, Siwen Xiu, Haitong Ge, Xiaolang Yan, Ahmed Amine Jerraya
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2010 A high efficient memory architecture for H.264/AVC motion compensation
abstract
In H.264/AVC decoding system, motion compensation operation occupies about 80% of the total memory access and becomes the system bottleneck. In this paper, a high efficient memory architecture for H.264/AVC motion compensation is proposed to extremely reduce external memory access bandwidth. A four-level hierarchical memory organization scheme is utilized to explore the reusability of neighboring blocks at an acceptable area cost. To improve the system processing throughput, five optimization techniques are adopted in motion compensation operation, which enable video decoder to achieve real-time decoding of HD 1080p video stream when operating at 110 MHz. Compared with the existing works, the proposed architecture is able to reduce the memory bandwidth requirement in motion compensation progress by 83.7% and performs better in the real-time application.
Chunshu Li, Kai Huang 0002, Xiaolang Yan, Jiong Feng, De Ma, Haitong Ge
ASAP2
2009 Simulink®-based heterogeneous multiprocessor SoC design flow for mixed hardware/software refinement and simulation
Sangil Han, Soo-Ik Chae, Lisane B. de Brisolara, Luigi Carro, Katalin Popovici, Xavier Guerin, Ahmed Amine Jerraya, Kai Huang 0002, Xiaolang Yan
Integr.8
2007 Simulink-Based MPSoC Design Flow: Case Study of Motion-JPEG and H.264
abstract
System-level design methodologies have been introduced as a solution to handle the design complexity of embedded multiprocessor SoC (MPSoC) systems. In this paper we describe a system-level design flow starting from Simulink specification, focusing on concurrent hardware and software design and verification at four different abstraction levels: Simulink Combined Algorithm and Architecture Model (CAAM), Virtual Architecture, Transaction-accurate Model and Virtual Prototype. We used two multimedia applications, Motion-JPEG and H.264, to evaluate this design flow. Experimental results show that our design flow can generate various MPSoC architectures from Simulink CAAM correctly and efficiently, allowing processor and task design space exploration at different abstraction levels.
Kai Huang 0002, Sangil Han, Katalin Popovici, Lisane B. de Brisolara, Xavier Guerin, Xiaolang Yan, Soo-Ik Chae, Luigi Carro, Ahmed Amine Jerraya
DAC1