Xiaowen Jiang 0001

dblp:192/3782-1 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-6283-2262ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Adaptive All-Digital Clock Data Calibration Circuit for Chiplet Interfaces
Zhenming Li, Pengkang Luo, Guowei Xing, Xiaowen Jiang 0001, Wei Xi 0001, Kai Huang 0002
ISCAS4
2026 An Efficient BIST and Fault Diagnosis Circuit for TSVs
Guowei Xing, Xiaowen Jiang 0001, Pengkang Luo, Zhenming Li, Wei Xi 0001, Kai Huang 0002
ISCAS2
2025 A Structure-Aware Irregular Blocking Method for Sparse LU Factorization
Dongliang Xiong, Kai Huang 0002, Changjun Wu, Xiaowen Jiang 0001
ICA3PP (3)5
2025 Quantitative Cost Model and Cost Optimization Methods for Multi-Technology-Node Architecture
abstract
In multi-chiplet design research, cost reduction is a key advantage of heterogeneous integration. Accurate system cost modeling is essential for making informed decisions early in the design process. However, previous cost models overlooked the impact of mixing different technology nodes on overall system cost. In this paper, we propose a quantitative cost model for estimating the cost of the multi-chiplet systems across multiple technology nodes and various chiplet partition granularities. In addition, three optimization strategies are introduced, focusing on partition granularity, technology node, and wafer metal layer, which work synergistically to improve the cost efficiency of 2.5D design. Case studies demonstrate that applying optimization strategies results in 18% to 31% cost savings compared to non-optimized 2.5D designs.
Qimin Yuan, Kai Huang 0002, Xiaowen Jiang 0001, Dongliang Xiong
ICCAD3
2023 Structured Term Pruning for Computational Efficient Neural Networks Inference
abstract
The state-of-the-art convolutional neural network accelerators are showing a growing interest in exploiting the bit-level sparsity and eliminating the ineffectual computations of zero bits. However, the excessive redundancy and the irregular distribution of nonzero bits limit the real speedup in the accelerators. To address this, we propose an algorithm-architecture codesign, named structured term pruning (STP), to boost the computation efficiency of neural networks inference. Specifically, we enhance the bit sparsity by guiding the weights toward the value with fewer power-of-two terms. Then, we structure the terms with layer-wise group budgets. Retraining is adopted to recover the accuracy drop. We also design the hardware of the group processing element and the fast signed-digital encoder for efficient implementation of STP networks. The system design of STP is realized with some easy alterations on an input stationary systolic array design. Extensive evaluation results demonstrate that STP can reduce significant inference computation costs, and achieve$2.35\times $computational energy saving for the ResNet18 network on the ImageNet dataset.
Kai Huang 0002, Bowen Li 0017, Siang Chen, Luc Claesen, Wei Xi 0001, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong, Xiaolang Yan
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2023 Efficient Halftoning via Deep Reinforcement Learning
abstract
Halftoning aims to reproduce a continuous-tone image with pixels whose intensities are constrained to two discrete levels. This technique has been deployed on every printer, and the majority of them adopt fast methods (e.g., ordered dithering, error diffusion) that fail to render structural details, which determine halftone's quality. Other prior methods of pursuing visual pleasure by searching for the optimal halftone solution, on the contrary, suffer from their high computational cost. In this paper, we propose a fast and structure-aware halftoning method via a data-driven approach. Specifically, we formulate halftoning as a reinforcement learning problem, in which each binary pixel's value is regarded as an action chosen by a virtual agent with a shared fully convolutional neural network (CNN) policy. In the offline phase, an effective gradient estimator is utilized to train the agents in producing high-quality halftones in one action step. Then, halftones can be generated online by one fast CNN inference. Besides, we propose a novel anisotropy suppressing loss function, which brings the desirable blue-noise property. Finally, we find that optimizing SSIM could result in holes in flat areas, which can be avoided by weighting the metric with the contone's contrast map. Experiments show that our framework can effectively train a light-weight CNN, which is 15x faster than previous structure-aware methods, to generate blue-noise halftones with satisfactory visual quality. We also present a prototype of deep multitoning to demonstrate the extensibility of our method.
Haitian Jiang, Dongliang Xiong, Xiaowen Jiang 0001, Kai Huang 0002
IEEE Trans. Image Process.3
2023 Structured Dynamic Precision for Deep Neural Networks Quantization
abstract
Deep Neural Networks (DNNs) have achieved remarkable success in various Artificial Intelligence applications. Quantization is a critical step in DNNs compression and acceleration for deployment. To further boost DNN execution efficiency, many works explore to leverage the input-dependent redundancy with dynamic quantization for different regions. However, the sensitive regions in the feature map are irregularly distributed, which restricts the real speed up for existing accelerators. To this end, we propose an algorithm-architecture co-design, named Structured Dynamic Precision (SDP). Specifically, we propose a quantization scheme in which the high-order bit part and the low-order bit part of data can be masked independently. And a fixed number of term parts are dynamically selected for computation based on the importance of each term in the group. We also present a hardware design to enable the algorithm efficiently with small overheads, whose inference time mainly scales with the precision proportionally. Evaluation experiments on extensive networks demonstrate that compared to the state-of-the-art dynamic quantization accelerator DRQ, our SDP can achieve 29% performance gain and 51% energy reduction for the same level of model accuracy.
Kai Huang 0002, Bowen Li 0017, Dongliang Xiong, Haitian Jiang, Xiaowen Jiang 0001, Xiaolang Yan, Luc Claesen, Dehong Liu, Junjian Chen, Zhili Liu
ACM Trans. Design Autom. Electr. Syst.5
2022 Halftoning with Multi-Agent Deep Reinforcement Learning
abstract
Deep neural networks have recently succeeded in digital halftoning using vanilla convolutional layers with high parallelism. However, existing deep methods fail to generate halftones with a satisfying blue-noise property and require complex training schemes. In this paper, we propose a halftoning method based on multi-agent deep reinforcement learning, called HALFTONERS, which learns a shared policy to generate high-quality halftone images. Specifically, we view the decision of each binary pixel value as an action of a virtual agent, whose policy is trained by a low-variance policy gradient. Moreover, the blue-noise property is achieved by a novel anisotropy suppressing loss function. Experiments show that our halftoning method produces high-quality halftones while staying relatively fast.
Haitian Jiang, Dongliang Xiong, Xiaowen Jiang 0001, Aiguo Yin, Kai Huang 0002
ICIP3
2022 Structured precision skipping: Accelerating convolutional neural networks with budget-aware dynamic precision selection
Kai Huang 0002, Siang Chen, Bowen Li 0017, Luc Claesen, Hao Yao, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong
J. Syst. Archit.7
2022 Acceleration-Aware Fine-Grained Channel Pruning for Deep Neural Networks via Residual Gating
abstract
Deep neural networks have achieved remarkable advancement in various intelligence tasks. However, the massive computation and storage consumption limit applications on resource-constrained devices. While channel pruning has been widely applied to compress models, it is challenging to reach very deep compressions for such a coarse-grained pruning structure without significant performance degradation. In this article, we propose an acceleration-aware fine-grained channel pruning (AFCP) framework for accelerating neural networks, which optimizes trainable gate parameters by estimating residual errors between pruned and original channels with hardware characteristics. Our fine-grained concept consists of both algorithm and structure levels. Different from existing methods that leverage a predefined pruning criterion, AFCP explicitly considers both zero-out and similar criteria for each channel, and adaptively selects the suitable one via residual gate parameters. For structure level, AFCP adopts a fine-grained channel pruning strategy for residual neural networks and a decomposition-based structure, which further extends the pruning optimization space. Moreover, instead of using theoretical computation costs, such as floating-point operations, we propose the hardware predictor that bridges the gap between realistic acceleration and pruning procedure to guide the learning of pruning, which improves the efficiency of model pruning when deployed on accelerators. Extensive evaluation results demonstrate that AFCP outperforms state-of-the-art methods, and achieves a favorable balance between model performance and computation cost.
Kai Huang 0002, Siang Chen, Bowen Li 0017, Luc Claesen, Hao Yao, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2021 Expected Energy Optimization for Real-Time Multiprocessor SoCs Running Periodic Tasks with Uncertain Execution Time
abstract
Energy optimization plays an increasingly critical role in designing an embedded real-time multiprocessor System on Chip (MPSoC). Dynamic Voltage Frequency Scaling (DVFS) and Dynamic Power Management (DPM) are preferable techniques to optimize energy consumption. However, previous DVFS and DPM algorithms were mostly designed for inter-task scheduling, without sufficient exploration on intra-task scheduling for further energy reduction. This paper presents a new intra-task scheduling approach considering the probabilistic distribution of task execution time, and it optimizes the mathematical expectation of power consumption (expected power consumption) for periodic dependent tasks with uncertain execution time running on MPSoCs using DVFS and DPM. The energy-efficient scheduling problem can be formulated by means of mixed integer linear programming (MILP) with the proposed technique. Moreover, we also propose a technique to compress the exploration space by reorganizing the probabilistic profiling information of all tasks. Our experimental results on synthetic and realistic benchmarks show that the proposed approach achieves up to 30 percent energy savings compared with other existing methods.
Kai Huang 0002, Ke Wang 0034, Dandan Zheng 0001, Xiaowen Jiang 0001, Rongjie Yan, Xiaolang Yan
IEEE Trans. Sustain. Comput.4
2019 A Scalable and Adaptable ILP-Based Approach for Task Mapping on MPSoC Considering Load Balance and Communication Optimization
abstract
Task mapping has been a hot topic in multiprocessor system-on-chip software design for decades. During the mapping process, load balance (LB) and communication optimization have been two important performance optimization factors. This paper studies the relations between LB, interprocessor communications, and communication pipeline technique during the mapping process, and proposes an integer linear programming (ILP)-based static task mapping approach, which considers both LB and communication optimization. The approach consists of an optimized ILP model for task mapping with fewer variables compared to previous ILP mapping works. Moreover, to enhance the scalability of the ILP task mapping, the task-processor-cluster algorithm is proposed to reduce the scale of the task graph and the number of processors and then solve the coarse-grained input by the ILP mapping. To increase the adaptability of the ILP task mapping, the improved augmented E-constraint method is further integrated with the ILP formulations to select the best mapping for different applications. Experimental results on a 2/4/8/16/24-CPU platform of both synthetic and real-life benchmarks demonstrate the efficiency of the proposed approach.
Kai Huang 0002, Dandan Zheng 0001, Min Yu 0006, Xiaowen Jiang 0001, Xiaolang Yan, Lisane B. de Brisolara, Ahmed Amine Jerraya
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2017 Providing Predictable Performance via a Slowdown Estimation Model
abstract
Interapplication interference at shared main memory slows down different applications differently. A few slowdown estimation models have been proposed to provide predictable performance by quantifying memory interference, but they have relatively low accuracy. Thus, we propose a more accurate slowdown estimation model called SEM at main memory. First, SEM unifies the slowdown estimation model by measuring IPC directly. Second, SEM uses the per-bank structure to monitor memory interference and improves estimation accuracy by considering write interference, row-buffer interference, and data bus interference. The evaluation results show that SEM has significantly lower slowdown estimation error (4.06%) compared to STFM (30.15%) and MISE (10.1%).
Dongliang Xiong, Kai Huang 0002, Xiaowen Jiang 0001, Xiaolang Yan
ACM Trans. Archit. Code Optim.3
2016 Memory Access Scheduling Based on Dynamic Multilevel Priority in Shared DRAM Systems
abstract
Interapplication interference at shared main memory severely degrades performance and increasing DRAM frequency calls for simple memory schedulers. Previous memory schedulers employ a per-application ranking scheme for high system performance or a per-group ranking scheme for low hardware cost, but few provide a balance. We propose DMPS, a memory scheduler based on dynamic multilevel priority. First, DMPS uses “memory occupancy” to measure interference quantitatively. Second, DMPS groups applications, favors latency-sensitive groups, and dynamically prioritizes applications by employing a per-level ranking scheme. The simulation results show that DMPS has 7.2% better system performance and 22% better fairness over FRFCFS at low hardware complexity and cost.
Dongliang Xiong, Kai Huang 0002, Xiaowen Jiang 0001, Xiaolang Yan
ACM Trans. Archit. Code Optim.3