EDBT 2026 Demo / reviewers in the wild / expert
Dongliang Xiong
dblp:192/3737
· DBLP profile ↗
15ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0002-4882-7504ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Structure-Aware Irregular Blocking Method for Sparse LU Factorization
Dongliang Xiong, Kai Huang 0002, Changjun Wu, Xiaowen Jiang 0001 |
ICA3PP (3) | 2 |
| 2025 | Quantitative Cost Model and Cost Optimization Methods for Multi-Technology-Node ArchitectureabstractIn multi-chiplet design research, cost reduction is a key advantage of heterogeneous integration. Accurate system cost modeling is essential for making informed decisions early in the design process. However, previous cost models overlooked the impact of mixing different technology nodes on overall system cost. In this paper, we propose a quantitative cost model for estimating the cost of the multi-chiplet systems across multiple technology nodes and various chiplet partition granularities. In addition, three optimization strategies are introduced, focusing on partition granularity, technology node, and wafer metal layer, which work synergistically to improve the cost efficiency of 2.5D design. Case studies demonstrate that applying optimization strategies results in 18% to 31% cost savings compared to non-optimized 2.5D designs. Qimin Yuan, Kai Huang 0002, Xiaowen Jiang 0001, Dongliang Xiong |
ICCAD | 4 |
| 2025 | ATEP: An Asynchronous Timing Error Prediction Circuit With Adaptive Voltage and Frequency ScalingabstractTiming error prediction circuits have demonstrated greater efficiency in reducing the worst case timing margins of conventional circuits. However, prior works of timing error prediction circuits have complicated the clock tree or introduced a substantial number of delay cells along data paths, resulting in a considerable increase in area overhead. This work introduces asynchronous timing error prediction circuit (ATEP), an ATEP that integrates timing error prediction technology with bundle-data asynchronous templates. The proposed circuit leverages delay lines in the request wire to generate the warning detection window (WDW) independent of clock signals, thereby reducing area overhead and streamlining the clock tree. In addition, we present an adaptive voltage and frequency scaling (AVFS) controller, which evaluates the likelihood of warnings or the quantity of warning paths based on path activation rates to determine when to cease adjustments. This strategy helps to identify the frequency closer to the point of first failure (PoFF). Furthermore, we propose a dynamic warning detector gating strategy to gate warning detectors based on the current environment, further decreasing power consumption. Implementing this circuit on an RISC-V processor, targeting 28-nm CMOS technology, yields up to a 56.8% performance improvement with only a 3.79% area cost and up to a 28.0% reduction in power consumption. Dongliang Xiong, Huibo Gao, Kai Huang 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | Blade-DA: Click-Based Asynchronous Resilient Circuit With Dynamic Delay AdjustmentabstractAsynchronous resilient circuits have demonstrated higher efficiency and robustness due to their flexible local control mechanism and ability to eliminate metastability, compared to the synchronous resilient circuits. This paper proposes the Blade-DA template, which is an asynchronous resilient circuit that incorporates a delay adjustment scheme. The Blade-DA template utilizes the Click template to simplify circuit design with a larger timing resilient window (TRW). Additionally, the Blade-DA template integrates a dynamic delay adjustment controller, which adaptively tunes the TRW based on real-time timing error rates. In this way, this controller reduces the timing error rate to improve performance under unfavorable Process, Voltage, and Temperature (PVT) corners. Moreover, we propose an optimization method of adding extra constraints to change the delay distribution, which further improves the performance of Blade-DA template. Compared to the previous controller with a large TRW, the proposed asynchronous controller offers a significant 81.65% improvement in maximum frequency and a 61.93% reduction in area while operating at the same frequency. Furthermore, the Blade-DA template is applied to a three-stage RISC-V processor, resulting in a 7.92%-14.78% performance improvement under unfavorable PVT conditions compared to the fixed TRW asynchronous template, while incurring only a 1.63% additional area cost. Dongliang Xiong, Kai Huang 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Structured Term Pruning for Computational Efficient Neural Networks InferenceabstractThe state-of-the-art convolutional neural network accelerators are showing a growing interest in exploiting the bit-level sparsity and eliminating the ineffectual computations of zero bits. However, the excessive redundancy and the irregular distribution of nonzero bits limit the real speedup in the accelerators. To address this, we propose an algorithm-architecture codesign, named structured term pruning (STP), to boost the computation efficiency of neural networks inference. Specifically, we enhance the bit sparsity by guiding the weights toward the value with fewer power-of-two terms. Then, we structure the terms with layer-wise group budgets. Retraining is adopted to recover the accuracy drop. We also design the hardware of the group processing element and the fast signed-digital encoder for efficient implementation of STP networks. The system design of STP is realized with some easy alterations on an input stationary systolic array design. Extensive evaluation results demonstrate that STP can reduce significant inference computation costs, and achieve$2.35\times $computational energy saving for the ResNet18 network on the ImageNet dataset. Kai Huang 0002, Bowen Li 0017, Siang Chen, Luc Claesen, Wei Xi 0001, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong, Xiaolang Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2023 | Efficient Halftoning via Deep Reinforcement LearningabstractHalftoning aims to reproduce a continuous-tone image with pixels whose intensities are constrained to two discrete levels. This technique has been deployed on every printer, and the majority of them adopt fast methods (e.g., ordered dithering, error diffusion) that fail to render structural details, which determine halftone's quality. Other prior methods of pursuing visual pleasure by searching for the optimal halftone solution, on the contrary, suffer from their high computational cost. In this paper, we propose a fast and structure-aware halftoning method via a data-driven approach. Specifically, we formulate halftoning as a reinforcement learning problem, in which each binary pixel's value is regarded as an action chosen by a virtual agent with a shared fully convolutional neural network (CNN) policy. In the offline phase, an effective gradient estimator is utilized to train the agents in producing high-quality halftones in one action step. Then, halftones can be generated online by one fast CNN inference. Besides, we propose a novel anisotropy suppressing loss function, which brings the desirable blue-noise property. Finally, we find that optimizing SSIM could result in holes in flat areas, which can be avoided by weighting the metric with the contone's contrast map. Experiments show that our framework can effectively train a light-weight CNN, which is 15x faster than previous structure-aware methods, to generate blue-noise halftones with satisfactory visual quality. We also present a prototype of deep multitoning to demonstrate the extensibility of our method. Haitian Jiang, Dongliang Xiong, Xiaowen Jiang 0001, Kai Huang 0002 |
IEEE Trans. Image Process. | 2 |
| 2023 | Structured Dynamic Precision for Deep Neural Networks QuantizationabstractDeep Neural Networks (DNNs) have achieved remarkable success in various Artificial Intelligence applications. Quantization is a critical step in DNNs compression and acceleration for deployment. To further boost DNN execution efficiency, many works explore to leverage the input-dependent redundancy with dynamic quantization for different regions. However, the sensitive regions in the feature map are irregularly distributed, which restricts the real speed up for existing accelerators. To this end, we propose an algorithm-architecture co-design, named Structured Dynamic Precision (SDP). Specifically, we propose a quantization scheme in which the high-order bit part and the low-order bit part of data can be masked independently. And a fixed number of term parts are dynamically selected for computation based on the importance of each term in the group. We also present a hardware design to enable the algorithm efficiently with small overheads, whose inference time mainly scales with the precision proportionally. Evaluation experiments on extensive networks demonstrate that compared to the state-of-the-art dynamic quantization accelerator DRQ, our SDP can achieve 29% performance gain and 51% energy reduction for the same level of model accuracy. Kai Huang 0002, Bowen Li 0017, Dongliang Xiong, Haitian Jiang, Xiaowen Jiang 0001, Xiaolang Yan, Luc Claesen, Dehong Liu, Junjian Chen, Zhili Liu |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2022 | Halftoning with Multi-Agent Deep Reinforcement LearningabstractDeep neural networks have recently succeeded in digital halftoning using vanilla convolutional layers with high parallelism. However, existing deep methods fail to generate halftones with a satisfying blue-noise property and require complex training schemes. In this paper, we propose a halftoning method based on multi-agent deep reinforcement learning, called HALFTONERS, which learns a shared policy to generate high-quality halftone images. Specifically, we view the decision of each binary pixel value as an action of a virtual agent, whose policy is trained by a low-variance policy gradient. Moreover, the blue-noise property is achieved by a novel anisotropy suppressing loss function. Experiments show that our halftoning method produces high-quality halftones while staying relatively fast. Haitian Jiang, Dongliang Xiong, Xiaowen Jiang 0001, Aiguo Yin, Kai Huang 0002 |
ICIP | 2 |
| 2022 | Structured precision skipping: Accelerating convolutional neural networks with budget-aware dynamic precision selection
Kai Huang 0002, Siang Chen, Bowen Li 0017, Luc Claesen, Hao Yao, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong |
J. Syst. Archit. | 9 |
| 2022 | Acceleration-Aware Fine-Grained Channel Pruning for Deep Neural Networks via Residual GatingabstractDeep neural networks have achieved remarkable advancement in various intelligence tasks. However, the massive computation and storage consumption limit applications on resource-constrained devices. While channel pruning has been widely applied to compress models, it is challenging to reach very deep compressions for such a coarse-grained pruning structure without significant performance degradation. In this article, we propose an acceleration-aware fine-grained channel pruning (AFCP) framework for accelerating neural networks, which optimizes trainable gate parameters by estimating residual errors between pruned and original channels with hardware characteristics. Our fine-grained concept consists of both algorithm and structure levels. Different from existing methods that leverage a predefined pruning criterion, AFCP explicitly considers both zero-out and similar criteria for each channel, and adaptively selects the suitable one via residual gate parameters. For structure level, AFCP adopts a fine-grained channel pruning strategy for residual neural networks and a decomposition-based structure, which further extends the pruning optimization space. Moreover, instead of using theoretical computation costs, such as floating-point operations, we propose the hardware predictor that bridges the gap between realistic acceleration and pruning procedure to guide the learning of pruning, which improves the efficiency of model pruning when deployed on accelerators. Extensive evaluation results demonstrate that AFCP outperforms state-of-the-art methods, and achieves a favorable balance between model performance and computation cost. Kai Huang 0002, Siang Chen, Bowen Li 0017, Luc Claesen, Hao Yao, Junjian Chen, Xiaowen Jiang 0001, Zhili Liu, Dongliang Xiong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2021 | DPOQ: Dynamic Precision Onion QuantizationabstractWith the development of deployment platforms and application scenarios for deep neural networks, traditional fixed network architectures cannot meet the requirements. Meanwhile the dynamic network inference becomes a new research trend. Many slimmable and scalable networks have been proposed to satisfy different resource constraints (e.g., storage, latency and energy). And a single network may support versatile architectural configurations including: depth, width, kernel size, and resolution. In this paper, we propose a novel network architecture reuse strategy enabling dynamic precision in parameters. Since our low-precision networks are wrapped in the high-precision networks like an onion, we name it dynamic precision onion quantization (DPOQ). We train the network by using the joint loss with scaled gradients. To further improve the performance and make different precision network compatible with each other, we propose the precision shift batch normalization (PSBN). And we also propose a scalable input-specific inference mechanism based on this architecture and make the network more adaptable. Experiments on the CIFAR and ImageNet dataset have shown that our DPOQ achieves not only better flexibility but also higher accuracy than the individual quantization. Bowen Li 0017, Kai Huang 0002, Siang Chen, Dongliang Xiong, Luc Claesen |
ACML | 4 |
| 2020 | DFQF: Data Free Quantization-aware Fine-tuningabstractData free deep neural network quantization is a practical challenge, since the original training data is often unavailable due to some privacy, proprietary or transmission issues. The existing methods implicitly equate data-free with training-free and quantize model manually through analyzing the weights’ distribution. It leads to a significant accuracy drop in lower than 6-bit quantization. In this work, we propose the data free quantization-aware fine-tuning (DFQF), wherein no real training data is required, and the quantized network is fine-tuned with generated images. Specifically, we start with training a generator from the pre-trained full-precision network with inception score loss, batch-normalization statistics loss and adversarial loss to synthesize a fake image set. Then we fine-tune the quantized student network with the full-precision teacher network and the generated images by utilizing knowledge distillation (KD). The proposed DFQF outperforms state-of-the-art post-train quantization methods, and achieve W4A4 quantization of ResNet20 on the CIFAR10 dataset within 1% accuracy drop. Bowen Li 0017, Kai Huang 0002, Siang Chen, Dongliang Xiong, Haitian Jiang, Luc Claesen |
ACML | 4 |
| 2020 | Fine-Grained Channel Pruning for Deep Residual Neural Networks
Siang Chen, Kai Huang 0002, Dongliang Xiong, Bowen Li 0017, Luc Claesen |
ICANN (2) | 3 |
| 2017 | Providing Predictable Performance via a Slowdown Estimation ModelabstractInterapplication interference at shared main memory slows down different applications differently. A few slowdown estimation models have been proposed to provide predictable performance by quantifying memory interference, but they have relatively low accuracy. Thus, we propose a more accurate slowdown estimation model called SEM at main memory. First, SEM unifies the slowdown estimation model by measuring IPC directly. Second, SEM uses the per-bank structure to monitor memory interference and improves estimation accuracy by considering write interference, row-buffer interference, and data bus interference. The evaluation results show that SEM has significantly lower slowdown estimation error (4.06%) compared to STFM (30.15%) and MISE (10.1%). Dongliang Xiong, Kai Huang 0002, Xiaowen Jiang 0001, Xiaolang Yan |
ACM Trans. Archit. Code Optim. | 1 |
| 2016 | Memory Access Scheduling Based on Dynamic Multilevel Priority in Shared DRAM SystemsabstractInterapplication interference at shared main memory severely degrades performance and increasing DRAM frequency calls for simple memory schedulers. Previous memory schedulers employ a per-application ranking scheme for high system performance or a per-group ranking scheme for low hardware cost, but few provide a balance. We propose DMPS, a memory scheduler based on dynamic multilevel priority. First, DMPS uses “memory occupancy” to measure interference quantitatively. Second, DMPS groups applications, favors latency-sensitive groups, and dynamically prioritizes applications by employing a per-level ranking scheme. The simulation results show that DMPS has 7.2% better system performance and 22% better fairness over FRFCFS at low hardware complexity and cost. Dongliang Xiong, Kai Huang 0002, Xiaowen Jiang 0001, Xiaolang Yan |
ACM Trans. Archit. Code Optim. | 1 |