EDBT 2026 Demo / reviewers in the wild / expert
Qiaosha Zou
dblp:23/10300
· DBLP profile ↗
18ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0001-6662-4316ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimization of IoUT Systems: A Hierarchical Federated Transfer Learning Approach Based on UAV Computation Offloading
Jian Hou 0002, Congcong Yang, Qiaosha Zou, Junkai Chen, Xiaotong Nie |
ICIC (4) | 3 |
| 2025 | Personalized Incentive Mechanism in Federated Learning via Variational Expectation Maximization
Xianyu Luo, Jian Hou 0002, Shuyun Luo, Qiaosha Zou |
ICIC (10) | 5 |
| 2025 | A Workload-Balance-Aware Accelerator Enabling Dense-to-Arbitrary-Sparse Neural NetworksabstractDeep neural networks (DNNs) have proved their great potential over various perceptual and cognitive tasks with the cost of ever-growing storage capacity and computation complexity. Sparse representations in neural networks have emerged as a compelling method to achieve substantial reductions in computational overhead and energy consumption. However, the introduction of sparsity presents challenges such as irregular memory accesses and wasted computation cycles. Traditional methods have attempted to address these challenges with varied success, unfortunately, often involving complex hardware designs or sacrificing sparsity’s benefits. In this article, we propose a software-hardware codesign solution, comprising an offline workload balancing encoding algorithm toward arbitrary sparsity in DNNs and a dedicated Processing Element array–based accelerator with a lightweight switch network. Extensive experiments are conducted to demonstrate the conclusion that our proposal is feasible to support various network structures and a wide range of sparsity ratios. With the encoding algorithm, a 1.16×–2.61× acceleration is achieved compared to the baseline. The system-on-chip measures 7.9 mm 2 , achieving an energy efficiency of 0.7 TOPS/W (dense), 2.1 TOPS/W (at 75% sparsity), and 10.3 TOPS/W (at 99.9% sparsity) at 0.9 V and 1,066-MHz clock frequency. Maohua Nie, Longfei Gou, Junmin He, Yongchuan Dong, Qiaosha Zou, Chuanjin Richard Shi |
ACM Trans. Embed. Comput. Syst. | 9 |
| 2024 | Frequency Adaptive Normalization For Non-stationary Time Series ForecastingabstractTime series forecasting typically needs to address non-stationary data with evolving trend and seasonal patterns. To address the non-stationarity, reversible instance normalization has been recently proposed to alleviate impacts from the trend with certain statistical measures, e.g., mean and variance. Although they demonstrate improved predictive accuracy, they are limited to expressing basic trends and are incapable of handling seasonal patterns. To address this limitation, this paper proposes a new instance normalization solution, called frequency adaptive normalization (FAN), which extends instance normalization in handling both dynamic trend and seasonal patterns. Specifically, we employ the Fourier transform to identify instance-wise predominant frequent components that cover most non-stationary factors.
Furthermore, the discrepancy of those frequency components between inputs and outputs is explicitly modeled as a prediction task with a simple MLP model. FAN is a model-agnostic method that can be applied to arbitrary predictive backbones. We instantiate FAN on four widely used forecasting models as the backbone and evaluate their prediction performance improvements on eight benchmark datasets. FAN demonstrates significant performance advancement, achieving 7.76\%$\sim$37.90\% average improvements in MSE. Our code is publicly available at http://github.com/icannotnamemyself/FAN. Weiwei Ye, Songgaojun Deng, Qiaosha Zou, Ning Gui |
NeurIPS | 3 |
| 2024 | BAP: Bilateral asymptotic pruning for optimizing CNNs on image tasks
Jingfei Chang, Liping Tao, Bo Lyu, Xiangming Zhu 0001, Shanyun Liu, Qiaosha Zou, Hongyang Chen 0001 |
Inf. Sci. | 6 |
| 2023 | AutoMap: Automatic Mapping of Neural Networks to Deep Learning Accelerators for Edge DevicesabstractEmerging deep neural networks (DNNs) have been emerging in applications (object detection, automatic speech recognition, etc.) deployed on edge devices. To improve the energy efficiency of edge devices, domain-specific deep learning accelerators (DLAs) are designed with limited on-chip resources. The manifold DLA designs and evolving DNN topologies bring challenges for applications mapping and scheduling on hardware resources. In this article, we propose an automatic DNN mapping framework named AutoMap, given the hardware backend information. First, a computational graph representation called extended directed weighted graph (EDWG) is proposed, which realizes unified expression for both spatial and temporal network interlayer connections. Second, an associated partitioner is implemented for splitting an EDWG into subEDWGs, which incorporates the on-chip memory constraint and facilitates weight data reuse on chip. Finally, a dynamic memory allocation strategy is utilized to alleviate the feature storing burden introduced by the multivarious network sizes and connections. Compared to the baseline mapping methods, experimental results show that our proposed automatic mapping framework can help to speedup the execution of several DNNs on state-of-the-art DLAs, ranging from$1.27\times $to$3.45\times $. The utilization of the PE array can increase from 20% to 64%. Maohua Nie, Qiaosha Zou, Chuanjin Richard Shi |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Augmenting aspect-level sentiment classification with distance-related local context input
Yongchuan Dong, Qiaosha Zou, Chuanjin Richard Shi |
J. Supercomput. | 2 |
| 2023 | Accelerating Distributed GNN Training by CodesabstractEmerging graph neural network (GNN) has recently attracted much attention and has been used extensively in many real-world applications thanks to its powerful expression ability of unstructured data. The real-world graph datasets are very large-scale, which can contain up to billions of nodes and tens of billions of edges. It usually requires distributed system to train GNN on such huge datasets. As a result, the data communication overheads between machines become the bottleneck of GNN computation. Our profiling results show that getting attributes from remote machines during sampling phase in GNN occupies$> $75% of the time of the training process. To address this issue, in this article, we propose Coded Neighbor Sampling (CNS) framework, which introduces codes technique to reduce the communication overheads of GNN. In the proposed CNS framework, the codes technique is coupled with GNN sampling method to exploit the data excess among different machines caused by unstructured nature of graph data. An analytical performance model is built for the proposed CNS framework, whose results are corroborated by the simulation and validate the benefit of the proposed CNS framework over both conventional GNN training method and conventional codes technique. Performance metrics, such as communication overheads, runtime, and throughput, of the proposed CNS framework are evaluated on a distributed GNN training simulation system implemented on MPI4py platform. The results show that, on average, the proposed CNS framework can save communication overhead by 40.6%, 35.5%, and 16.5%, reduce the runtime by 12.1%, 17.0%, and 10.0%, and improve the throughput by 16.2%, 24.4%, and 11.2%, respectively, when training GNN models with Cora, PubMed, and Large Taobao. Tianchan Guan, Dimin Niu, Qiaosha Zou, Hongzhong Zheng, Chuanjin Richard Shi, Yuan Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | SAIL: A Deep-Learning-Based System for Automatic Gait Assessment From TUG VideosabstractGait disorders are common in the elderly people, seriously hinder patients’ mobility and sometimes indicate underlying severe neurological diseases. Timely and automatic diagnosis of gait disorders is greatly desired. Existing methods with wearable devices put burdens on patients. We establish a video-based algorithm namedSAILto perform contactless gait assessment automatically. The SAIL contains three parts, namely,skeleton detector,parameter extractor, andgait classifier. Using a pose estimation algorithm, the skeleton detector converts RGB videos to a human skeleton sequence. Then, the parameter extractor extracts gait parameters from skeletons with a signal detection technique. Finally, a trained Support vector machine is used as a gait classifier to detect abnormal gait. The SAIL achieves 86.2% sensitivity and 98.5% specificity for abnormal gait detection on ourSAIL-TUGdataset, outperforming general clinic doctors with 76.4% and 97.4%, respectively. Nine gait parameters and the binary gait classification result are included in the final gait report. We implement an automatic gait assessment system based on SAIL and deployed the user-interface software in more than 60 hospitals for practical applications. More than 30 000 gait reports have been automatically generated. Moreover, we establish a publicly available dataset namedSAIL-TUGincluding 404 annotated Timed “Up & Go” videos. Qiaosha Zou, Yanmin Tang, Chuanjin Richard Shi |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2021 | Systolic-Array Deep-Learning Acceleration Exploring Pattern-Indexed Coordinate-Assisted Sparsity for Real-Time On-Device Speech ProcessingabstractThis paper presents a hardware-software co-design for efficient sparse deep neural networks (DNNs) implementation in a regular systolic array for real-time on-device speech processing. The weight pruning format, exploring pattern-based coordinate-assisted (PICA) sparsity, expands the pattern-based pruning into both convolutional neural networks (CNNs) and recurrent neural networks (RNNs). It reduces the index storage overhead as well as avoids accuracy degradation. The proposed systolic accelerator leverages the intrinsic data reuse and locality to accommodate the PICA-based sparsity without using complex data distribution networks. It also supports DNNs with different topologies. By reducing the model size by 16x, PICA sparsification reduces 6.02x index storage overhead while still achieving 20.7% WER in TIMIT dataset. For the pruned WaveNet and LSTM, the accelerator achieves 0.62 and 2.69 TOPS/W energy efficiency, 1.7x to 10x higher than the state-of-the-art. Shiwei Liu 0002, Qiaosha Zou, Chuanjin Richard Shi |
ACM Great Lakes Symposium on VLSI | 4 |
| 2017 | Thermomechanical Stress-Aware Management for 3-D IC DesignsabstractThermal characteristics have been considered as one of the most challenging problems in 3-D integrated circuits (3-D ICs). Due to the thermal expansion coefficient mismatch between through-silicon vias (TSVs) and the silicon substrate, and the presence of elevated thermal gradients, thermomechanical stress issues are exacerbated in 3-D ICs. In this brief, we propose a solution that combines design-time and run-time techniques to reduce thermomechanical stress and the associated reliability issues. A TSV stress-aware floorplan policy is proposed to minimize the possibility of wafer cracking and interfacial delamination. In addition, a run-time thermal management scheme effectively eliminates large thermal gradients between layers. Experimental results show that the reliability of 3-D design can be significantly improved due to the reduced TSV thermal load and the elimination of mechanical damaging thermal cycling pattern. Moreover, impacts of thermal characteristics in TSVs and thermal vias insertion are explored. Qiaosha Zou, Eren Kursun, Yuan Xie 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Pinatubo: a processing-in-memory architecture for bulk bitwise operations in emerging non-volatile memoriesabstractProcessing-in-memory (PIM) provides high bandwidth, massive parallelism, and high energy efficiency by implementing computations in main memory, therefore eliminating the overhead of data movement between CPU and memory. While most of the recent work focused on PIM in DRAM memory with 3D die-stacking technology, we propose to leverage the unique features of emerging non-volatile memory (NVM), such as resistance-based storage and current sensing, to enable efficient PIM design in NVM. We propose Pinatubo1, a Processing In Non-volatile memory ArchiTecture for bUlk Bitwise Operations. Instead of integrating complex logic inside the cost-sensitive memory, Pinatubo redesigns the read circuitry so that it can compute the bitwise logic of two or more memory rows very efficiently, and support one-step multi-row operations. The experimental results on data intensive graph processing and database applications show that Pinatubo achieves a ~500× speedup, ~28000× energy saving on bitwise operations, and 1.12× overall speedup, 1.11× overall energy saving over the conventional processor. Shuangchen Li, Cong Xu 0002, Qiaosha Zou, Jishen Zhao, Yuan Xie 0001 |
DAC | 3 |
| 2015 | Heterogeneous architecture design with emerging 3D and non-volatile memory technologiesabstractEnergy becomes the primary concern in nowadays multi-core architecture designs. Moore's law predicts that the exponentially increasing number of cores can be packed into a single chip every two years, however, the increasing power density is the obstacle to continuous performance gains. Recent studies show that heterogeneous multi-core is a competitive promising solution to optimize performance per watt. In this paper, different types of heterogeneous architecture are discussed. For each type, current challenges and latest solutions are briefly introduced. Preliminary analyses are performed to illustrate the scalability of the heterogeneous system and the potential benefits towards future application requirements. Moreover, we demonstrate the advantages of leveraging three-dimensional (3D) integration on heterogeneous architectures. With 3D die stacking, disparate technologies can be integrated on the same chip, such as the CMOS logic and emerging non-volatile memory, enabling a new paradigm of architecture design.1 Qiaosha Zou, Matthew Poremba, Junfeng Zhao 0003, Yuan Xie 0001 |
ASP-DAC | 1 |
| 2014 | 3DLAT: TSV-based 3D ICs crosstalk minimization utilizing Less Adjacent Transition codeabstract3D integration is one of the promising solutions to overcome the interconnect bottleneck with vertical interconnect through-silicon vias (TSVs). This paper investigates the crosstalk in 3D IC designs, especially the capacitive crosstalk in TSV interconnects. We propose a novel ω-LAT coding scheme to reduce the capacitive crosstalk and minimize the power consumption overhead in the TSV array. Combining with the Transition Signaling, the LAT coding scheme restricts the number of transitions in every transmission cycle to minimize the crosstalk and power consumption. Compared to other 3D crosstalk minimization coding schemes, the proposed coding can provide the same delay reduction with more affordable overhead. The performance and power analysis show that when ω is 4, the proposed LAT coding scheme can achieve 38% interconnect crosstalk delay reduction compared to the data transmission without coding. By reducing the value of ω, further reduction can be achieved1. Qiaosha Zou, Dimin Niu, Yuan Xie 0001 |
ASP-DAC | 1 |
| 2014 | TSV power supply array electromigration lifetime analysis in 3D ICSabstractElectromigration (EM) can cause severe reliability issues in contemporary integrated circuits. For the emerging three-dimensional integrated circuits (3D ICs), the introduction of through-silicon vias (TSVs) as the vertical signal carrier complicates the electromigration analysis. In particular, an accurate EM analysis on TSV arrays that are used in the power supply network is critical since the large current going through those TSVs can accelerate their degradation. In this work, we propose a novel EM analysis framework that focuses on TSV arrays in the power supply network, under the circumstance of uneven current distribution. The impacts of various design factors on the EM lifetime are discussed in detail. Our results reveal that the predicted TSV array lifetime is largely biased without proper current distribution analysis, resulting in an unexpected early failure. Qiaosha Zou, Tao Zhang 0032, Cong Xu 0002, Yuan Xie 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2013 | Thermomechanical stress-aware management for 3D IC designsabstractThe thermomechanical stress has been considered as one of the most challenging problems in three-dimensional integration circuits (3D ICs), due to the thermal expansion coefficient mismatch between the through-silicon vias (TSVs) and silicon substrate, and the presence of elevated thermal gradients. To address the stress issue, we propose a thorough solution that combines design-time and run-time techniques for the relief of thermomechanical stress and the associated reliability issues. A sophisticated TSV stress-aware floorplan policy is proposed to minimize the possibility of wafer cracking and interfacial delamination. In addition, the run-time thermal management scheme effectively eliminates large thermal gradients between layers. The experimental results show that the reliability of 3D design can be significantly improved due to the reduced TSV thermal load and the elimination of mechanical damaging thermal cycling pattern. Qiaosha Zou, Tao Zhang 0032, Eren Kursun, Yuan Xie 0001 |
DATE | 1 |
| 2013 | Low power multi-level-cell resistive memory design with incomplete data mappingabstractPhase change memory (PCM) has been widely studied as a potential DRAM alternative. The multi-level cell (MLC) can further increase the memory density and reduce the fabrication cost by storing multiple bits in a single cell. Nevertheless, large write power, high write latency, as well as reliability issue resulted from the resistance drift, bring in challenges for MLC PCM based memory design. In contrast, the emerging Resistive Random Access Memory (ReRAM), which has similar MLC property as PCM, demonstrates better performance and energy efficiency compared to PCM. In addition, due to the physical switching behaviors of ReRAM cell, the resistance drift phenomenon does not exist. In this paper, we propose a low power MLC ReRAM design. We first study the programming method of MLC ReRAM and identify that programming latency and energy are highly dependent on the data pattern written to the cell. Based on this observation, we propose incomplete data mapping (IDM), which maps an eight-level-cell into six states to prevent the time/energy consuming data patterns from appearing in the cell. Furthermore, in order to improve endurance of MLC RAM, which is much smaller than single-level cell (SLC) ReRM due to the complex programming method, we propose Dynamic Data ReMapping (DDRM) to selectively regulate memory blocks from IDM state back to complete data mapping (CDM) state. We demonstrate that the proposed design can work effectively with existing error-correction schemes but requires much smaller space overhead. Experimental results show that, IDM can reduce the energy performance by at most 15% with negligible performance overhead. By combining the DDRM with existing error-correction scheme, DDRM can improve the memory lifetime by 2.75× compared with conventional memory architectures. Dimin Niu, Qiaosha Zou, Cong Xu 0002, Yuan Xie 0001 |
ICCD | 2 |
| 2012 | 3DHLS: Incorporating high-level synthesis in physical planning of three-dimensional (3D) ICsabstractThree-dimensional (3D) circuit integration is a promising technology to alleviate performance and power related issues raised by interconnects in nanometer CMOS. Physical planning of three-dimensional integrated circuits is substantially different from that of traditional planar integrated circuits, due to the presence of multiple layers of dies. To realize the full potential offered by three-dimensional integration, it is necessary to take physical information into consideration at higher-levels of the design abstraction for 3D ICs. This paper proposes an incremental system-level synthesis framework that tightly integrates behavioral synthesis of modules into the layer assignment and floorplan- ning stage of 3D IC design. Behavioral synthesis is implemented as a sub-routine to be called to adjust delay/power/variability/area of circuit modules during the physical planning process. Experimental results show that with the proposed synthesis-during-planning methodology, the overall timing yield is improved by 8%, and the chip peak temperature reduced by 6.6°C, compared to the conventional planning-after-synthesis approach. Guangyu Sun 0003, Qiaosha Zou, Yuan Xie 0001 |
DATE | 3 |