EDBT 2026 Demo / reviewers in the wild / expert
Xiaotian Zhao
dblp:251/1103
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DARE: Enriching Physical Dataflow Awareness for Macro Placement OptimizationabstractPhysical dataflow, which defines the detailed connections among cells and macros, is a critical yet underexplored factor in automatic macro placement. It becomes increasingly important for enabling intelligent design automation to minimize manual intervention and reduce design iterations. Existing macro or mixed-size placers with dataflow awareness primarily focus on intrinsic relationships among macros, overlooking the crucial influence of standard cell clusters on macro placement. To address this, we propose DARE, which extracts hidden connections between macros and standard cells and incorporates a series of algorithms to enrich dataflow awareness, integrating them into placement constraints for improved macro placement. To further optimize placement results, we introduce two fine-tuning steps: (1) congestion optimization by taking macro area into consideration, and (2) flipping decisions to determine the optimal macro orientation based on the extracted dataflow information. By integrating enhanced dataflow awareness into placement constraints and applying these fine-tuning steps, the proposed approach achieves an average 7.9% improvement in half-perimeter wirelength (HPWL) across multiple widely used benchmark designs compared to a state-of-the-art dataflow-aware macro placer. Additionally, it significantly improves congestion, reducing overflow by an average of 82.5%, and achieves improvements of 36.97% in Worst Negative Slack (WNS) and 59.44% in Total Negative Slack (TNS). The approach also maintains efficient runtime throughout the entire placement, incurring less than a 1.5% runtime overhead. These results show that the proposed dataflow-driven methodology, combined with the fine-tuning steps, provides an effective foundation for macro placement within the OpenROAD flow and can be further extended to other design flows in the future to enhance placement quality. Xiaotian Zhao, Yichen Cai 0004, Yushan Pan, Xinfei Guo |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2024 | Standard Cells Do Matter: Uncovering Hidden Connections for High-Quality Macro PlacementabstractIt becomes increasingly critical for an intelligent macro placer to be able to uncover a good dataflow for a large-scale chip to reduce churns from manual trials and errors. Existing macro placers or mixed-size placement engines that were equipped with dataflow awareness mostly focused on extracting intrinsic relations among macros only, ignoring the fact that standard cell clusters play an essential role in determining the location of macros. In this paper, we identify the necessity of macro-cell connection awareness for high-quality macro placement and propose a novel methodology to extract all “hidden” relationships efficiently among macros and cell clusters. By integrating the discovered connections as part of the placement constraints, the proposed methodology achieves an average of 2.8% and 5.5% half perimeter wire length (HPWL) improvement for considering one-hop macro-cell and two-hop macro-cell-cell dataflow connections respectively, when compared against a recently proposed dataflow-aware macro placer. A maximum of 9.7% HPWL improvement has been achieved, incurring only less than 1% runtime penalty. In addition, the congestion has been improved significantly by the proposed method, yielding an average of 62.9% and 73.4% overflow reduction for one-hop and two-hop dataflow considerations. The proposed dataflow connection extraction methodology has been demonstrated to be a significant starting point for macro placement and can be integrated into the existing design flows while delivering better design quality. Xiaotian Zhao, Tianju Wang, Run Jiao, Xinfei Guo |
DATE | 1 |
| 2024 | CINEMA: A Configurable Binary Segmentation Based Arithmetic Module for Mixed-Precision In-Memory AccelerationabstractThe emergence of mixed-precision quantization (MPQ) applied to edge AI models highlights the critical need for hardware support. It is a promising model compression approach with minimal accuracy loss but poses a notable hardware design challenge in the intricate balance required between computing reconfigurability and the resulting area or energy overheads. As the Compute-in-Memory (CiM) paradigm becomes prevalent for accelerating edge inference and demonstrates promising results in enhancing energy efficiency by eliminating overloaded data traffic, efficient in-memory computing circuitry becomes paramount to maximize these advantages. While the memory cell itself cannot handle complex arithmetic logic, especially in the context of MPQ-supporting mult-precision computing, the peripheral pairing module is required to be modular, portable, and scalable. In this paper, we propose a novel bit-precision configurable arithmetic module based on binary segmentation, supporting fine-grained precision ranging from 2 to 8 bits. This module achieves a peak throughput of 16 GOPS and maximum energy efficiency of 8.56 TOPS/W, occupying only 1778 μm2in 28nm technology node. The throughput and energy efficiency achieve 1.18x and 6.58x improvements compared with baseline work. Its portability enables seamless integration with various memory technologies, enhancing support for MPQ efficiently. Runxi Wang, Ruge Xu, Xiaotian Zhao, Kai Jiang 0007, Xinfei Guo |
ISCAS | 3 |
| 2023 | Design Space Exploration of Layer-Wise Mixed-Precision Quantization with Tightly Integrated Edge Inference UnitsabstractLayer-wise mixed-precision quantization (MPQ) has become prevailing for edge inference since it strikes a better balance between accuracy and efficiency compared to the uniform quantization scheme. Existing MPQ strategies either lacked hardware awareness or incurred huge computation costs, which gated their deployment at the edge. In this work, we propose a novel MPQ search algorithm that obtains an optimal scheme by "sampling" layer-wise sensitivity with respect to a newly proposed metric that incorporates both accuracy and proxy of hardware cost. To further efficiently deploy post-training MPQ on edge chips, we propose to tightly integrate the quantized inference units as part of the processor pipeline through micro-architecture and Instruction Set Architecture (ISA) co-design. Evaluation results show that the proposed search algorithm achieves 3% ~ 11% higher inference accuracy with similar hardware cost compared to the state-of-the-art MPQ strategies. In addition, the tightly integrated MPQ units achieve speedup of 15.13x ~ 29.65x compared to a baseline RISC-V processor. Xiaotian Zhao, Yimin Gao, Vaibhav Verma, Ruge Xu, Mircea R. Stan, Xinfei Guo |
ACM Great Lakes Symposium on VLSI | 1 |
| 2019 | Transfer Learning Methods for Spoken Language UnderstandingabstractIn this paper, we present a series of methods to improve the performance of spoken language understanding in the 1st Chinese Audio-Textual Spoken Language Understanding Challenge (CATSLU 2019) which is aimed to improve the robustness for automatic speech recognition (ASR) errors and to solve the problem of not enough labeled data in new domains. We combine word information and char information to improve the performance of the semantic parser. We also use some transfer learning methods like correlation alignments to improve the robustness of the spoken language understanding system. Then we merge the rule method and the neural network method to raise system output performance. In video and weather domains with few training data, we use both the transfer learning model trained on multi-domain data and the rule-based approach. Our approaches achieve F1 scores of 86.83%, 92.84%, 94.16%, and 93.04% on the test sets of map, music, video and weather domains. Chengda Tang, Xiaotian Zhao, Xuancai Li, Zhuolin Jin, Dequan Zheng, Tiejun Zhao |
ICMI | 3 |