VLDB 2026 Research / reviewers in the wild / expert
Zixiao Wang 0001
dblp:141/1943-1 · also Zi-Xiao Wang 0001
· DBLP profile ↗
16ranked-venue papers
7as first author
16since 2021 · last 2026
0009-0000-8179-0996ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCLOG: Don't Cares-based Logic Optimization using Pre-training Graph Neural NetworksabstractLogic rewriting serves as a robust optimization technique that enhances Boolean networks by substituting small segments with more effective implementations. The incorporation of don’t cares in this process often yields superior optimization results. Nevertheless, the calculation of don’t cares within a Boolean network can be resourceintensive. Therefore, it is crucial to develop effective strategies that mitigate the computational costs associated with don’t cares while simultaneously facilitating the exploration of improved optimization outcomes. To address these challenges, this paper proposes DCLOG, a don’t cares-based logic optimization framework, to efficiently and effectively optimize a given Boolean network. DCLOG leverages a pretrained graph neural network model to filter out cuts without don’t cares and then performs an incremental window simulation to calculate don’t cares for each cut. Experimental results demonstrate the effectiveness and efficiency of DCLOG on large Boolean networks, specifically average size reductions of 15.64 % and 1.44 % while requiring less than 23.84 % and $44.70 \%$ of the average runtime compared with state-of-the-art methods for the majority-inverter graph (MIG), respectively. Rongliang Fu, Libo Shen, Ziyi Wang 0010, Zhengxing Lei, Zixiao Wang 0001, Junying Huang, Bei Yu 0001, Tsung-Yi Ho |
ASP-DAC | 5 |
| 2026 | Video-based Visible-Event Cross-modal Person Re-identification for Edge AI Surveillance SystemsabstractVideo-based cross-modal person re-identification (ReID) is a critical task for video surveillance and security systems, particularly in resource-constrained edge AI environments. While existing crossmodal ReID methods primarily focus on thermal-visible matching, event cameras, with their low power consumption, high temporal resolution, and sparse data representation, offer significant advantages for edgebased surveillance systems by reducing data processing overhead and enabling robust performance under challenging lighting conditions. In this paper, we introduce a novel task: video-based visible-event person re-identification (VE ReID), which aims to match identities across RGB and event camera modalities. To the best of our knowledge, this is the first work to systematically define and investigate this cross-modal task in the context of event-driven edge AI. Specifically, we curate evaluation benchmarks from existing RGB-event datasets and synthesize a new RGB-event dataset, explicitly adapting them to the cross-modal ReID setting to enable a more comprehensive evaluation of VE ReID. Extensive experiments reveal that existing cross-modal state-of-the-art (SOTA) methods fail to effectively address the unique challenges posed by event data, highlighting the importance of tailored solutions for this task. To this end, we propose a novel method that constructs auxiliary modalities using frequency information from RGB and event tracklets, aligning them effectively through a fine-grained metric learning loss. Our approach not only achieves significant accuracy improvements over existing methods but also demonstrates the potential of event cameras for efficient and scalable edge AI surveillance applications. All source code and benchmarks are publicly available at https://github.com/yxgnahz/ASPDAC26-Event-RGBReID. Xinyun Zhang 0001, Zixiao Wang 0001, Yurui Kuang, Bei Yu 0001 |
ASP-DAC | 2 |
| 2026 | FastRW: An Efficient Random Walk Method for Steady-State Thermal AnalysisabstractThermal simulation is increasingly critical in modern IC design and manufacturing. Random walk methods based on the Feynman-Kac formula enable efficient local temperature estimation without computing the full temperature field. However, in practical scenarios without Dirichlet boundary conditions, these methods often require excessively long paths and heuristic truncation rules. In this work, we revisit Feynman-Kac sampling and derive an exact characterization of the truncation error: the expected residual contribution is a simple scalar multiple of the temperature at the truncation point. This insight leads to FastRW, a random-walk framework that safely applies aggressive truncation. FastRW uses a cheap, noisy prior temperature field to approximate the residual term and shorten individual paths, and further exploits cross-relations among query points through a Bayesian posterior update to reduce the number of required walks. Experiments on 3DIC steady-state thermal benchmarks show that FastRW achieves over 6× speedup over prior Feynman-Kac-based methods with better accuracy. Zixiao Wang 0001, Tianshu Hou, Zhen Zhuang, Tsung-Yi Ho, Farzan Farnia, Bei Yu 0001 |
DATE | 1 |
| 2026 | DiffResist: Physics-Constrained Diffusion for Photoresist ModelingabstractAccurate and efficient 3D photoresist simulation is essential for optical lithography at advanced technology nodes. Existing methods that predict 3D resist profiles from aerial images either rely on analytical reaction-diffusion solvers, which are slow, or on high-capacity 3D generative models, which are costly to train and deploy. We instead formulate 3D resist prediction as a depth-wise 2D generation task conditioned on the aerial image. DiffResist introduces a physics-constrained diffusion model whose reverse steps are aligned with resist exposure physics: a two-stage noise schedule connects physically meaningful layers to a Gaussian prior, and boundary conditions at the resist-air interface are injected to suppress error propagation. Combined with a lightweight super-resolution module, DiffResist achieves state-of-the-art accuracy on a public benchmark with over 10 × faster inference than 3D diffusion baselines. Zixiao Wang 0001, Jieya Zhou, Xinyun Zhang 0001, Shoubo Hu, Farzan Farnia, Bei Yu 0001 |
DATE | 1 |
| 2026 | Node2Node: Node Adaptation with Transformer for Cross-Node Hotspot DetectionabstractAs semiconductor manufacturing advances to smaller process nodes, hotspot detection has become critical for ensuring the manufacturability and reliability of integrated circuit (IC) layouts. However, existing detection methods rely heavily on labeled data tailored to specific nodes, resulting in poor generalizability across nodes due to variations in layout geometries and fabrication processes. Labeling new data at advanced nodes is also costly and time-consuming. To overcome these challenges, we propose Node2Node, the first adaptation framework explicitly designed for cross-node hotspot detection. Node2Node integrates a novel node-invariant encoder with a node-specific encoder to jointly capture transferable and node-dependent features. To further improve robustness, we introduce a bidirectional center alignment strategy, which refines pseudo-labels by leveraging a small amount of labeled data from the target node. Additionally, a cross-node distribution loss is introduced to explicitly align feature distributions between nodes. Extensive experiments demonstrate that Node2Node substantially improves cross-node generalization and achieves state-of-the-art hotspot detection performance. Silin Chen, Yibo Huang 0009, Xinyun Zhang 0001, Zixiao Wang 0001, Bei Yu 0001, Ningmu Zou |
DATE | 5 |
| 2025 | FlexPose: Pose Distribution Adaptation with Limited GuidanceabstractNumerous well-annotated human key-point datasets are publicly available to date. However, annotating human poses for newly collected images is still a costly and time-consuming progress. Pose distributions from different datasets share similar pose hinge-structure priors with different geometric transformations, such as pivot orientation, joint rotation, and bone length ratio. The difference between Pose distributions is essentially the difference between the transformation distributions. Inspired by this fact, we propose a method to calibrate a pre-trained pose generator in which the pose prior has already been learned to an adapted one following a new pose distribution. We treat the representation of human pose joint coordinates as skeleton image and transfer a pre-trained pose annotation generator with only a few annotation guidance. By fine-tuning a limited number of linear layers that closely related to the pose transformation, the adapted generator is able to produce any number of pose annotations that are similar to the target poses. We evaluate our proposed method, FlexPose, on several cross-dataset settings both qualitatively and quantitatively, which demonstrates that our approach achieves state-of-the-art performance compared to the existing generative-model-based transfer learning methods when given limited annotation guidance. Zixiao Wang 0001, Junwu Weng, Bei Yu 0001 |
AAAI | 1 |
| 2025 | SDM-PEB: Spatial-Depthwise Mamba for Enhanced Post-Exposure Bake SimulationabstractThe post-exposure bake (PEB) process is a critical step in semiconductor lithography, directly impacting resist profile accuracy and circuit pattern fidelity. Precise modeling of PEB is essential for controlling photoacid diffusion and inhibitor reactions. In this paper, we introduce SDM-PEB, an advanced modeling framework designed to enhance the accuracy of PEB simulations by capturing both intra-layer spatial dependencies and inter-layer depthwise interactions. Leveraging a unique hierarchical feature extractor with overlapped patch merging and efficient self-attention, our approach effectively captures both coarse and fine features at multiple scales. The spatial-depthwise Mamba-based attention unit, centered on a customized selective scan and structured state space model, efficiently captures spatial and depthwise dependencies, enabling precise 3D PEB simulation. Additionally, a PEB focal loss and differential depth divergence regularization term improve the sensitivity to both spatial and depthwise variations, addressing inherent data imbalances in 3D PEB simulations. Our framework is validated with commercial rigorous model, and experimental results demonstrate that the SDM-PEB outperforms previous methods in accuracy and efficiency. Ziyang Yu 0001, Peng Xu 0052, Zixiao Wang 0001, Binwu Zhu, Qipan Wang, Yibo Lin, Runsheng Wang, Bei Yu 0001, Martin D. F. Wong |
DAC | 3 |
| 2025 | DiffPattern-Flex: Efficient Layout Pattern Generation via Discrete Diffusion
Zixiao Wang 0001, Wenqian Zhao 0002, Yunheng Shen, Guojin Chen, Farzan Farnia, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | BAQE: Backend-Adaptive DNN Deployment via Synchronous Bayesian Quantization and Hardware Configuration ExplorationabstractEfficiently deploying deep learning (DL) algorithms on different hardware backends has become a time-consuming challenge. Achieving ultimate inference efficiency on hardware requires both algorithm-level model compression techniques, such as model quantization, and hardware-level optimization, such as operation reconfiguration and scheduling. In this article, we propose BAQE, a unified deployment framework that bridges the gap between algorithm-level and backend-level optimization. By constructing a global search space, we can synchronously optimize both the model quantization settings and backend configuration parameters. To accelerate this laborious and time-consuming process, we propose a searching strategy based on multiobjective Bayesian optimization (BO) using a Gaussian model with deep kernel learning as the surrogate model. More importantly, BAQE can easily adapt to various backends with different hardware resources efficiently and effectively. Each inner step of the optimization process is aware of the genuine hardware resources, ensuring that all accuracy/latency metrics and historical knowledge/feedback are evaluated directly on the device within each iteration. Empirical results demonstrate that our approach achieves both superior inference time and accuracy with a faster optimization process. Wenqian Zhao 0002, Zixiao Wang 0001, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | HAPE: Hardware-Aware LLM Pruning For Efficient On-Device Inference OptimizationabstractOver the past few years, large language models (LLMs) have demonstrated remarkable performance and versatility across a variety of complex tasks. However, their deployment has been challenged by their substantial model size and computational requirements. Pruning is a effective approach to make the model parameters sparse, thereby acquire inference acceleration. While not everyone requires training or fine-tuning large models, the diverse range of applications necessitates the deployment of LLMs on different devices. Model pruning and compression have emerged as areas of deep research interest to address these challenges. In consideration of versatility and practicality, we have designed a hardware-aware pruning process for general-purpose hardware/edge devices to enable efficient deployment and inference of LLMs. Instead of considering sparse ratio alone, we are motivated to design a pruning framework that incorporates genuine inference speed-up sensitivity from each pruning structure. Moreover, our framework breaks the layer-by-layer pruning setting and fuse several layers into one pruning stage to allow cross-layer optimization. Apart from that, we hold pragmatism by conducting compilation optimization during pruning. This step is critical because most sparsity patterns barely show distinct speed acceleration with corresponding dataflow and memory optimization. Our process operates within a post-training framework, obviating the need for additional training and thereby reducing resource requirements, while ensuring diverse inference speed and accuracy requirements on hardware. Wenqian Zhao 0002, Lancheng Zou, Zixiao Wang 0001, Xufeng Yao, Bei Yu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | ChatPattern: Layout Pattern Customization via Natural LanguageabstractExisting works focus on fixed-size layout pattern generation, while the more practical free-size pattern generation receives limited attention. In this paper, we propose ChatPattern, a novel Large-Language-Model (LLM) powered framework for flexible pattern customization. ChatPattern utilizes a two-part system featuring an expert LLM agent and a highly controllable layout pattern generator. The LLM agent can interpret natural language requirements and operate design tools to meet specified needs, while the generator excels in conditional layout generation, pattern modification, and memory-friendly patterns extension. Experiments on challenging pattern generation setting shows the ability of ChatPattern to synthesize high-quality large-scale patterns. Zixiao Wang 0001, Yunheng Shen, Xufeng Yao, Wenqian Zhao 0002, Farzan Farnia, Bei Yu 0001 |
DAC | 1 |
| 2024 | GTCO: Graph and Tensor Co-Design for Transformer-Based Image Recognition on Tensor CoresabstractDeep learning frameworks or compilers optimize the operators in computation graph using fixed templates via significant engineering efforts, which may miss potential optimizations such as operator fusion. Therefore, automatically implementing and optimizing the emerging new combinations of operators on a specific hardware accelerator is of importance. In this article, we introduce GTCO, a tensor compilation system designed to accelerate transformer-based vision models’ inference on GPUs. GTCO tackles the operator fusion techniques in the transformer-based model using a novel dynamic programming algorithm and proposes a search policy with new sketch generation rules for the fused batch matrix multiplication and softmax operators. Tensor programs are sampled from an effective search space, and a hardware abstraction with hierarchical mapping from tensor computation to domain-specific accelerators (Tensor Cores) is formally defined. Finally, our framework can map and transform tensor expression into efficient CUDA kernels with hardware intrinsics on GPU. Our experimental results demonstrate that GTCO improves the end-to-end execution performance by up to$1.73\times $relative to the cutting-edge deep learning library TensorRT on NVIDIA GPUs with Tensor Cores. Xufeng Yao, Qi Sun 0002, Wenqian Zhao 0002, Shixin Chen, Zixiao Wang 0001, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Ultrafast Source Mask Optimization via Conditional Discrete DiffusionabstractSource mask optimization (SMO) is vital for mitigating lithography imaging distortions caused by shrinking critical dimensions in integrated circuit fabrication. However, the computational intensity of SMO, involving multiple integrals in Abbe’s theory, hinders its widespread adoption and advancement. In this paper, we present Diff-SMO, a highly efficient and accurate SMO framework with a primary emphasis on enhancing source optimization techniques. Previous research was confined to mask optimization acceleration due to the constraints of the academia lithography model. Diff-SMO extends the scope of optimization by concurrently refining the intricate interplay between the source and mask. We first develop a GPU-accelerated lithography simulator grounded in Abbe’s theory, enabling full GPU acceleration throughout the SMO process. Furthermore, we propose a discrete diffusion model for generating quasi-optimal sources, significantly improving computational efficiency. Our experimental results demonstrate exceptional imaging fidelity, surpassing the state-of-the-art, with over 200 times higher throughput compared to traditional SMO methods. Guojin Chen, Zixiao Wang 0001, Bei Yu 0001, David Z. Pan, Martin D. F. Wong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Truncate-Split-Contrast: A Framework for Learning from Mislabeled VideosabstractLearning with noisy label is a classic problem that has been extensively studied for image tasks, but much less for video in the literature. A straightforward migration from images to videos without considering temporal semantics and computational cost is not a sound choice. In this paper, we propose two new strategies for video analysis with noisy labels: 1) a lightweight channel selection method dubbed as Channel Truncation for feature-based label noise detection. This method selects the most discriminative channels to split clean and noisy instances in each category. 2) A novel contrastive strategy dubbed as Noise Contrastive Learning, which constructs the relationship between clean and noisy instances to regularize model training. Experiments on three well-known benchmark datasets for video classification show that our proposed truNcatE-split-contrAsT (NEAT) significantly outperforms the existing baselines. By reducing the dimension to 10% of it, our method achieves over 0.4 noise detection F1-score and 5% classification accuracy improvement on Mini-Kinetics dataset under severe noise (symmetric-80%). Thanks to Noise Contrastive Learning, the average classification accuracy improvement on Mini-Kinetics and Sth-Sth-V1 is over 1.6%. Zixiao Wang 0001, Junwu Weng, Chun Yuan 0003, Jue Wang 0001 |
AAAI | 1 |
| 2023 | DiffPattern: Layout Pattern Generation via Discrete DiffusionabstractDeep generative models dominate the existing literature in layout pattern generation. However, leaving the guarantee of legality to an inexplicable neural network could be problematic in several applications. In this paper, we propose DiffPattern to generate reliable layout patterns. DiffPattern introduces a novel diverse topology generation method via a discrete diffusion model with compute-efficiently lossless layout pattern representation. Then a white-box pattern assessment is utilized to generate legal patterns given desired design rules. Our experiments on several benchmark settings show that DiffPattern significantly outperforms existing baselines and is capable of synthesizing reliable layout patterns. Zixiao Wang 0001, Yunheng Shen, Wenqian Zhao 0002, Guojin Chen, Farzan Farnia, Bei Yu 0001 |
DAC | 1 |
| 2023 | ATFormer: A Learned Performance Model with Transfer Learning Across Devices for Deep Learning Tensor ProgramsabstractThe training and inference efficiency of everlarger deep neural networks highly rely on the performance of tensor operators on specific hardware platforms.Therefore, a compilationbased optimization flow with automatic tensor generation and parameter tuning is necessary for efficient model deployment.While compilation-based methods with performance models can provide dynamic and suitable code optimization, they suffer from a large design space exploration with rough measurement accuracy and poor transferability among different hardware platforms.This paper presents ATFormer, a simple yet efficient design with attention-inspired modules to accurately predict the performance of optimized operators by capturing global and long-range dependencies within a complete scheduling space.Compared with state-of-the-arts, ATFormer can predict the optimal implementation of tensor operators to reduce inference time with minimal effort on modern DNN benchmarks.Furthermore, ATFormer with pre-trained parameters can quickly adapt to different workloads and hardware via transfer learning. Wenqian Zhao 0002, Zixiao Wang 0001, Bei Yu 0001 |
EMNLP | 4 |