VLDB 2026 Research / reviewers in the wild / expert
Xinyun Zhang 0001
dblp:150/9539-1
· DBLP profile ↗
20ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-7763-7507ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 4 first-author · 13 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Video-based Visible-Event Cross-modal Person Re-identification for Edge AI Surveillance SystemsabstractVideo-based cross-modal person re-identification (ReID) is a critical task for video surveillance and security systems, particularly in resource-constrained edge AI environments. While existing crossmodal ReID methods primarily focus on thermal-visible matching, event cameras, with their low power consumption, high temporal resolution, and sparse data representation, offer significant advantages for edgebased surveillance systems by reducing data processing overhead and enabling robust performance under challenging lighting conditions. In this paper, we introduce a novel task: video-based visible-event person re-identification (VE ReID), which aims to match identities across RGB and event camera modalities. To the best of our knowledge, this is the first work to systematically define and investigate this cross-modal task in the context of event-driven edge AI. Specifically, we curate evaluation benchmarks from existing RGB-event datasets and synthesize a new RGB-event dataset, explicitly adapting them to the cross-modal ReID setting to enable a more comprehensive evaluation of VE ReID. Extensive experiments reveal that existing cross-modal state-of-the-art (SOTA) methods fail to effectively address the unique challenges posed by event data, highlighting the importance of tailored solutions for this task. To this end, we propose a novel method that constructs auxiliary modalities using frequency information from RGB and event tracklets, aligning them effectively through a fine-grained metric learning loss. Our approach not only achieves significant accuracy improvements over existing methods but also demonstrates the potential of event cameras for efficient and scalable edge AI surveillance applications. All source code and benchmarks are publicly available at https://github.com/yxgnahz/ASPDAC26-Event-RGBReID. Xinyun Zhang 0001, Zixiao Wang 0001, Yurui Kuang, Bei Yu 0001 |
ASP-DAC | 1 |
| 2026 | DiffResist: Physics-Constrained Diffusion for Photoresist ModelingabstractAccurate and efficient 3D photoresist simulation is essential for optical lithography at advanced technology nodes. Existing methods that predict 3D resist profiles from aerial images either rely on analytical reaction-diffusion solvers, which are slow, or on high-capacity 3D generative models, which are costly to train and deploy. We instead formulate 3D resist prediction as a depth-wise 2D generation task conditioned on the aerial image. DiffResist introduces a physics-constrained diffusion model whose reverse steps are aligned with resist exposure physics: a two-stage noise schedule connects physically meaningful layers to a Gaussian prior, and boundary conditions at the resist-air interface are injected to suppress error propagation. Combined with a lightweight super-resolution module, DiffResist achieves state-of-the-art accuracy on a public benchmark with over 10 × faster inference than 3D diffusion baselines. Zixiao Wang 0001, Jieya Zhou, Xinyun Zhang 0001, Shoubo Hu, Farzan Farnia, Bei Yu 0001 |
DATE | 3 |
| 2026 | Node2Node: Node Adaptation with Transformer for Cross-Node Hotspot DetectionabstractAs semiconductor manufacturing advances to smaller process nodes, hotspot detection has become critical for ensuring the manufacturability and reliability of integrated circuit (IC) layouts. However, existing detection methods rely heavily on labeled data tailored to specific nodes, resulting in poor generalizability across nodes due to variations in layout geometries and fabrication processes. Labeling new data at advanced nodes is also costly and time-consuming. To overcome these challenges, we propose Node2Node, the first adaptation framework explicitly designed for cross-node hotspot detection. Node2Node integrates a novel node-invariant encoder with a node-specific encoder to jointly capture transferable and node-dependent features. To further improve robustness, we introduce a bidirectional center alignment strategy, which refines pseudo-labels by leveraging a small amount of labeled data from the target node. Additionally, a cross-node distribution loss is introduced to explicitly align feature distributions between nodes. Extensive experiments demonstrate that Node2Node substantially improves cross-node generalization and achieves state-of-the-art hotspot detection performance. Silin Chen, Yibo Huang 0009, Xinyun Zhang 0001, Zixiao Wang 0001, Bei Yu 0001, Ningmu Zou |
DATE | 4 |
| 2026 | RATuner: Retrieval-Augmented VLSI Flow Design Parameter Tuning Framework
Peng Xu 0052, Ziyang Yu 0001, Yuan Pu 0001, Xinyun Zhang 0001, Donger Luo, Hao Geng, Tsung-Yi Ho, Bei Yu 0001 |
DATE | 4 |
| 2026 | HyDAS: Hybrid Domain Deformed Attention for Selective Hotspot DetectionabstractTechnology node scaling is challenged in many aspects, including pitch reduction, patterning flexibility, and lithography process variability during manufacturing. Without exception, layout hotspot detection, one of the critical steps to achieving design closure, also requires upgrading the associated techniques. With the rapid development of deep learning techniques, the detector exploiting convolutional neural network (CNN) is superior to ones based on pattern matching and classical machine learning algorithms. However, due to the local nature of CNN, the traditional CNN-based detector fails to model the relationship between the patterns in a large-sized layout, resulting in ignoring the impact of light propagation and some optical effects during photolithography. Even worse, another challenge arises from the fact that engineers cannot fully trust the results of learning model-based detectors, especially when handling some complicated layout patterns in practice. This makes it very difficult to deploy the detectors. Observing the facts, we propose a vision transformer (ViT) model-based layout hotspot detector with a deformed attention mechanism, where the training paradigm is inspired by the large pre-trained foundation model (e.g., OpenAI’s GPT-n series) and fine-tuning. Considering the light diffraction during photolithography, the hybrid domain (i.e., spatial and spectral domain) layout inputs via multi-channel are leveraged. Besides, our proposed detector integrates a selective option where the model can choose to do prediction or send to engineers based on the misclassification risk level. Experimental results on the ICCAD2012 metal layer benchmarks and ICCAD2020 via layer benchmarks demonstrate the effectiveness and efficiency of our approach. We have made the ICCAD2020 dataset publicly available to support further research in hotspot detection, enable benchmarking across different process nodes and layout types, and facilitate reproducibility in the field. The dataset is accessible at https: //github.com/shadowior/ICCAD2020. Qi Sun 0002, Su Zheng, Xinyun Zhang 0001, Bei Yu 0001, Hao Geng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | MM-GRADE: A Multi-Modal EDA Tool Documentation QA Framework Leveraging Retrieval Augmented GenerationabstractThe complexity of EDA tools necessitates the development of advanced documentation query answering systems to enhance user efficiency and reduce the associated learning curve. Recent innovations in the use of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) for EDA tool documentation have demonstrated significant progress; however, these approaches typically lack the multi-modal capabilities required to effectively handle visual data, such as circuit layout and GUI screenshots provided through user input. To address the concern, we introduce a multi-modal RAG system that incorporates two domain-customized modules: a multi-modal retriever model finetuned by the customized bilevel hard negative mining (BHNM) strategy, and a vision large language model (VLLM) finetuned using a tailored extract-score-answer pipeline. Moreover, we have manually curated ORD-MMBench, a multi-modal QA benchmark comprising 120 high-quality question-document-answer triplets based on OpenROAD documentation. Experimental results demonstrate that our customized RAG framework outperforms state-of-the-art multi-modal RAG flows and models on ORD-MMBench. Yuan Pu 0001, Zhuolun He, Shutong Lin, Jiajun Qin, Xinyun Zhang 0001, Hairuo Han, Haisheng Zheng, Cheng Zhuo, Qi Sun 0002, David Z. Pan, Bei Yu 0001 |
ICCAD | 5 |
| 2025 | Prerouting Timing Prediction Across Different Technology NodesabstractIn the domain of very-large-scale integration (VLSI) design, the accuracy of prerouting timing prediction is of paramount importance for ensuring the performance and reliability of integrated circuits. Traditional methods based on machine learning necessitate the availability of extensive and high-quality datasets. However, this requirement poses significant challenges for advanced technology nodes due to the laborious and time-intensive nature of data preparation. To address this critical issue, we introduce a novel transfer learning framework that leverages data from preceding technology nodes to facilitate learning and prediction on the target node. Our methodology commences with the disentanglement and alignment of timing path features across different nodes, ensuring the preservation and effective translation of intrinsic timing path properties. Subsequently, we employ a Bayesian-based model to predict the arrival times of individual timing paths. This model is particularly adept at managing the high-variability inherent in arrival times and exhibits strong generalization capabilities to novel design scenarios. Moreover, we propose a new algorithm to reweight the preceding node data during training by estimating their transferability through the cell type distribution. We validate the efficacy of our proposed framework through comprehensive experimental evaluations, demonstrating successful transfer learning from 130 or 45 to 7-nm technology nodes. The results underscore the potential of our approach to significantly mitigate the dependency on extensive data preparation while maintaining high accuracy in timing prediction for cutting-edge VLSI designs. Xinyun Zhang 0001, Binwu Zhu, Fangzhou Liu 0005, Jiaxi Jiang, Ziyi Wang 0010, Peng Xu 0052, Hong Xu 0001, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | p-Laplacian Adaptation for Generative Pre-trained Vision-Language ModelsabstractVision-Language models (VLMs) pre-trained on large corpora have demonstrated notable success across a range of downstream tasks. In light of the rapidly increasing size of pre-trained VLMs, parameter-efficient transfer learning (PETL) has garnered attention as a viable alternative to full fine-tuning. One such approach is the adapter, which introduces a few trainable parameters into the pre-trained models while preserving the original parameters during adaptation. In this paper, we present a novel modeling framework that recasts adapter tuning after attention as a graph message passing process on attention graphs, where the projected query and value features and attention matrix constitute the node features and the graph adjacency matrix, respectively. Within this framework, tuning adapters in VLMs necessitates handling heterophilic graphs, owing to the disparity between the projected query and value space. To address this challenge, we propose a new adapter architecture, p-adapter, which employs p-Laplacian message passing in Graph Neural Networks (GNNs). Specifically, the attention weights are re-normalized based on the features, and the features are then aggregated using the calibrated attention matrix, enabling the dynamic exploitation of information with varying frequencies in the heterophilic attention graphs. We conduct extensive experiments on different pre-trained VLMs and multi-modal tasks, including visual question answering, visual entailment, and image captioning. The experimental results validate our method's significant superiority over other PETL methods. Our code is available at https://github.com/wuhy68/p-Adapter/. Haoyuan Wu, Xinyun Zhang 0001, Peng Xu 0052, Peiyu Liao, Xufeng Yao, Bei Yu 0001 |
AAAI | 2 |
| 2024 | Progressively Knowledge Distillation via Re-parameterizing Diffusion Reverse ProcessabstractKnowledge distillation aims at transferring knowledge from the teacher model to the student one by aligning their distributions. Feature-level distillation often uses L2 distance or its variants as the loss function, based on the assumption that outputs follow normal distributions. This poses a significant challenge when distribution gaps are substantial since this loss function ignores the variance term. To address the problem, we propose to decompose the transfer objective into small parts and optimize it progressively. This process is inspired by diffusion models from which the noise distribution is mapped to the target distribution step by step. However, directly employing diffusion models is impractical in the distillation scenario due to its heavy reverse process. To overcome this challenge, we adopt the structural re-parameterization technique to generate multiple student features to approximate the teacher features sequentially. The multiple student features are combined linearly in inference time without extra cost. We present extensive experiments performed on various transfer scenarios, such as CNN-to-CNN and Transformer-to-CNN, that validate the effectiveness of our approach. Xufeng Yao, Fanbin Lu, Yuechen Zhang, Xinyun Zhang 0001, Wenqian Zhao 0002, Bei Yu 0001 |
AAAI | 4 |
| 2024 | Fracturing-aware Curvilinear ILT via Circular E-beam Mask WriterabstractInverse lithography technology (ILT) plays a crucial role in optical proximity correction, tending to generate curvilinear masks for optimal process windows. Traditional curvilinear mask manufacturing involves fracturing into rectangles, requiring expensive mask write times. A novel E-beam mask writer that writes variable radius circles per shot significantly reduces the shot count for curvilinear masks. To exploit this mask writer's benefits, we present two methods to generate circular fracturing-aware masks. The first one converts pixel-based masks from existing ILT methods into circle-based masks using predefined rules. The second one integrates circular constraints into the ILT process, generating circle-based masks directly via optimization. Extensive experimental results validate both approaches' effectiveness. Xinyun Zhang 0001, Su Zheng, Guojin Chen, Binwu Zhu, Hong Xu 0001, Bei Yu 0001 |
DAC | 1 |
| 2024 | Disentangle, Align and Generalize: Learning A Timing Predictor from Different Technology NodesabstractIn VLSI design, accurate pre-routing timing prediction is paramount. Traditional machine learning-based methods require extensive data, posing challenges for advanced technology nodes due to the time-consuming data preparation. To mitigate this issue, we propose a novel transfer learning framework that uses data from previous nodes for learning on the target node. Our method initially disentangles and aligns timing path features across different nodes, then predicts each path's arrival time employing a Bayesian-based model capable of handling highly variable arrival time and generalizing to new designs. Experimental results on transfer learning from 130nm to 7nm nodes validate our method's effectiveness. Xinyun Zhang 0001, Binwu Zhu, Fangzhou Liu 0005, Ziyi Wang 0010, Peng Xu 0052, Hong Xu 0001, Bei Yu 0001 |
DAC | 1 |
| 2024 | ChatEDA: A Large Language Model Powered Autonomous Agent for EDAabstractThe integration of a complex set of Electronic Design Automation (EDA) tools to enhance interoperability is a critical concern for circuit designers. Recent advancements in large language models (LLMs) have showcased their exceptional capabilities in natural language processing and comprehension, offering a novel approach to interfacing with EDA tools. This research paper introduces ChatEDA, an autonomous agent for EDA empowered by a large language model, AutoMage, complemented by EDA tools serving as executors. ChatEDA streamlines the design flow from the Register-Transfer Level (RTL) to the Graphic Data System Version II (GDSII) by effectively managing task decomposition, script generation, and task execution. Through comprehensive experimental evaluations, ChatEDA has demonstrated its proficiency in handling diverse requirements, and our fine-tuned AutoMage model has exhibited superior performance compared to GPT-4 and other similar LLMs. Haoyuan Wu, Zhuolun He, Xinyun Zhang 0001, Xufeng Yao, Su Zheng, Haisheng Zheng, Bei Yu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Reconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting TranscriptsabstractMeetings typically involve multiple participants and lengthy conversations, resulting in redundant and trivial content.To overcome these challenges, we propose a two-step framework, Reconstruct before Summarize (RbS), for effective and efficient meeting summarization.RbS first leverages a self-supervised paradigm to annotate essential contents by reconstructing the meeting transcripts.Secondly, we propose a relative positional bucketing (RPB) algorithm to equip (conventional) summarization models to generate the summary.Despite the additional reconstruction process, our proposed RPB significantly compressed the input, leading to faster processing and reduced memory consumption compared to traditional summarization methods.We validate the effectiveness and efficiency of our method through extensive evaluations and analysis.On two meeting summarization datasets, AMI and ICSI, our approach outperforms previous state-of-the-art approaches without relying on large-scale pretraining or expert-grade annotating tools. Haochen Tan, Han Wu 0004, Wei Shao 0009, Xinyun Zhang 0001, Mingjie Zhan, Zhaohui Hou, Ding Liang, Linqi Song |
EMNLP | 4 |
| 2023 | Efficient Deep Space Filling CurveabstractSpace-filling curves (SFCs) act as a linearization approach to map data in higher dimensional space to lower dimensional space, which is used comprehensively in computer vision, such as image/point cloud compression, hashing and etc. Currently, researchers formulate the problem of searching for an optimal SFC to the problem of finding a single Hamiltonian circuit on the image grid graph. Existing methods adopt graph neural networks (GNN) for SFC search. By modeling the pixel grid as a graph, they first adopt GNN to predict the edge weights and then generate a minimum spanning tree (MST) based on the predictions, which is further used to construct the SFC. However, GNN-based methods suffer from high computational costs and memory footprint usage. Besides, MST generation is un-differentiable, which is infeasible to optimize via gradient descent. To remedy these issues, we propose a GNN-based SFC-search framework with a tailored algorithm that largely reduces computational cost of GNN. Additionally, we propose a siamese network learning scheme to optimize DNN-based models in an end-to-end fashion. Extensive experiments show that our proposed method outperforms both DNN-based methods and traditional SFCs, e.g. Hilbert curve, by a large margin on various benchmarks. Xufeng Yao, Xinyun Zhang 0001, Bei Yu 0001 |
ICCV | 3 |
| 2023 | DRC-SG 2.0: Efficient Design Rule Checking Script Generation via Key Information ExtractionabstractDesign Rule Checking (DRC) is a critical step in integrated circuit design. DRC requires formatted scripts as the input to design rule checkers. However, these scripts are manually generated in the foundry, which is tedious and error prone for generation of thousands of rules in advanced technology nodes. To mitigate this issue, we propose the first DRC script generation framework, leveraging a deep learning-based key information extractor to automatically identify essential arguments from rules and a script translator to organize the extracted arguments into executable DRC scripts. We further enhance the performance of the extractor with three specific design rule generation techniques and a multi-task learning-based rule classification module. Experimental results demonstrate that the framework can generate a single rule script in 5.46 ms on average, with the extractor achieving 91.1% precision and 91.8% recall on the key information extraction. Compared with the manual generation, our framework can significantly reduce the turnaround time and speed up process design closure. Binwu Zhu, Xinyun Zhang 0001, Yibo Lin, Bei Yu 0001, Martin D. F. Wong |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2022 | Context-Based Contrastive Learning for Scene Text RecognitionabstractPursuing accurate and robust recognizers has been a long-lasting goal for scene text recognition (STR) researchers. Recently, attention-based methods have demonstrated their effectiveness and achieved impressive results on public benchmarks. The attention mechanism enables models to recognize scene text with severe visual distortions by leveraging contextual information. However, recent studies revealed that the implicit over-reliance of context leads to catastrophic out-of-vocabulary performance. On the contrary to the superior accuracy of the seen text, models are prone to misrecognize unseen text even with good image quality. We propose a novel framework, Context-based contrastive learning (ConCLR), to alleviate this issue. Our proposed method first generates characters with different contexts via simple image concatenation operations and then optimizes contrastive loss on their embeddings. By pulling together clusters of identical characters within various contexts and pushing apart clusters of different characters in embedding space, ConCLR suppresses the side-effect of overfitting to specific contexts and learns a more robust representation. Experiments show that ConCLR significantly improves out-of-vocabulary generalization and achieves state-of-the-art performance on public benchmarks together with attention-based recognizers. Xinyun Zhang 0001, Binwu Zhu, Xufeng Yao, Qi Sun 0002, Ruiyu Li, Bei Yu 0001 |
AAAI | 1 |
| 2022 | PCL: Proxy-based Contrastive Learning for Domain GeneralizationabstractDomain generalization refers to the problem of training a model from a collection of different source domains that can directly generalize to the unseen target domains. A promising solution is contrastive learning, which attempts to learn domain-invariant representations by exploiting rich semantic relations among sample-to-sample pairs from different domains. A simple approach is to pull positive sample pairs from different domains closer while pushing other negative pairs further apart. In this paper, we find that directly applying contrastive-based methods (e.g., supervised contrastive learning) are not effective in domain generalization. We argue that aligning positive sample-to-sample pairs tends to hinder the model generalization due to the significant distribution gaps between different domains. To address this issue, we propose a novel proxy-based contrastive learning method, which replaces the original sample-to-sample relations with proxy-to-sample relations, significantly alleviating the positive alignment issue. Experiments on the four standard benchmarks demonstrate the effectiveness of the proposed method. Furthermore, we also consider a more complex scenario where no ImageNet pre-trained models are provided. Our method consistently shows better performance. Xufeng Yao, Xinyun Zhang 0001, Yuechen Zhang, Qi Sun 0002, Ran Chen 0001, Ruiyu Li, Bei Yu 0001 |
CVPR | 3 |
| 2022 | GTuner: tuning DNN computations on GPU via graph attention networkabstractIt is an open problem to compile DNN models on GPU and improve the performance. A novel framework, GTuner, is proposed to jointly learn from the structures of computational graphs and the statistical features of codes to find the optimal code implementations. A Graph ATtention network (GAT) is designed as the performance estimator in GTuner. In GAT, graph neural layers are used to propagate the information in the graph and a multi-head self-attention module is designed to learn the complicated relationships between the features. Under the guidance of GAT, the GPU codes are generated through auto-tuning. Experimental results demonstrate that our method outperforms the previous arts remarkably. Qi Sun 0002, Xinyun Zhang 0001, Hao Geng, Yuxuan Zhao 0001, Haisheng Zheng, Bei Yu 0001 |
DAC | 2 |
| 2021 | Hotspot Detection via Multi-task Learning and Transformer EncoderabstractWith the rapid development of semiconductors and the continuous scaling-down of circuit feature size, hotspot detection has become much more challenging and crucial as a critical step in the physical verification flow. In recent years, advanced deep learning techniques have spawned many frameworks for hotspot detection. However, most existing hotspot detectors can only detect defects arising in the central region of small clips, making the whole detection process time-consuming on large layouts. Some advanced hotspot detectors can detect multiple hotspots in a large area but need to propose potential defect regions, and a refinement step is required to locate the hotspot precisely. To simplify the procedure of multi-stage detectors, an end - to-end single-stage hotspot detector is proposed to identify hotspots on large scales without refining potential regions. Besides, multiple tasks are developed to learn various pattern topological features. Also, a feature aggregation module based on Transformer Encoder is designed to globally capture the relationship between different features, further enhancing the feature representation ability. Experimental results show that our proposed framework achieves higher accuracy over prior methods with faster inference speed. Binwu Zhu, Ran Chen 0001, Xinyun Zhang 0001, Fan Yang 0001, Xuan Zeng 0001, Bei Yu 0001, Martin D. F. Wong |
ICCAD | 3 |
| 2021 | Fast and Efficient DNN Deployment via Deep Gaussian Transfer LearningabstractDeep neural networks (DNNs) have been widely used recently while their hardware deployment optimizations are very time-consuming and the historical deployment knowledge is not utilized efficiently. In this paper, to accelerate the optimization process and find better deployment configurations, we propose a novel transfer learning method based on deep Gaussian processes (DGPs). Firstly, a deep Gaussian process (DGP) model is built on the historical data to learn empirical knowledge. Secondly, to transfer knowledge to a new task, a tuning set is sampled for the new task under the guidance of the DGP model. Then DGP is tuned according to the tuning set via maximum-a-posteriori (MAP) estimation to accommodate for the new task and finally used to guide the deployments of the task. The experiments show that our method achieves the best inference latencies of convolutions while accelerating the optimization process significantly, compared with previous arts. Qi Sun 0002, Tinghuan Chen, Hao Geng, Xinyun Zhang 0001, Bei Yu 0001 |
ICCV | 5 |