Zhen Wang 0030

dblp:78/6727-30 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
20since 2021 · last 2025
0000-0001-5350-4569ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 6 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 LMM-IR: Large-Scale Netlist-Aware Multimodal Framework for Static IR-Drop Prediction
abstract
Static IR drop analysis is a fundamental and critical task in the field of chip design. Nevertheless, this process can be quite time-consuming, potentially requiring several hours. Moreover, addressing IR drop violations frequently demands iterative analysis, thereby causing the computational burden. Therefore, fast and accurate IR drop prediction is vital for reducing the overall time invested in chip design. In this paper, we firstly propose a novel multimodal approach that efficiently processes SPICE files through large-scale netlist transformer (LNT). Our key innovation is representing and processing netlist topology as 3D point cloud representations, enabling efficient handling of netlist with up to hundreds of thousands to millions nodes. All types of data, including netlist files and image data, are encoded into latent space as features and fed into the model for static voltage drop prediction. This enables the integration of data from multiple modalities for complementary predictions. Experimental results demonstrate that our proposed algorithm can achieve the best F1 score and the lowest MAE among the winning teams of the ICCAD 2023 contest and the state-of-theart algorithms.
Zhen Wang 0030, Hongquan He, Qi Xu 0004, Tinghuan Chen, Hao Geng
DAC2
2025 LMLitho: A Large Vision Model-Driven Lithography Simulation Framework
abstract
As IC fabrication advances toward smaller process nodes, design technology co-optimization (DTCO) has emerged as a critical enabler of chip performance advancements. Lithography simulation, vital for bridging design and manufacturing, now plays an indispensable role in designing litho-friendly layouts/masks and developing resolution enhancement techniques (RETs). While academia and industry have explored statistical techniques and machine learning models for simulators, the computing paradigm and hardware prevent these solutions from efficiently and accurately simulating the complicated optical imaging coupled with resist film imaging. In this paper, we propose a new simulation paradigm: LMLitho (large vision model-driven lithography simulator), trained on circa one hundred thousand triplets of illumination maps, masks, and resist images. The cross-attention mechanism in our simulator inherently captures diffraction patterns akin to light wave interference within mask features, while hierarchical attention layers enable the modeling of long-range diffraction effects (e.g., proximity effects). A comprehensive dataset encompassing diverse classical types of source and mask patterns, including both metal-1 and via layers, is generated to meet the requirements of training our large vision model-based simulator1. The experimental results demonstrate that our simulator achieves over 120× speedup compared to existing commercial solutions while preserving comparable high fidelity, and exhibits superior generalization to advanced process nodes. When deployed in inverse lithography technology (ILT)-guided mask optimization workflows, masks of higher quality are generated than existing solutions.
Zhen Wang 0030, Hongquan He, Xuming He 0001, Qi Sun 0002, Cheng Zhuo, Bei Yu 0001, Jingyi Yu 0001, Hao Geng
ICCAD1
2025 When Semi-Supervised LVM Meets Frequency-Based Critical Layout Pattern Selection
abstract
Critical pattern selection is an essential foundation for lithography technologies such as full-chip Source Mask Optimization (SMO) and Optical Proximity Correction (OPC) model calibration. Traditional methods are mostly based on image feature extraction and clustering, without integrating lithography domain knowledge, and they suffer from high time complexity and excessive computational resource consumption. As a result, they fail to fully capture the optical characteristics of complex patterns and lack subsequent verification of the "critical" nature of the patterns. Moreover, past work often relied on manually annotated data, and the generalization ability needs further validation. To overcome these limitations, this paper proposes a semi-supervised critical pattern selection method based on frequency domain and LVM collaboration. This method defines the critical patterns by describing the optical characteristics of the patterns in the frequency domain and uses the ViT model to extract the nonlinear features of the patterns, ultimately selecting the critical patterns efficiently. This approach meets the demands of high-precision and high-efficiency lithography techniques for advanced processes. A series of simulations and comparisons performed using commercial tool, verify the validity of the proposed method.
Liuke Wang, Shenshuo Yao, Shihan Wang 0007, Zhen Wang 0030, Jingyi Yu 0001, Hao Geng
ICCAD4
2025 LithoSim: A Large, Holistic Lithography Simulation Benchmark for AI-Driven Semiconductor Manufacturing
abstract
Lithography orchestrates a symphony of light, mask and photochemicals to transfer the integrated circuit patterns onto the wafer. Lithography simulation serves as the critical nexus between circuit design and manufacturing, where its speed and accuracy fundamentally govern the optimization quality of downstream resolution enhancement techniques (RET). While machine learning promises to circumvent computational limitations of lithography process through data-driven or physics-informed approximations of computational lithography, existing simulators suffer from inadequate lithographic awareness due to insufficient training data capturing essential process variations and mask correction rules. We present LithoSim, the most comprehensive lithography simulation benchmark to date, featuring over $4$ million high-resolution input-output pairs with rigorous physical correspondence. The dataset systematically incorporates alterable optical source distributions, metal and via mask topologies with optical proximity correction (OPC) variants, and process windows reflecting fab-realistic variations. By integrating domain-specific metrics spanning AI performance and lithographic fidelity, LithoSim establishes a unified evaluation framework for data-driven and physics-informed computational lithography. The data (https://huggingface.co/datasets/grandiflorum/LithoSim), code (https://dw-hongquan.github.io/LithoSim), and pre-trained models (https://huggingface.co/grandiflorum/LithoSim) are released openly to support the development of hybrid ML-based and high-fidelity lithography simulation for the benefit of semiconductor manufacturing.
Hongquan He, Zhen Wang 0030, Jingya Wang 0001, Xuming He 0001, Bei Yu 0001, Jingyi Yu 0001, Hao Geng
NeurIPS2
2025 A Comprehensive Survey of Data Augmentation in Visual Reinforcement Learning
Guozheng Ma, Zhen Wang 0030, Zhecheng Yuan, Xueqian Wang 0001, Bo Yuan 0003, Dacheng Tao
Int. J. Comput. Vis.2
2024 Revisiting Plasticity in Visual Reinforcement Learning: Data, Modules and Training Stages
abstract
Plasticity, the ability of a neural network to evolve with new data, is crucial for high-performance and sample-efficient visual reinforcement learning (VRL). Although methods like resetting and regularization can potentially mitigate plasticity loss, the influences of various components within the VRL framework on the agent's plasticity are still poorly understood. In this work, we conduct a systematic empirical exploration focusing on three primary underexplored facets and derive the following insightful conclusions: (1) data augmentation is essential in maintaining plasticity; (2) the critic's plasticity loss serves as the principal bottleneck impeding efficient training; and (3) without timely intervention to recover critic's plasticity in the early stages, its loss becomes catastrophic. These insights suggest a novel strategy to address the high replay ratio (RR) dilemma, where exacerbated plasticity loss hinders the potential improvements of sample efficiency brought by increased reuse frequency. Rather than setting a static RR for the entire training process, we propose Adaptive RR, which dynamically adjusts the RR based on the critic’s plasticity level. Extensive evaluations indicate that Adaptive RR not only avoids catastrophic plasticity loss in the early stages but also benefits from more frequent reuse in later phases, resulting in superior sample efficiency.
Guozheng Ma, Sen Zhang 0006, Zixuan Liu 0002, Zhen Wang 0030, Yixin Chen 0001, Li Shen 0008, Xueqian Wang 0001, Dacheng Tao
ICLR5
2023 Adaptive Self-Supervised Continual Learning
abstract
Continual Learning (CL) studies the problem of developing a robust model that can learn new tasks while retaining previously learned knowledge. However, the current CL methods exclusively focus on data with annotations, disregarding that unlabelled data is the mainstream in real-world applications. To close this research gap, this study concentrates on continual self-supervised learning, which is plagued by challenges of memory over-fitting and class imbalance. Besides, these challenges are exacerbated throughout incremental training. Aimed at addressing these challenges from both loss and data perspectives, we introduce a framework, Adaptive Self-supervised Continual Learning (ASCL). Specifically, we devise an Adaptive Sharpness-Aware Minimization (ASAM) module responsible for identifying flatter local minima in the loss landscape with a smaller memory over-fitting risk. Additionally, we design an Adaptive Memory Enhancement (AME) module responsible for rebalancing self-supervised loss with new and old tasks from a data perspective. Finally, the adaptive mechanisms in AME and ASAM modules dynamically adjust the loss landscape sharpness and memory enhancement strength with the feedback of intermediate training results. The results of our extensive experiments demonstrate the state-of-the-art performance of our methods in continual self-supervised learning scenarios across multiple datasets.
Lilei Wu, Zhen Wang 0030, Jie Liu 0002
ECAI2
2023 DeWave: Discrete Encoding of EEG Waves for EEG to Text Translation
abstract
The translation of brain dynamics into natural language is pivotal for brain-computer interfaces (BCIs), a field that has seen substantial growth in recent years. With the swift advancement of large language models, such as ChatGPT, the need to bridge the gap between the brain and languages becomes increasingly pressing. Current methods, however, require eye-tracking fixations or event markers to segment brain dynamics into word-level features, which can restrict the practical application of these systems. These event markers may not be readily available or could be challenging to acquire during real-time inference, and the sequence of eye fixations may not align with the order of spoken words. To tackle these issues, we introduce a novel framework, DeWave, that integrates discrete encoding sequences into open-vocabulary EEG-to-text translation tasks. DeWave uses a quantized variational encoder to derive discrete codex encoding and align it with pre-trained language models. This discrete codex representation brings forth two advantages: 1) it alleviates the order mismatch between eye fixations and spoken words by introducing text-EEG contrastive alignment training, and 2) it minimizes the interference caused by individual differences in EEG waves through an invariant discrete codex. Our model surpasses the previous baseline (40.1 and 31.7) by 3.06% and 6.34\%, respectively, achieving 41.35 BLEU-1 and 33.71 Rouge-F on the ZuCo Dataset. Furthermore, this work is the first to facilitate the translation of entire EEG signal periods without the need for word-level order markers (e.g., eye fixations), scoring 20.5 BLEU-1 and 29.5 Rouge-1 on the ZuCo Dataset, respectively.
Yiqun Duan, Charles Chau, Zhen Wang 0030, Yu-Kai Wang, Chin-Teng Lin
NeurIPS3
2023 Learning Better with Less: Effective Augmentation for Sample-Efficient Visual Reinforcement Learning
abstract
Data augmentation (DA) is a crucial technique for enhancing the sample efficiency of visual reinforcement learning (RL) algorithms. Notably, employing simple observation transformations alone can yield outstanding performance without extra auxiliary representation tasks or pre-trained encoders. However, it remains unclear which attributes of DA account for its effectiveness in achieving sample-efficient visual RL. To investigate this issue and further explore the potential of DA, this work conducts comprehensive experiments to assess the impact of DA's attributes on its efficacy and provides the following insights and improvements: (1) For individual DA operations, we reveal that both ample spatial diversity and slight hardness are indispensable. Building on this finding, we introduce Random PadResize (Rand PR), a new DA operation that offers abundant spatial diversity with minimal hardness. (2) For multi-type DA fusion schemes, the increased DA hardness and unstable data distribution result in the current fusion schemes being unable to achieve higher sample efficiency than their corresponding individual operations. Taking the non-stationary nature of RL into account, we propose a RL-tailored multi-type DA fusion scheme called Cycling Augmentation (CycAug), which performs periodic cycles of different DA operations to increase type diversity while maintaining data distribution consistency. Extensive evaluations on the DeepMind Control suite and CARLA driving simulator demonstrate that our methods achieve superior sample efficiency compared with the prior state-of-the-art methods.
Guozheng Ma, Linrui Zhang, Haoyu Wang 0018, Zilin Wang 0002, Zhen Wang 0030, Li Shen 0008, Xueqian Wang 0001, Dacheng Tao
NeurIPS6
2023 Cross-domain multi-style merge for image captioning
Yiqun Duan, Zhen Wang 0030, Yi Li 0050, Jingya Wang 0001
Comput. Vis. Image Underst.2
2023 Trust-Region Adaptive Frequency for Online Continual Learning
abstract
Abstract In the paradigm of online continual learning, one neural network is exposed to a sequence of tasks, where the data arrive in an online fashion and previously seen data are not accessible. Such online fashion causes insufficient learning and severe forgetting on past tasks issues, preventing a good stability-plasticity trade-off, where ideally the network is expected to have high plasticity to adapt to new tasks well and have the stability to prevent forgetting on old tasks simultaneously. To solve these issues, we propose a trust-region adaptive frequency approach, which alternates between standard-process and intra-process updates. Specifically, the standard-process replays data stored in a coreset and interleaves the data with current data, and the intra-process updates the network parameters based on the coreset. Furthermore, to improve the unsatisfactory performance stemming from online fashion, the frequency of the intra-process is adjusted based on a trust region, which is measured by the confidence score of current data. During the intra-process, we distill the dark knowledge to retain useful learned knowledge. Moreover, to store more representative data in the coreset, a confidence-based coreset selection is presented in an online manner. The experimental results on standard benchmarks show that the proposed method significantly outperforms state-of-art continual learning algorithms.
Yajing Kong, Liu Liu 0014, Maoying Qiao, Zhen Wang 0030, Dacheng Tao
Int. J. Comput. Vis.4
2023 Cross task neural architecture search for EEG signal recognition
Yiqun Duan, Zhen Wang 0030, Yi Li 0050, Jianhang Tang, Yu-Kai Wang, Chin-Teng Lin
Neurocomputing2
2023 SIN: Semantic Inference Network for Few-Shot Streaming Label Learning
abstract
Streaming label learning aims to model newly emerged labels for multilabel classification systems, which requires plenty of new label data for training. However, in changing environments, only a small amount of new label data can practically be collected. In this work, we formulate and study few-shot streaming label learning (FSLL), which models emerging new labels with only a few annotated examples by utilizing the knowledge learned from past labels. We propose a meta-learning framework, semantic inference network (SIN), which can learn and infer the semantic correlation between new labels and past labels to adapt FSLL tasks from a few examples effectively. SIN leverages label semantic representation to regularize the output space and acquires labelwise meta-knowledge based on gradient-based meta-learning. Moreover, SIN incorporates a novel label decision module with a meta-threshold loss to find the optimal confidence thresholds for each new label. Theoretically, we illustrate that the proposed semantic inference mechanism could constrain the complexity of hypotheses space to reduce the risk of overfitting and achieve better generalizability. Experimentally, extensive empirical results and ablation studies demonstrate the performance of SIN is superior to the prior state-of-the-art methods on FSLL.
Zhen Wang 0030, Liu Liu 0014, Yiqun Duan, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2022 Continual Learning through Retrieval and Imagination
abstract
Continual learning is an intellectual ability of artificial agents to learn new streaming labels from sequential data. The main impediment to continual learning is catastrophic forgetting, a severe performance degradation on previously learned tasks. Although simply replaying all previous data or continuously adding the model parameters could alleviate the issue, it is impractical in real-world applications due to the limited available resources. Inspired by the mechanism of the human brain to deepen its past impression, we propose a novel framework, Deep Retrieval and Imagination (DRI), which consists of two components: 1) an embedding network that constructs a unified embedding space without adding model parameters on the arrival of new tasks; and 2) a generative model to produce additional (imaginary) data based on the limited memory. By retrieving the past experiences and corresponding imaginary data, DRI distills knowledge and rebalances the embedding space to further mitigate forgetting. Theoretical analysis demonstrates that DRI can reduce the loss approximation error and improve the robustness through retrieval and imagination, bringing better generalizability to the network. Extensive experiments show that DRI performs significantly better than the existing state-of-the-art continual learning methods and effectively alleviates catastrophic forgetting.
Zhen Wang 0030, Liu Liu 0014, Yiqun Duan, Dacheng Tao
AAAI1
2022 Continual Learning with Lifelong Vision Transformer
abstract
Continual learning methods aim at training a neural network from sequential data with streaming labels, relieving catastrophic forgetting. However, existing methods are based on and designed for convolutional neural networks (CNNs), which have not utilized the full potential of newly emerged powerful vision transformers. In this paper, we propose a novel attention-based framework Lifelong Vision Transformer (LVT), to achieve a better stability-plasticity trade-off for continual learning. Specifically, an inter-task attention mechanism is presented in LVT, which implicitly absorbs the previous tasks' information and slows down the drift of important attention between previous tasks and the current task. LVT designs a dual-classifier structure that independently injects new representation to avoid catas-trophic interference and accumulates the new and previous knowledge in a balanced manner to improve the overall performance. Moreover, we develop a confidence-aware memory update strategy to deepen the impression of the previous tasks. The extensive experimental results show that our approach achieves state-of-the-art performance with even fewer parameters on continual learning benchmarks.
Zhen Wang 0030, Liu Liu 0014, Yiqun Duan, Yajing Kong, Dacheng Tao
CVPR1
2022 Balancing Stability and Plasticity Through Advanced Null Space in Continual Learning
Yajing Kong, Liu Liu 0014, Zhen Wang 0030, Dacheng Tao
ECCV (26)3
2022 Online Continual Learning with Contrastive Vision Transformer
Zhen Wang 0030, Liu Liu 0014, Yajing Kong, Jiaxian Guo, Dacheng Tao
ECCV (20)1
2022 Input Enhanced Logarithmic Factorization Network for CTR Prediction
Xianzhuang Li, Zhen Wang 0030, Xuesong Wu 0003, Bo Yuan 0003, Xueqian Wang 0001
PAKDD (3)2
2022 Position-aware image captioning with spatial relation
Yiqun Duan, Zhen Wang 0030, Jingya Wang 0001, Yu-Kai Wang, Chin-Teng Lin
Neurocomputing2
2021 Multi-label Few-shot Learning with Semantic Inference (Student Abstract)
abstract
Few-shot learning can adapt the classification model to new labels with only a few labeled examples. Previous studies mainly focus on the scenario of a single category label per example but have not solved the more challenging multi-label scenario with exponential-sized output space and low-data effectively. In this paper, we propose a semantic-aware meta-learning model for multi-label few-shot learning. Our approach can learn and infer the semantic correlation between unseen labels and historical labels to quickly adapt multi-label tasks from only a few examples. Specifically, features can be mapped into the semantic embedding space via label word vectors to explore and exploit the label correlation, and thus cope with the challenge on the overwhelming size of the output space. Then a novel semantic inference mechanism is designed for leveraging prior knowledge learned from historical labels, which will produce good generalization performance on new labels to alleviate the low-data problem. Finally, extensive empirical results show that the proposed method significantly outperforms the existing state-of-the-art methods on the multi-label few-shot learning tasks.
Zhen Wang 0030, Yiqun Duan, Liu Liu 0014, Dacheng Tao
AAAI1
2020 Deep Streaming Label Learning
abstract
In multi-label learning, each instance can be associated with multiple and non-exclusive labels. Previous studies assume that all the labels in the learning process are fixed and static; however, they ignore the fact that the labels will emerge continuously in changing environments. In order to fill in these research gaps, we propose a novel deep neural network (DNN) based framework, Deep Streaming Label Learning (DSLL), to classify instances with newly emerged labels effectively. DSLL can explore and incorporate the knowledge from past labels and historical models to understand and develop emerging new labels. DSLL consists of three components: 1) a streaming label mapping to extract deep relationships between new labels and past labels with a novel label-correlation aware loss; 2) a streaming feature distillation propagating feature-level knowledge from the historical model to a new model; 3) a senior student network to model new labels with the help of knowledge learned from the past. Theoretically, we prove that DSLL admits tight generalization error bounds for new labels in the DNN framework. Experimentally, extensive empirical results show that the proposed method performs significantly better than the existing state-of-the-art multi-label learning methods to handle the continually emerging new labels.
Zhen Wang 0030, Liu Liu 0014, Dacheng Tao
ICML1
2020 A Dual Input-aware Factorization Machine for CTR Prediction
abstract
Factorization Machines (FMs) refer to a class of general predictors working with real valued feature vectors, which are well-known for their ability to estimate model parameters under significant sparsity and have found successful applications in many areas such as the click-through rate (CTR) prediction. However, standard FMs only produce a single fixed representation for each feature across different input instances, which may limit the CTR model’s expressive and predictive power. Inspired by the success of Input-aware Factorization Machines (IFMs), which aim to learn more flexible and informative representations of a given feature according to different input instances, we propose a novel model named Dual Input-aware Factorization Machines (DIFMs) that can adaptively reweight the original feature representations at the bit-wise and vector-wise levels simultaneously. Furthermore, DIFMs strategically integrate various components including Multi-Head Self-Attention, Residual Networks and DNNs into a unified end-to-end model. Comprehensive experiments on two real-world CTR prediction datasets show that the DIFM model can outperform several state-of-the-art models consistently.
Wantong Lu, Yongzhe Chang, Zhen Wang 0030, Bo Yuan 0003
IJCAI4
2019 DBSVEC: Density-Based Clustering Using Support Vector Expansion
abstract
DBSCAN is a popular clustering algorithm that can discover clusters of arbitrary shapes with broad applications. However, DBSCAN is computationally expensive, as it performs range queries for all the points to determine their neighbors and grow the clusters. To address this problem, we propose a novel approximate density-based clustering algorithm named DBSVEC. DBSVEC introduces support vectors into density-based clustering, which allows performing range queries only on a small subset of points called the core support vectors. This technique significantly improves the efficiency while retaining high-quality cluster results. We evaluate the performance of DBSVEC via extensive experiments on real and synthetic datasets. The results show that DBSVEC is up to three orders of magnitude faster than DBSCAN. Compared with the state-of-the-art approximate density-based clustering methods, DBSVEC is up to two orders of magnitude faster, and the clustering results of DBSVEC are more similar to those of DBSCAN.
Zhen Wang 0030, Rui Zhang 0003, Jianzhong Qi 0001, Bo Yuan 0003
ICDE1
2019 An Input-aware Factorization Machine for Sparse Prediction
abstract
Factorization machines (FMs) are a class of general predictors working effectively with sparse data, which represents features using factorized parameters and weights. However, the accuracy of FMs can be adversely affected by the fixed representation trained for each feature, as the same feature is usually not equally predictive and useful in different instances. In fact, the inaccurate representation of features may even introduce noise and degrade the overall performance. In this work, we improve FMs by explicitly considering the impact of individual input upon the representation of features. We propose a novel model named \textit{Input-aware Factorization Machine} (IFM), which learns a unique input-aware factor for the same feature in different instances via a neural network. Comprehensive experiments on three real-world recommendation datasets are used to demonstrate the effectiveness and mechanism of IFM. Empirical results indicate that IFM is significantly better than the standard FM model and consistently outperforms four state-of-the-art deep learning based methods.
Zhen Wang 0030, Bo Yuan 0003
IJCAI2