Yihan Wen

dblp:314/2717 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Mamba-HTEPA: A multi-branch structure framework for multimodal grading of meningiomas using Mamba
Zhuo Zhang 0020, Yihan Wen, Wendi Liang, Guanchong Niu, Quanfeng Ma
Expert Syst. Appl.2
2026 Rule-Semantic Generative Calibration Blur Detection for UAV Imagery
Yihan Wen, Zhuo Zhang 0020, Xianping Ma, Peipei Zhu, Jinglei Li, Guanchong Niu, Qiguang Miao
IEEE Trans. Circuits Syst. Video Technol.1
2025 MOSS: Multi-Modal Representation Learning on Sequential Circuits
abstract
Deep learning has significantly advanced Electronic Design Automation (EDA), with circuit representation learning emerging as a key area for modeling the relationship between a circuit’s structure and functionality. Existing methods primarily use either Large Language Models (LLMs) for Register Transfer Level (RTL) code analysis or Graph Neural Networks (GNNs) for netlist modeling. While LLMs excel at high-level functional understanding, they struggle with detailed netlist behavior. GNNs, however, face challenges when scaling to larger sequential circuits due to long-range information dependencies and insufficient functional supervision, leading to decreased accuracy and limited generalization. To address these challenges, we propose MOSS, a multimodal framework that integrates GNNs with LLMs for sequential circuit modeling. By enhancing D-type Flip-Flop (DFF) node features with embeddings from fine-tuned LLMs on RTL code, we focus the GNN on critical anchor points, reducing reliance on long-range dependencies. The LLM also provides global circuit embeddings, offering efficient supervision for functionality-related tasks. Additionally, MOSS introduces an adaptive aggregation method and a two-phase propagation mechanism in the GNN to better model signal propagation and sequential feedback within the circuit. Experimental results demonstrate that MOSS significantly improves the accuracy of functionality and performance predictions for sequential circuits compared to existing methods, particularly in larger circuits where previous models struggle. Specifically, MOSS achieves a $\mathbf{9 5. 2 \%}$ accuracy in arrival time prediction.
Jianan Mu, Tianmeng Yang, Silin Liu, Yihan Wen, Hui Wang 0152, Zhiteng Chao, Husheng Han, Zizhen Liu, Shengwen Liang, Jing Ye 0001, Bei Yu 0001, Xiaowei Li 0001, Huawei Li 0001
DAC9
2025 EPICS: Efficient Parallel Pattern Fault Simulation for Sequential Circuits via Strongly Connected Components
abstract
As functional safety of electronic chips gains importance in autonomous vehicles and aerospace, standards like ISO 26262 mandate high diagnostic coverage, requiring extensive gate-level fault simulations. However, for large-scale industrial sequential circuits, these simulations are time-consuming, creating a significant bottleneck in chip development. Prior approaches have focused on reducing computational complexity and optimizing CPU hardware usage by minimizing redundant computations during fault propagation and leveraging bit-level parallel processing capabilities. Techniques like parallel-pattern and event-driven simulations have improved performance in combinational circuits but face limitations in sequential circuits due to timing dependencies within loops. The challenge lies in parallelizing simulations across different cycles without violating these dependencies, which is exacerbated by the complex feedback structures in SCCs. In this work, we propose a novel parallel-pattern fault simulation framework that combines loop fusion with efficient event traversal to accelerate sequential circuit simulations. By compiling simple loops into larger nodes, we reduce the number of feedback events without introducing excessive redundancy. For larger SCCs, we develop specialized algorithms for selecting loop entrance nodes based on indegree analysis and implement the lazy propagation strategy for internal nodes. This approach minimizes simulation events caused by inaccurate predictions and reduces overhead associated with false event propagation. We integrate these techniques into our simulation framework, EPICS, which strategically mixes compiled and event-driven simulations to optimize performance. Experimental results demonstrate that EPICS achieves a $5.94 \times$ speedup over state-of-the-art commercial tool while maintaining the same fault coverage.
Hui Wang 0152, Jianan Mu, Yihan Wen, Zizhen Liu, Shengwen Liang, Jing Ye 0001, Xiaowei Li 0001, Huawei Li 0001
DAC6
2025 VIRTUAL: Vector-based Dynamic Power Estimation via Decoupled Multi-Modality Learning
abstract
Dynamic power analysis in digital integrated circuits (ICs) conventionally relies on gate-level synthesis and simulation, creating a critical bottleneck in iterative design flows. We propose VIRTUAL, a multi-modality learning framework for rapid post-synthesis dynamic power estimation directly from Register-Transfer Level (RTL) implementations and input waveform vectors, eliminating the need for gate-level synthesis and extensive simulation. By decoupling features from input port waveforms and RTL implementations, VIRTUAL employs a transformer-based encoder to extract temporal patterns from input port waveforms and a graph neural network (GNN) to capture structural and functional dependencies within RTL implementations. Through self-supervised contrastive learning across sequential and graph modalities, the framework learns robust power-relevant representations with minimal labeled data. Subsequently, VIRTUAL refines the multi-modality embeddings using a lightweight fusion module and a power prediction head, enabling dynamic power estimation for fixed clock periods within seconds or minutes. Experimental evaluations demonstrate approximately a Pearson correlation coefficient (PCC) of 0.842 and a mean absolute percentage error (MAPE) of 23.43%, while achieving 14.27×speedup compared to traditional gate-level power analysis workflows. Experimental results across diverse RTL designs and input port waveforms validate that our proposed learning-based approach maintains the accuracy while significantly reducing design iteration time, transforming hours of synthesis and simulation into minutes of direct prediction.
Yuntao Lu, Yihan Wen, Jianan Mu, Huawei Li 0001, Bei Yu 0001
ICCAD3
2025 Bridging Layout and RTL: Knowledge Distillation based Timing Prediction
abstract
Accurate and efficient timing prediction at the register-transfer level (RTL) remains a fundamental challenge in electronic design automation (EDA), particularly in striking a balance between accuracy and computational efficiency. While static timing analysis (STA) provides high-fidelity results through comprehensive physical parameters, its computational overhead makes it impractical for rapid design iterations. Conversely, existing RTL-level approaches sacrifice accuracy due to the limited physical information available. We propose RTLDistil, a novel cross-stage knowledge distillation framework that bridges this gap by transferring precise physical characteristics from a layout-aware teacher model (Teacher GNN) to an efficient RTL-level student model (Student GNN), both implemented as graph neural networks (GNNs). RTLDistil efficiently predicts key timing metrics, such as arrival time (AT), and employs a multi-granularity distillation strategy that captures timing-critical features at node, subgraph, and global levels. Experimental results demonstrate that RTLDistil achieves significant improvement in RTL-level timing prediction error reduction, compared to state-of-the-art prediction models. This framework enables accurate early-stage timing prediction, advancing EDA’s “left-shift” paradigm while maintaining computational efficiency. Our code and dataset will be publicly available at https://github.com/sklp-eda-lab/RTLDistil.
Yihan Wen, Jianan Mu, Jing Ye 0001, Bei Yu 0001, Huawei Li 0001
ICML2
2025 SAW: Semantic-Aware WebRTC Transmission Using Diffusion-Based Scalable Video Coding
abstract
As video transmission systems expand into various complex scenarios, real-time video coding methods are essential for maintaining low latency and high perceptual quality across varying network conditions. In this work, we propose service-aware Web real-time communication (WebRTC), a semantic-assisted WebRTC system built on scalable video coding (SVC). Specifically, this system is structured with three layers: 1)$\mathcal {L}_{1}$extracts and down-samples semantic information at the encoder, employing a novel super-resolution (SR) method named BUS-DDIM at the decoder to enhance the transmission efficiency and machine vision recognition rate; 2)$\mathcal {L}_{2}$adaptively compresses high-quality video by discarding frames with little motion at the encoder to address latency issues under poor network conditions, and utilize the adjacent frame-guided denoised interpolation model called the adjacent frame-guided denoised diffusion implicit model for restoring the video; and 3)$\mathcal {L}_{3}$transmits high-quality video tailored for users with high-definition video requirements and favorable network conditions. These layers dynamically enhance the visual experience and ensure low latency across various network environments. Experiments are conducted on diverse videos to validate the effectiveness of the proposed framework. The performance evaluation under real-time scenarios indicates significant enhancements in video quality and transmission efficiency, showcasing compatibility and versatility across various applications.
Yihan Wen, Jinglei Li, Chung Shue Chen, Guanchong Niu
IEEE Internet Things J.1
2024 Semantic-Based Motion Detection Method for Unmanned Aerial Vehicle Data Transmission
abstract
Unmanned Aerial Vehicles (UAVs) are crucial for wireless network transmissions, particularly in challenging en-vironments of regular inspection. However, transmitting high-resolution video data from UAV s poses challenges due to limited resources and significant data volumes. Traditional video compression methods, removing redundant information with a single frame, suffer from quality loss as compression rates increase. To address these issues, we propose a novel framework, namely Semantic-based Motion Detection Compression (SMDC) to perform the video compression with high-quality resolution. The proposed framework incorporates Generative Diffusion Change Detection (TransC-GD-CD), a robust semantic-based change detection method, to accurately detect motion between adjacent video frames. Specifically, frames with slight motion are eliminated, thereby reducing network bandwidth requirements for the UAV inspection. Furthermore, a neural network-based interpolation technique is integrated to restore the information loss and ensure smooth playback. Experimental results show that SMDC outperforms traditional compression methods based on H.264, achieving higher video quality at matched bitrates. The exceptional performance of SMDC promises its potential as an effective solution for high-resolution video transmission in scenarios with limited bandwidth.
Yihan Wen, Feixiang Liu, Qi Cao 0001, Guanchong Niu
ICC1
2024 Peak Power and Dynamic IR-drop Assessment via Waveform Augmenting
abstract
Pre-silicon power and IR-drop estimation are crucial parts of the chip design process. Vector-based and vectorless assessments are commonly employed to estimate the worst peak power and dynamic IR-drop of the design. However, with rapid growth in chip scale and complexity, the waveform-driven vector-based assessment encounters the coverage challenge due to the difficulty in generating test waveforms that encompass all potential worst-case scenarios. Additionally, Vectorless assessments consistently yield overly pessimistic estimations which may lead to significant overdesign. This paper proposes a semi-vector-based assessment flow aimed at offering a more reasonable estimation of worst-case peak power and IR-drop. In the proposed assessment, functionally independent modules are identified through the analysis of module toggle activity correlation (MTAC) on the existing waveform. By making these functionally independent modules toggle simultaneously, an augmented waveform approximating the worst peak power and dynamic IR-drop scenario is actively generated while preserving a similar MTAC to the existing waveform. By applying dynamic power and IR-drop analysis to the augmented waveform, previously unaddressed weaknesses are identified. Experimental results on an industrial design indicate that the proposed worst-case assessment result is 3× and 2.5× more accurate than the vector-based result for worst peak power and dynamic IR-drop. Similarly, it is 33× and 6× more accurate than the vectorless result.
Yihan Wen, Bei Yu 0001
ICCAD1
2024 Enhanced Facial Restoration with Misinformation-Filtered Guide-Denoising Diffusion Probabilistic Models
abstract
Most of the existing generation models encounter notable challenges in complex scenes, particularly with inaccuracies in facial organs and textures that do not align with actual conditions. Traditional face restoration methods, heavily dependent on facial geometry and reference priors, often generate incorrect facial images that contribute misleading prior information. In this study, a Misinformation-Flitered GuideDenoising Diffusion Probabilistic Models (MF-GDDPM) is proposed to address these issue. Specifically, MF-GDDPM employs low-pass filtering to remove high-frequency details that contain misleading prior information. This process results in filtered low-dimensional facial contours that guide the diffusion model in generating high-quality facial images. To further enhance the fidelity of the generated results, a dualstream encoder within the Denoising Unet is constructed to process facial contours and high-dimensional details separately, while the Attention Feature Fusion (AFF) attention mechanism ensures the fidelity of image restoration. We have also incorporated the Natural Image Quality Evaluator (NIQE), a deep learning-based image quality assessment tool, into our framework as a novel loss function to crucially ensure the naturalness of restored images. Overall, the proposed method marks a significant improvement in generating accurate and clear facial images using diffusion models.
Wendi Liang, Yihan Wen, Jianuo Jiang, Tat-Ming Lok, Guanchong Niu
ICIP2
2024 KDSMALL: A lightweight small object detection algorithm based on knowledge distillation
Xiaodon Wang, Yusheng Fan, Yishuai Yang, Yihan Wen, Zhengyuan Lin, Langlang Chen, Shizhou Yao, Liu Zequn, Jianqing Wang
Comput. Commun.5
2024 GCD-DDPM: A Generative Change Detection Model Based on Difference-Feature-Guided DDPM
abstract
Deep learning (DL)-based methods have recently shown great promise in bitemporal change detection (CD). Existing discriminative methods based on convolutional neural networks (CNNs) and Transformers rely on discriminative representation learning for change recognition while struggling with exploring local and long-range contextual dependencies. As a result, it is still challenging to obtain fine-grained and robust CD maps in diverse ground scenes. To cope with this challenge, this work proposes a generative CD model called GCD-DDPM to directly generate CD maps by exploiting the denoising diffusion probabilistic model (DDPM), instead of classifying each pixel into changed or unchanged categories. Furthermore, the difference conditional encoder (DCE), is designed to guide the generation of CD maps by exploiting multilevel difference features. Leveraging the variational inference (VI) procedure, GCD-DDPM can adaptively recalibrate the CD results through an iterative inference process, while accurately distinguishing subtle and irregular changes in diverse scenes. Finally, a noise suppression-based semantic enhancer (NSSE) is specifically designed to mitigate noise in the current step’s change-aware feature representations from the CD Encoder. This refinement, serving as an attention map, can guide subsequent iterations while enhancing CD accuracy. Extensive experiments on four high-resolution CD datasets (CDD) confirm the superior performance of the proposed GCD-DDPM. The code for this work will be available athttps://github.com/udrs/GCD.
Yihan Wen, Xianping Ma, Man-On Pun
IEEE Trans. Geosci. Remote. Sens.1
2023 Risk Propagation Based Vector Profiling for High Coverage Dynamic IR-Drop Analysis
abstract
Vector-based dynamic IR-drop analysis is a crucial aspect for enhancing yield in chip fabrication since it provides accurate IR-drop simulation with real waveform. To evaluate waveforms with a large duration from numerous working scenarios, vector profiling is widely used to increase scalability. In real cases, only a few windows selected by vector profiling are assessed by dynamic IR-drop analysis, rather than the whole waveform. Therefore, the coverage of vector profiling methods becomes a major concern, especially in EUV process node. The IR-drop locality effect on multi-pattern layers makes traditional vector profiling methods less robust. The real worst-case waveform window which may lead to silicon failure is frequently missed, which ultimately impacts the coverage of profiling. This paper proposes a novel risk propagation-based vector profiling method that achieves better estimation of IR-drop risk by considering the locality through examining not only the self-power-induced IR-drop but also the drop propagated from surrounding regions. The experimental results have shown that the proposed vector profiling achieved 4.3 times greater probability of covering the worst IR-drop window compared to traditional profiling. The proposed profiling also discovered additional IR-drop risky regions which were missed by traditional profiling.
Yihan Wen
ICCAD1