Xudong Lu 0004

dblp:15/3008-4 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0008-9082-4084ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 FabThink: A Wafer Analysis Multimodal LLM via Chain-of-Thought-Driven Retrieval Augmentation
abstract
Retrieval-Augmented Generation (RAG) incorporates external knowledge to support Large Language Models (LLMs) in generating more accurate, fact-based answers. However, standard RAG methods applied in LLMs lack adaptation to specific domains, limiting their effectiveness in handling complex and specialized knowledge in wafer manufacturing, such as defect root cause analysis, which leads to lower retrieval accuracy and increased large model hallucinations. We propose a wafer-domain-tailored Multimodal Large Language Model (MLLM), FabThink, which aims to optimize the RAG process through a unique multimodal Chain-of-Thought (CoT) framework to address the above issues. Specifically, we propose a "logical decomposition, cross-modal integration, multi-turn retrieval" strategy to refine the process of solving complex queries and enhance the precision of document retrieval. In addition, we introduce an adaptive weighted ranking for critical document selection and fine-tune a text generator to answer wafer-related questions. Experimental results on fab data show that FabThink excels in detection, retrieval, and generation tasks, strongly supporting defect analysis in the integrated circuits (IC) domain.
Xudong Lu 0004, Jinyuan Deng, Hao Geng, Hanming Wu, Qi Sun 0002, Cheng Zhuo
ICCAD3
2024 FabGPT: An Efficient Large Multimodal Model for Complex Wafer Defect Knowledge Queries
abstract
Intelligence is key to advancing integrated circuit (IC) fabrication. Recent breakthroughs in Large Multimodal Models (LMMs) have unlocked extraditionary abilities in understanding images and text, fostering intelligent fabrication. Leveraging the power of LMMs, we introduce FabGPT, a customized IC fabrication large multimodal model for wafer defect knowledge query. FabGPT manifests expertise in conducting defect detection in Scanning Electron Microscope (SEM) images, performing root cause analysis, and providing expert Q&A on fabrication processes. FabGPT matches enhanced multimodal features to automatically detect minute defects under complex wafer backgrounds and reduce the subjectivity of manual threshold settings. Besides, the proposed modulation module and interactive corpus training strategy embed wafer defect knowledge into the pre-trained model, effectively balancing Q&A queries related to defect knowledge and original knowledge and mitigating the modality bias issues. Experiments on in-house fab data show that FabGPT achieves significant performance improvement in wafer defect detection and knowledge querying.
Xudong Lu 0004, Qi Sun 0002, Hanming Wu, Cheng Zhuo
ICCAD2
2024 SEM-CLIP: Precise Few-Shot Learning for Nanoscale Defect Detection in Scanning Electron Microscope Image
abstract
In the field of integrated circuit manufacturing, the detection and classification of nanoscale wafer defects are critical for subsequent root cause analysis and yield enhancement. The complex background patterns observed in scanning electron microscope (SEM) images and the diverse textures of the defects pose significant challenges. Traditional methods usually suffer from insufficient data, labels, and poor transferability. In this paper, we propose a novel few-shot learning approach, SEM-CLIP, for accurate defect classification and segmentation. SEM-CLIP customizes the Contrastive Language-Image Pretraining (CLIP) model to better focus on defect areas and minimize background distractions, thereby enhancing segmentation accuracy. We employ text prompts enriched with domain knowledge as prior information to assist in precise analysis. Additionally, our approach incorporates feature engineering with textual guidance to categorize defects more effectively. SEM-CLIP requires little annotated data, substantially reducing labor demands in the semiconductor industry. Extensive experimental validation demonstrates that our model achieves impressive classification and segmentation results under few-shot learning scenarios.
Xudong Lu 0004, Yining Chen 0001, Qi Sun 0002, Cheng Zhuo
ICCAD3
2024 DCAFuse: Dual-Branch Diffusion-CNN Complementary Feature Aggregation Network for Multi-Modality Image Fusion
abstract
Multi-modality image fusion (MMIF) aims to integrate the complementary features of source images into the fused image, including target saliency and texture specifics. Recently, image fusion methods leveraging diffusion models have demonstrated commendable results. Despite their strengths, diffusion models reduce the capability to perceive local features. Additionally, their inherent working mechanism, introducing noise to the inputs, consequently leads to a loss of original information. To overcome this problem, we propose a novel Diffusion-CNN feature Aggregation Fusion (DCAFuse) network that can extract complementary features from the dual branches and aggregate them effectively. Specifically, we utilize the denoising diffusion probabilistic model (DDPM) in the diffusion-based branch to construct global information, and multi-scale convolutional kernels in the CNN-based branch to extract local detailed features. Afterward, we design a novel complementary feature aggregation module (CFAM). By constructing coordinate attention maps for features, CFAM captures long-range dependencies in both horizontal and vertical directions, thereby dynamically guiding the aggregation weights of branches. In addition, to further improve the complementarity of dual-branch features, we introduce a novel loss function based on cosine similarity and a unique denoising timestep selection strategy. Extensive experimental results show that our proposed DCAFuse outperforms other state-of-the-art methods in multiple image fusion tasks, including infrared and visible image fusion (IVF) and medical image fusion (MIF).
Xudong Lu 0004, Haiwen Hong, Qi Sun 0002, Cheng Zhuo
ACM Multimedia1