VLDB 2026 Research / reviewers in the wild / expert
Boyang Wang 0003
dblp:05/11538-3
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2025
0009-0002-2885-2814ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Component-Aware Unsupervised Logical Anomaly Generation for Industrial Anomaly DetectionabstractAnomaly detection is critical in industrial manufacturing for ensuring product quality and improving efficiency in automated processes. The scarcity of anomalous samples limits traditional detection methods, making anomaly generation essential for expanding the data repository. However, recent generative models often produce unrealistic anomalies increasing false positives, or require real-world anomaly samples for training. In this work, we treat anomaly generation as a compositional problem and propose ComGEN, a component-aware and unsupervised framework that addresses the gap in logical anomaly generation. Our method comprises a multi-component learning strategy to disentangle visual components, followed by subsequent generation editing procedures. Disentangled text-to-component pairs, revealing intrinsic logical constraints, conduct attention-guided residual mapping and model training with iteratively matched references across multiple scales. Experiments on the MVTecLOCO dataset confirm the efficacy of ComGEN, achieving the best AUROC score of$\mathbf{9 1. 2 \%}$. Additional experiments on the real-world scenario of Diesel Engine and widelyused MVTecAD dataset demonstrate significant performance improvements when integrating simulated anomalies generated by ComGEN into automated production workflows. Xuan Tong, Yang Chang, Qing Zhao 0007, Jiawen Yu, Boyang Wang 0003, Junxiong Lin, Yuxuan Lin 0001, Xinji Mai, Haoran Wang 0006, Zeng Tao, Yan Wang 0068 |
ICRA | 5 |
| 2025 | Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial EnvironmentsabstractAnomaly detection and localization in automated industrial manufacturing can significantly enhance production efficiency and product quality. Existing methods are capable of detecting surface defects in pre-defined or controlled imaging environments. However, accurately detecting workpiece defects in complex and unstructured industrial environments with varying views, poses and illumination remains challenging. We propose a novel anomaly detection and localization method specifically designed to handle inputs with perturbative patterns. Our approach introduces a new framework based on a collaborative distillation heterogeneous teacher network (HetNet), an adaptive local-global feature fusion module, and a local multivariate Gaussian noise generation module. HetNet can learn to model the complex feature distribution of normal patterns using limited information about local disruptive changes. We conducted extensive experiments on mainstream benchmarks. HetNet demonstrates superior performance with approximately 10% improvement across all evaluation metrics on MSC-AD under industrial conditions, while achieving state-of-the-art results on other datasets, validating its resilience to environmental fluctuations and its capability to enhance the reliability of industrial anomaly detection systems across diverse scenarios. Tests in real-world environments further confirm that HetNet can be effectively integrated into production lines to achieve robust and real-time anomaly detection. Codes, images and videos are published on the project website at: https://zihuatanejoyu.github.io/HetNet/ Jiawen Yu, Jieji Ren, Yang Chang, Qiaojun Yu, Xuan Tong, Boyang Wang 0003, Xinji Mai |
IROS | 6 |
| 2025 | Observe finer to select better: Learning key frame extraction via semantic coherence for dynamic facial expression recognition in the wild
Shaoqi Yan, Yan Wang 0068, Xinji Mai, Zeng Tao, Wei Song 0007, Qing Zhao 0007, Boyang Wang 0003, Haoran Wang 0006, Shuyong Gao |
Inf. Sci. | 7 |
| 2024 | Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution
Junxiong Lin, Yan Wang 0068, Zeng Tao, Boyang Wang 0003, Qing Zhao 0007, Haorang Wang, Xuan Tong, Xinji Mai, Yuxuan Lin 0001, Wei Song 0007, Jiawen Yu, Shaoqi Yan |
ECCV (52) | 4 |
| 2024 | FD-UAD: Unsupervised Anomaly Detection Platform Based on Defect Autonomous Imaging and Enhancement
Yang Chang, Yuxuan Lin 0001, Boyang Wang 0003, Qing Zhao 0007, Yan Wang 0068 |
IJCAI | 3 |
| 2024 | Suppressing Uncertainties in Degradation Estimation for Blind Super-Resolution
Junxiong Lin, Zen Tao, Xuan Tong, Xinji Mai, Haoran Wang 0006, Boyang Wang 0003, Yan Wang 0068, Qing Zhao 0007, Jiawen Yu, Yuxuan Lin 0001, Shaoqi Yan, Shuyong Gao |
ACM Multimedia | 6 |
| 2024 | All rivers run into the sea: Unified Modality Brain-Inspired Emotional Central MechanismabstractIn the field of affective computing, fully leveraging information from a variety of sensory modalities is essential for the comprehensive understanding and processing of human emotions. Inspired by the process through which the human brain handles emotions and the theory of cross-modal plasticity, we propose UMBEnet, a brain-like unified modal affective processing network. The primary design of UMBEnet includes a Dual-Stream (DS) structure that fuses inherent prompts with a Prompt Pool and a Sparse Feature Fusion (SFF) module. The design of the Prompt Pool is aimed at integrating information from different modalities, while inherent prompts are intended to enhance the system's predictive guidance capabilities and effectively manage knowledge related to emotion classification. Moreover, considering the sparsity of effective information across different modalities, the SSF module aims to make full use of all available sensory data through the sparse integration of modality fusion prompts and inherent prompts, maintaining high adaptability and sensitivity to complex emotional states. Extensive experiments on the largest benchmark datasets in the Dynamic Facial Expression Recognition (DFER) field, including DFEW, FERV39k, and MAFW, have proven that UMBEnet consistently outperforms the current state-of-the-art methods. Notably, in scenarios of Modality Missingness and multimodal contexts, UMBEnet significantly surpasses the leading current methods, demonstrating outstanding performance and adaptability in tasks that involve complex emotional understanding with rich multimodal information. Code can be obtained at https://github.com/Xinji-Mai/UMBEnet. Xinji Mai, Junxiong Lin, Haoran Wang 0006, Zeng Tao, Yan Wang 0068, Shaoqi Yan, Xuan Tong, Jiawen Yu, Boyang Wang 0003, Ziheng Zhou 0005, Qing Zhao 0007, Shuyong Gao |
ACM Multimedia | 9 |
| 2024 | Mixed noise-guided mutual constraint framework for unsupervised anomaly detection in smart industries
Qing Zhao 0007, Yan Wang 0068, Yuxuan Lin 0001, Shaoqi Yan, Wei Song 0007, Boyang Wang 0003, Yang Chang, Lizhe Qi |
Comput. Commun. | 6 |
| 2024 | MSC-AD: A Multiscene Unsupervised Anomaly Detection Dataset for Small Defect Detection of Casting SurfaceabstractIntelligent detection of product surface defects in the industrial scene is the key to ensuring product quality. On general benchmarks, current unsupervised anomaly detection techniques have achieved significant success. When used in complex industrial environments (e.g., large industrial components with small defects), the model needs to be able to adapt to different imaging scenarios (e.g., illumination and resolution) and accurately detect and localize anomalies, but its performance is still far from satisfactory. Besides, the complex and unstable optical lighting environment for collecting such data poses major challenges in establishing unified benchmarks for optical lighting and imaging resolution in defect detection. To fill this gap, we build a standard imaging system-based multiscene unsupervised anomaly detection dataset, coined as MSC-AD. In particular, it provides 12 imaging scenes, i.e., a cross combination of low-to-high three illuminations and 150 × 150 to 600 × 600 four resolutions, in which six types of large casting surfaces with different structures include five kinds of small defects with sample-level and pixel-level precise ground truth. We systematically investigate representative baseline methods and empirical analysis on this dataset to obtain a number of interesting findings, e.g., how to detach from distinctly different imaging scenes, and how to distinguish between subtly normal–anomaly classes. To the best of our knowledge, MSC-AD is the first multi-illumination, multiresolution, multisurface, and multidefect dataset built in a standard imaging system. Qing Zhao 0007, Yan Wang 0068, Boyang Wang 0003, Junxiong Lin, Shaoqi Yan, Wei Song 0007, Antonio Liotta, Jiawen Yu, Shuyong Gao |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Freq-HD: An Interpretable Frequency-based High-Dynamics Affective Clip Selection Method for in-the-Wild Facial Expression Recognition in VideosabstractThe in-the-wild dynamic facial expression recognition (DFER) has been challenging due to several high-dynamics factors such as limited dynamic expression-related frames and variable non-expression noise in facial expression sequences. To provide more expression-related clips for DFER models, we propose a novel and interpretable frequency-based method (Freq-HD) for high-dynamics affective clip selection. It can select clips containing pure expression changes from sequences and aid different DFER network structures in recognizing in-the-wild dynamic facial expressions more accurately and efficiently. We first design a novel spatial-temporal frequency analysis (STFA) module to compute the dynamics values of each clip by using sliding windows and spatial-temporal frequency analysis. Moreover, we propose a multi-band complementary selection (MBC) module to amend the inappropriate reaction of the dynamics values of different spatial frequency bands in STFA when expression-irrelevant noise occurs. Specifically, the MBC uses an ingenious mapping method to generate the inhibitory factors to complement and separate the dynamics of expressions and non-expressions in different frequency bands. The Freq-HD can select the most expression-correlated clips and the consisting frames, which could be incorporated into any existing DFER models. We extensively evaluate the Freq-HD on two in-the-wild datasets and four DFER baselines, showing that our method significantly improves the subsequent network performance while using fewer input frames and reducing computation cost. More ablation studies and visualization analysis provide further empirical evidence of the effectiveness of our method. Zeng Tao, Yan Wang 0068, Zhaoyu Chen 0001, Boyang Wang 0003, Shaoqi Yan, Kaixun Jiang, Shuyong Gao |
ACM Multimedia | 4 |
| 2023 | A Capture to Registration Framework for Realistic Image Super-Resolution in the Industry EnvironmentabstractThe acquisition and processing of visual data in industrial environments are of paramount importance. High-resolution (HR) images offer superior clarity and richer textural detail compared to low-resolution (LR) images. On the one hand, owing to the incorporation of richer information, HR images demonstrate substantially enhanced performance compared to LR images in downstream applications, such as anomaly detection. On the other hand, they provide valuable insights to designers and quality inspectors who require a detailed understanding of the images. Currently, the majority of research on super-resolution focuses on natural scenes such as cities and fields, however, the development of datasets for industrial scenes is still in its infancy. To address the image distortion in building realistic LR-HR image pairs in the industry environment, we design a capture to registration framework. It consists of the standard imaging system, physical calibration of the imaging system, as well as the rigid to elastic registration of the LR-HR image pairs. Thus, we build the first realistic industrial sence super-resolution dataset (IndSR), comprises of 50 sets of calibrated images with three scale factors and five typical defects. To benchmark IndSR, we employ quantitative, qualitative, and task-oriented studies to evaluate the representative super-resolution and anomaly detection methods. Besides, we systematically investigate and discuss the performances and results of the existing SISR methods to advance research in the field of super-resolution in industry environment. The IndSR dataset can be available from https://byw4ng.github.io/IndSR/. Boyang Wang 0003, Yan Wang 0068, Qing Zhao 0007, Junxiong Lin, Zeng Tao, Pinxue Guo, Zhaoyu Chen 0001, Kaixun Jiang, Shaoqi Yan, Shuyong Gao |
ACM Multimedia | 1 |