EDBT 2026 Demo / reviewers in the wild / expert
Jiawen Yu
dblp:258/5034
· DBLP profile ↗
17ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Joint Analysis of Localization and CoMP Transmission Performance in Integrated Sensing and Communication Networks
Muyu Mei, Jiawen Yu, Li Feng 0003, Chunhui Feng, Baoyi Xu, Xu Bao 0001, Mingwu Yao |
WCNC | 2 |
| 2025 | OUS: Bridging Scene Context and Facial Features to Overcome the Rigid Cognitive ProblemabstractDynamic Facial Expression Recognition (DFER) is crucial for affective computing but often overlooks the impact of scene context. We have identified a significant issue in current DFER tasks: human annotators typically integrate emotions from various angles, including environmental cues and body language, whereas existing DFER methods tend to consider the scene as noise that needs to be filtered out, focusing solely on facial information. We refer to this as the Rigid Cognitive Problem. The Rigid Cognitive Problem can lead to discrepancies between the cognition of annotators and models in some samples. To align more closely with the human cognitive paradigm of emotions, we propose an Overall Understanding of the Scene DFER method (OUS). OUS effectively integrates scene and facial features, combining scene-specific emotional knowledge for DFER. Extensive experiments on the two largest datasets in the DFER field, DFEW and FERV39k, demonstrate that OUS significantly outperforms existing methods. By analyzing the Rigid Cognitive Problem, OUS successfully understands the complex relationship between scene context and emotional expression, closely aligning with human emotional understanding in real-world scenarios. Xinji Mai, Haoran Wang 0006, Zeng Tao, Junxiong Lin, Shaoqi Yan, Yan Wang 0068, Jiawen Yu, Xuan Tong |
AAAI | 7 |
| 2025 | D2SP: Dynamic Dual-Stage Purification Framework for Dual Noise Mitigation in Vision-based Affective RecognitionabstractThe current advancements in Dynamic Facial Expression Recognition (DFER) methods mainly focus on better capturing the spatial and temporal features of facial expressions. However, DFER datasets contain a substantial amount of noisy samples, and few have addressed the issue of handling this noise. We identified two types of noise: one is caused by low-quality data resulting from factors such as occlusion, dim lighting, and blurriness; the other arises from mislabeled data due to annotation bias by annotators. Addressing the two types of noise, we have meticulously crafted a Dynamic Dual-Stage Purification (D2SP) Framework. This initiative aims to dynamically purify the DFER datasets of these two types of noise, ensuring that only high-quality and correctly labeled data is used in the training process. To mitigate low-quality samples, we introduce the Coarse-Grained Pruning (CGP) stage, which computes sample weights and prunes those low-weight samples. After CGP, the Fine-Grained Correction (FGC) stage evaluates prediction stability to correct mislabeled data. Moreover, D2SP is conceived as a general, plug-and-play framework, tailored to integrate seamlessly with prevailing DFER methods. Extensive experiments covering prevalent DFER datasets and deploying multiple benchmark methods have substantiated D2SP’s ability to enhance performance metrics. Haoran Wang 0006, Xinji Mai, Zeng Tao, Xuan Tong, Junxiong Lin, Yan Wang 0068, Jiawen Yu, Shaoqi Yan, Ziheng Zhou 0005 |
CVPR | 7 |
| 2025 | HSS-IAD: A Heterogeneous Same-Sort Industrial Anomaly Detection DatasetabstractMulti-class Unsupervised Anomaly Detection algorithms (MUAD) are receiving increasing attention due to their relatively low deployment costs and improved training efficiency. However, the real-world effectiveness of MUAD methods is questioned due to limitations in current Industrial Anomaly Detection (IAD) datasets. These datasets contain numerous classes that are unlikely to be produced by the same factory and fail to cover multiple structures or appearances. Additionally, the defects do not reflect real-world characteristics. Therefore, we introduce the Heterogeneous Same-Sort Industrial Anomaly Detection (HSS-IAD) dataset, which contains 8,580 images of metallic-like industrial parts and precise anomaly annotations. These parts exhibit variations in structure and appearance, with subtle defects that closely resemble the base materials. We also provide foreground images for synthetic anomaly generation. Finally, we evaluate popular IAD methods on this dataset under multi-class and class-separated settings, demonstrating its potential to bridge the gap between existing datasets and real factory conditions. The dataset is available at https://github.com/Qiqigeww/HSS-IAD-Dataset. Qishan Wang 0002, Shuyong Gao, Jiawen Yu, Xuan Tong |
ICME | 4 |
| 2025 | Component-Aware Unsupervised Logical Anomaly Generation for Industrial Anomaly DetectionabstractAnomaly detection is critical in industrial manufacturing for ensuring product quality and improving efficiency in automated processes. The scarcity of anomalous samples limits traditional detection methods, making anomaly generation essential for expanding the data repository. However, recent generative models often produce unrealistic anomalies increasing false positives, or require real-world anomaly samples for training. In this work, we treat anomaly generation as a compositional problem and propose ComGEN, a component-aware and unsupervised framework that addresses the gap in logical anomaly generation. Our method comprises a multi-component learning strategy to disentangle visual components, followed by subsequent generation editing procedures. Disentangled text-to-component pairs, revealing intrinsic logical constraints, conduct attention-guided residual mapping and model training with iteratively matched references across multiple scales. Experiments on the MVTecLOCO dataset confirm the efficacy of ComGEN, achieving the best AUROC score of$\mathbf{9 1. 2 \%}$. Additional experiments on the real-world scenario of Diesel Engine and widelyused MVTecAD dataset demonstrate significant performance improvements when integrating simulated anomalies generated by ComGEN into automated production workflows. Xuan Tong, Yang Chang, Qing Zhao 0007, Jiawen Yu, Boyang Wang 0003, Junxiong Lin, Yuxuan Lin 0001, Xinji Mai, Haoran Wang 0006, Zeng Tao, Yan Wang 0068 |
ICRA | 4 |
| 2025 | Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial EnvironmentsabstractAnomaly detection and localization in automated industrial manufacturing can significantly enhance production efficiency and product quality. Existing methods are capable of detecting surface defects in pre-defined or controlled imaging environments. However, accurately detecting workpiece defects in complex and unstructured industrial environments with varying views, poses and illumination remains challenging. We propose a novel anomaly detection and localization method specifically designed to handle inputs with perturbative patterns. Our approach introduces a new framework based on a collaborative distillation heterogeneous teacher network (HetNet), an adaptive local-global feature fusion module, and a local multivariate Gaussian noise generation module. HetNet can learn to model the complex feature distribution of normal patterns using limited information about local disruptive changes. We conducted extensive experiments on mainstream benchmarks. HetNet demonstrates superior performance with approximately 10% improvement across all evaluation metrics on MSC-AD under industrial conditions, while achieving state-of-the-art results on other datasets, validating its resilience to environmental fluctuations and its capability to enhance the reliability of industrial anomaly detection systems across diverse scenarios. Tests in real-world environments further confirm that HetNet can be effectively integrated into production lines to achieve robust and real-time anomaly detection. Codes, images and videos are published on the project website at: https://zihuatanejoyu.github.io/HetNet/ Jiawen Yu, Jieji Ren, Yang Chang, Qiaojun Yu, Xuan Tong, Boyang Wang 0003, Xinji Mai |
IROS | 1 |
| 2025 | ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich ManipulationabstractVision-Language-Action (VLA) models have advanced general-purpose robotic manipulation by leveraging pretrained visual and linguistic representations. However, they struggle with contact-rich tasks that require fine-grained control involving force, especially under visual occlusion or dynamic uncertainty. To address these limitations, we propose \textbf{ForceVLA}, a novel end-to-end manipulation framework that treats external force sensing as a first-class modality within VLA systems. ForceVLA introduces \textbf{FVLMoE}, a force-aware Mixture-of-Experts fusion module that dynamically integrates pretrained visual-language embeddings with real-time 6-axis force feedback during action decoding. This enables context-aware routing across modality-specific experts, enhancing the robot's ability to adapt to subtle contact dynamics. We also introduce \textbf{ForceVLA-Data}, a new dataset comprising synchronized vision, proprioception, and force-torque signals across five contact-rich manipulation tasks. ForceVLA improves average task success by 23.2\% over strong $\pi_0$-based baselines, achieving up to 80\% success in tasks such as plug insertion. Our approach highlights the importance of multimodal integration for dexterous manipulation and sets a new benchmark for physically intelligent robotic control. Code and data will be released at https://sites.google.com/view/forcevla2025/. Jiawen Yu, Hairuo Liu, Qiaojun Yu, Jieji Ren, Ce Hao, Haitong Ding, Guangyu Huang, Guofan Huang, Panpan Cai, Cewu Lu |
NeurIPS | 1 |
| 2025 | Joint Sensing-Communication Performance Analysis of ISAC-Enabled VCNabstractIntegrated sensing and communication (ISAC) is emerging as a key technology and research focus for future vehicular communication networks (VCN). It achieves efficient reuse of wireless infrastructure and spectrum resources through the collaborative design of sensing and communication functionalities. However, this integration leads to inevitable mutual interference, and the complexity of channel conditions further complicates the coordination of these functionalities. This paper primarily focuses on the joint sensing-communication performance analysis of ISAC-enabled VCN. Specifically, we model the spatial distribution of roads using a Poisson line process, while the locations of vehicles and roadside units (RSUs) are represented by two one-dimensional Poisson point processes. We characterize the dynamic interference distribution caused by RSUs during the sensing and communication phases and calculate the probability of successful perception (PSP) for a typical pair. Furthermore, for this typical pair, we meticulously derive the communication coverage probability based on the derived PSP for such a pair. To provide a detailed analysis of the interaction between communication and sensing functionalities, we evaluate their trade-off relationship and derive the joint probability of ISAC coverage. Moreover, we perform comprehensive simulations to verify the theoretical results. Additionally, the numerical results demonstrate how different parameters impact the network performance, providing guidance for network deployment and resource allocation under certain performance requirements. Jiawen Yu, Muyu Mei, Li Feng 0003, Xu Bao 0001, Lijuan Xu 0002, Baoyi Xu, Mingwu Yao |
IEEE Trans. Commun. | 1 |
| 2024 | Adaptive Multi-modal Fusion of Spatially Variant Kernel Refinement with Diffusion Model for Blind Image Super-Resolution
Junxiong Lin, Yan Wang 0068, Zeng Tao, Boyang Wang 0003, Qing Zhao 0007, Haorang Wang, Xuan Tong, Xinji Mai, Yuxuan Lin 0001, Wei Song 0007, Jiawen Yu, Shaoqi Yan |
ECCV (52) | 11 |
| 2024 | Suppressing Uncertainties in Degradation Estimation for Blind Super-Resolution
Junxiong Lin, Zen Tao, Xuan Tong, Xinji Mai, Haoran Wang 0006, Boyang Wang 0003, Yan Wang 0068, Qing Zhao 0007, Jiawen Yu, Yuxuan Lin 0001, Shaoqi Yan, Shuyong Gao |
ACM Multimedia | 9 |
| 2024 | All rivers run into the sea: Unified Modality Brain-Inspired Emotional Central MechanismabstractIn the field of affective computing, fully leveraging information from a variety of sensory modalities is essential for the comprehensive understanding and processing of human emotions. Inspired by the process through which the human brain handles emotions and the theory of cross-modal plasticity, we propose UMBEnet, a brain-like unified modal affective processing network. The primary design of UMBEnet includes a Dual-Stream (DS) structure that fuses inherent prompts with a Prompt Pool and a Sparse Feature Fusion (SFF) module. The design of the Prompt Pool is aimed at integrating information from different modalities, while inherent prompts are intended to enhance the system's predictive guidance capabilities and effectively manage knowledge related to emotion classification. Moreover, considering the sparsity of effective information across different modalities, the SSF module aims to make full use of all available sensory data through the sparse integration of modality fusion prompts and inherent prompts, maintaining high adaptability and sensitivity to complex emotional states. Extensive experiments on the largest benchmark datasets in the Dynamic Facial Expression Recognition (DFER) field, including DFEW, FERV39k, and MAFW, have proven that UMBEnet consistently outperforms the current state-of-the-art methods. Notably, in scenarios of Modality Missingness and multimodal contexts, UMBEnet significantly surpasses the leading current methods, demonstrating outstanding performance and adaptability in tasks that involve complex emotional understanding with rich multimodal information. Code can be obtained at https://github.com/Xinji-Mai/UMBEnet. Xinji Mai, Junxiong Lin, Haoran Wang 0006, Zeng Tao, Yan Wang 0068, Shaoqi Yan, Xuan Tong, Jiawen Yu, Boyang Wang 0003, Ziheng Zhou 0005, Qing Zhao 0007, Shuyong Gao |
ACM Multimedia | 8 |
| 2024 | MSC-AD: A Multiscene Unsupervised Anomaly Detection Dataset for Small Defect Detection of Casting SurfaceabstractIntelligent detection of product surface defects in the industrial scene is the key to ensuring product quality. On general benchmarks, current unsupervised anomaly detection techniques have achieved significant success. When used in complex industrial environments (e.g., large industrial components with small defects), the model needs to be able to adapt to different imaging scenarios (e.g., illumination and resolution) and accurately detect and localize anomalies, but its performance is still far from satisfactory. Besides, the complex and unstable optical lighting environment for collecting such data poses major challenges in establishing unified benchmarks for optical lighting and imaging resolution in defect detection. To fill this gap, we build a standard imaging system-based multiscene unsupervised anomaly detection dataset, coined as MSC-AD. In particular, it provides 12 imaging scenes, i.e., a cross combination of low-to-high three illuminations and 150 × 150 to 600 × 600 four resolutions, in which six types of large casting surfaces with different structures include five kinds of small defects with sample-level and pixel-level precise ground truth. We systematically investigate representative baseline methods and empirical analysis on this dataset to obtain a number of interesting findings, e.g., how to detach from distinctly different imaging scenes, and how to distinguish between subtly normal–anomaly classes. To the best of our knowledge, MSC-AD is the first multi-illumination, multiresolution, multisurface, and multidefect dataset built in a standard imaging system. Qing Zhao 0007, Yan Wang 0068, Boyang Wang 0003, Junxiong Lin, Shaoqi Yan, Wei Song 0007, Antonio Liotta, Jiawen Yu, Shuyong Gao |
IEEE Trans. Ind. Informatics | 8 |
| 2024 | Cluster-based two-branch framework for point cloud attribute compression
Longhua Sun, Jin Wang 0023, Qing Zhu 0004, Jiaying Liu 0015, Jiawen Yu |
Vis. Comput. | 5 |
| 2022 | Blockchain-based Crowd-sensing Trust Management Mechanism for Crowd EvacuationabstractWith the continuous urban expansion, the safety of public places has attracted more and more attention. Once an accident occurs and people distrust their surroundings, it will cause congestion, trampling and other events. Individual trust value is particularly important in crowd evacuation. However, it is very difficult to evaluate the individual credibility with existing methods. We propose a blockchain-based crowd-sensing trust management mechanism (B-CSTM) to address this problem. First, we provide a message credibility evaluation mechanism that can evaluate the credibility of messages sent by individuals on spot. Second, we developed a trust calculation method that uses a distributed Evacuation Perception Unit (EPU) to calculate individuals’ trust values by perception scores. Finally, we build a blockchain-based trust management mechanism model that uses blockchain to store and query trust values. Through experimental analysis, we visualize the trust results in the evacuation scenario. It is shown that the mechanism is feasible for the management of trust values (collection, calculation and storage, etc.) in crowd evacuation. Jiawen Yu, Guijuan Zhang, Dianjie Lu, Hong Liu 0013 |
CSCWD | 1 |
| 2022 | Traffic Sign Recognition Using Ulam's GameabstractIn this paper, we propose a conditional early exiting framework with Ulam's Game for traffic sign recognition. Since the traffic sign recognition system has extremely high requirements on dynamic performance, we pays more attention to improving the detection efficiency, hoping to obtain results in a shorter time. In our system, we use a modified ResNet-50 as backbone network to do feature extraction and use a Pooling module to accumulate feature. Then, we have a Gate module to determine whether the feature have accumulated enough to begin Ulam's Game. A classifier is used to get candidate results, which are used to run Ulam's Game and get the final prediction. The model shows good detection accuracy and dynamic performance in multiple data sets (Mini-Kinetics, ActivityNet, Lisa). Haofeng Zheng, Ruikang Luo, Yaofeng Song, Jiawen Yu |
ICARCV | 5 |
| 2022 | An Improved Topology Prediction of Alpha-Helical Transmembrane Protein Based on Deep Multi-Scale Convolutional Neural NetworkabstractAlpha-helical proteins ( αTMPs) are essential in various biological processes. Despite their tertiary structures are crucial for revealing complex functions, experimental structure determination remains challenging and costly. In the past decades, various sequence-based topology prediction methods have been developed to bridge the gap between the sequences and structures by characterizing the structural features, but significant improvements are still required. Deep learning brings a great opportunity for its powerful representation learning capability from limited original data. In this work, we improved our αTMP topology prediction method DMCTOP using deep learning, which composed of two deep convolutional blocks to simultaneously extract local and global contextual features. Consequently, the inputs were simplified to reflect the original features of the sequence, including a protein sequence feature and an evolutionary conservation feature. DMCTOP can efficiently and accurately identify all topological types and the N-terminal orientation for an αTMP sequence. To validate the effectiveness of our method, we benchmarked DMCTOP against 13 peer methods according to the whole sequence, the transmembrane segment and the traditional criterion in testing experiments. All the results reveal that our method achieved the highest prediction accuracy and outperformed all the previous methods. The method is available at https://icdtools.nenu.edu.cn/dmctop. Jiawen Yu, Zhe Liu 0030, Han Wang 0028, Zhiqiang Ma 0003, Dong Xu 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2019 | DMCTOP: Topology Prediction of Alpha-Helical Transmembrane Protein Based on Deep Multi-Scale Convolutional Neural NetworkabstractAlpha-helical transmembrane proteins ($\alpha \text{TMPs}$) belong to an important category of integral membranes. Their structures are highly valuable in relevant research, but costly to solve experimentally. Sequence-based topology prediction provides a practical computational approach to characterize the structure features. Although much progress had been made in the past decade, there is significant room for improvement in predicting the topology structure. Deep learning brings a great opportunity for its capability of mining new features from data. In this work, we propose a novel$\alpha \text{TMP}$topology prediction method DMCTOP using a Deep Multi-Scale Convolutional Neural Network (DMCNN), which composes of two deep convolutional blocks to extract local and global contextual features. Consequently, the inputs of DMCTOP is simplified to a protein sequence feature and an evolutionary conservation feature. DMCTOP can efficiently and accurately identify all topological types and the N-terminal orientation for an$\alpha \text{TMP}$sequence. In the testing experiments, the prediction accuracy was calculated according to the whole sequence, the transmembrane segments and the traditional criterion. Our state-of-the-art method achieved the highest prediction accuracy compared to all the previous methods. The standalone tool is available at https://github.com/NENUBioCompute/DMCTOP. Han Wang 0028, Jiawen Yu, Dong Xu 0002 |
BIBM | 3 |