Yanfu Yan

dblp:74/8828 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0008-2475-6802ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution Detection
abstract
Numerous machine learning (ML) models have been developed, including those for software engineering (SE) tasks, under the assumption that training and testing data come from the same distribution. However, training and testing distributions often differ, as training datasets rarely encompass the entire distribution, while testing distribution tends to shift over time. Hence, when confronted with out-of-distribution (OOD) instances that differ from the training data, a reliable and trustworthy SE ML model must be capable of detecting them to either abstain from making predictions, or potentially forward these OODs to appropriate models handling other categories or tasks. In this paper, we develop two types of SE-specific OOD detection models, unsupervised and weakly-supervised OOD detection for code. The unsupervised OOD detection approach is trained solely on in-distribution samples while the weakly-supervised approach utilizes a tiny number of OOD samples to further enhance the detection performance in various OOD scenarios. Extensive experimental results demonstrate that our proposed methods significantly outperform the baselines in detecting OOD samples from four different scenarios simultaneously and also positively impact a main code understanding task.
Yanfu Yan, Viet Duong, Huajie Shao, Denys Poshyvanyk
ICSE1
2024 Semantic GUI Scene Learning and Video Alignment for Detecting Duplicate Video-based Bug Reports
abstract
Video-based bug reports are increasingly being used to document bugs for programs centered around a graphical user interface (GUI). However, developing automated techniques to manage video-based reports is challenging as it requires identifying and understanding often nuanced visual patterns that capture key information about a reported bug. In this paper, we aim to overcome these challenges by advancing the bug report management task of duplicate detection for video-based reports. To this end, we introduce a new approach, called Janus, that adapts the scene-learning capabilities of vision transformers to capture subtle visual and textual patterns that manifest on app UI screens --- which is key to differentiating between similar screens for accurate duplicate report detection. Janus also makes use of a video alignment technique capable of adaptive weighting of video frames to account for typical bug manifestation patterns. In a comprehensive evaluation on a benchmark containing 7,290 duplicate detection tasks derived from 270 video-based bug reports from 90 Android app bugs, the best configuration of our approach achieves an overall mRR/mAP of 89.8%/84.7%, and for the large majority of duplicate detection tasks, outperforms prior work by ≈9% to a statistically significant degree. Finally, we qualitatively illustrate how the scene-learning capabilities provided by Janus benefits its performance.
Yanfu Yan, Nathan Cooper, Oscar Chaparro, Kevin Moran, Denys Poshyvanyk
ICSE1
2023 ACER: An AST-based Call Graph Generator Framework
abstract
We introduce ACER, an AST-based call graph generator framework. ACER leverages tree-sitter to interface with any language. We opted to focus on generators that operate on abstract syntax trees (ASTs) due to their speed and simplicitly in certain scenarios; however, a fully quantified intermediate representation usually provides far better information at the cost of requiring compilation. To evaluate our framework, we created two context-insensitive Java generators and compared them to existing open-source Java generators.Code: https://github.com/WM-SEMERU/ACER
Yanfu Yan, Denys Poshyvanyk
SCAM2
2022 Light Attention Embedding for Facial Expression Recognition
abstract
Facial expression recognition is important for human–computer interaction and other applications. Several facial expression datasets have been published in recent decades and have enabled improvements in algorithms for classifying emotions. However, recognition of realistic expressions in real-world conditions is still challenging because of uncontrolled conditions, such as lighting, brightness, pose, and occlusion. In this paper, we propose a light attention embedding network based on the spatial attention mechanism (LAENet-SA), which can focus on locations in an image that are relevant to emotion. LAENet-SA allows a small number of attention modules to be embedded and can be constructed from typical convolutional neural networks. The performance of LAENet-SA on facial expression recognition has been validated on three facial expression datasets, including a lab-controlled dataset and two in-the-wild datasets. Experimental results show that LAENet-SA improved the performance on each dataset, compared with state-of-the-art methods, and achieved better generalization when tested on facial images with occlusion.
Jian Xue 0002, Ke Lu 0002, Yanfu Yan
IEEE Trans. Circuits Syst. Video Technol.4
2022 Fine-Grained Categorization From RGB-D Images
abstract
In the field of computer vision, fine-grained visual categorization has attracted a lot of attention and made great progress due to convolutional neural networks and a large number of publicly available datasets. With next-generation sensing technology, RGB-D cameras can provide high-quality synchronized RGB and depth images for solving many computer vision problems. Although RGB-D cameras have been used in the context of multi-view object category detection and scene understanding, they have not been widely used in fine-grained classification. In this paper, we introduce a multiview RGB-D dataset RGBD-FG for fine-grained categorization. Currently, the dataset contains 93 051 RGB-D images covering 19 super-categories and 50 sub-categories of common vegetables and fruit, and is organized in a hierarchical manner. We provide extensive experimental results to establish state-of-the-art benchmarks for our dataset, illustrating its diversity and scope for improvement through future work. We also propose a novel modality-specific multimodal network called FS-Multimodal network, which can solve two limitations of multimodal networks trained based on fine-tuning techniques: over-fitting and lack of effective depth-specific features. We hope that our study lays the foundations for fine-grained categorization of RGB-D data.
Yanhao Tan, Mohammad Muntasir Rahman, Yanfu Yan, Jian Xue 0002, Ling Shao 0001, Ke Lu 0002
IEEE Trans. Multim.3
2019 A Markerless Body Motion Capture System for Character Animation Based on Multi-view Cameras
abstract
A novel application system is proposed in this paper to achieve the generation of 3D character animation driven by markerless human body motion capture. The whole pipeline of the system consists of four parts: capturing motion data by multiple cameras, detecting 2D human body joints and estimating 3D joints, calculating bone transformation matrices, and generating character animation. Its main objective is to generate 3D skeleton and animation for 3D characters from multi-view images captured by ordinary cameras. The computation complexity of 3D skeleton reconstruction based on 3D vision is reduced accordingly to achieve the frame-by-frame motion capture. The experimental results show that our system is effective and efficient for capturing human action and animating 3D cartoon characters simultaneously.
Jinbao Wang 0001, Ke Lyu, Yanfu Yan
ICASSP5
2019 Dense Attention Network for Facial Expression Recognition in the Wild
abstract
Recognizing facial expression is significant for human-computer interaction system and other applications. A certain number of facial expression datasets have been published in recent decades and helped with the improvements for emotion classification algorithms. However, recognition of the realistic expressions in the wild is still challenging because of uncontrolled lighting, brightness, pose, occlusion, etc. In this paper, we propose an attention mechanism based module which can help the network focus on the emotion-related locations. Furthermore, we produce two network structures named DenseCANet and DenseSANet by using the attention modules based on the backbone of DenseNet. Then these two networks and original DenseNet are trained on wild dataset AffectNet and lab-controlled dataset CK+. Experimental results show that the DenseSANet has improved the performance on both datasets comparing with the state-of-the-art methods.
Ke Lu 0002, Jian Xue 0002, Yanfu Yan
MMAsia4
2018 A fast Cascade Shape Regression Method based on CNN-based Initialization
abstract
Cascade shape regression (CSR) methods predict facial landmarks by iteratively updating an initial shape and are state-of-the-art. The initial shape always limits the result and causes local optimum, which is usually obtained from the average face or by randomly picking a face from the training set. In this paper, we propose a CNN-based initial method for CSR. Convolution neural network provides a highly robust initial shape estimation, while the following CSR algorithm fine-tunes the initialization rapidly to achieve higher accuracy. Furthermore, CNN-based initial approach is proposed to get 68-point initial shape, which is calculated from convolutional network 5-point result by the radial basis function interpolation with thin-plate splines (RBF-TPS). Extensive experiments demonstrate that CSR methods are sensitive to the initialization and proposed approach gets favorable results compared to state-of-the-art algorithms and achieves real-time performance.
Jian Xue 0002, Ke Lu 0002, Yanfu Yan
ICPR4