Yang Liu 0182

dblp:51/3710-182 · DBLP profile ↗
← Back
19ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0003-2157-0080ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 APGFusion : Adaptive PoolFormer and CNN medical image fusion network based on convolutional gated linear units
Danhua Lu, Yang Liu 0182, Feng Yang 0014
Expert Syst. Appl.3
2026 SORT-LFR: Revisiting SORT for Multi-Object Tracking in Low-Frame-Rate Videos
abstract
For certain applications like highway surveillance systems, only low-frame-rate videos are recorded, which presents a huge challenge to existing trackers, as objects tend to undergo far more abrupt changes in location, motion, and appearance between successive frames compared to normal frame rates. To handle the above challenges, we propose a novel approach, namely$\mathbb {SORT}$-$\mathbb {LFR}$, for$\mathbb {S}$imple$\mathbb {O}$nline and$\mathbb {R}$ealtime$\mathbb {T}$racking in$\mathbb {L}$ow-$\mathbb {F}$rame-$\mathbb {R}$ate videos, which consists of following techniques: 1) A feature-prior association strategy to improve the capability to track new objects with significant displacements; 2) A Kalman filter using acceleration in state space (accel-fused Kalman filter) to improve the motion estimation capability for non-constant velocity moving objects; 3) A detection-guided adaptive exponential moving average (DG-AEMA) feature update mechanism to enhance feature temporal modeling capability for tracked objects; 4) A trajectory-covariance threshold tuning (TCTT) method to filter out incorrect association results. Through these techniques, the proposed SORT achieves 91.8 HOTA, 92.6 MOTA and 93.9 IDF1, which surpass all state-of-the-art trackers on the public CityFlow and our private HighwayTrack datasets under the low-frame-rate setting.
Yawen Huang, Yubei Lin, Ziwei Zhu 0005, Xingming Zhang 0001, Yang Liu 0182, Yuexiang Li, Yefeng Zheng 0001
IEEE Trans. Multim.6
2025 Gender Fairness of Machine Learning Algorithms for Pain Detection
abstract
Automated pain detection through machine learning (ML) and deep learning (DL) algorithms holds significant potential in healthcare, particularly for patients unable to self-report pain levels. However, the accuracy and fairness of these algorithms across different demographic groups (e.g., gender) remain under-researched. This paper investigates the gender fairness of ML and DL models trained on the UNBC-McMaster Shoulder Pain Expression Archive Database, evaluating the performance of various models in detecting pain based solely on the visual modality of participants’ facial expressions. We compare traditional ML algorithms, Linear Support Vector Machine (L SVM) and Radial Basis Function SVM (RBF SVM), with DL methods, Convolutional Neural Network (CNN) and Vision Transformer (ViT), using a range of performance and fairness metrics. While ViT achieved the highest accuracy and a selection of fairness metrics, all models exhibited gender-based biases. These findings highlight the persistent trade-off between accuracy and fairness, emphasising the need for fairness-aware techniques to mitigate biases in automated healthcare systems.
Yuting Shang, Jiaee Cheong, Yang Liu 0182, Hatice Gunes
FG4
2025 Diffusion Model and Class-Balanced Adaptive Threshold for Federated Semi-supervised Non-IID Image Classification
Guirong Liang, Yang Liu 0182, Feng Yang 0014
ICIC (11)2
2025 Dynamic class-balanced threshold Federated Semi-Supervised Learning by exploring diffusion model and all unlabeled data
Yang Liu 0182, Guirong Liang, Feng Yang 0014
Future Gener. Comput. Syst.2
2025 Multi-consistency for semi-supervised medical image segmentation via diffusion models
Yunzhu Chen, Yang Liu 0182, Manti Lu, Liyao Fu, Feng Yang 0014
Pattern Recognit.2
2024 Unified Video and Image Representation for Boosted Video Face Forgery Detection
abstract
Face forgery detection is crucial in preserving the security and integrity of facial data amidst the rapid developments in face manipulation techniques and deep generative models. Existing methods for video face forgery detection typically assume that all frames in a forged video are manipulated, while identifying partially forged videos with only a subset of altered frames is still a challenge to be solved. To address this issue, we propose a novel framework, i.e., the UVIF, that utilizes additional annotated images to provide fine-grained supervision for detecting partial forgeries in videos. The UVIF integrates a unified encoder and a multi-task learning paradigm to model both facial videos and images for boosted video face forgery detection. A 2D backbone with temporal fusion modules is employed for the unified encoder. A pseudo labeling process is also designed for facial video frames to bridge the representation of individual video frames and static images. Extensive experiments on benchmark datasets demonstrate the effectiveness of our framework, outperforming state-of-the-art methods in detecting partially forged videos while introducing no additional computational overhead. Our code is available at https://github.com/haotianll/UVIF.
Chenhui Pan, Yang Liu 0182, Guoying Zhao 0001
ECAI3
2024 Benchmarking deep Facial Expression Recognition: An extensive protocol with balanced dataset in the wild
abstract
Facial expression recognition (FER) is crucial in enhancing human-computer interaction. While current FER methods, leveraging various open-source deep learning models and training techniques, have shown promising accuracy and generalizability, their efficacy often diminishes in real-world scenarios that are not extensively studied. Addressing this gap, we introduce a novel in-the-wild balanced testing facial expression dataset designed for cross-domain validation, called BTFER. We rigorously evaluated widely utilized networks and self-designed architectures, adhering to a standardized protocol. Additionally, we explored different configurations, including input resolutions, class balance management, and pre-trained strategies, to ascertain their impact on performance. Through comprehensive testing across three major FER datasets and our in-depth cross-validation, we have ranked these network architectures and formulated a series of practical guidelines for implementing deep learning-based FER solutions in real-life applications. This paper also delves into the ethical considerations, privacy concerns, and regulatory aspects relevant to the deployment of FER technologies in sectors such as marketing, education, entertainment, and healthcare, aiming to foster responsible and effective use. The BTFER dataset and the implementation code are available in Kaggle and Github, respectively.
Gianmarco Ipinze Tutuianu, Yang Liu 0182, Ari Alamäki, Janne Kauttonen
Eng. Appl. Artif. Intell.2
2024 Exploring contactless techniques in multimodal emotion recognition: insights into diverse applications, challenges, solutions, and prospects
abstract
Abstract In recent years, emotion recognition has received significant attention, presenting a plethora of opportunities for application in diverse fields such as human–computer interaction, psychology, and neuroscience, to name a few. Although unimodal emotion recognition methods offer certain benefits, they have limited ability to encompass the full spectrum of human emotional expression. In contrast, Multimodal Emotion Recognition (MER) delivers a more holistic and detailed insight into an individual's emotional state. However, existing multimodal data collection approaches utilizing contact-based devices hinder the effective deployment of this technology. We address this issue by examining the potential of contactless data collection techniques for MER. In our tertiary review study, we highlight the unaddressed gaps in the existing body of literature on MER. Through our rigorous analysis of MER studies, we identify the modalities, specific cues, open datasets with contactless cues, and unique modality combinations. This further leads us to the formulation of a comparative schema for mapping the MER requirements of a given scenario to a specific modality combination. Subsequently, we discuss the implementation of Contactless Multimodal Emotion Recognition (CMER) systems in diverse use cases with the help of the comparative schema which serves as an evaluation blueprint. Furthermore, this paper also explores ethical and privacy considerations concerning the employment of contactless MER and proposes the key principles for addressing ethical and privacy concerns. The paper further investigates the current challenges and future prospects in the field, offering recommendations for future research and development in CMER. Our study serves as a resource for researchers and practitioners in the field of emotion recognition, as well as those intrigued by the broader outcomes of this rapidly progressing technology.
Umair Ali Khan, Qianru Xu, Yang Liu 0182, Altti Lagstedt, Ari Alamäki, Janne Kauttonen
Multim. Syst.3
2024 Interactions for Socially Shared Regulation in Collaborative Learning: An Interdisciplinary Multimodal Dataset
abstract
Socially shared regulation plays a pivotal role in the success of collaborative learning. However, evaluating socially shared regulation of learning (SSRL) proves challenging due to the dynamic and infrequent cognitive and socio-emotional interactions, which constitute the focal point of SSRL. To address this challenge, this article gathers interdisciplinary researchers to establish a multimodal dataset with cognitive and socio-emotional interactions for SSRL study. Firstly, to induce cognitive and socio-emotional interactions, learning science researchers designed a special collaborative learning task with regulatory trigger events among triadic people for the SSRL study. Secondly, this dataset includes various modalities like video, Kinect data, audio, and physiological data (accelerometer, EDA, heart rate) from 81 high school students in 28 groups, offering a comprehensive view of the SSRL process. Thirdly, three-level verbal interaction annotations and nonverbal interactions including facial expression, eye gaze, gesture, and posture are provided, which could further contribute to interdisciplinary fields such as computer science, sociology, and education. In addition, comprehensive analysis verifies the dataset’s effectiveness. As far as we know, this is the first multimodal dataset for studying SSRL among triadic group members.
Yante Li, Yang Liu 0182, Andy Nguyen, Henglin Shi, Eija Vuorenmaa, Sanna Järvelä, Guoying Zhao 0001
ACM Trans. Interact. Intell. Syst.2
2024 Uncertain Facial Expression Recognition via Multi-Task Assisted Correction
abstract
Deep models for facial expression recognition achieve high performance by training on large-scale labeled data. However, publicly available datasets contain uncertain facial expressions caused by ambiguous annotations or confusing emotions, which could severely decline the robustness. Previous studies usually follow the bias elimination method in general tasks without considering the uncertainty problem from the perspective of different corresponding sources. This article proposes a novel method of multi-task assisted correction in addressing uncertain facial expression recognition called MTAC. Specifically, a confidence estimation block and a weighted regularization module are applied to highlight solid samples and suppress uncertain samples in every batch. In addition, two auxiliary tasks, i.e., action unit detection and valence-arousal measurement, are introduced to learn semantic distributions from a data-driven AU graph and mitigate category imbalance based on latent dependencies between discrete and continuous emotions, respectively. Moreover, a re-labeling strategy guided by feature-level similarity constraint further generates new labels for identified uncertain samples to promote model learning. The proposed method can flexibly combine with existing frameworks in a fully-supervised or weakly-supervised manner. Experiments on five popular benchmarks demonstrate that the MTAC substantially improves over baselines when facing synthetic and real uncertainties and outperforms the state-of-the-art methods.
Yang Liu 0182, Xingming Zhang 0001, Janne Kauttonen, Guoying Zhao 0001
IEEE Trans. Multim.1
2023 Exploring Effective Knowledge Distillation for Tiny Object Detection
abstract
Detecting tiny objects is a long-standing and critical problem in object detection, with broad real-world applications such as autonomous driving, surveillance, and medical diagnosis. Recent studies for tiny object detection often cause extra computational costs during inference due to introducing feature maps with increased resolution or additional network modules. This scarifies the inference speed for better detection accuracy and may heavily limit their availability to real-world applications. Therefore, this paper turns to knowledge distillation to improve the representation learning of a small model regarding both superior detection accuracy and fast inference speed. The masked scale-aware feature distillation and local attention distillation are proposed to address the critical issues in the distillation of tiny objects. Experimental results on two tiny benchmarks indicate that our method can bring noticeable performance gains to different detectors while keeping their original inference speeds. Our method also shows competitive performance compared to state-of-the-art methods for tiny object detection. Our code is available at https://github.com/haotianll/TinyKD.
Qing Liu 0003, Yang Liu 0182, Yixiong Liang, Guoying Zhao 0001
ICIP3
2023 Graph-Based Facial Affect Analysis: A Review
abstract
As one of the most important affective signals, facial affect analysis (FAA) is essential for developing human-computer interaction systems. Early methods focus on extracting appearance and geometry features associated with human affects while ignoring the latent semantic information among individual facial changes, leading to limited performance and generalization. Recent work attempts to establish a graph-based representation to model these semantic relationships and develop frameworks to leverage them for various FAA tasks. This paper provides a comprehensive review of graph-based FAA, including the evolution of algorithms and their applications. First, the FAA background knowledge is introduced, especially on the role of the graph. We then discuss approaches widely used for graph-based affective representation in literature and show a trend towards graph construction. For the relational reasoning in graph-based FAA, existing studies are categorized according to their non-deep or deep learning methods, emphasizing the latest graph neural networks. Performance comparisons of the state-of-the-art graph-based FAA methods are also summarized. Finally, we discuss the challenges and potential directions. As far as we know, this is the first survey of graph-based FAA methods. Our findings can serve as a reference for future research in this field.
Yang Liu 0182, Xingming Zhang 0001, Yante Li, Jinzhao Zhou, Xin Li 0116, Guoying Zhao 0001
IEEE Trans. Affect. Comput.1
2022 Uncertain Label Correction via Auxiliary Action Unit Graphs for Facial Expression Recognition
abstract
High-quality annotated images are significant to deep facial expression recognition (FER) methods. However, uncertain labels, mostly existing in large-scale public datasets, often mislead the training process. In this paper, we achieve uncertain label correction of facial expressions using auxiliary action unit (AU) graphs, called ULC-AG. Specifically, a weighted regularization module is introduced to highlight valid samples and suppress category imbalance in every batch. Based on the latent dependency between emotions and AUs, an auxiliary branch using graph convolutional layers is added to extract the semantic information from graph topologies. Finally, a re-labeling strategy corrects the ambiguous annotations by comparing their feature similarities with semantic templates. Experiments show that our ULC-AG achieves 89.31% and 61.57% accuracy on RAF-DB and AffectNet datasets, respectively, outperform the baseline and state-of-the-art methods.
Yang Liu 0182, Xingming Zhang 0001, Janne Kauttonen, Guoying Zhao 0001
ICPR1
2022 Deep Learning for Micro-Expression Recognition: A Survey
abstract
Micro-expressions (MEs) are involuntary facial movements revealing people's hidden feelings in high-stake situations and have practical importance in various fields. Early methods for Micro-expression Recognition (MER) are mainly based on traditional features. Recently, with the success of Deep Learning (DL) in various tasks, neural networks have received increasing interest in MER. Different from macro-expressions, MEs are spontaneous, subtle, and rapid facial movements, leading to difficult data collection and annotation, thus publicly available datasets are usually small-scale. Currently, various DL approaches have been proposed to solve the ME issues and improve MER performance. In this survey, we provide a comprehensive review of deep MER and define a new taxonomy for the field encompassing all aspects of MER based on DL, including datasets, each step of the deep MER pipeline, and performance comparisons of the most influential methods. The basic approaches and advanced developments are summarized and discussed for each aspect. Additionally, we conclude the remaining challenges and potential directions for the design of robust MER systems. Finally, ethical considerations in MER are discussed. To the best of our knowledge, this is the first survey of deep MER methods, and this survey can serve as a reference point for future MER research.
Yante Li, Jinsheng Wei, Yang Liu 0182, Janne Kauttonen, Guoying Zhao 0001
IEEE Trans. Affect. Comput.3
2021 SG-DSN: A Semantic Graph-based Dual-Stream Network for facial expression recognition
Yang Liu 0182, Xingming Zhang 0001, Jinzhao Zhou, Lunkai Fu
Neurocomputing1
2021 Facial expression recognition using frequency multiplication network with uniform rectangular features
Jinzhao Zhou, Xingming Zhang 0001, Yubei Lin, Yang Liu 0182
J. Vis. Commun. Image Represent.4
2020 Facial Expression Recognition Using Spatial-Temporal Semantic Graph Network
abstract
Motions of facial components convey significant information of facial expressions. Although remarkable advancement has been made, the dynamic of facial topology has not been fully exploited. In this paper, a novel facial expression recognition (FER) algorithm called Spatial Temporal Semantic Graph Network (STSGN) is proposed to automatically learn spatial and temporal patterns through end-to-end feature learning from facial topology structure. The proposed algorithm not only has greater discriminative power to capture the dynamic patterns of facial expression and stronger generalization capability to handle different variations but also higher interpretability. Experimental evaluation on two popular datasets, CK+ and Oulu-CASIA, shows that our algorithm achieves more competitive results than other state-of-the-art methods.
Jinzhao Zhou, Xingming Zhang 0001, Yang Liu 0182, Xiangyuan Lan
ICIP3
2020 Learning the Connectivity: Situational Graph Convolution Network for Facial Expression Recognition
abstract
Previous studies recognizing expressions with facial graph topology mostly use a fixed facial graph structure established by the physical dependencies among facial landmarks. However, the static graph structure inherently lacks flexibility in non-standardized scenarios. This paper proposes a dynamic-graph-based method for effective and robust facial expression recognition. To capture action-specific dependencies among facial components, we introduce a link inference structure, called the Situational Link Generation Module (SLGM). We further propose the Situational Graph Convolution Network (SGCN) to automatically detect and recognize facial expression in various conditions. Experimental evaluations on two lab-constrained datasets, CK+ and Oulu, along with an in-the-wild dataset, AFEW, show the superior performance of the proposed method. Additional experiments on occluded facial images further demonstrate the robustness of our strategy.
Jinzhao Zhou, Xingming Zhang 0001, Yang Liu 0182
VCIP3