VLDB 2026 Research / reviewers in the wild / expert
Kangning Yang
dblp:220/4120
· DBLP profile ↗
11ranked-venue papers
5as first author
7since 2021 · last 2025
0000-0002-7106-0022ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | F2T2-HIT: A U-Shaped FFT Transformer and Hierarchical Transformer for Reflection RemovalabstractSingle Image Reflection Removal (SIRR) technique plays a crucial role in image processing by eliminating unwanted reflections from the background. These reflections, often caused by photographs taken through glass surfaces, can significantly degrade image quality. SIRR remains a challenging problem due to the complex and varied reflections encountered in real-world scenarios. These reflections vary significantly in intensity, shapes, light sources, sizes, and coverage areas across the image, posing challenges for most existing methods to effectively handle all cases. To address these challenges, this paper introduces a U-shaped Fast Fourier Transform Transformer and Hierarchical Transformer (F2T2-HiT) architecture, an innovative Transformer-based design for SIRR. Our approach uniquely combines Fast Fourier Transform (FFT) Transformer blocks and Hierarchical Transformer blocks within a UNet framework. The FFT Transformer blocks leverage the global frequency domain information to effectively capture and separate reflection patterns, while the Hierarchical Transformer blocks utilize multi-scale feature extraction to handle reflections of varying sizes and complexities. Extensive experiments conducted on three publicly available testing datasets demonstrate state-of-the-art performance, validating the effectiveness of our approach. Jie Cai 0001, Kangning Yang, Ling Ouyang, Lan Fu, Jiaming Ding, Huiming Sun, Chiu Man Ho, Zibo Meng |
ICIP | 2 |
| 2025 | OpenRR-1k: A Scalable Dataset for Real-World Reflection RemovalabstractReflection removal technology plays a crucial role in photography and computer vision applications. However, existing techniques are hindered by the lack of high-quality in-the-wild datasets. In this paper, we propose a novel paradigm for collecting reflection datasets from a fresh perspective. Our approach is convenient, cost-effective, and scalable, while ensuring that the collected data pairs are of high quality, perfectly aligned, and represent natural and diverse scenarios. Following this paradigm, we collect a Real-world, Diverse, and Pixel-aligned dataset (named OpenRR-1k dataset), which contains 1,000 high-quality transmission-reflection image pairs collected in the wild. Through the analysis of several reflection removal methods and benchmark evaluation experiments on our dataset, we demonstrate its effectiveness in improving robustness in challenging real-world environments. Our dataset is available at https://github.com/caijie0620/OpenRR-1k. Kangning Yang, Ling Ouyang, Huiming Sun, Jie Cai 0001, Lan Fu, Jiaming Ding, Chiu Man Ho, Zibo Meng |
ICIP | 1 |
| 2023 | Survey on Emotion Sensing Using Mobile DevicesabstractThe rapid development and ubiquity of mobile and wearable devices promises to enable researchers to monitor users’ granular emotional data in a less intrusive manner. Researchers have used a wide variety of mobile and wearable devices for this purpose, and have proposed various approaches to sense users’ emotional states. In this survey, we utilise three established digital libraries (ACM Digital Library,IEEE Xplore Digital Library, andSpringer Nature). We analysed and critically assessed the different approaches used in the three stages (perception, learning, inference) of a typical mobile emotion sensing framework, following a structured paper selection process. The contribution of this survey is three-fold; first, we document all the latest relevant literature on mobile emotion sensing research; second, we describe how mobile and wearable devices use their sensing and computing capabilities to monitor human emotions; third, we discuss challenges and opportunities of mobile emotion sensing to demonstrate the potential of this thriving field of research. Kangning Yang, Benjamin Tag, Chaofan Wang 0001, Zhanna Sarsenbayeva, Tilman Dingler, Greg Wadley, Jorge Gonçalves 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2023 | Behavioral and Physiological Signals-Based Deep Multimodal Approach for Mobile Emotion RecognitionabstractWith the rapid development of mobile and wearable devices, it is increasingly possible to access users’ affective data in a more unobtrusive manner. On this basis, researchers have proposed various systems to recognize user’s emotional states. However, most of these studies rely on traditional machine learning techniques and a limited number of signals, leading to systems that either do not generalize well or would frequently lack sufficient information for emotion detection in realistic scenarios. In this paper, we propose a novel attention-based LSTM system that uses a combination of sensors from a smartphone (front camera, microphone, touch panel) and a wristband (photoplethysmography, electrodermal activity, and infrared thermopile sensor) to accurately determine user’s emotional states. We evaluated the proposed system by conducting a user study with 45 participants. Using collected behavioral (facial expression, speech, keystroke) and physiological (blood volume, electrodermal activity, skin temperature) affective responses induced by visual stimuli, our system was able to achieve an average accuracy of 89.2 percent for binary positive and negative emotion classification under leave-one-participant-out cross-validation. Furthermore, we investigated the effectiveness of different combinations of data signals to cover different scenarios of signal availability. Kangning Yang, Chaofan Wang 0001, Zhanna Sarsenbayeva, Benjamin Tag, Tilman Dingler, Greg Wadley, Jorge Gonçalves 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2022 | Hand Hygiene Quality Assessment Using Image-to-Image Translation
Chaofan Wang 0001, Kangning Yang, Weiwei Jiang 0001, Jing Wei 0002, Zhanna Sarsenbayeva, Jorge Gonçalves 0001, Vassilis Kostakos |
MICCAI (8) | 2 |
| 2022 | Mobile Emotion Recognition via Multiple Physiological Signals using Convolution-augmented TransformerabstractRecognising and monitoring emotional states play a crucial role in mental health and well-being management. Importantly, with the widespread adoption of smart mobile and wearable devices, it has become easier to collect long-term and granular emotion-related physiological data passively, continuously, and remotely. This creates new opportunities to help individuals manage their emotions and well-being in a less intrusive manner using off-the-shelf low-cost devices. Pervasive emotion recognition based on physiological signals is, however, still challenging due to the difficulty to efficiently extract high-order correlations between physiological signals and users' emotional states. In this paper, we propose a novel end-to-end emotion recognition system based on a convolution-augmented transformer architecture. Specifically, it can recognise users' emotions on the dimensions of arousal and valence by learning both the global and local fine-grained associations and dependencies within and across multimodal physiological data (including blood volume pulse, electrodermal activity, heart rate, and skin temperature). We extensively evaluated the performance of our model using the K-EmoCon dataset, which is acquired in naturalistic conversations using off-the-shelf devices and contains spontaneous emotion data. Our results demonstrate that our approach outperforms the baselines and achieves state-of-the-art or competitive performance. We also demonstrate the effectiveness and generalizability of our system on another affective dataset which used affect inducement and commercial physiological sensors. Kangning Yang, Benjamin Tag, Chaofan Wang 0001, Tilman Dingler, Greg Wadley, Jorge Gonçalves 0001 |
ICMR | 1 |
| 2021 | Benchmarking commercial emotion detection systems using realistic distortions of facial image datasets
Kangning Yang, Chaofan Wang 0001, Zhanna Sarsenbayeva, Benjamin Tag, Tilman Dingler, Greg Wadley, Jorge Gonçalves 0001 |
Vis. Comput. | 1 |
| 2020 | Does Smartphone Use Drive our Emotions or vice versa? A Causal AnalysisabstractIn this paper, we demonstrate the existence of a bidirectional causal relationship between smartphone application use and user emotions. In a two-week long in-the-wild study with 30 participants we captured 502,851 instances of smartphone application use in tandem with corresponding emotional data from facial expressions. Our analysis shows that while in most cases application use drives user emotions, multiple application categories exist for which the causal effect is in the opposite direction. Our findings shed light on the relationship between smartphone use and emotional states. We furthermore discuss the opportunities for research and practice that arise from our findings and their potential to support emotional well-being. Zhanna Sarsenbayeva, Gabriele Marini, Niels van Berkel, Chu Luo, Weiwei Jiang 0001, Kangning Yang, Greg Wadley, Tilman Dingler, Vassilis Kostakos, Jorge Gonçalves 0001 |
CHI | 6 |
| 2018 | Multimodal Affective Analysis Using Hierarchical Attention Strategy with Word-Level AlignmentabstractMultimodal affective computing, learning to recognize and interpret human affect and subjective information from multiple data sources, is still challenging because:(i) it is hard to extract informative features to represent human affects from heterogeneous inputs; (ii) current fusion strategies only fuse different modalities at abstract levels, ignoring time-dependent interactions between modalities. Addressing such issues, we introduce a hierarchical multimodal architecture with attention and word-level fusion to classify utterance-level sentiment and emotion from text and audio data. Our introduced model outperforms state-of-the-art approaches on published datasets, and we demonstrate that our model's synchronized attention over modalities offers visual interpretability. Kangning Yang, Shiyu Fu, Shuhong Chen, Xinyu Li 0003, Ivan Marsic |
ACL (1) | 2 |
| 2018 | Hybrid Attention based Multimodal Network for Spoken Language ClassificationabstractWe examine the utility of linguistic content and vocal characteristics for multimodal deep learning in human spoken language understanding. We present a deep multimodal network with both feature attention and modality attention to classify utterance-level speech data. The proposed hybrid attention architecture helps the system focus on learning informative representations for both modality-specific feature extraction and model fusion. The experimental results show that our system achieves state-of-the-art or competitive results on three published multimodal datasets. We also demonstrated the effectiveness and generalization of our system on a medical speech dataset from an actual trauma scenario. Furthermore, we provided a detailed comparison and analysis of traditional approaches and deep learning methods on both feature extraction and fusion. Kangning Yang, Shiyu Fu, Shuhong Chen, Xinyu Li 0003, Ivan Marsic |
COLING | 2 |
| 2018 | Human Conversation Analysis Using Attentive Multimodal Networks with Hierarchical Encoder-DecoderabstractHuman conversation analysis is challenging because the meaning can be expressed through words, intonation, or even body language and facial expression. We introduce a hierarchical encoder-decoder structure with attention mechanism for conversation analysis. The hierarchical encoder learns word-level features from video, audio, and text data that are then formulated into conversation-level features. The corresponding hierarchical decoder is able to predict different attributes at given time instances. To integrate multiple sensory inputs, we introduce a novel fusion strategy with modality attention. We evaluated our system on published emotion recognition, sentiment analysis, and speaker trait analysis datasets. Our system outperformed previous state-of-the-art approaches in both classification and regressions tasks on three datasets. We also outperformed previous approaches in generalization tests on two commonly used datasets. We achieved comparable performance in predicting co-existing labels using the proposed model instead of multiple individual models. In addition, the easily-visualized modality and temporal attention demonstrated that the proposed attention mechanism helps feature selection and improves model interpretability. Xinyu Li 0003, Kaixiang Huang, Shiyu Fu, Kangning Yang, Shuhong Chen, Moliang Zhou, Ivan Marsic |
ACM Multimedia | 5 |