VLDB 2026 Research / reviewers in the wild / expert
Guanghao Yin
dblp:247/1052
· DBLP profile ↗
9ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Noise-Resistant Multimodal Transformer for Emotion Recognition
Yuanyuan Liu 0004, Haoyu Zhang 0001, Yibing Zhan, Zijing Chen, Guanghao Yin, Zhe Chen 0013 |
Int. J. Comput. Vis. | 5 |
| 2024 | Token-disentangling Mutual Transformer for multimodal emotion recognition
Guanghao Yin, Yuanyuan Liu 0004, Haoyu Zhang 0001, Fang Fang 0008, Chang Tang, Liangxiao Jiang |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | Online Streaming Video Super-Resolution With Convolutional Look-Up TableabstractOnline video streaming has fundamental limitations on the transmission bandwidth and computational capacity and super-resolution is a promising potential solution. However, applying existing video super-resolution methods to online streaming is non-trivial. Existing video codecs and streaming protocols (e.g., WebRTC) dynamically change the video quality both spatially and temporally, which leads to diverse and dynamic degradations. Furthermore, online streaming has a strict requirement for latency that most existing methods are less applicable. As a result, this paper focuses on the rarely exploited problem setting of online streaming video super resolution. To facilitate the research on this problem, a new benchmark dataset named LDV-WebRTC is constructed based on a real-world online streaming system. Leveraging the new benchmark dataset, we propose a novel method specifically for online video streaming, which contains a convolution and Look-Up Table (LUT) hybrid model to achieve better performance-latency trade-off. To tackle the changing degradations, we propose a mixture-of-expert-LUT module, where a set of LUT specialized in different degradations are built and adaptively combined to handle different degradations. Experiments show our method achieves 720P video SR around 100 FPS, while significantly outperforms existing LUT-based methods and offers competitive performance compared to efficient CNN-based methods. Code is available at https://github.com/quzefan/ConvLUT. Guanghao Yin, Zefan Qu, Xinyang Jiang, Zhenhua Han, Ningxin Zheng, Huan Yang 0005, Xiaohong Liu 0001, Yuqing Yang 0001, Dongsheng Li 0002, Lili Qiu |
IEEE Trans. Image Process. | 1 |
| 2023 | Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment AnalysisabstractThough Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved.To alleviate this, we present Adaptive Language-guided Multimodal Transformer (ALMT), which incorporates an Adaptive Hyper-modality Learning (AHL) module to learn an irrelevance/conflict-suppressing representation from visual and audio features under the guidance of language features at different scales.With the obtained hypermodality representation, the model can obtain a complementary and joint representation through multimodal fusion for effective MSA.In practice, ALMT achieves state-of-the-art performance on several popular datasets (e.g., MOSI, MOSEI and CH-SIMS) and an abundance of ablation demonstrates the validity and necessity of our irrelevance/conflict suppression mechanism. Haoyu Zhang 0001, Yu Wang 0246, Guanghao Yin, Kejun Liu, Yuanyuan Liu 0004, Tianshu Yu 0001 |
EMNLP | 3 |
| 2022 | Content-Variant Reference Image Quality Assessment via Knowledge DistillationabstractGenerally, humans are more skilled at perceiving differences between high-quality (HQ) and low-quality (LQ) images than directly judging the quality of a single LQ image. This situation also applies to image quality assessment (IQA). Although recent no-reference (NR-IQA) methods have made great progress to predict image quality free from the reference image, they still have the potential to achieve better performance since HQ image information is not fully exploited. In contrast, full-reference (FR-IQA) methods tend to provide more reliable quality evaluation, but its practicability is affected by the requirement for pixel-level aligned reference images. To address this, we firstly propose the content-variant reference method via knowledge distillation (CVRKD-IQA). Specifically, we use non-aligned reference (NAR) images to introduce various prior distributions of high-quality images. The comparisons of distribution differences between HQ and LQ images can help our model better assess the image quality. Further, the knowledge distillation transfers more HQ-LQ distribution difference information from the FR-teacher to the NAR-student and stabilizing CVRKD-IQA performance. Moreover, to fully mine the local-global combined information, while achieving faster inference speed, our model directly processes multiple image patches from the input with the MLP-mixer. Cross-dataset experiments verify that our model can outperform all NAR/NR-IQA SOTAs, even reach comparable performance than FR-IQA methods on some occasions. Since the content-variant and non-aligned reference HQ images are easy to obtain, our model can support more IQA applications with its robustness to content variations. Our code is available: https://github.com/guanghaoyin/CVRKD-IQA. Guanghao Yin, Wei Wang 0009, Zehuan Yuan, Chuchu Han, Wei Ji 0008, Shouqian Sun, Changhu Wang |
AAAI | 1 |
| 2022 | MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the WildabstractDynamic facial expression recognition (FER) databases provide important data support for affective computing and applications. However, most FER databases are annotated with several basic mutually exclusive emotional categories and contain only one modality, e.g., videos. The monotonous labels and modality cannot accurately imitate human emotions and fulfill applications in the real world. In this paper, we propose MAFW, a large-scale multi-modal compound affective database with 10,045 video-audio clips in the wild. Each clip is annotated with a compound emotional category and a couple of sentences that describe the subjects' affective behaviors in the clip. For the compound emotion annotation, each clip is categorized into one or more of the 11 widely-used emotions, i.e., anger, disgust, fear, happiness, neutral, sadness, surprise, contempt, anxiety, helplessness, and disappointment. To ensure high quality of the labels, we filter out the unreliable annotations by an Expectation Maximization (EM) algorithm, and then obtain 11 single-label emotion categories and 32 multi-label emotion categories. To the best of our knowledge, MAFW is the first in-the-wild multi-modal database annotated with compound emotion annotations and emotion-related captions. Additionally, we also propose a novel Transformer-based expression snippet feature learning method to recognize the compound emotions leveraging the expression-change relations among different emotions and modalities. Extensive experiments on MAFW database show the advantages of the proposed method over other state-of-the-art methods for both uni- and multi-modal FER. Our MAFW database is publicly available from https://mafw-database.github.io/MAFW. Yuanyuan Liu 0004, Chuanxu Feng, Wenbin Wang 0001, Guanghao Yin, Jiabei Zeng, Shiguang Shan |
ACM Multimedia | 5 |
| 2022 | Conditional Hyper-Network for Blind Super-Resolution With Multiple DegradationsabstractAlthough the single-image super-resolution (SISR) methods have achieved great success on the single degradation, they still suffer performance drop with multiple degrading effects in real scenarios. Recently, some blind and non-blind models for multiple degradations have been explored. However, these methods usually degrade significantly for distribution shifts between the training and test data. Towards this end, we propose a novel conditional hyper-network framework for super-resolution with multiple degradations (named CMDSR), which helps the SR framework learn how to adapt to changes in the degradation distribution of input. We extract degradation prior at the task-level with the proposed ConditionNet, which will be used to adapt the parameters of the basic SR network (BaseNet). Specifically, the ConditionNet of our framework first learns the degradation prior from a support set, which is composed of a series of degraded image patches from the same task. Then the adaptive BaseNet rapidly shifts its parameters according to the conditional features. Moreover, in order to better extract degradation prior, we propose a task contrastive loss to shorten the inner-task distance and enlarge the cross-task distance between task-level features. Without predefining degradation maps, our blind framework can conduct one single parameter update to yield considerable improvement in SR results. Extensive experiments demonstrate the effectiveness of CMDSR over various blind, and even several non-blind methods. The flexible BaseNet structure also reveals that CMDSR can be a general framework for a large series of SISR models. Our code is available at https://github.com/guanghaoyin/CMDSR. Guanghao Yin, Wei Wang 0009, Zehuan Yuan, Wei Ji 0008, Dongdong Yu, Shouqian Sun, Tat-Seng Chua, Changhu Wang |
IEEE Trans. Image Process. | 1 |
| 2022 | A Multimodal Framework for Large-Scale Emotion Recognition by Fusing Music and Electrodermal Activity SignalsabstractConsiderable attention has been paid to physiological signal-based emotion recognition in the field of affective computing. For reliability and user-friendly acquisition, electrodermal activity (EDA) has a great advantage in practical applications. However, EDA-based emotion recognition with large-scale subjects is still a tough problem. The traditional well-designed classifiers with hand-crafted features produce poorer results because of their limited representation abilities. And the deep learning models with auto feature extraction suffer the overfitting drop-off because of large-scale individual differences. Since music has a strong correlation with human emotion, static music can be involved as the external benchmark to constrain various dynamic EDA signals. In this article, we make an attempt by fusing the subject’s individual EDA features and the external evoked music features. And we propose an end-to-end multimodal framework, the one-dimensional residual temporal and channel attention network (RTCAN-1D). For EDA features, the channel-temporal attention mechanism for EDA-based emotion recognition is first involved in mine the temporal and channel-wise dynamic and steady features. The comparisons with single EDA-based SOTA models on DEAP and AMIGOS datasets prove the effectiveness of RTCAN-1D to mine EDA features. For music features, we simply process the music signal with the open-source toolkit openSMILE to obtain external feature vectors. We conducted systematic and extensive evaluations. The experiments on the current largest music emotion dataset PMEmo validate that the fusion of EDA and music is a reliable and efficient solution for large-scale emotion recognition. Guanghao Yin, Shouqian Sun |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | User Independent Emotion Recognition with Residual Signal-Image NetworkabstractUser independent emotion recognition with large scale physiological signals is a tough problem. There exist many advanced methods but they are conducted under relatively small datasets with dozens of subjects. Here, we propose Res-SIN, a novel end-to-end framework using Electrodermal Activity(EDA) signal images to classify human emotion. We first apply convex optimization-based EDA (cvxEDA) to decompose signals and mine the static and dynamic emotion changes. Then, we transform decomposed signals to images so that they can be effectively processed by CNN frameworks. The Res-SIN combines individual emotion features and external emotion benchmarks to accelerate convergence. We evaluate our approach on the PMEmo dataset, the currently largest emotional dataset containing music and EDA signals. To the best of author's knowledge, our method is the first attempt to classify large scale subject-independent emotion with 7962 pieces of EDA signals from 457 subjects. Experimental results demonstrate the reliability of our model and the binary classification accuracy of 73.65% and 73.43% on arousal and valence dimension can be used as a baseline. Guanghao Yin, Shouqian Sun, Hui Zhang 0064, Ning Zou |
ICIP | 1 |