EDBT 2026 Demo / reviewers in the wild / expert
Taoyue Wang
dblp:276/3243
· DBLP profile ↗
12ranked-venue papers
1as first author
9since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | You Only Need One Stage: Novel-View Synthesis from a Single Blind Face ImageabstractWe propose a novel one-stage method, NVB-Face, for generating consistent Novel-View images directly from a single Blind Face image. Existing approaches to novel-view synthesis for objects or faces typically require a high-resolution RGB image as input. When dealing with degraded images, the conventional pipeline follows a two-stage process: first restoring the image to high resolution, then synthesizing novel views from the restored result. However, this approach is highly dependent on the quality of the restored image, often leading to inaccuracies and inconsistencies in the final output. To address this limitation, we extract single-view features directly from the blind face image and introduce a feature manipulator that transforms these features into 3D-aware, multi-view latent representations. Leveraging the powerful generative capacity of a diffusion model, our framework synthesizes high-quality, consistent novel-view face images. Experimental results show that our method significantly outperforms traditional two-stage approaches in both consistency and fidelity. Taoyue Wang, Xiang Zhang 0030, Huiyuan Yang, Lijun Yin 0001 |
AAAI | 1 |
| 2026 | ReMiX-MAE: Learning Missing-Channel Cross-Modal Representations from RGB-Only Clinical Facial Videos for Sympathetic-Mediated Pain Assessment
Nan Bi, Taoyue Wang, Lijun Yin 0001 |
FG | 2 |
| 2026 | Inter-Stance: A Dyadic Multimodal Corpus for Conversational Stance Analysis
Xiang Zhang 0030, Taoyue Wang, Nan Bi, Cody Zhou, Zoie Wang, Yuming Su, Jeffrey F. Cohn, Lijun Yin 0001 |
FG | 3 |
| 2024 | Enhancing Face Recognition in Low-Quality Images Based on Restoration and 3D Multiview GenerationabstractFace recognition in low-quality images presents a significant challenge, particularly when images or videos are captured under adverse conditions such as atmosphere turbulence due to long distance capture causing noise, blurring, and distortion in low resolution videos. This paper proposes a novel and comprehensive approach to address the challenges in atmosphere turbulence, leveraging a multi-stage process that includes weak-strong image restoration using Generative Adversarial Network (GAN) based model and Stable Diffusion (SD) based model, 3D face view generation with Neural Radiance Fields (NeRF), and adaptive face recognition on fine-augmented datasets. Our methodology shows substantial improvements in face recognition accuracy on average from 55.6% to 74.7% on the level-3 turbulence images of LFW, CFP, CALFW, as well as improvements on datasets of TinyFace, the BRIAR BRC1, and the BRIAR BGC, demonstrating the performance enhancement in extremely challenging conditions. Xiang Zhang 0030, Taoyue Wang, Lijun Yin 0001 |
IJCB | 3 |
| 2024 | Multimodal Channel-Mixing: Channel and Spatial Masked AutoEncoder on Facial Action Unit DetectionabstractRecent studies have focused on utilizing multi-modal data to develop robust models for facial Action Unit (AU) detection. However, the heterogeneity of multi-modal data poses challenges in learning effective representations. One such challenge is extracting relevant features from multiple modalities using a single feature extractor. Moreover, previous studies have not fully explored the potential of multi-modal fusion strategies. In contrast to the extensive work on late fusion, there are limited investigations on early fusion for channel information exploration. This paper presents a novel multi-modal reconstruction network, named Multimodal Channel-Mixing (MCM), as a pre-trained model to learn robust representation for facilitating multi-modal fusion. The approach follows an early fusion setup, integrating a Channel-Mixing module, where two out of five channels are randomly dropped. The dropped channels then are reconstructed from the remaining channels using masked autoencoder. This module not only reduces channel redundancy, but also facilitates multi-modal learning and reconstruction capabilities, resulting in robust feature learning. The encoder is fine-tuned on a downstream task of automatic facial action unit detection. Pre-training experiments were conducted on BP4D+, followed by fine-tuning on BP4D and DISFA to assess the effectiveness and robustness of the proposed framework. The results demonstrate that our method meets and surpasses the performance of state-of-the-art baseline methods. Xiang Zhang 0030, Huiyuan Yang, Taoyue Wang, Lijun Yin 0001 |
WACV | 3 |
| 2024 | Disagreement Matters: Exploring Internal Diversification for Redundant Attention in Generic Facial Action AnalysisabstractThis paper demonstrates the effectiveness of a diversification mechanism for building a more robust multi-attention system in generic facial action analysis. While previous multi-attention (e.g., visual attention and self-attention) research on facial expression recognition (FER) and Action Unit (AU) detection have been thoroughly studied to focus on ”external attention diversification”, where attention branches localize different facial areas, we delve into the realm of ”internal attention diversification” and explore the impact of diverse attention patterns within the same Region of Interest (RoI). Our experiments reveal that variability in attention patterns significantly impacts model performance, indicating that unconstrained multi-attention plagued by redundancy and over-parameterization, leading to sub-optimal results. To tackle this issue, we propose a compact module that guides the model to achieve self-diversified multi-attention. Our method is applied to both CNN-based and Transformer-based models, benchmarked on popular databases such as BP4D and DISFA for AU detection, as well as CK+, MMI, BU-3DFE, and BP4D+ for facial expression recognition. We also evaluate the mechanism on Self-attention and Channel-wise attention designs for improving their adaptive capabilities in multi-modal feature fusion tasks. The multi-modal evaluation is conducted on BP4D, BP4D+, and our newly developed large-scale comprehensive emotion database BP4D++, which contains well-synchronized and aligned sensor modalities, addressing the scarcity of annotations and identities in human affective computing. We plan to release the new database to the research community, fostering further advancements in this field. Zheng Zhang 0023, Xiang Zhang 0030, Taoyue Wang, Huiyuan Yang, Umur A. Ciftci, Jeffrey F. Cohn, Lijun Yin 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2023 | ReactioNet: Learning High-order Facial Behavior from Universal Stimulus-Reaction by Dyadic Relation ReasoningabstractDiverse visual stimuli can evoke various human affective states, which are usually manifested in an individual’s muscular actions and facial expressions. In lab-controlled emotion datasets, such a critical component (i.e., stimulus) was commonly designed in a limited way, making researchers incapable of generalizing the universal correlation and causation of stimulus-reaction as well as predicting possible emotions from context, timing, and relation. In this paper, we collected a large-scale spontaneous facial behavior database ReactioNet, which contains 1.1 million coupled stimulus-reaction tuples (visual/audio/caption from both stimuli and subjects). We introduce a new facial behavior detection scenario, Dyadic Relation Reasoning (DRR), which aims to detect facial actions by reasoning their relations with stimuli. By aggregating the dyadic information, our method essentially forms a relation prototype Universal Stimulus Reaction (U-SR), which encodes the low-order and high-order relationships between stimulus agents and facial reactions. A framework with both non-graph and graph modules is further developed to evaluate DRR-based facial action unit detection, facial expression recognition, and scene classification. Specifically, to learn "what" arouses a facial reaction, the non-graph module associates and projects the fine-grained stimulus-reaction features into common subspaces using cross-domain contrastive learning. To learn "how" stimulus-reaction pairs are mutually related, the graph module adopts Graph Convolution Network to represent, converge, and infer the dyadic U-SR relation under two relation assumptions (i.e., homophily and heterophily [68]). Extensive experiments demonstrate the effectiveness of the proposed work. The new dataset will be available for the research community. Taoyue Wang, Geran Zhao, Xiang Zhang 0030, Xi Kang, Lijun Yin 0001 |
ICCV | 2 |
| 2023 | Knowledge-Spreader: Learning Semi-Supervised Facial Action Dynamics by Consistifying Knowledge GranularityabstractRecent studies on dynamic facial action unit (AU) detection have extensively relied on dense annotations. However, manual annotations are difficult, time-consuming, and costly. The canonical semi-supervised learning (SSL) methods ignore the consistency, extensibility, and adaptability of structural knowledge across spatial-temporal domains. Furthermore, the reliance on offline design and excessive parameters hinder the efficiency of the learning process. To remedy these issues, we propose a lightweight and online semi-supervised framework, a so-called Knowledge-Spreader (KS), to learn AU dynamics with sparse annotations. By formulating SSL as a Progressive Knowledge Distillation (PKD) problem, we aim to infer cross-domain information, specifically from spatial to temporal domains, by consistifying knowledge granularity within Teacher-Students Network. Specifically, KS employs sparsely annotated key-frames to learn AU dependencies as the privileged knowledge. Then, the model spreads the learned knowledge to their unlabeled neighbours by jointly applying knowledge distillation and pseudo-labeling, and completes the temporal information as the expanded knowledge. We term the progressive knowledge distillation as "Knowledge Spreading", which allows our model to learn spatial-temporal knowledge from video clips with only one label allocated. Extensive experiments demonstrate that KS achieves competitive performance as compared to the state of the arts under the circumstances of using only 2% labels on BP4D and 5% labels on DISFA. In addition, we have tested it on our newly developed large-scale comprehensive emotion database BP4D++, which contains considerable samples across well-synchronized and aligned sensor modalities for alleviating the scarcity issue of annotations and identities. Xiang Zhang 0030, Taoyue Wang, Lijun Yin 0001 |
ICCV | 3 |
| 2023 | Weakly-Supervised Text-driven Contrastive Learning for Facial Behavior UnderstandingabstractContrastive learning has shown promising potential for learning robust representations by utilizing unlabeled data. However, constructing effective positive-negative pairs for contrastive learning on facial behavior datasets remains challenging. This is because such pairs inevitably encode the subject-ID information, and the randomly constructed pairs may push similar facial images away due to the limited number of subjects in facial behavior datasets. To address this issue, we propose to utilize activity descriptions, coarse-grained information provided in some datasets, which can provide high-level semantic information about the image sequences but is often neglected in previous studies. More specifically, we introduce a two-stage Contrastive Learning with Text-Embeded framework for Facial behavior understanding (CLEF). The first stage is a weakly-supervised contrastive learning method that learns representations from positive-negative pairs constructed using coarse-grained activity information. The second stage aims to train the recognition of facial expressions or facial action units by maximizing the similarity between the image and the corresponding text label names. The proposed CLEF achieves state-of-the-art performance on three in-the-lab datasets for AU recognition and three in-the-wild datasets for facial expression recognition. Xiang Zhang 0030, Taoyue Wang, Huiyuan Yang, Lijun Yin 0001 |
ICCV | 2 |
| 2020 | Set Operation Aided Network for Action Units DetectionabstractAs a large number of parameters exist in deepmodel based methods, training such models usually requires many fully AU-annotated facial images. This is true with regard to the number of frames in two widely used datasets: BP4D[31] and DISFA [18], while those frames were captured from a small number of subjects (41, 27 respectively). This is problematic, as subjects produce highly consistent facial muscle movements, adding more frames per subject would only adds more close points in the feature space, and thus the classifier does not benefit from those extra frames. Data augmentation methods can be applied to alleviate the problem to a certain degree, but they fail to augment new subjects. We propose a novel Set Operation Aided Network (SO-Net) for action units detection. Specifically, new features and the corresponding labels are generated by adding set operations to both the feature and label spaces. The generated new features can be treated as a representation of a hypothetical image. As a result, we can implicitly obtain training examples beyond what was originally observed in the dataset. Therefore, the deep model is forced to learn subject-independent features, and is generalizable to unseen subjects. SO-Net is end-to-end trainable, and can be flexibly plugged in any CNN model during training. We evaluate the proposed method on two public datasets, BP4D and DISFA. The experiment shows a state-of-the-art performance, demonstrating the effectiveness of the proposed method. Huiyuan Yang, Taoyue Wang, Lijun Yin 0001 |
FG | 2 |
| 2020 | Region of Interest Based Graph Convolution: A Heatmap Regression Approach for Action Unit DetectionabstractMachine vision of human facial expressions has been studied for decades, from prototypical expressions to Action Units (AUs), from hand-crafted to deep features, from multi-class to multi-label classifications. Since the widely adopted deep networks lack interpretation on learnt representations, human prior knowledge cannot be effectively imposed and examined. On the other hand, AU is a human defined concept. In order to align with this idea, a finer level of network design is desired. In this paper, we first extend the heatmaps to ROI maps, encoding the location of both positive and negative occurred AUs, then employ a well-designed backbone network to regress it. In this way, AU detection is performed in two stages, key regions localization and occurrence classification. To prompt the spatial dependency among ROIs, we utilize graph convolution for feature refinement. The decomposition of similarity matrix is supervised by AU labels. This novel framework is evaluated on two benchmark databases (BP4D and DISFA) for AU detection. The experimental results are superior to the state-of-the-art algorithms and baseline models, demonstrating the effectiveness of our proposed method. Zheng Zhang 0023, Taoyue Wang, Lijun Yin 0001 |
ACM Multimedia | 2 |
| 2020 | Adaptive Multimodal Fusion for Facial Action Units RecognitionabstractMultimodal facial action units (AU) recognition aims to build models that are capable of processing, correlating, and integrating information from multiple modalities (i.e., 2D images from a visual sensor, 3D geometry from 3D imaging, and thermal images from an infrared sensor). Although the multimodel data can provide rich information, there are two challenges that have to be addressed when learning from multimodal data: 1) the model must capture the complex cross-modal interactions in order to utilize the additional and mutual information effectively; 2) the model must be robust enough in the circumstance of unexpected data corruptions during testing, in case of a certain modality missing or being noisy. In this paper, we propose a novel A daptive M ultimodal F usion method (AMF ) for AU detection, which learns to select the most relevant feature representations from different modalities by a re-sampling procedure conditioned on a feature scoring module. The feature scoring module is designed to allow for evaluating the quality of features learned from multiple modalities. As a result, AMF is able to adaptively select more discriminative features, thus increasing the robustness to missing or corrupted modalities. In addition, to alleviate the over-fitting problem and make the model generalize better on the testing data, a cut-switch multimodal data augmentation method is designed, by which a random block is cut and switched across multiple modalities. We have conducted a thorough investigation on two public multimodal AU datasets, BP4D and BP4D+, and the results demonstrate the effectiveness of the proposed method. Ablation studies on various circumstances also show that our method remains robust to missing or noisy modalities during tests. Huiyuan Yang, Taoyue Wang, Lijun Yin 0001 |
ACM Multimedia | 2 |