EDBT 2026 Demo / reviewers in the wild / expert
Yuanhao Yu
dblp:00/10782
· DBLP profile ↗
19ranked-venue papers
1as first author
18since 2021 · last 2026
0000-0001-8176-9716ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 11 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ETR: Entropy Trend Reward for Efficient Chain-of-Thought ReasoningabstractChain-of-thought (CoT) reasoning improves large language model performance on complex tasks, but often produces excessively long and inefficient reasoning traces.Existing methods shorten CoTs using length penalties or global entropy reduction, implicitly assuming that low uncertainty is desirable throughout reasoning.We show instead that reasoning efficiency is governed by the trajectory of uncertainty.CoTs with dominant downward entropy trends are substantially shorter.Motivated by this insight, we propose Entropy Trend Reward (ETR), a trajectory-aware objective that encourages progressive uncertainty reduction while allowing limited local exploration.We integrate ETR into Group Relative Policy Optimization (GRPO) and evaluate it across multiple reasoning models and challenging benchmarks.ETR consistently achieves a superior accuracy-efficiency trade-off, improving DeepSeek-R1-Distill-7B by +9.9% accuracy while reducing CoT length by 67% across four benchmarks. Xuan Xiong, Huan Liu 0014, Li Gu, Zhixiang Chi, Yuanhao Yu, Yang Wang 0003 |
ACL (1) | 6 |
| 2025 | Extended Loss: Incorporating Long Context into Training Models when using Short Audio Frames
Quang Minh Dinh, Hoda Rezaee Kaviani, Mehrdad Hosseinzadeh, Yuanhao Yu |
INTERSPEECH | 4 |
| 2025 | Multi-level optimizing of parameters in stochastic configuration networks based on cloud model and nutcracker optimization algorithm
Yuanhao Yu, Kun Li 0011 |
Inf. Sci. | 2 |
| 2024 | Test-Time Personalization with Meta Prompt for Gaze EstimationabstractDespite the recent remarkable achievement in gaze estimation, efficient and accurate personalization of gaze estimation without labels is a practical problem but rarely touched on in the literature. To achieve efficient personalization, we take inspiration from the recent advances in Natural Language Processing (NLP) by updating a negligible number of parameters, "prompts", at the test time. Specifically, the prompt is additionally attached without perturbing original network and can contain less than 1% of a ResNet-18's parameters. Our experiments show high efficiency of the prompt tuning approach. The proposed one can be 10 times faster in terms of adaptation speed than the methods compared. However, it is non-trivial to update the prompt for personalized gaze estimation without labels. At the test time, it is essential to ensure that the minimizing of particular unsupervised loss leads to the goals of minimizing gaze estimation error. To address this difficulty, we propose to meta-learn the prompt to ensure that its updates align with the goal. Our experiments show that the meta-learned prompt can be effectively adapted even with a simple symmetry loss. In addition, we experiment on four cross-dataset validations to show the remarkable advantages of the proposed method. Huan Liu 0014, Julia Qi, Mohammad Hassanpour, Yang Wang 0003, Konstantinos N. Plataniotis, Yuanhao Yu |
AAAI | 7 |
| 2024 | Adapting to Distribution Shift by Visual Domain Prompt GenerationabstractIn this paper, we aim to adapt a model at test-time using a few unlabeled data to address distribution shifts.
To tackle the challenges of extracting domain knowledge from a limited amount of data, it is crucial to utilize correlated information from pre-trained backbones and source domains. Previous studies fail to utilize recent foundation models with strong out-of-distribution generalization. Additionally, domain-centric designs are not flavored in their works. Furthermore, they employ the process of modelling source domains and the process of learning to adapt independently into disjoint training stages. In this work, we propose an approach on top of the pre-computed features of the foundation model. Specifically, we build a knowledge bank to learn the transferable knowledge from source domains. Conditioned on few-shot target data, we introduce a domain prompt generator to condense the knowledge bank into a domain-specific prompt. The domain prompt then directs the visual features towards a particular domain via a guidance module. Moreover, we propose a domain-aware contrastive loss and employ meta-learning to facilitate domain knowledge extraction. Extensive experiments are conducted to validate the domain knowledge extraction. The proposed method outperforms previous work on 5 large-scale benchmarks including WILDS and DomainNet. Zhixiang Chi, Li Gu, Tao Zhong 0003, Huan Liu 0014, Yuanhao Yu, Konstantinos N. Plataniotis, Yang Wang 0003 |
ICLR | 5 |
| 2024 | Visually Guided Audio Source Separation with Meta Consistency LearningabstractIn this paper, we tackle the problem of visually guided audio source separation in the context of both known and unknown objects (e.g., musical instruments). Recent successful end-to-end deep learning approaches adopt a single network with fixed parameters to generalize across unseen test videos. However, it can be challenging to generalize in cases where the distribution shift between training and test videos is higher as they fail to utilize internal information of unknown test videos. Based on this observation, we introduce a meta-consistency driven test time adaptation scheme that enables the pretrained model to quickly adapt to known and unknown test music videos in order to bring substantial improvements. In particular, we design a self-supervised audio-visual consistency objective as an auxiliary task that learns the synchronization between audio and its corresponding visual embedding. Concretely, we apply a meta-consistency training scheme to further optimize the pretrained model for effective and faster test time adaptation. We obtain substantial performance gains with only a smaller number of gradient updates and without any additional parameters for the task of audio source separation. Extensive experimental results across datasets demonstrate the effectiveness of our proposed method. Md. Amirul Islam, Seyed Shahabeddin Nabavi, Irina Kezele, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005 |
WACV | 5 |
| 2024 | HARWE: A multi-modal large-scale dataset for context-aware human activity recognition in smart working environments
Alireza Esmaeilzehi, Ensieh Khazaei, Kai Wang 0068, Navjot Kaur Kalsi, Pai Chet Ng, Huan Liu 0014, Yuanhao Yu, Dimitrios Hatzinakos, Konstantinos N. Plataniotis |
Pattern Recognit. Lett. | 7 |
| 2023 | Meta-Auxiliary Learning for Future Depth Prediction in VideosabstractWe consider a new problem of future depth prediction in videos. Given a sequence of observed frames in a video, the goal is to predict the depth map of a future frame that has not been observed yet. Depth estimation plays a vital role for scene understanding and decision-making in intelligent systems. Predicting future depth maps can be valuable for autonomous vehicles to anticipate the behaviours of their surrounding objects. Our proposed model for this problem has a two-branch architecture. One branch is for the primary task of future depth prediction. The other branch is for an auxiliary task of image reconstruction. The auxiliary branch can act as a regularization. Inspired by some recent work on test-time adaption, we use the auxiliary task during testing to adapt the model to a specific test video. We also propose a novel meta-auxiliary learning that learns the model specifically for the purpose of effective test-time adaptation. Experimental results demonstrate that our proposed approach outperforms other alternative methods. Huan Liu 0014, Zhixiang Chi, Yuanhao Yu, Yang Wang 0003, Jun Chen 0005, Jin Tang 0005 |
WACV | 3 |
| 2023 | Test-Time Adaptation for Optical Flow Estimation Using Motion VectorsabstractDue to the prohibitive cost as well as technical challenges in annotating ground-truth optical flow for large-scale realistic video datasets, the existing deep learning models for optical flow estimation mostly rely on synthetic data for training, which in turn may lead to significant performance degradation under test-data distribution shift in real-world environments. In this work, we propose the methodology to tackle this important problem. We design a self-supervised learning task for adjusting the optical flow estimation model at test time. We exploit the fact that most videos are stored in compressed formats, from which compact information on motion, in the form of motion vectors and residuals, can be made readily available. We formulate the self-supervised task as motion vector prediction, and link this task to optical flow estimation. To the best of our knowledge, our Test-Time Adaption guided with Motion Vectors (TTA-MV), is the first work to perform such adaptation for optical flow. The experimental results demonstrate that TTA-MV can improve the generalization capability of various well-known deep learning methods for optical flow estimation, such as FlowNet, PWCNet, and RAFT. Seyed Mehdi Ayyoubzadeh, Irina Kezele, Yuanhao Yu, Xiaolin Wu 0001, Yang Wang 0003, Jin Tang 0005 |
IEEE Trans. Image Process. | 4 |
| 2023 | Stress Detection Through Wrist-Based Electrodermal Activity Monitoring and Machine LearningabstractStress is an inevitable part of modern life. While stress can negatively impact a person's life and health, positive and under-controlled stress can also enable people to generate creative solutions to problems encountered in their daily lives. Although it is hard to eliminate stress, we can learn to monitor and control its physical and psychological effects. It is essential to provide feasible and immediate solutions for more mental health counselling and support programs to help people relieve stress and improve their mental health. Popular wearable devices, such as smartwatches with several sensing capabilities, including physiological signal monitoring, can alleviate the problem. This work investigates the feasibility of using wrist-based electrodermal activity (EDA) signals collected from wearable devices to predict people's stress status and identify possible factors impacting stress classification accuracy. We use data collected from wrist-worn devices to examine the binary classification discriminating stress from non-stress. For efficient classification, five machine learning-based classifiers were examined. We explore the classification performance on four available EDA databases under different feature selections. According to the results, Support Vector Machine (SVM) outperforms the other machine learning approaches with an accuracy of 92.9 for stress prediction. Additionally, when the subject classification included gender information, the performance analysis showed significant differences between males and females. We further examine a multimodal approach for stress classifications. The results indicate that wearable devices with EDA sensors have a great potential to provide helpful insight for improved mental health monitoring. Lili Zhu, Petros Spachos, Pai Chet Ng, Yuanhao Yu, Yang Wang 0003, Konstantinos N. Plataniotis, Dimitrios Hatzinakos |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Fast Human Pose Estimation in Compressed VideosabstractCurrent approaches for human pose estimation in videos can be categorized into per-frame and warping-based methods. Both approaches have their pros and cons. For example, per-frame methods are generally more accurate, but they are often slow. Warping-based approaches are more efficient, but the performance is usually not good. To bridge the gap, in this paper, we propose a novel fast framework for human pose estimation to meet the real-time inference with controllable accuracy degradation in compressed video domain. Our approach takes advantage of the motion representation (called “motion vector”) that is readily available in a compressed video. Pose joints in a frame are obtained by directly warping the pose joints from the previous frame using the motion vectors. We also propose modules to correct possible errors introduced by the pose warping when needed. Extensive experimental results demonstrate the effectiveness of our proposed framework for accelerating the speed of top-down human pose estimation in videos. Huan Liu 0014, Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jun Chen 0005, Jin Tang 0005 |
IEEE Trans. Multim. | 5 |
| 2022 | MetaFSCIL: A Meta-Learning Approach for Few-Shot Class Incremental LearningabstractIn this paper, we tackle the problem of few-shot class incremental learning (FSCIL). FSCIL aims to incrementally learn new classes with only a few samples in each class. Most existing methods only consider the incremental steps at test time. The learning objective of these methods is often hand-engineered and is not directly tied to the objective (i.e. incrementally learning new classes) during testing. Those methods are sub-optimal due to the misalignment between the training objectives and what the methods are expected to do during evaluation. In this work, we proposed a bi-level optimization based on meta-learning to directly optimize the network to learn how to incrementally learn in the setting of FSCIL. Concretely, we propose to sample sequences of incremental tasks from base classes for training to simulate the evaluation protocol. For each task, the model is learned using a meta-objective such that it is capable to perform fast adaptation without forgetting. Furthermore, we propose a bi-directional guided modulation, which is learned to automatically modulate the activations to reduce catastrophic forgetting. Extensive experimental results demonstrate that the proposed method outperforms the baseline and achieves the state-of-the-art results on CIFARIOO, MiniImageNet, and CUB200 datasets. Zhixiang Chi, Li Gu, Huan Liu 0014, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005 |
CVPR | 5 |
| 2022 | Few-Shot Class-Incremental Learning via Entropy-Regularized Data-Free Replay
Huan Liu 0014, Li Gu, Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jun Chen 0005, Jin Tang 0005 |
ECCV (24) | 5 |
| 2022 | Hierarchical Deep Learning Model with Inertial and Physiological Sensors Fusion for Wearable-Based Human Activity RecognitionabstractThis paper presents a human activity recognition (HAR) system with wearable devices. While various approaches have been suggested for HAR, most of them focus on either 1) the inertial sensors to capture the physical movement or 2) subject-dependent evaluations that are less practical to real world cases. To this end, our work integrates sensing in-puts from physiological sensors to compensate the limitation of inertial sensors in capturing the human activities with less physical movements. Physiological sensors can capture physiological responses reflecting human behaviors in executing daily activities. To simulate a realistic application, three different evaluation scenarios are considered, namely All-access, Cross-subject and Cross-activity. Lastly, we propose a Hierarchical Deep Learning (HDL) model, which improves the accuracy and stability of HAR, compared to conventional models. Our proposed HDL with fusion of inertial and physiological sensing inputs achieves 97.16%, 92.23%, 90.18% average accuracy in All-access, Cross-subject, Cross-activity scenarios, which confirms the effectiveness of our approach. Dae Yon Hwang, Pai Chet Ng, Yuanhao Yu, Yang Wang 0003, Petros Spachos, Dimitrios Hatzinakos, Konstantinos N. Plataniotis |
ICASSP | 3 |
| 2022 | Feasibility Study of Stress Detection with Machine Learning through EDA from Wearable DevicesabstractThe recent pandemic has brought tremendous changes to everyone’s life, causing stress about losing loved ones, losing jobs, and having changes in sleep or eating habits. This study investigates the feasibility of utilizing Electrodermal Activity (EDA) collected from wearable devices to detect people’s stress. EDA can quantify the changes in sympathetic dynamics by measuring sweat produced by our sweat glands. Currently, the adoption of EDA sensors to commercially off-the-shelf smart-watches is still in the infancy stage, and only a few brands have the EDA sensors implemented into their smartwatch. To facilitate our feasibility study, we need the datasets that contain the EDA signals collected from wearable devices. This paper uses two publicly available datasets containing the EDA signals collected from research-grade wearable devices. We cast the stress detection problem as a binary classification problem and trained the classifiers with three popular machine learning methods: K-Nearest Neighbor, Logistic Regression, and Random Forests. According to experimental results, Random Forests achieves an accuracy of 85.7% to classify stress from non-stress status. The results verified that wearable devices with EDA sensors have the potential to predict stress status. Lili Zhu, Pai Chet Ng, Yuanhao Yu, Yang Wang 0003, Petros Spachos, Dimitrios Hatzinakos, Konstantinos N. Plataniotis |
ICC | 3 |
| 2022 | Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-ExpertsabstractIn this paper, we tackle the problem of domain shift. Most existing methods perform training on multiple source domains using a single model, and the same trained model is used on all unseen target domains. Such solutions are sub-optimal as each target domain exhibits its own specialty, which is not adapted. Furthermore, expecting single-model training to learn extensive knowledge from multiple source domains is counterintuitive. The model is more biased toward learning only domain-invariant features and may result in negative knowledge transfer. In this work, we propose a novel framework for unsupervised test-time adaptation, which is formulated as a knowledge distillation process to address domain shift. Specifically, we incorporate Mixture-of-Experts (MoE) as teachers, where each expert is separately trained on different source domains to maximize their specialty. Given a test-time target domain, a small set of unlabeled data is sampled to query the knowledge from MoE. As the source domains are correlated to the target domains, a transformer-based aggregator then combines the domain knowledge by examining the interconnection among them. The output is treated as a supervision signal to adapt a student prediction network toward the target domain. We further employ meta-learning to enforce the aggregator to distill positive knowledge and the student network to achieve fast adaptation. Extensive experiments demonstrate that the proposed method outperforms the state-of-the-art and validates the effectiveness of each proposed component. Our code is available at https://github.com/n3il666/Meta-DMoE. Tao Zhong 0003, Zhixiang Chi, Li Gu, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005 |
NeurIPS | 5 |
| 2021 | Test-Time Fast Adaptation for Dynamic Scene Deblurring via Meta-Auxiliary LearningabstractIn this paper, we tackle the problem of dynamic scene deblurring. Most existing deep end-to-end learning approaches adopt the same generic model for all unseen test images. These solutions are sub-optimal, as they fail to utilize the internal information within a specific image. On the other hand, a self-supervised approach, SelfDeblur, enables internal training within a test image from scratch, but it does not fully take advantage of large external datasets. In this work, we propose a novel self-supervised meta-auxiliary learning to improve the performance of deblurring by integrating both external and internal learning. Concretely, we build a self-supervised auxiliary reconstruction task that shares a portion of the network with the primary deblurring task. The two tasks are jointly trained on an external dataset. Furthermore, we propose a meta-auxiliary training scheme to further optimize the pretrained model as a base learner, which is applicable for fast adaptation at test time. During training, the performance of both tasks is coupled. Therefore, we are able to exploit the internal information at test time via the auxiliary task to enhance the performance of deblurring. Extensive experimental results across evaluation datasets demonstrate the effectiveness of test-time adaptation of the proposed method. Zhixiang Chi, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005 |
CVPR | 3 |
| 2021 | Toward Personalized Emotion Recognition: A Face Recognition Based Attention Method for Facial Emotion RecognitionabstractThis paper aims to address the subject-dependent challenge of the facial emotion recognition (FER) task. To accomplish this, we propose a novel face recognition based attention FER (FRA-FER) framework which propagates subtle face recognition (FR) features through the FER network. Particularly, first a spatial attention map from the feature maps of an FR convolutional neural network (CNN) is created and then it is fused into the FER-CNN. By doing this FR feature propagation, the FER network is personalized as it takes the advantage of the FR features learned from large-scale face recognition datasets. Experiments on the two challenging datasets AffectNet and AFEW demonstrate the superiority of our proposed FRA-FER network to the state-of-the-art work. Mostafa Shahabinejad, Yang Wang 0003, Yuanhao Yu, Jin Tang 0005 |
FG | 3 |
| 2018 | Robust discriminative tracking via structured prior regularization
Yuanhao Yu, Qingsong Wu, Thia Kirubarajan, Yasuo Uehara |
Image Vis. Comput. | 1 |