Yu Guan 0001

dblp:86/6151-1 · DBLP profile ↗
← Back
26ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-1283-3806ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Towards Mitigating Modality Bias in Vision-Language Models for Temporal Action Localization
abstract
Jiaqi Li, Guangming Wang, Shuntian Zheng, Minzhe Ni, Xiaoman Lu, Guanghui Ye, Yu Guan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiaqi Li 0008, Guangming Wang 0001, Shuntian Zheng, Minzhe Ni, Xiaoman Lu, Guanghui Ye, Yu Guan 0001
ACL (1)7
2025 VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video Questions
abstract
Complex video question-answering (VQA) requires in-depth understanding of video contents including object and action recognition as well as video classification and summarization, which exhibits great potential in emerging applications in education and entertainment, etc. Multimodal large language models (MLLMs) may accomplish this task by grasping the intention of a question and decomposing it to a series of visual recognition sub-tasks to find out the answer with the help of an agent. To tackle this task, we first collect a new dedicated Complex VQA dataset named CVQA and then propose VQAGuider, an innovative framework planning a few atomic visual recognition tools by video-related API matching. VQAGuider facilitates a deep engagement with video content and precise responses to complex video-related questions by MLLMs, which is beyond aligning visual and language features for simple VQA tasks. Our experiments demonstrate VQAGuider is capable of navigating the complex VQA tasks by MLLMs and improves the accuracy by 29.6% and 17.2% on CVQA and the existing VQA datasets, respectively, highlighting its potential in advancing MLLMs’s capabilities in video understanding.
Yuyan Chen, Jiyuan Jia, Yu Guan 0001, Ming Yang 0007, Qingpei Guo
ACL (1)5
2025 DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
abstract
The challenge in LLM-based video understanding lies in preserving visual and semantic information in long videos while maintaining a memory-affordable token count. However, redundancy and correspondence in videos have hindered the performance potential of existing methods. Through statistical learning on current datasets, we observe that redundancy occurs in both repeated and answer-irrelevant frames, and the corresponding frames vary with different questions. This suggests the possibility of adopting dynamic encoding to balance detailed video information preservation with token budget reduction. To this end, we propose a dynamic cooperative network, DynFocus, for memory-efficient video encoding in this paper. Specifically, i) a Dynamic Event Prototype Estimation (DPE) module to dynamically select meaningful frames for question answering; (ii) a Compact Cooperative Encoding (CCE) module that encodes meaningful frames with detailed visual appearance and the remaining frames with sketchy perception separately. We evaluate our method on five publicly available benchmarks, and experimental results consistently demonstrate that our method achieves competitive performance. Code is available at https://github.com/Simon98-AI/ DynFocus
Qingpei Guo, Liyuan Pan, Liu Liu 0009, Yu Guan 0001, Ming Yang 0007
CVPR5
2025 Engage for All: Making Ordinary Image Descriptions Appealing Again!
Yuyan Chen, Jinghan Cao, Yu Guan 0001, Ming-Hsuan Yang 0001, Qing Guo 0005
ICCV5
2025 DSleepNet: Disentanglement Learning for Personal Attribute-Agnostic Three-Stage Sleep Classification Using Wearable Sensing Data
abstract
Long-term non-invasive sleep stage monitoring is instrumental in comprehending the progression of sleep disorders, cardiovascular diseases, and the interplay between sleep, type 2 diabetes, and neurodegenerative diseases. However, the conventional deep learning approach is susceptible to personal attributes (PAs) such as age, Body Mass Index, and severity of sleep apnea existing in the training dataset, potentially hindering its generalisation capacity to unseen cohorts. This paper introduces DSleepNet, a novel approach that disentangles the feature space into PA-specific and PA-agnostic components using two probabilistic encoders. The PA-agnostic features, designed to remain unaffected by personal attributes, outperformed the baseline CNN, improving the mean F1 score by up to 8.7% (baseline: 60.3) and Cohen's Kappa by 4.7% (baseline: 55.5), especially in reducing the impact of sleep apnea. DSleepNet functions without the need for target cohort data during training. It operates without the need to acquire PA data during inference, nor does it require fine-tuning. A novel Independent Excitation mechanism is incorporated into the latent feature space to remove correlations between the two types of features. Comprehensive testing in various PA settings has demonstrated its efficacy in improving the model's robustness.
Bing Zhai, Haoran Duan 0001, Yu Guan 0001, Huy Phan, Wai Lok Woo
IEEE J. Biomed. Health Informatics3
2023 Part-aware Prototypical Graph Network for One-shot Skeleton-based Action Recognition
abstract
In this paper, we study the problem of one-shot skeleton-based action recognition, which poses unique challenges in learning transferable representation from base classes to novel classes, particularly for fine-grained actions. Existing meta-learning frameworks typically rely on the body-level representations in spatial dimension, which limits the generalisation to capture subtle visual differences in the fine-grained label space. To overcome the above limitation, we propose a part-aware prototypical representation for one-shot skeleton-based action recognition. Our method captures skeleton motion patterns at two distinctive spatial levels, one for global contexts among all body joints, referred to as body level, and the other attends to local spatial regions of body parts, referred to as the part level. We also devise a class-agnostic attention mechanism to highlight important parts for each action class. Specifically, we develop a part-aware prototypical graph network consisting of three modules: a cascaded embedding module for our dual-level modelling, an attention-based part fusion module to fuse parts and generate part-aware prototypes, and a matching module to perform classification with the part-aware representations. We demonstrate the effectiveness of our method on two public skeleton-based action recognition datasets: NTU RGB+D 120 and NW-UCLA.
Tailin Chen, Desen Zhou, Jian Wang 0066, Qian He 0001, Chuanyang Hu, Errui Ding, Yu Guan 0001, Xuming He 0001
FG8
2023 Designing compact features for remote stroke rehabilitation monitoring using wearable accelerometers
abstract
Abstract Stroke is known as a major global health problem, and for stroke survivors it is key to monitor the recovery levels. However, traditional stroke rehabilitation assessment methods (such as the popular clinical assessment) can be subjective and expensive, and it is also less convenient for patients to visit clinics in a high frequency. To address this issue, in this work based on wearable sensing and machine learning techniques, we develop an automated system that can predict the assessment score in an objective manner. With wrist-worn sensors, accelerometer data is collected from 59 stroke survivors in free-living environments for a duration of 8 weeks, and we map the week-wise accelerometer data (3 days per week) to the assessment score by developing signal processing and predictive model pipeline. To achieve this, we propose two types of new features, which can encode the rehabilitation information from both paralysed and non-paralysed sides while suppressing the high-level noises such as irrelevant daily activities. Based on the proposed features, we further develop the longitudinal mixed-effects model with Gaussian process prior (LMGP), which can model the random effects caused by different subjects and time slots (during the 8 weeks). Comprehensive experiments are conducted to evaluate our system on both acute and chronic patients, and the promising results suggest its effectiveness.
Yu Guan 0001, Jian Qing Shi, Xiu-Li Du, Janet A. Eyre
CCF Trans. Pervasive Comput. Interact.2
2022 Action Quality Assessment with Temporal Parsing Transformer
Yang Bai 0011, Desen Zhou, Songyang Zhang 0001, Jian Wang 0066, Errui Ding, Yu Guan 0001, Yang Long 0001, Jingdong Wang 0001
ECCV (4)6
2021 Discriminative Latent Semantic Graph for Video Captioning
abstract
Video captioning aims to automatically generate natural language sentences that can describe the visual contents of a given video. Existing generative models like encoder-decoder frameworks cannot explicitly explore the object-level interactions and frame-level information from complex spatio-temporal data to generate semantic-rich captions. Our main contribution is to identify three key problems in a joint framework for future video summarization tasks. 1) Enhanced Object Proposal: we propose a novel Conditional Graph that can fuse spatio-temporal information into latent object proposal. 2) Visual Knowledge: Latent Proposal Aggregation is proposed to dynamically extract visual words with higher semantic levels. 3) Sentence Validation: A novel Discriminative Language Validator is proposed to verify generated captions so that key semantic concepts can be effectively preserved. Our experiments on two public datasets (MVSD and MSR-VTT) manifest significant improvements over state-of-the-art approaches on all metrics, especially for BLEU-4 and CIDEr. Our code is available at https://github.com/baiyang4/D-LSG-Video-Caption.
Yang Bai 0011, Junyan Wang 0001, Yang Long 0001, Bingzhang Hu, Yang Song 0001, Maurice Pagnucco, Yu Guan 0001
ACM Multimedia7
2021 Learning Multi-Granular Spatio-Temporal Graph Network for Skeleton-based Action Recognition
abstract
The task of skeleton-based action recognition remains a core challenge in human-centred scene understanding due to the multiple granularities and large variation in human motion. Existing approaches typically employ a single neural representation for different motion patterns, which has difficulty in capturing fine-grained action classes given limited training data. To address the aforementioned problems, we propose a novel multi-granular spatio-temporal graph network for skeleton-based action classification that jointly models the coarse- and fine-grained skeleton motion patterns. To this end, we develop a dual-head graph network consisting of two interleaved branches, which enables us to extract features at two spatio-temporal resolutions in an effective and efficient manner. Moreover, our network utilises a cross-head communication strategy to mutually enhance the representations of both heads. We conducted extensive experiments on three large-scale datasets, namely NTU RGB+D 60, NTU RGB+D 120, and Kinetics-Skeleton, and achieves the state-of-the-art performance on all the benchmarks, which validates the effectiveness of our method1.
Tailin Chen, Desen Zhou, Jian Wang 0066, Yu Guan 0001, Xuming He 0001, Errui Ding
ACM Multimedia5
2021 Invariant Deep Compressible Covariance Pooling for Aerial Scene Categorization
abstract
Learning discriminative and invariant feature representation is the key to visual image categorization. In this article, we propose a novel invariant deep compressible covariance pooling (IDCCP) to solve nuisance variations in aerial scene categorization. We consider transforming the input image according to a finite transformation group that consists of multiple confounding orthogonal matrices, such as the D4 group. Then, we adopt a Siamese-style network to transfer the group structure to the representation space, where we can derive a trivial representation that is invariant under the group action. The linear classifier trained with trivial representation will also be possessed with invariance. To further improve the discriminative power of representation, we extend the representation to the tensor space while imposing orthogonal constraints on the transformation matrix to effectively reduce feature dimensions. We conduct extensive experiments on the publicly released aerial scene image data sets and demonstrate the superiority of this method compared with state-of-the-art methods. In particular, with using ResNet architecture, our IDCCP model can reduce the dimension of the tensor representation by about 98% without sacrificing accuracy (i.e., < 0.5%).
Yi Ren 0001, Gerard P. Parr, Yu Guan 0001, Ling Shao 0001
IEEE Trans. Geosci. Remote. Sens.4
2020 Query Twice: Dual Mixture Attention Meta Learning for Video Summarization
abstract
Video summarization aims to select representative frames to retain high-level information, which is usually solved by predicting the segment-wise importance score via a softmax function. However, softmax function suffers in retaining high-rank representations for complex visual or sequential information, which is known as the Softmax Bottleneck problem. In this paper, we propose a novel framework named Dual Mixture Attention (DMASum) model with Meta Learning for video summarization that tackles the softmax bottleneck problem, where the Mixture of Attention layer (MoA) effectively increases the model capacity by employing twice self-query attention that can capture the second-order changes in addition to the initial query-key attention, and a novel Single Frame Meta Learning rule is then introduced to achieve more generalization to small datasets with limited training sources. Furthermore, the DMASum significantly exploits both visual and sequential attention that connects local key-frame and global attention in an accumulative way. We adopt the new evaluation protocol on two public datasets, SumMe, and TVSum. Both qualitative and quantitative experiments manifest significant improvements over the state-of-the-art methods.
Junyan Wang 0001, Yang Bai 0011, Yang Long 0001, Bingzhang Hu, Zhenhua Chai, Yu Guan 0001, Xiaolin Wei
ACM Multimedia6
2020 Dual reference age synthesis
Yuan Zhou 0023, Bingzhang Hu, Jun He 0006, Yu Guan 0001, Ling Shao 0001
Neurocomputing4
2020 Multi-Granularity Canonical Appearance Pooling for Remote Sensing Scene Classification
abstract
Recognising remote sensing scene images remains challenging due to large visual-semantic discrepancies. These mainly arise due to the lack of detailed annotations that can be employed to align pixel-level representations with high-level semantic labels. As the tagging process is labour-intensive and subjective, we hereby propose a novel Multi-Granularity Canonical Appearance Pooling (MG-CAP) to automatically capture the latent ontological structure of remote sensing datasets. We design a granular framework that allows progressively cropping the input image to learn multi-grained features. For each specific granularity, we discover the canonical appearance from a set of pre-defined transformations and learn the corresponding CNN features through a maxout-based Siamese style architecture. Then, we replace the standard CNN features with Gaussian covariance matrices and adopt the proper matrix normalisations for improving the discriminative power of features. Besides, we provide a stable solution for training the eigenvalue-decomposition function (EIG) in a GPU and demonstrate the corresponding back-propagation using matrix calculus. Extensive experiments have shown that our framework can achieve promising results in public remote sensing scene datasets.
Yu Guan 0001, Ling Shao 0001
IEEE Trans. Image Process.2
2019 Order Matters: Shuffling Sequence Generation for Video Prediction
Junyan Wang 0001, Bingzhang Hu, Yang Long 0001, Yu Guan 0001
BMVC4
2019 Unsupervised Machine Learning for Card Payment Fraud Detection
Mario Parreño Centeno, Mohammed Aamir Ali, Yu Guan 0001, Aad P. A. van Moorsel
CRiSIS3
2019 Remote monitoring of stroke patients' rehabilitation using wearable accelerometers
abstract
In this work, we outline an automated system for continuous assessment of patient recovery levels generically in movement-related disorders using wrist-worn accelerometers, with visualisation of aspects of recovery trends. We demonstrate our model on assessment of recoveries from hemiparetic stroke.
Shane Halloran, Yu Guan 0001, Jian Qing Shi, Janet A. Eyre
UbiComp3
2019 Generic compact representation through visual-semantic ambiguity removal
Yang Long 0001, Yu Guan 0001, Ling Shao 0001
Pattern Recognit. Lett.2
2019 Triple Verification Network for Generalized Zero-Shot Learning
abstract
Conventional Zero-shot Learning approaches often suffer from severe performance degradation in the Generalised Zero-shot Learning (GZSL) scenario, i.e. to recognise test images that are from both seen and unseen classes. This paper studies the Class-level Over-fitting (CO) and empirically shows its effects to GZSL. We then address ZSL as a Triple Verification problem and propose a unified optimisation of regression and compatibility functions, i.e. two main streams of existing ZSL approaches. The complementary losses mutually regularise the same model to mitigate the CO problem. Furthermore, we implement a deep extension paradigm to linear models and significantly outperforms state-of-the-art methods in both GZSL and ZSL scenarios on the four standard benchmarks.
Haofeng Zhang 0001, Yang Long 0001, Yu Guan 0001, Ling Shao 0001
IEEE Trans. Image Process.3
2018 Towards Universal Representation for Unseen Action Recognition
abstract
Unseen Action Recognition (UAR) aims to recognise novel action categories without training examples. While previous methods focus on inner-dataset seen/unseen splits, this paper proposes a pipeline using a large-scale training source to achieve a Universal Representation (UR) that can generalise to a more realistic Cross-Dataset UAR (CDUAR) scenario. We first address UAR as a Generalised Multiple-Instance Learning (GMIL) problem and discover 'building-blocks' from the large-scale ActivityNet dataset using distribution kernels. Essential visual and semantic components are preserved in a shared space to achieve the UR that can efficiently generalise to new datasets. Predicted UR exemplars can be improved by a simple semantic adaptation, and then an unseen action can be directly recognised using UR during the test. Without further training, extensive experiments manifest significant improvements over the UCF101 and HMDB51 benchmarks.
Yi Zhu 0001, Yang Long 0001, Yu Guan 0001, Shawn D. Newsam, Ling Shao 0001
CVPR3
2018 Remote Cloud-Based Automated Stroke Rehabilitation Assessment Using Wearables
abstract
We outline a system enabling accurate remote assessment of stroke rehabilitation levels using wrist worn accelerometer time series data. The system is built based on features generated from clustering models across sliding windows in the data and makes use of computation in the cloud. Predictive models are built using advanced machine learning techniques.
Shane Halloran, Jian Qing Shi, Yu Guan 0001, Michael Dunne-Willows, Janet A. Eyre
eScience3
2018 Inference of a compact representation of sensor fingerprint for source camera identification
abstract
Sensor pattern noise (SPN) is an inherent fingerprint of imaging devices, which provides an effective way for source camera identification (SCI). Although SPNs extracted from large image blocks usually yield high identification accuracy, their high dimensionality would incur a high computational cost in the matching stage, consequently hindering many applications that require efficient camera matchings. In this work, we employ and evaluate the concept of principal component analysis (PCA) de-noising in SCI tasks. Based on this concept, we present a framework that formulates a compact SPN representation. To enhance the de-noising effect, we introduce a training set construction procedure that minimizes the impact of various interfering artifacts, which is especially useful in some challenging cases, e.g., when only textured reference images are available. To further boost the SCI performance, a novel approach based on linear discriminant analysis (LDA) is adopted to extract more discriminant SPN features. To evaluate our methods, extensive experiments are conducted on the Dresden image database. The results indicate that the proposed framework can serve as an effective post-processing procedure, which not only boosts the performance, but also greatly reduces the computational cost in the matching phase.
Ruizhe Li 0002, Chang-Tsun Li, Yu Guan 0001
Pattern Recognit.3
2016 Enhanced SVD for Collaborative Filtering
Xin Guan 0002, Chang-Tsun Li, Yu Guan 0001
PAKDD (2)3
2015 A compact representation of sensor fingerprint for camera identification and fingerprint matching
abstract
Sensor Pattern Noise (SPN) has been proved as an effective fingerprint of imaging devices to link pictures to the cameras that acquired them. In practice, forensic investigators usually extract this camera fingerprint from large image block to improve the matching accuracy because large image blocks tend to contain more SPN information. As a result, camera fingerprints usually have a very high dimensionality. However, the high dimensionality of fingerprint will incur a costly computation in the matching phase, thus hindering many interesting applications which require an efficient real-time camera matching. To solve this problem, an effective feature extraction method based on PCA and LDA is proposed in this work to compress the dimensionality of camera fingerprint. Our experimental results show that the proposed feature extraction algorithm could greatly reduce the size of fingerprint and enhance the performance in term of Receiver Operating Characteristic (ROC) curve of several existing methods.
Ruizhe Li 0002, Chang-Tsun Li, Yu Guan 0001
ICASSP3
2015 Incremental update of feature extractor for camera identification
abstract
Sensor Pattern Noise (SPN) is an inherent fingerprint of imaging devices, which has been widely used in the tasks of digital camera identification, image classification and forgery detection. In our previous work, a feature extraction method based on PCA denoising concept was applied to extract a set of principal components from the original noise residual. However, this algorithm is inefficient when query cameras are continuously received. To solve this problem, we propose an extension based on Candid Covariance-free Incremental PCA (CCIPCA) and two modifications to incrementally update the feature extractor according to the received cameras. Experimental results show that the PCA and CCIPCA based features both outperform their original features on the ROC performance, and CCIPCA is more efficient on camera updating.
Ruizhe Li 0002, Chang-Tsun Li, Yu Guan 0001
ICIP3
2015 On Reducing the Effect of Covariate Factors in Gait Recognition: A Classifier Ensemble Method
abstract
Robust human gait recognition is challenging because of the presence of covariate factors such as carrying condition, clothing, walking surface, etc. In this paper, we model the effect of covariates as an unknown partial feature corruption problem. Since the locations of corruptions may differ for different query gaits, relevant features may become irrelevant when walking condition changes. In this case, it is difficult to train one fixed classifier that is robust to a large number of different covariates. To tackle this problem, we propose a classifier ensemble method based on the random subspace Method (RSM) and majority voting (MV). Its theoretical basis suggests it is insensitive to locations of corrupted features, and thus can generalize well to a large number of covariates. We also extend this method by proposing two strategies, i.e, local enhancing (LE) and hybrid decision-level fusion (HDF) to suppress the ratio of false votes to true votes (before MV). The performance of our approach is competitive against the most challenging covariates like clothing, walking surface, and elapsed time. We evaluate our method on the USF dataset and OU-ISIR-B dataset, and it has much higher performance than other state-of-the-art algorithms.
Yu Guan 0001, Chang-Tsun Li, Fabio Roli
IEEE Trans. Pattern Anal. Mach. Intell.1