Kun Hu 0008

dblp:61/9177-8 · DBLP profile ↗
← Back
43ranked-venue papers
8as first author
39since 2021 · last 2026
0000-0002-6891-8059ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 2 first-author · 30 since 2021Artificial intelligence and machine learning · 24 · 4 first-author · 22 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 2 since 2021
YearPublicationVenuePosition
2026 TWINFUZZ: Dual-Model Fuzzing for Robustness Generalization in Deep Learning
abstract
Deep learning (DL) models are increasingly deployed in safety-critical applications such as face recognition, autonomous driving, and medical diagnosis. Despite their impressive accuracy, they remain vulnerable to adversarial examples - subtle perturbations that can cause incorrect predictions, i.e., the robustness issues. While adversarial training improves robustness against known attacks, it often fails to generalize to unseen or stronger threats, revealing a critical gap in robustness generalization. In this work, we propose a dual-model fuzzing framework to enhance generalized robustness in DL models. Central to our method is a lightweight metric, the Lagrangian Information Bottleneck (LIB), which guides entropy-based mutation toward semantically meaningful and high-risk regions of the input space. The executor uses a resistant model and a more error-prone vulnerable model; their prediction consistency forms the basis of agreement mining, a label-free oracle for isolating decision-boundary samples. To ensure fuzzing effectiveness, we further introduce a task-driven seed selection strategy (e.g., SSIM for vision) that filters out low-quality inputs. We implement a prototype, TWINFUZZ, and evaluate it on six benchmark datasets and nine DL models. Compared with state-of-the-art testing approaches, TWINFUZZ achieves superior improvements in both training-specific and generalized robustness.
Enze Dai, Wentao Mo, Kun Hu 0008, Xiaogang Zhu 0001, Xi Xiao 0001, Sheng Wen, Shaohua Wang 0002, Yang Xiang 0001
AAAI3
2026 Pb4U-GNet: Resolution-Adaptive Garment Simulation via Propagation-before-Update Graph Network
abstract
Garment simulation is fundamental to various applications in computer vision and graphics, from virtual try-on to digital human modelling. However, conventional physics-based methods remain computationally expensive, hindering their application in time-sensitive scenarios. While graph neural networks (GNNs) offer promising acceleration, existing approaches exhibit poor cross-resolution generalisation, demonstrating significant performance degradation on higher-resolution meshes beyond the training distribution. This stems from two key factors: (1) existing GNNs employ fixed message-passing depth that fails to adapt information aggregation to mesh density variation, and (2) vertex-wise displacement magnitudes are inherently resolution-dependent in garment simulation. To address these issues, we introduce Propagation-before-Update Graph Network (Pb4U-GNet), a resolution-adaptive framework that decouples message propagation from feature updates. Pb4U-GNet incorporates two key mechanisms: (1) dynamic propagation depth control, adjusting message-passing iterations based on mesh resolution, and (2) geometry-aware update scaling, which scales predictions according to local mesh characteristics. Extensive experiments show that even trained solely on low-resolution meshes, Pb4U-GNet exhibits strong generalisability across diverse mesh resolutions, addressing a fundamental challenge in neural garment simulation.
Aoran Liu, Kun Hu 0008, Clinton Mo, Qiuxia Wu, Wenxiong Kang, Zhiyong Wang 0001
AAAI2
2026 DuoCast: Duo-Probabilistic Diffusion for Precipitation Nowcasting
abstract
Accurate short-term precipitation forecasting is critical for weather-sensitive decision-making in agriculture, transportation, and disaster response. Existing deep learning approaches often struggle to balance global structural consistency with local detail preservation, especially under complex meteorological conditions. We propose DuoCast, a dual-diffusion framework that decomposes precipitation forecasting into low- and high-frequency components modeled in orthogonal latent subspaces. We theoretically prove that this frequency decomposition reduces prediction error compared to conventional single branch U-Net diffusion models. In DuoCast, the low-frequency model captures large-scale trends via convolutional encoders conditioned on weather front dynamics, while the high-frequency model refines fine-scale variability using a self-attention-based architecture. Experiments on four benchmark radar datasets show that DuoCast consistently outperforms state-of-the-art baselines, achieving superior accuracy in both spatial detail and temporal evolution.
Penghui Wen, Mengwei He, Patrick Filippi, Thomas F. A. Bishop, Zhiyong Wang 0001, Kun Hu 0008
AAAI8
2026 KANMultiSign: Multi-scale sequence-based pose animation from sign language notation with Kolmogorov-Arnold networks
abstract
Sign language production from symbolic notation offers a scalable route to accessible sign animation. We present KANMultiSign, a multi-scale sequence generator that translates HamNoSys notation into two-dimensional human pose sequences. Our framework makes two complementary contributions. First, we introduce a coarse-to-fine generation strategy with multi-scale supervision: the model is first guided by an intermediate body–hand–face scaffold to encourage global structural coherence, and then refines fine-grained hand articulation to improve finger-level detail. Second, we investigate integrating Kolmogorov–Arnold Network modules into a Transformer backbone, using learnable univariate function primitives to model the highly non-linear mapping from discrete phonological symbols to continuous body kinematics with a compact parameterization. Experiments on multiple public corpora spanning Polish, German, Greek, and French sign languages show consistent reductions in dynamic time warping based joint error compared with a strong notation-to-pose baseline, while using substantially fewer parameters. Controlled ablations further indicate that KAN-based variants substantially reduce parameter count while maintaining competitive performance when coupled with multi-scale supervision, rather than serving as the main driver of accuracy gains. These findings position multi-scale supervision as the key mechanism for improving notation-conditioned pose generation, with KAN offering a compact alternative for efficient modeling. Our code will be publicly available.
Guanyi Du, Lintao Wang 0002, Kun Hu 0008
Neurocomputing3
2026 Keyframe selection from motion capture data with dual-agent reinforcement learning
abstract
Animation production workflows centred around motion capture techniques require animators to edit motions based on a set of keyframes. However, most existing keyframe selection methods are optimisation-based, which suffer from the issues of flexibility and efficiency. In this paper, a novel deep reinforcement learning method with dual agents are proposed for unsupervised keyframe selection. First, an S-Agent and an R-Agent evaluate the actions of selection and refinement, respectively. A deep spatio-temporal network, namely graph keyframe evaluation network (GKEN), is proposed for the agents. Then, an animation specified reward is devised based on reconstruction, which fulfills three important properties of the animation workflow: incremental reward, order insensitivity and non-diminishing returns. During the inference, it is no longer necessary to compute the reconstruction, which significantly decreases the run-time latency. Experiments on the CMU MoCap dataset demonstrate the efficiency of the proposed method without clearly compromising the effectiveness compared with the state-of-the-art methods. • A deep reinforcement learning with dual-agent to identify motion keyframes. • A spatio-temporal deep agent with graph convolutions and transformers. • Comprehensive experiments and human demonstrations for MoCap keyframing.
Kun Hu 0008, Clinton Mo, Mingyang Ma 0004, Shaohui Mei, Zhiyong Wang 0001
Pattern Recognit.1
2025 RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation Learning
abstract
Masked point modeling methods have recently achieved great success in self-supervised learning for point cloud data. However, these methods are sensitive to rotations and often exhibit sharp performance drops when encountering rotational variations. In this paper, we propose a novel Rotation-Invariant Masked AutoEncoders (RI-MAE) to address two major challenges: 1) achieving rotation-invariant latent representations, and 2) facilitating self-supervised reconstruction in a rotation-invariant manner. For the first challenge, we introduce RI-Transformer, which features disentangled geometry content, rotation-invariant relative orientation and position embedding mechanisms for constructing rotation-invariant point cloud latent space. For the second challenge, a novel dual-branch student-teacher architecture is devised. It enables the self-supervised learning via the reconstruction of masked patches within the learned rotation-invariant latent space. Each branch is based on an RI-Transformer, and they are connected with an additional RI-Transformer predictor. The teacher encodes all point patches, while the student solely encodes unmasked ones. Finally, the predictor predicts the latent features of the masked patches using the output latent embeddings from the student, supervised by the outputs from the teacher. Extensive experiments demonstrate that our method is robust to rotations, achieving the state-of-the-art performance on various downstream tasks.
Kunming Su, Qiuxia Wu, Panpan Cai, Xiaogang Zhu 0001, Xuequan Lu, Zhiyong Wang 0001, Kun Hu 0008
AAAI7
2025 DC-PCN: Point Cloud Completion Network with Dual-Codebook Guided Quantization
abstract
Point cloud completion aims to reconstruct complete 3D shapes from partial 3D point clouds. With advancements in deep learning techniques, various methods for point cloud completion have been developed. Despite achieving encouraging results, a significant issue remains: these methods often overlook the variability in point clouds sampled from a single 3D object surface. This variability can lead to ambiguity and hinder the achievement of more precise completion results. Therefore, in this study, we introduce a novel point cloud completion network, namely Dual-Codebook Point Completion Network (DC-PCN), following an encder-decoder pipeline. The primary objective of DC-PCN is to formulate a singular representation of sampled point clouds originating from the same 3D surface. DC-PCN introduces a dual-codebook design to quantize point-cloud representations from a multilevel perspective. It consists of an encoder-codebook and a decoder-codebook, designed to capture distinct point cloud patterns at shallow and deep levels. Additionally, to enhance the information flow between these two codebooks, we devise an information exchange mechanism. This approach ensures that crucial features and patterns from both shallow and deep levels are effectively utilized for completion. Extensive experiments on the PCN, ShapeNet_Part, and ShapeNet34 datasets demonstrate the state-of-the-art performance of our method.
Qiuxia Wu, Kunming Su, Zhiyong Wang 0001, Kun Hu 0008
AAAI5
2025 B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
abstract
Recently, Vision Large Language Models (VLLMs) integrated with vision encoders have shown promising performance in vision understanding. The key of VLLMs is to encode visual content into sequences of visual tokens, enabling VLLMs to simultaneously process both visual and textual content. However, understanding videos, especially long videos, remain a challenge to VLLMs as the number of visual tokens grows rapidly when encoding videos, resulting in the risk of exceeding the context window of VLLMs and introducing heavy computation burden. To restrict the number of visual tokens, existing VLLMs either: (1) uniformly downsample videos into a fixed number of frames or (2) reducing the number of visual tokens encoded from each frame. We argue the former solution neglects the rich temporal cue in videos and the later overlooks the spatial details in each frame. In this work, we present Balanced-VLLM (B-VLLM): a novel VLLM framework that aims to effectively leverage task relevant spatio-temporal cues while restricting the number of visual tokens under the VLLM context window length. At the core of our method, we devise a text-conditioned adaptive frame selection module to identify frames relevant to the visual understanding task. The selected frames are then de-duplicated using a temporal frame token merging technique. The visual tokens of the selected frames are processed through a spatial token sampling module and an optional spatial token merging strategy to achieve precise control over the token count. Experimental results show that B-VLLM is effective in balancing the number of frames and visual tokens in video understanding, yielding superior performance on various video understanding benchmarks. Our code is available at https://github.com/zhuqiangLu/B-VLLM.
Zhuqiang Lu, Zhenfei Yin, Mengwei He, Zhihui Wang 0001, Zhiyong Wang 0001, Kun Hu 0008
ICCV7
2025 PUMPS: Skeleton-Agnostic Point-Based Universal Motion Pre-Training for Synthesis in Human Motion Tasks
Clinton Mo, Kun Hu 0008, Chengjiang Long, Dong Yuan 0001, Wan-Chi Siu, Zhiyong Wang 0001
ICCV2
2025 Extended Short- and Long-Range Mesh Learning for Fast and Generalised Garment Simulation
abstract
3D garment simulation is a critical component for producing cloth-based graphics. Recent advancements in graph neural networks (GNNs) offer a promising approach for efficient garment simulation. However, GNNs require extensive message-passing to propagate information such as physical forces and maintain contact awareness across the entire garment mesh, which becomes computationally inefficient at higher resolutions. To address this, we devise a novel GNN-based mesh learning framework with two key components to extend the message-passing range with minimal overhead, namely the Laplacian-Smoothed Dual Message-Passing (LSDMP) and the Geodesic Self-Attention (GSA) modules. LSDMP enhances message-passing with a Laplacian features smoothing process, which efficiently propagates the impact of each vertex to nearby vertices. Concurrently, GSA introduces geodesic distance embeddings to represent the spatial relationship between vertices and utilises attention mechanisms to capture global mesh information. The two modules operate in parallel to ensure both short- and long-range mesh modelling. Extensive experiments demonstrate the state-of-the-art performance of our method, requiring fewer layers and lower inference latency.1
Aoran Liu, Kun Hu 0008, Clinton Mo, ChangYang Li, Zhiyong Wang 0001
ICME2
2025 PTSR: A Unified Patch Tokenization, Selection and Representation Framework for Efficient Micro-expression Recognition
abstract
Micro-expression recognition is a challenging task of identifying hidden emotion, as micro-expressions have brief durations and involve small-scale facial muscle movements. Although deep learning-based methods, especially transformer-based methods, have achieved impressive performance in this task, these methods exhibit high computational complexity and struggle to learn effective representations in the context of typically small-scale micro-expression datasets, due to the excess of tokens in the multi-head self-attention. Moreover, most existing methods do not differentiate the importance of local features, especially in micro-expression recognition with subtle changes. Therefore, we propose a novel unified Patch Tokenization, Selection and Representation framework (PTSR) with vision Transformer for micro-expression recognition. Specifically, PTSR first presents a dual norm shifted patch tokenization (DNSPT) module to learn spatial relations between neighboring pixels of the face region, which is implemented by elaborating spatial transformation and dual norm projection. Then, we employ a local-global attention module (LAM) to extract the local-global image feature, incorporating a dynamic token selection module (DTSM) to select important patches/tokens, thereby capturing more discriminative representations for the input clip. Extensive experiments are conducted on 4 widely used public datasets, i.e., CASME II, SAMM, SMIC, CAS(ME)3, and the experimental results indicate that our method can achieve clear performance improvements over the state-of-the-art methods, such as 8.37% improvement on the CAS(ME)3 dataset in terms of UF1 and 3.1% improvement on the SMIC dataset in terms of UAR metric.
Liangyu Fu, Junbo Wang 0003, Qiangguo Jin, Yining Zhu, Hongsong Wang 0001, Kun Hu 0008
ICMR8
2025 MirrorDiff: Learning Mirror Diffusion for Image Captioning via Regeneration
abstract
Recently, diffusion models which have achieved promising progress in text-to-image generation generally have also been generally explored for image captioning. However, these diffusion-based image captioning methods usually suffer from semantic inconsistency between image content and textual description, thus producing lagging results compared with Auto-Regressive (AR) ones. To this end, in this paper, we propose a novel dual diffusion-based framework namely MirrorDiff, to achieve semantic consistency with a symmetric image-to-text-to-image generation model, which acts like a mirror that maps the original input image into a regenerated image via the generated caption. Specifically, it first utilizes both pre-trained image encoder and text encoder to obtain image representation and textual representation respectively, then forwards the image representation and the noisy textual representation into a continuous diffusion model to output an intermediate sentence. To semantically align the intermediate sentence with the input image, a diffusion-based visual regenerator is employed to regenerate the input image conditioned on the intermediate sentence, resulting in a proposed visual regeneration loss. Different from most existing image captioning methods, MirrorDiff is a plug-and-play framework which can be plugged into many previous image captioning methods, and further evaluate the generated sentence via the visual similarity between the input image and the regenerated image. Extensive experiments on the MS COCO dataset show that our method achieves obvious improvements over state-of-the-art diffusion-based methods, up to 127.9 on CIDEr, and achieves competitive performance on multiple evaluation metrics over the auto-regressive methods trained on larger-scale datasets.
Junbo Wang 0003, Liangyu Fu, Yining Zhu, Qiangguo Jin, Hongsong Wang 0001, Kun Hu 0008
ICMR8
2025 Graph traverse reference network for sign language corpus retrieval in the wild
abstract
Sign languages are the primary languages of the deaf community as well as hearing individuals who are unable to speak, which engage the visual-manual modality to convey meanings. In recent years, there has been an explosive growth of sign language videos available from video streaming and social media service platforms. Given the size of these corpora, sign language users often face significant challenges in effectively acquiring the information they need. Therefore, we propose a novel deep learning architecture, namely Graph Traverse Reference Network (GTRN), allowing visual signing queries to retrieve relevant sign language videos (documents) from a large corpus. GTRN introduces a traverse graph, which provides coarse-to-fine reference information in a hierarchical manner from frame-level to body-part-level observations. A reference-based attention is devised to obtain the embedding for a visual input of each level, which allows the computations to be allocated and processed at difference locations regarding local devices and central servers. A contrastive learning strategy optimizes GTRN in pursuit of a joint latent space for the queries and the documents by their meanings. Moreover, GTRN is compatible to leverage existing general visual representation foundation models, by which their resulted embeddings are used as the frame-level reference of GTRN. To the best of our knowledge, it is one of the first studies on using visual signing queries for retrieving sign language videos in a real-world setting and comprehensive experiments were conducted which demonstrated the effectiveness of our proposed method. • A deep network formulates sign language embeddings for video queries and documents. • A reference-based attention method and contrastive strategy for embedding alignment. • Comprehensive experiments and analyses on a real-world sign language corpus.
Kun Hu 0008, Fengxiang He, Adam Schembri, Zhiyong Wang 0001
Neurocomputing1
2025 Music source separation via hybrid waveform and spectrogram based generative adversarial network
Qiuxia Wu, Haipeng Deng, Kun Hu 0008, Zhiyong Wang 0001
Multim. Tools Appl.3
2024 Terrain Diffusion Network: Climatic-Aware Terrain Generation with Geological Sketch Guidance
abstract
Sketch-based terrain generation seeks to create realistic landscapes for virtual environments in various applications such as computer games, animation and virtual reality. Recently, deep learning based terrain generation has emerged, notably the ones based on generative adversarial networks (GAN). However, these methods often struggle to fulfill the requirements of flexible user control and maintain generative diversity for realistic terrain. Therefore, we propose a novel diffusion-based method, namely terrain diffusion network (TDN), which actively incorporates user guidance for enhanced controllability, taking into account terrain features like rivers, ridges, basins, and peaks. Instead of adhering to a conventional monolithic denoising process, which often compromises the fidelity of terrain details or the alignment with user control, a multi-level denoising scheme is proposed to generate more realistic terrains by taking into account fine-grained details, particularly those related to climatic patterns influenced by erosion and tectonic activities. Specifically, three terrain synthesisers are designed for structural, intermediate, and fine-grained level denoising purposes, which allow each synthesiser concentrate on a distinct terrain aspect. Moreover, to maximise the efficiency of our TDN, we further introduce terrain and sketch latent spaces for the synthesizers with pre-trained terrain autoencoders. Comprehensive experiments on a new dataset constructed from NASA Topology Images clearly demonstrate the effectiveness of our proposed method, achieving the state-of-the-art performance. Our code is available at https://github.com/TDNResearch/TDN.
Zexin Hu, Kun Hu 0008, Clinton Mo, Zhiyong Wang 0001
AAAI2
2024 Autoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image Generation
abstract
A 360-degree (omni-directional) image provides an all-encompassing spherical view of a scene. Recently, there has been an increasing interest in synthesising 360-degree images from conventional narrow field of view (NFoV) images captured by digital cameras and smartphones, for providing immersive experiences in various scenarios such as virtual reality. Yet, existing methods typically fall short in synthesizing intricate visual details or ensure the generated images align consistently with user-provided prompts. In this study, autoregressive omni-aware generative network (AOG-Net) is proposed for 360-degree image generation by outpainting an incomplete 360-degree image progressively with NFoV and text guidances joinly or individually. This autoregressive scheme not only allows for deriving finer-grained and text-consistent patterns by dynamically generating and adjusting the process but also offers users greater flexibility to edit their conditions throughout the generation process. A global-local conditioning mechanism is devised to comprehensively formulate the outpainting guidance in each autoregressive step. Text guidances, omni-visual cues, NFoV inputs and omni-geometry are encoded and further formulated with cross-attention based transformers into a global stream and a local stream into a conditioned generative backbone model. As AOG-Net is compatible to leverage large-scale models for the conditional encoder and the generative prior, it enables the generation to use extensive open-vocabulary text guidances. Comprehensive experiments on two commonly used 360-degree image datasets for both indoor and outdoor settings demonstrate the state-of-the-art performance of our proposed method. Our code is available at https://github.com/zhuqiangLu/AOG-NET-360.
Zhuqiang Lu, Kun Hu 0008, Lei Bai 0001, Zhiyong Wang 0001
AAAI2
2024 SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation
abstract
The Segment Anything Model (SAM) is a powerful foundation model that has revolutionised image segmentation. To apply SAM to surgical instrument segmentation, a common approach is to locate precise points or boxes of instruments and then use them as prompts for SAM in a zero-shot manner. However, we observe two problems with this naive pipeline: (1) the domain gap between natural objects and surgical instruments leads to inferior generalisation of SAM; and (2) SAM relies on precise point or box locations for accurate segmentation, requiring either extensive manual guidance or a well-performing specialist detector for prompt preparation, which leads to a complex multi-stage pipeline. To address these problems, we introduce SurgicalSAM, a novel end-to-end efficient-tuning approach for SAM to effectively integrate surgical-specific information with SAM’s pre-trained knowledge for improved generalisation. Specifically, we propose a lightweight prototype-based class prompt encoder for tuning, which directly generates prompt embeddings from class prototypes and eliminates the use of explicit prompts for improved robustness and a simpler pipeline. In addition, to address the low inter-class variance among surgical instrument categories, we propose contrastive prototype learning, further enhancing the discrimination of the class prototypes for more accurate class prompting. The results of extensive experiments on both EndoVis2018 and EndoVis2017 datasets demonstrate that SurgicalSAM achieves state-of-the-art performance while only requiring a small number of tunable parameters. The source code is available at https://github.com/wenxi-yue/SurgicalSAM.
Wenxi Yue, Jing Zhang 0037, Kun Hu 0008, Yong Xia 0001, Jiebo Luo 0001, Zhiyong Wang 0001
AAAI3
2024 Motion Keyframe Interpolation for Any Human Skeleton via Temporally Consistent Point Cloud Sampling and Reconstruction
Clinton Mo, Kun Hu 0008, Chengjiang Long, Dong Yuan 0001, Zhiyong Wang 0001
ECCV (82)2
2024 Identity-Consistent Diffusion Network for Grading Knee Osteoarthritis Progression in Radiographic Imaging
Wenhua Wu 0005, Kun Hu 0008, Wenxi Yue, Wei Li 0058, Milena Simic, ChangYang Li, Wei Xiang 0001, Zhiyong Wang 0001
ECCV (84)2
2024 Radio Frequency Signal based Human Silhouette Segmentation: A Sequential Diffusion Approach
abstract
Radio frequency (RF) signals have been proved to be flexible for human silhouette segmentation (HSS) under complex environments. Existing studies are mainly based on a one-shot approach, which lacks a coherent projection ability from the RF domain. Additionally, the spatio-temporal patterns have not been fully explored for human motion dynamics in HSS. Therefore, we propose a two-stage Sequential Diffusion Model (SDM) to progressively synthesize high-quality segmentation jointly with the considerations on motion dynamics. Cross-view transformation blocks are devised to guide the diffusion model in a multi-scale manner for comprehensively characterizing human related patterns in an individual frame such as directional projection from signal planes. Moreover, spatio-temporal blocks are devised to fine-tune the frame-level model to incorporate spatio-temporal contexts and motion dynamics, enhancing the consistency of the segmentation maps. Comprehensive experiments on a public benchmark - HIBER demonstrate the state-of-the-art performance of our method with an IoU 0.732. Our code is available at https://github.com/ph-w2000/SDM.
Penghui Wen, Kun Hu 0008, Dong Yuan 0001, ChangYang Li, Zhiyong Wang 0001
ICME2
2024 Bridging the Gap: Sketch-Aware Interpolation Network for High-Quality Animation Sketch Inbetweening
abstract
Hand-drawn 2D animation workflow is typically initiated with the creation of sketch keyframes. Subsequent manual inbetweens are crafted for smoothness, which is a labor-intensive process and the prospect of automatic animation sketch interpolation has become highly appealing. Yet, common frame interpolation methods are generally hindered by two key issues: 1) limited texture and colour details in sketches, and 2) exaggerated alterations between two sketch keyframes. To overcome these issues, we propose a novel deep learning method - Sketch-Aware Interpolation Network (SAIN). This approach incorporates multi-level guidance that formulates region-level correspondence, stroke-level correspondence and pixel-level dynamics. A multi-stream U-Transformer is then devised to characterize sketch inbetweening patterns using these multi-level guides through the integration of self / cross-attention mechanisms. Additionally, to facilitate future research on animation sketch inbetweening, we constructed a large-scale dataset - STD-12K, comprising 30 sketch animation series in diverse artistic styles. Comprehensive experiments on this dataset convincingly show that our proposed SAIN surpasses the state-of-the-art interpolation methods. Our code and dataset are avaliable in https://github.com/none-master/SAIN.
Kun Hu 0008, Wei Bao 0001, Chang Wen Chen, Zhiyong Wang 0001
ACM Multimedia2
2024 SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization
Lintao Wang 0002, Xiaogang Zhu 0001, Xuequan Lu, Zhiyong Wang 0001, Kun Hu 0008
MMAsia6
2024 CFRL: Coarse-Fine Decoupled Representation Learning For Long-Tailed Recognition
abstract
Data often faces a severe class imbalance issue in the real world, meaning that the number of instances within classes varies greatly, following a long-tailed distribution.In this case, the direct application of supervised learning yields poor performance.Existing long-tailed recognition (LTR) methods often heavily rely on the label information to enhance tail classes' accuracy at the expense of head class by an image-level end-to-end resampling strategy to address data distribution imbalance.Nevertheless, they neglect label bias, which can severely affect the LTR model's accuracy.In this paper, we propose a novel approach, namely Coarse-Fine Decoupled Representation Learning (CFRL) for LTR.Our core idea is to decouple data representations from the classifier and decompose representation learning into two stages: image-level and patch-level.Specifically, in the image-level stage, we leverage unsupervised learning on image-level information to reduce the impact of label bias caused by imbalanced datasets.In the patch-level stage, we introduce patch-level rotation augmentation as negative samples, forcing the model to acquire more comprehensive information.Our theoretical and empirical analyses demonstrate that the approach does not sacrifice the accuracy of head classes while significantly reducing the overfitting of tail classes, improving both of them.We showcase state-of-the-art results on CIFAR, ImageNet, and iNaturalist datasets.Furthermore, we illustrate that this training methodology can be combined with various existing Long-Tailed Recognition (LTR) methods, further enhancing their performance.
Yiran Song, Qianyu Zhou 0001, Kun Hu 0008, Lizhuang Ma, Xuequan Lu
MMAsia3
2024 T2QRM: Text-Driven Quadruped Robot Motion Generation
Kun Hu 0008, Zhiyong Wang 0001, Wenxiong Kang
MMAsia4
2024 TLDW: Extreme Multimodal Summarization of News Videos
abstract
Multimodal summarisation with multimodal output is drawing increasing attention due to the rapid growth of multimedia data. While several methods have been proposed to summarise visual-text contents, their multimodal outputs are not succinct enough at an extreme level to address the information overload issue. To the end of extreme multimodal summarisation, we introduce a new task, eXtreme Multimodal Summarisation with Multimodal Output (XMSMO) for the scenario of TL;DW - Too Long; Didn’t Watch, akin to TL;DR. XMSMO aims to summarise a video-document pair into a summary with an extremely short length, which consists of one cover frame as the visual summary and one sentence as the textual summary. We propose a novel unsupervised Hierarchical Optimal Transport Network (HOT-Net) consisting of three components: hierarchical multimodal encoder, hierarchical multimodal fusion decoder, and optimal transport solver. Our method is trained, without using reference summaries, by optimising the visual and textual coverage from the perspectives of the distance between the semantic distributions under optimal transport plans. To facilitate the study on this task, we constructed a large-scale dataset, XMSMO-News, by harvesting 4,891 video-document pairs. The experimental results show that our method achieves promising performance in terms of ROUGE and IoU metrics. Our dataset and source code will be publicly available in GitHub.
Peggy Tang, Kun Hu 0008, Lei Zhang 0001, Jiebo Luo 0001, Zhiyong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Soft Multiprototype Clustering Algorithm via Two-Layer Semi-NMF
abstract
This article proposes a novel soft multiprototype clustering algorithm (SMP) for high-dimensional data clustering with noisy and complex structural patterns. SMP integrates dimensionality reduction, multiprototype clustering, and multiprototype merge clustering under a two-layer seminonnegative matrix factorization (semi-NMF) architecture. Specifically, the first semi-NMF layer performs multiprototype clustering, which solves the problem that a single prototype cannot represent complex data structures. Meanwhile, the multiprototype fuzzy clustering constraints ensure that the multiprototypes better characterize the original data structure. The second semi-NMF layer performs multiprototype merge clustering to mitigate the issues of heavy computation burden and poor antinoise performance of the spectral clustering algorithm. The introduction of the Laplace graph matrix regularization constraint in this layer assists SMP in completing the merging of multiprototypes with complex data structures. Comprehensive experiments demonstrate that the proposed method outperforms the state-of-the-art algorithms.
Shan Zeng, Xiangjun Duan, Kun Hu 0008, Yuan Yan Tang
IEEE Trans. Fuzzy Syst.5
2024 Higher Order Polynomial Transformer for Fine-Grained Freezing of Gait Detection
abstract
Freezing of Gait (FoG) is a common symptom of Parkinson's disease (PD), manifesting as a brief, episodic absence, or marked reduction in walking, despite a patient's intention to move. Clinical assessment of FoG events from manual observations by experts is both time-consuming and highly subjective. Therefore, machine learning-based FoG identification methods would be desirable. In this article, we address this task as a fine-grained human action recognition problem based on vision inputs. A novel deep learning architecture, namely, higher order polynomial transformer (HP-Transformer), is proposed to incorporate pose and appearance feature sequences to formulate fine-grained FoG patterns. In particular, a higher order self-attention mechanism is proposed based on higher order polynomials. To this end, linear, bilinear, and trilinear transformers are formulated in pursuit of discriminative fine-grained representations. These representations are treated as multiple streams and further fused by a cross-order fusion strategy for FoG detection. Comprehensive experiments on a large in-house dataset collected during clinical assessments demonstrate the effectiveness of the proposed method, and an area under the receiver operating characteristic (ROC) curve (AUC) of 0.92 is achieved for detecting FoG.
Renfei Sun, Kun Hu 0008, Kaylena A. Ehgoetz Martens, Markus Hagenbuchner, Ah Chung Tsoi, Mohammed Bennamoun, Simon J. G. Lewis, Zhiyong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.2
2023 Multi-Scale Control Signal-Aware Transformer for Motion Synthesis without Phase
abstract
Synthesizing controllable motion for a character using deep learning has been a promising approach due to its potential to learn a compact model without laborious feature engineering. To produce dynamic motion from weak control signals such as desired paths, existing methods often require auxiliary information such as phases for alleviating motion ambiguity, which limits their generalisation capability. As past poses often contain useful auxiliary hints, in this paper, we propose a task-agnostic deep learning method, namely Multi-scale Control Signal-aware Transformer (MCS-T), with an attention based encoder-decoder architecture to discover the auxiliary information implicitly for synthesizing controllable motion without explicitly requiring auxiliary information such as phase. Specifically, an encoder is devised to adaptively formulate the motion patterns of a character's past poses with multi-scale skeletons, and a decoder driven by control signals to further synthesize and predict the character's state by paying context-specialised attention to the encoded past motion patterns. As a result, it helps alleviate the issues of low responsiveness and slow transition which often happen in conventional methods not using auxiliary information. Both qualitative and quantitative experimental results on an existing biped locomotion dataset, which involves diverse types of motion transitions, demonstrate the effectiveness of our method. In particular, MCS-T is able to successfully generate motions comparable to those generated by the methods using auxiliary information.
Lintao Wang 0002, Kun Hu 0008, Lei Bai 0001, Wanli Ouyang, Zhiyong Wang 0001
AAAI2
2023 Continuous Intermediate Token Learning with Implicit Motion Manifold for Keyframe Based Motion Interpolation
abstract
Deriving sophisticated 3D motions from sparse keyframes is a particularly challenging problem, due to continuity and exceptionally skeletal precision. The action features are often derivable accurately from the full series of keyframes, and thus, leveraging the global context with transformers has been a promising data-driven embedding approach. However, existing methods are often with inputs of interpolated intermediate frame for continuity using basic interpolation methods with keyframes, which result in a trivial local minimum during training. In this paper, we propose a novel framework to formulate latent motion manifolds with keyframe-based constraints, from which the continuous nature of intermediate token representations is considered. Particularly, our proposed framework consists of two stages for identifying a latent motion subspace, i.e., a keyframe encoding stage and an intermediate token generation stage, and a subsequent motion synthesis stage to extrapolate and compose motion data from manifolds. Through our extensive experiments conducted on both the LaFAN1 and CMU Mocap datasets, our proposed method demonstrates both superior interpolation accuracy and high visual similarity to ground truth motions.
Clinton Mo, Kun Hu 0008, Chengjiang Long, Zhiyong Wang 0001
CVPR2
2023 Material-Aware Self-Supervised Network for Dynamic 3D Garment Simulation
abstract
Dynamic 3D garment simulation has various applications in many domains. Recently, self-supervised learning for this task has been studied to reduce annotation costs. However, different material characteristics of garments have been rarely explored, limiting the generalization and flexibility of existing methods. Therefore, in this paper, a novel self-supervised deep learning architecture is proposed, namely Material-aware Self-supervised Network (MSN), as a material-aware approach for dynamically simulating garments with different materials. Specifically, a material-aware parameterized regressor is introduced based on the observation that material characteristics change continuously regarding the fabric parameters. As a result, MSN realises real-time garment simulation with various material properties without model re-training. Moreover, to simulate garments of different categories (e.g., t-shirts vs. dresses), a sampling-based linear skinning strategy is studied in MSN. Comprehensive experiments on the widely used AMASS dataset demonstrated the effectiveness of MSN both quantitatively and qualitatively.
Aoran Liu, Kun Hu 0008, Wenxi Yue, Qiuxia Wu, Zhiyong Wang 0001
ICME2
2023 Robust Audio Anti-Spoofing with Fusion-Reconstruction Learning on Multi-Order Spectrograms
abstract
Robust audio anti-spoofing has been increasingly challenging due to the recent advancements on deepfake techniques. While spectrograms have demonstrated their capability for anti-spoofing, complementary information presented in multi-order spectral patterns have not been well explored, which limits their effectiveness for varying spoofing attacks. Therefore, we propose a novel deep learning method with a spectral fusion-reconstruction strategy, namely S2pecNet, to utilise multi-order spectral patterns for robust audio anti-spoofing representations. Specifically, spectral patterns up to second-order are fused in a coarse-to-fine manner and two branches are designed for the fine-level fusion from the spectral and temporal contexts. A reconstruction from the fused representation to the input spectrograms further reduces the potential fused information loss. Our method achieved the state-of-the-art performance with an EER of 0.77% on a widely used dataset: ASVspoof2019 LA Challenge.
Penghui Wen, Kun Hu 0008, Wenxi Yue, Sen Zhang 0006, Wanlei Zhou 0001, Zhiyong Wang 0001
INTERSPEECH2
2023 TopicCAT: Unsupervised Topic-Guided Co-Attention Transformer for Extreme Multimodal Summarisation
abstract
The exponential growth of multimedia data has sparked a surge of interest in multimodal summarisation with multimodal output (MSMO). A relatively unexplored but essential task within this field is extreme multimodal summarisation, a process that involves creating extremely concise multimodal summaries to further address the issue of multimedia information overload. In this study, we propose a novel Unsupervised Topic-guided Co-Attention Transformer (TopicCAT) neural network to produce extreme multimodal summaries for video-document pairs. The approach consists of two learning stages for a comprehensive multimodal understanding, guided by topic-based insights: a unimodal learning stage and a cross-modal learning stage, in which a cross-modal topic model is devised to capture the overarching themes present in both documents and videos. To achieve unsupervised learning, eliminating the need for resource-expensive collection of ground-truth multimodal summaries, we propose an optimal transport-based optimisation scheme to evaluate summary coverage from a semantic distribution perspective at the topic-level. Comprehensive experiments demonstrate the effectiveness of our proposed TopicCAT method on a multimodal news dataset, achieving a BERTScore of 84.46 and an accuracy of 0.60.
Peggy Tang, Kun Hu 0008, Lei Zhang 0001, Junbin Gao, Jiebo Luo 0001, Zhiyong Wang 0001
ACM Multimedia2
2023 Federated Unsupervised Cluster-Contrastive learning for person Re-identification: A coarse-to-fine approach
abstract
Person Re-identification (ReID) has attracted considerable interests in recent years, largely driven by the escalating demand for public safety measures. However, the acquisition and handling of sensitive personal data can trigger significant privacy concerns. Federated learning has been introduced as a potential solution to this problem, with the goal of limiting the exposure of sensitive data across different participating entities (clients). Existing methods often depend on labor-intensive data annotations and face difficulties in maintaining cross-domain uniformity. To tackle these challenges, we propose a Federated Unsupervised Cluster-Contrastive (FedUCC) method based on deep learning for Person ReID that follows a generic-to-specific learning strategy. First, FedUCC procures generic knowledge from a conventional federated learning scheme to aggregate and distribute parameters across local clients. Second, specialized knowledge is explored to facilitate client personalization by disentangling client-specific knowledge from generic knowledge through parameter localization. Third, to further enhance effective fine-grained patterns instead of overfitting on specialized client knowledge, we investigate two key aspects: patch-level feature alignment and camera-invariant learning. Comprehensive experiments on eight public benchmark datasets demonstrate the state-of-the-art performance of our proposed method.
Jian F. Weng, Kun Hu 0008, Jingya Wang 0001, Zhiyong Wang 0001
Comput. Vis. Image Underst.2
2023 Region Assisted Sketch Colorization
abstract
Automatic sketch colorization is a challenging task that aims to generate a color image from a sketch, primarily due to its inherently ill-posed nature. While many approaches have shown promising results, two significant challenges remain: limited color patterns and a wide range of artifacts such as color bleeding and semantic inconsistencies among relevant regions. These issues stem from the operation of traditional convolutional structures, which capture structural features in a pixel-wise manner, resulting in inadequate utilization of regional information within the sketch. Therefore, we propose the Region-Assisted Sketch Coloring (RASC) method, which introduces an intermediate representation called the 'Region Map' to explicitly characterize the regional information of the sketch. This Region Map is derived from the input sketch and is effectively formulated by our RASC architecture, enhancing the perception of region-wise features beyond the original pixel-wise features. Specifically, we start by employing the sketch encoder to extract hierarchical feature maps from the input sketches. Subsequently, we introduce a coarse-to-fine decoder comprising a series of Region-based Modulation (RM) blocks. This decoder modulates features that combine the modulation results of its previous block and the sketch features of the corresponding encoder block with our Region Formulation module. Each module explicitly formulates the sketch features in a region-wise manner. This accurately captures both the inner-region local style and inter-region global context dependency, resulting in various color patterns and fewer synthesis artifacts. Our experimental results show that our proposed method surpasses state-of-the-art methods in both synthetic and real sketch datasets.
Ning Wang 0025, Muyao Niu, Zhihui Wang 0001, Kun Hu 0008, Bin Liu 0040, Zhiyong Wang 0001
IEEE Trans. Image Process.4
2023 Multi-Level Adversarial Spatio-Temporal Learning for Footstep Pressure Based FoG Detection
abstract
Freezing of gait (FoG) is one of the most common symptoms of Parkinson's disease, which is a neurodegenerative disorder of the central nervous system impacting millions of people around the world. To address the pressing need to improve the quality of treatment for FoG, devising a computer-aided detection and quantification tool for FoG has been increasingly important. As a non-invasive technique for collecting motion patterns, the footstep pressure sequences obtained from pressure sensitive gait mats provide a great opportunity for evaluating FoG in the clinic and potentially in the home environment. In this study, FoG detection is formulated as a sequential modelling task and a novel deep learning architecture, namely Adversarial Spatio-temporal Network (ASTN), is proposed to learn FoG patterns across multiple levels. ASTN introduces a novel adversarial training scheme with a multi-level subject discriminator to obtain subject-independent FoG representations, which helps to reduce the over-fitting risk due to the high inter-subject variance. As a result, robust FoG detection can be achieved for unseen subjects. The proposed scheme also sheds light on improving subject-level clinical studies from other scenarios as it can be integrated with many existing deep architectures. To the best of our knowledge, this is one of the first studies of footstep pressure-based FoG detection and the approach of utilizing ASTN is the first deep neural network architecture in pursuit of subject-independent representations. In our experiments on 393 trials collected from 21 subjects, the proposed ASTN achieved an AUC 0.85, clearly outperforming conventional learning methods.
Kun Hu 0008, Shaohui Mei, Wei Wang 0115, Kaylena A. Ehgoetz Martens, Liang Wang 0001, Simon J. G. Lewis, David Dagan Feng, Zhiyong Wang 0001
IEEE J. Biomed. Health Informatics1
2023 Graph Fusion Network-Based Multimodal Learning for Freezing of Gait Detection
abstract
Freezing of gait (FoG) is identified as a sudden and brief episode of movement cessation despite the intention to continue walking. It is one of the most disabling symptoms of Parkinson's disease (PD) and often leads to falls and injuries. Many computer-aided FoG detection methods have been proposed to use data collected from unimodal sources, such as motion sensors, pressure sensors, and video cameras. However, there are limited efforts of multimodal-based methods to maximize the value of all the information collected from different modalities in clinical assessments and improve the FoG detection performance. Therefore, in this study, a novel end-to-end deep architecture, namely graph fusion neural network (GFN), is proposed for multimodal learning-based FoG detection by combining footstep pressure maps and video recordings. GFN constructs multimodal graphs by treating the encoded features of each modality as vertex-level inputs and measures their adjacency patterns to construct complementary FoG representations, thus reducing the representation redundancy among different modalities. In addition, since GFN is devised to process multimodal graphs of arbitrary structures, it is expected to achieve superior performance with inputs containing missing modalities, compared to the alternative unimodal methods. A multimodal FoG dataset was collected, which included clinical assessment videos and footstep pressure sequences of 340 trials from 20 PD patients. Our proposed GFN demonstrates a great promise of multimodal FoG detection with an area under the curve (AUC) of 0.882. To the best of our knowledge, this is one of the first studies to utilize multimodal learning for automated FoG detection, which offers significant opportunities for better patient assessments and clinical trials in the future.
Kun Hu 0008, Zhiyong Wang 0001, Kaylena A. Ehgoetz Martens, Markus Hagenbuchner, Mohammed Bennamoun, Ah Chung Tsoi, Simon J. G. Lewis
IEEE Trans. Neural Networks Learn. Syst.1
2022 Sign Language Translation with Hierarchical Spatio-Temporal Graph Neural Network
abstract
Sign language translation (SLT), which generates text in a spoken language from visual content in a sign language, is important to assist the hard-of-hearing community for their communications. Inspired by neural machine translation (NMT), most existing SLT studies adopted a general sequence to sequence learning strategy. However, SLT is significantly different from general NMT tasks since sign languages convey messages through multiple visual-manual aspects. Therefore, in this paper, these unique characteristics of sign languages are formulated as hierarchical spatio-temporal graph representations, including high-level and fine-level graphs of which a vertex characterizes a specified body part and an edge represents their interactions. Particularly, high-level graphs represent the patterns in the regions such as hands and face, and fine-level graphs consider the joints of hands and landmarks of facial regions. To learn these graph patterns, a novel deep learning architecture, namely hierarchical spatio-temporal graph neural network (HST-GNN), is proposed. Graph convolutions and graph self-attentions with neighborhood context are proposed to characterize both the local and the global graph properties. Experimental results on benchmark datasets demonstrated the effectiveness of the proposed method.
Jichao Kan, Kun Hu 0008, Markus Hagenbuchner, Ah Chung Tsoi, Mohammed Bennamoun, Zhiyong Wang 0001
WACV2
2022 Adversarial Evolving Neural Network for Longitudinal Knee Osteoarthritis Prediction
abstract
Knee osteoarthritis (KOA) as a disabling joint disease has doubled in prevalence since the mid-20th century. Early diagnosis for the longitudinal KOA grades has been increasingly important for effective monitoring and intervention. Although recent studies have achieved promising performance for baseline KOA grading, longitudinal KOA grading has been seldom studied and the KOA domain knowledge has not been well explored yet. In this paper, a novel deep learning architecture, namely adversarial evolving neural network (A-ENN), is proposed for longitudinal grading of KOA severity. As the disease progresses from mild to severe level, ENN involves the progression patterns for accurately characterizing the disease by comparing an input image it to the template images of different KL grades using convolution and deconvolution computations. In addition, an adversarial training scheme with a discriminator is developed to obtain the evolution traces. Thus, the evolution traces as fine-grained domain knowledge are further fused with the general convolutional image representations for longitudinal grading. Note that ENN can be applied to other learning tasks together with existing deep architectures, in which the responses characterize progressive representations. Comprehensive experiments on the Osteoarthritis Initiative (OAI) dataset were conducted to evaluate the proposed method. An overall accuracy was achieved as 62.7%, with the baseline, 12-month, 24-month, 36-month, and 48-month accuracy as 64.6%, 63.9%, 63.2%, 61.8% and 60.2%, respectively.
Kun Hu 0008, Wenhua Wu 0005, Wei Li 0058, Milena Simic, Albert Y. Zomaya, Zhiyong Wang 0001
IEEE Trans. Medical Imaging1
2021 Keyframe Extraction from Motion Capture Sequences with Graph based Deep Reinforcement Learning
abstract
Animation production workflows centred around motion capture techniques often require animators to edit the motion for various artistic and technical reasons. This process generally uses a set of keyframes. Unsupervised keyframe selection methods for motion capture sequences are highly demanded to reduce the laborious annotations. However, most existing methods are optimization-based, which cause the issues of flexibility and efficiency and eventually constrains the interactions and controls with animators. To address these limitations, we propose a novel graph based deep reinforcement learning method for efficient unsupervised keyframe selection. First, a reward function is devised in terms of reconstruction difference by comparing the original sequence and the interpolated sequence produced by the keyframes. The reward complies with the requirements of the animation pipeline satisfying: 1) incremental reward to evaluate the interpolated keyframes immediately; 2) order insensitivity for consistent evaluation; and 3) non-diminishing return for comparable rewards between optimal and sub-optimal solutions. Then by representing each skeleton frame as a graph, a graph-based deep agent is guided to heuristically select keyframes to maximize the reward. During the inference it is no longer necessary to estimate the reconstruction difference, and the evaluation time can be reduced significantly. The experimental results on the CMU Mocap dataset demonstrate that our proposed method is able to select keyframes at a high efficiency without clearly compromising the quality in comparison with the state-of-the-art methods.
Clinton Mo, Kun Hu 0008, Shaohui Mei, Zhiyong Wang 0001
ACM Multimedia2
2020 Speaker-Aware Monaural Speech Separation
Jiahao Xu 0002, Kun Hu 0008, Tran Duc Chung, Zhiyong Wang 0001
INTERSPEECH2
2020 Graph Sequence Recurrent Neural Network for Vision-Based Freezing of Gait Detection
abstract
Freezing of gait (FoG) is one of the most common symptoms of Parkinson's disease (PD), a neurodegenerative disorder which impacts millions of people around the world. Accurate assessment of FoG is critical for the management of PD and to evaluate the efficacy of treatments. Currently, the assessment of FoG requires well-trained experts to perform time-consuming annotations via vision-based observations. Thus, automatic FoG detection algorithms are needed. In this study, we formulate vision-based FoG detection, as a fine-grained graph sequence modelling task, by representing the anatomic joints in each temporal segment with a directed graph, since FoG events can be observed through the motion patterns of joints. A novel deep learning method is proposed, namely graph sequence recurrent neural network (GS-RNN), to characterize the FoG patterns by devising graph recurrent cells, which take graph sequences of dynamic structures as inputs. For the cases of which prior edge annotations are not available, a data-driven based adjacency estimation method is further proposed. To the best of our knowledge, this is one of the first studies on vision-based FoG detection using deep neural networks designed for graph sequences of dynamic structures. Experimental results on more than 150 videos collected from 45 patients demonstrated promising performance of the proposed GS-RNN for FoG detection with an AUC value of 0.90.
Kun Hu 0008, Zhiyong Wang 0001, Wei Wang 0115, Kaylena A. Ehgoetz Martens, Liang Wang 0001, Tieniu Tan, Simon J. G. Lewis, David Dagan Feng
IEEE Trans. Image Process.1
2020 Vision-Based Freezing of Gait Detection With Anatomic Directed Graph Representation
abstract
Parkinson's disease significantly impacts the life quality of millions of people around the world. While freezing of gait (FoG) is one of the most common symptoms of the disease, it is time consuming and subjective to assess FoG for well-trained experts. Therefore, it is highly desirable to devise computer-aided FoG detection methods for the purpose of objective and time-efficient assessment. In this paper, in line with the gold standard of FoG clinical assessment, which requires video or direct observation, we propose one of the first vision-based methods for automatic FoG detection. To better characterize FoG patterns, instead of learning an overall representation of a video, we propose a novel architecture of graph convolution neural network and represent each video as a directed graph where FoG related candidate regions are the vertices. A weakly-supervised learning strategy and a weighted adjacency matrix estimation layer are proposed to eliminate the resource expensive data annotation required for fully supervised learning. As a result, the interference of visual information irrelevant to FoG, such as gait motion of supporting staff involved in clinical assessments, has been reduced to improve FoG detection performance by identifying the vertices contributing to FoG events. To further improve the performance, the global context of a clinical video is also considered and several fusion strategies with graph predictions are investigated. Experimental results on more than 100 videos collected from 45 patients during a clinical assessment demonstrated promising performance of our proposed method with an AUC of 0.887.
Kun Hu 0008, Zhiyong Wang 0001, Shaohui Mei, Kaylena A. Ehgoetz Martens, Simon J. G. Lewis, David Dagan Feng
IEEE J. Biomed. Health Informatics1
2018 Vision-Based Freezing of Gait Detection with Anatomic Patch Based Representation
Kun Hu 0008, Zhiyong Wang 0001, Kaylena A. Ehgoetz Martens, Simon J. G. Lewis
ACCV (1)1