Zhiyong Wang 0001

dblp:62/234-1 · DBLP profile ↗
← Back
165ranked-venue papers
8as first author
90since 2021 · last 2026
0000-0002-8043-0312ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 102 · 7 first-author · 51 since 2021Artificial intelligence and machine learning · 62 · 2 first-author · 44 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 13 since 2021Databases, data management, data science and information retrieval · 10 · 7 since 2021Computer networks · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Pb4U-GNet: Resolution-Adaptive Garment Simulation via Propagation-before-Update Graph Network
abstract
Garment simulation is fundamental to various applications in computer vision and graphics, from virtual try-on to digital human modelling. However, conventional physics-based methods remain computationally expensive, hindering their application in time-sensitive scenarios. While graph neural networks (GNNs) offer promising acceleration, existing approaches exhibit poor cross-resolution generalisation, demonstrating significant performance degradation on higher-resolution meshes beyond the training distribution. This stems from two key factors: (1) existing GNNs employ fixed message-passing depth that fails to adapt information aggregation to mesh density variation, and (2) vertex-wise displacement magnitudes are inherently resolution-dependent in garment simulation. To address these issues, we introduce Propagation-before-Update Graph Network (Pb4U-GNet), a resolution-adaptive framework that decouples message propagation from feature updates. Pb4U-GNet incorporates two key mechanisms: (1) dynamic propagation depth control, adjusting message-passing iterations based on mesh resolution, and (2) geometry-aware update scaling, which scales predictions according to local mesh characteristics. Extensive experiments show that even trained solely on low-resolution meshes, Pb4U-GNet exhibits strong generalisability across diverse mesh resolutions, addressing a fundamental challenge in neural garment simulation.
Aoran Liu, Kun Hu 0008, Clinton Mo, Qiuxia Wu, Wenxiong Kang, Zhiyong Wang 0001
AAAI6
2026 DuoCast: Duo-Probabilistic Diffusion for Precipitation Nowcasting
abstract
Accurate short-term precipitation forecasting is critical for weather-sensitive decision-making in agriculture, transportation, and disaster response. Existing deep learning approaches often struggle to balance global structural consistency with local detail preservation, especially under complex meteorological conditions. We propose DuoCast, a dual-diffusion framework that decomposes precipitation forecasting into low- and high-frequency components modeled in orthogonal latent subspaces. We theoretically prove that this frequency decomposition reduces prediction error compared to conventional single branch U-Net diffusion models. In DuoCast, the low-frequency model captures large-scale trends via convolutional encoders conditioned on weather front dynamics, while the high-frequency model refines fine-scale variability using a self-attention-based architecture. Experiments on four benchmark radar datasets show that DuoCast consistently outperforms state-of-the-art baselines, achieving superior accuracy in both spatial detail and temporal evolution.
Penghui Wen, Mengwei He, Patrick Filippi, Thomas F. A. Bishop, Zhiyong Wang 0001, Kun Hu 0008
AAAI7
2026 Keyframe selection from motion capture data with dual-agent reinforcement learning
abstract
Animation production workflows centred around motion capture techniques require animators to edit motions based on a set of keyframes. However, most existing keyframe selection methods are optimisation-based, which suffer from the issues of flexibility and efficiency. In this paper, a novel deep reinforcement learning method with dual agents are proposed for unsupervised keyframe selection. First, an S-Agent and an R-Agent evaluate the actions of selection and refinement, respectively. A deep spatio-temporal network, namely graph keyframe evaluation network (GKEN), is proposed for the agents. Then, an animation specified reward is devised based on reconstruction, which fulfills three important properties of the animation workflow: incremental reward, order insensitivity and non-diminishing returns. During the inference, it is no longer necessary to compute the reconstruction, which significantly decreases the run-time latency. Experiments on the CMU MoCap dataset demonstrate the efficiency of the proposed method without clearly compromising the effectiveness compared with the state-of-the-art methods. • A deep reinforcement learning with dual-agent to identify motion keyframes. • A spatio-temporal deep agent with graph convolutions and transformers. • Comprehensive experiments and human demonstrations for MoCap keyframing.
Kun Hu 0008, Clinton Mo, Mingyang Ma 0004, Shaohui Mei, Zhiyong Wang 0001
Pattern Recognit.7
2026 Harnessing Text Insights With Visual Alignment for Medical Image Segmentation
abstract
Pre-trained vision-language models (VLMs) and language models (LMs) have recently garnered significant attention due to their remarkable ability to represent textual concepts, opening up new avenues in vision tasks. In medical image segmentation, efforts are being made to integrate text and image data using VLMs and LMs. However, current text-enhanced approaches face several challenges. First, using separate pre-trained vision and text models to encode image and text data can result in semantic shifts. Second, while VLMs can establish the correspondence between visual and textual features when pre-trained on paired image-text data, this alignment often deteriorates during segmentation tasks due to misalignment between the text and vision components in ongoing learning. In this paper, we propose TeViA, a novel approach that seamlessly integrates with various vision and text models, irrespective of their pre-training relationships. This integration is achieved through a segmentation-specific text-to-vision alignment design, ensuring both information gain and semantic consistency. Specifically, for each training data, a foreground visual representation is extracted from the segmentation head and used to supervise projection layers, thereby adjusting the textual features to better contribute to the segmentation task. Additionally, a historic visual prototype is created by aggregating target semantics from all training data and is updated using a momentum-based manner. This prototype aims to enhance the visual representation of each data instance by establishing feature-level connections, which in turn refines the textual features. The superiority of TeViA is validated on five public datasets, exhibiting over 6% Dice improvements compared to vision-only methods. Code is available at: https://github.com/jgfiuuuu/TeViA.
Qingjie Zeng, Zilin Lu, Yutong Xie 0001, Zhiyong Wang 0001, Yanning Zhang 0001, Yong Xia 0001
IEEE Trans. Medical Imaging5
2026 UniAlign: A Universal Cross-Modality Knowledge Alignment Framework for Fine-Grained Action Recognition
abstract
The key to fine-grained video action recognition is identifying subtle differences between action categories. Relying solely on visual features supervised by action labels makes it challenging to characterize robust and discriminative action dynamics from videos. With significant advancements in human pose estimation and the powerful capabilities of Vision-Language Models (VLMs), obtaining reliable and cost-free human pose data and textual semantics has become increasingly feasible, enabling their effective use in fine-grained action recognition. However, the inherent disparities in feature representations across different modalities necessitate a robust alignment strategy to achieve opti mal fusion. To address this, we propose a universal cross-modality knowledge alignment framework, namely UniAlign, to transfer the knowledge from such pre-trained multi-modal models into action recognition models. Specifically, UniAlign introduces two additional branches to extract pose features and textual semantics with the pre-trained pose encoder and VLM. To align the action relevant cues among video features, pose features, and textual semantics, we propose a Cross-Modality Similarity Aggregation module (CMSA) that utilizes the importance of different modal cues while aggregating cross-modal similarities. Additionally, we adopt a fine-tuning mechanism similar to Exponential Moving Average (EMA) to refine the textual semantics, ensuring that the semantic representations encoded by VLMs are preserved while being optimized towards the specific task preferences. Extensive experiments on widely used fine-grained action recognition benchmarks (e.g., FineGym, NTURGB-D, Diving48) and coarse-grained K400 dataset demonstrate the effectiveness of the proposed UniAlign method.
Yihan Wang 0011, Baoli Sun, Xinzhu Ma, Zhihui Wang 0001, Zhiyong Wang 0001
IEEE Trans. Multim.6
2026 Toward an Effective Action-Region Tracking Framework for Fine-Grained Video Action Recognition
abstract
Fine-grained action recognition (FGAR) aims to identify subtle and distinctive differences among fine-grained action categories. However, current recognition methods often capture coarse-grained motion patterns but struggle to identify subtle details in local regions evolving over time. In this work, we introduce the action-region tracking (ART) framework, a novel solution leveraging a query-response mechanism to discover and track the dynamics of distinctive local details, enabling distinguishing similar actions effectively. Specifically, we propose a region-specific semantic activation module that employs discriminative and text-constrained semantics serve as queries to capture the most action-related region responses in each video frame, facilitating interaction among spatial and temporal dimensions with corresponding video features. The captured region responses are then organized into action tracklets, which characterize the region-based action dynamics by linking related responses across different video frames in a coherent sequence. The text-constrained queries are designed to expressly encode nuanced semantic representations derived from the textual descriptions of action labels, as extracted by the language branches within visual language models. To optimize generated action tracklets, we design a multilevel tracklet contrastive constraint among multiple region responses at spatial and temporal levels, which can effectively distinguish individual region responses in each video frame (spatial level) and establish the correlation of similar region responses between adjacent video frames (temporal level). In addition, we implement a task-specific fine-tuning mechanism to refine textual semantics during training. This ensures that the semantic representations encoded by vision language models (VLMs) are not only preserved but also optimized for specific task preferences. Comprehensive experiments on several widely used action recognition benchmarks, i.e., FineGym, Diving48, NTURGB-D, Kinetics, and Something-Something, clearly demonstrate the superiority to previous state-of-the-art baselines.
Baoli Sun, Yihan Wang 0011, Xinzhu Ma, Zhihui Wang 0001, Kun Lu 0003, Zhiyong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.6
2025 RI-MAE: Rotation-Invariant Masked AutoEncoders for Self-Supervised Point Cloud Representation Learning
abstract
Masked point modeling methods have recently achieved great success in self-supervised learning for point cloud data. However, these methods are sensitive to rotations and often exhibit sharp performance drops when encountering rotational variations. In this paper, we propose a novel Rotation-Invariant Masked AutoEncoders (RI-MAE) to address two major challenges: 1) achieving rotation-invariant latent representations, and 2) facilitating self-supervised reconstruction in a rotation-invariant manner. For the first challenge, we introduce RI-Transformer, which features disentangled geometry content, rotation-invariant relative orientation and position embedding mechanisms for constructing rotation-invariant point cloud latent space. For the second challenge, a novel dual-branch student-teacher architecture is devised. It enables the self-supervised learning via the reconstruction of masked patches within the learned rotation-invariant latent space. Each branch is based on an RI-Transformer, and they are connected with an additional RI-Transformer predictor. The teacher encodes all point patches, while the student solely encodes unmasked ones. Finally, the predictor predicts the latent features of the masked patches using the output latent embeddings from the student, supervised by the outputs from the teacher. Extensive experiments demonstrate that our method is robust to rotations, achieving the state-of-the-art performance on various downstream tasks.
Kunming Su, Qiuxia Wu, Panpan Cai, Xiaogang Zhu 0001, Xuequan Lu, Zhiyong Wang 0001, Kun Hu 0008
AAAI6
2025 DC-PCN: Point Cloud Completion Network with Dual-Codebook Guided Quantization
abstract
Point cloud completion aims to reconstruct complete 3D shapes from partial 3D point clouds. With advancements in deep learning techniques, various methods for point cloud completion have been developed. Despite achieving encouraging results, a significant issue remains: these methods often overlook the variability in point clouds sampled from a single 3D object surface. This variability can lead to ambiguity and hinder the achievement of more precise completion results. Therefore, in this study, we introduce a novel point cloud completion network, namely Dual-Codebook Point Completion Network (DC-PCN), following an encder-decoder pipeline. The primary objective of DC-PCN is to formulate a singular representation of sampled point clouds originating from the same 3D surface. DC-PCN introduces a dual-codebook design to quantize point-cloud representations from a multilevel perspective. It consists of an encoder-codebook and a decoder-codebook, designed to capture distinct point cloud patterns at shallow and deep levels. Additionally, to enhance the information flow between these two codebooks, we devise an information exchange mechanism. This approach ensures that crucial features and patterns from both shallow and deep levels are effectively utilized for completion. Extensive experiments on the PCN, ShapeNet_Part, and ShapeNet34 datasets demonstrate the state-of-the-art performance of our method.
Qiuxia Wu, Kunming Su, Zhiyong Wang 0001, Kun Hu 0008
AAAI4
2025 XAI for In-Hospital Mortality Prediction via Multimodal ICU Data
abstract
Predicting in-hospital mortality for ICU patients is critical for improving clinical outcomes. Deep learning models have achieved remarkable accuracy but often lack explainability. To address this, we propose$X-M M P$, an eXplainable Multimodal Mortality Predictor based on transformers for heterogeneous ICU data. X-MMP integrates tabular time-series, vital signs, and clinical notes, and introduces a Layer-Wise Relevance Propagation (LRP) method to generate interpretable saliency maps across modalities. Evaluated on the multimodal MIMICIII dataset, X-MMP achieves competitive predictive performance and generates clinically meaningful explanations. Code: https://github.com/lixingqiao/XAI-ICU.
Xingqiao Li, Jindong Gu, Zhiyong Wang 0001, Yancheng Yuan, Fengxiang He, Bo Du 0001
BIBM3
2025 HGCLIP: Exploring Vision-Language Models with Graph Representations for Hierarchical Understanding
abstract
Object categories are typically organized into a multi-granularity taxonomic hierarchy. When classifying categories at different hierarchy levels, traditional uni-modal approaches focus primarily on image features, revealing limitations in complex scenarios. Recent studies integrating Vision-Language Models (VLMs) with class hierarchies have shown promise, yet they fall short of fully exploiting the hierarchical relationships. These efforts are constrained by their inability to perform effectively across varied granularity of categories. To tackle this issue, we propose a novel framework (HGCLIP) that effectively combines CLIP with a deeper exploitation of the Hierarchical class structure via Graph representation learning. We explore constructing the class hierarchy into a graph, with its nodes representing the textual or image features of each category. After passing through a graph encoder, the textual features incorporate hierarchical structure information, while the image features emphasize class-aware features derived from prototypes through the attention mechanism. Our approach demonstrates significant improvements on 11 diverse visual recognition benchmarks. Our codes are fully available at https://github.com/richard-peng-xia/HGCLIP.
Peng Xia 0005, Xingtong Yu, Lie Ju, Zhiyong Wang 0001, Peibo Duan, ZongYuan Ge
COLING5
2025 ReSo: A Reward-driven Self-organizing LLM-based Multi-Agent System for Reasoning Tasks
abstract
Multi-agent systems have emerged as a promising approach for enhancing the reasoning capabilities of large language models in complex problem-solving.However, current MAS frameworks are limited by poor flexibility and scalability, with underdeveloped optimization strategies.To address these challenges, we propose ReSo, which integrates task graph generation with a reward-driven two-stage agent selection process.The core of ReSo is the proposed Collaborative Reward Model, which can provide fine-grained reward signals for MAS cooperation for optimization.We also introduce an automated data synthesis framework for generating MAS benchmarks, without human annotations.Experimentally, ReSo matches or outperforms existing methods.ReSo achieves 33.7% and 32.3% accuracy on Math-MAS and SciBench-MAS SciBench, while other methods completely fail.The code and data are available at Reso.
Hejia Geng, Xiangyuan Xue, Yiran Qin, Zhiyong Wang 0001, Zhenfei Yin, Lei Bai 0001
EMNLP6
2025 B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
abstract
Recently, Vision Large Language Models (VLLMs) integrated with vision encoders have shown promising performance in vision understanding. The key of VLLMs is to encode visual content into sequences of visual tokens, enabling VLLMs to simultaneously process both visual and textual content. However, understanding videos, especially long videos, remain a challenge to VLLMs as the number of visual tokens grows rapidly when encoding videos, resulting in the risk of exceeding the context window of VLLMs and introducing heavy computation burden. To restrict the number of visual tokens, existing VLLMs either: (1) uniformly downsample videos into a fixed number of frames or (2) reducing the number of visual tokens encoded from each frame. We argue the former solution neglects the rich temporal cue in videos and the later overlooks the spatial details in each frame. In this work, we present Balanced-VLLM (B-VLLM): a novel VLLM framework that aims to effectively leverage task relevant spatio-temporal cues while restricting the number of visual tokens under the VLLM context window length. At the core of our method, we devise a text-conditioned adaptive frame selection module to identify frames relevant to the visual understanding task. The selected frames are then de-duplicated using a temporal frame token merging technique. The visual tokens of the selected frames are processed through a spatial token sampling module and an optional spatial token merging strategy to achieve precise control over the token count. Experimental results show that B-VLLM is effective in balancing the number of frames and visual tokens in video understanding, yielding superior performance on various video understanding benchmarks. Our code is available at https://github.com/zhuqiangLu/B-VLLM.
Zhuqiang Lu, Zhenfei Yin, Mengwei He, Zhihui Wang 0001, Zhiyong Wang 0001, Kun Hu 0008
ICCV6
2025 PUMPS: Skeleton-Agnostic Point-Based Universal Motion Pre-Training for Synthesis in Human Motion Tasks
Clinton Mo, Kun Hu 0008, Chengjiang Long, Dong Yuan 0001, Wan-Chi Siu, Zhiyong Wang 0001
ICCV6
2025 RobAVA: A Large-Scale Dataset and Baseline Towards Video Based Robotic Arm Action Understanding
Baoli Sun, Xinzhu Ma, Anqi Zou, Chuixuan Fan, Zhihui Wang 0001, Kun Lu 0003, Zhiyong Wang 0001
ICCV9
2025 VLIPP: Towards Physically Plausible Video Generation with Vision and Language Informed Physical Prior
abstract
Video diffusion models (VDMs) have advanced significantly in recent years, enabling the generation of highly realistic videos and drawing the attention of the community in their potential as world simulators. However, despite their capabilities, VDMs often fail to produce physically plausible videos due to an inherent lack of understanding of physics, resulting in incorrect dynamics and event sequences. To address this limitation, we propose a novel two-stage image-to-video generation framework that explicitly incorporates physics with vision and language informed physical prior. In the first stage, we employ a Vision Language Model (VLM) as a coarse-grained motion planner, integrating chain-of-thought and physics-aware reasoning to predict a rough motion trajectories/changes that approximate real-world physical dynamics while ensuring the inter-frame consistency. In the second stage, we use the predicted motion trajectories/changes to guide the video generation of a VDM. As the predicted motion trajectories/changes are rough, noise is added during inference to provide freedom to the VDM in generating motion with more fine details. Extensive experimental results demonstrate that our framework can produce physically plausible motion, and comparative evaluations highlight the notable superiority of our approach over existing methods. More video results are available on our Project Page: https://madaoer.github.io/projects/physically_plausible_video_generation.
Xindi Yang, Baolu Li 0001, Zhenfei Yin, Lei Bai 0001, Liqian Ma, Zhiyong Wang 0001, Jianfei Cai 0001, Tien-Tsin Wong, Huchuan Lu, Xu Jia 0012
ICCV7
2025 Diffusing to the Top: Boost Graph Neural Networks with Minimal Hyperparameter Tuning
abstract
Graph Neural Networks (GNNs) are proficient in graph representation learning and achieve promising performance on versatile tasks such as node classification and link prediction. Usually, a comprehensive hyperparameter tuning is essential for fully unlocking GNN's top performance, especially for complicated tasks such as node classification on large graphs and long-range graphs. This is usually associated with high computational and time costs and careful design of appropriate search spaces. This work introduces a graph-conditioned latent diffusion framework (GNN-Diff) to generate high-performing GNNs based on the model checkpoints of sub-optimal hyperparameters selected by a light-tuning coarse search. We validate our method through 166 experiments across four graph tasks: node classification on small, large, and long-range graphs, as well as link prediction. Our experiments involve 10 classic and state-of-the-art target models and 20 publicly available datasets. The results consistently demonstrate that GNN-Diff: (1) boosts the performance of GNNs with efficient hyperparameter tuning; and (2) presents high stability and generalizability on unseen data across multiple generation runs. The code is available at https://github.com/lequanlin/GNN-Diff.
Lequan Lin, Dai Shi, Andi Han, Zhiyong Wang 0001, Junbin Gao
ICLR4
2025 When Graph Neural Networks Meet Dynamic Mode Decomposition
abstract
Graph Neural Networks (GNNs) have emerged as fundamental tools for a wide range of prediction tasks on graph-structured data. Recent studies have drawn analogies between GNN feature propagation and diffusion processes, which can be interpreted as dynamical systems. In this paper, we delve deeper into this perspective by connecting the dynamics in GNNs to modern Koopman theory and its numerical method, Dynamic Mode Decomposition (DMD). We illustrate how DMD can estimate a low-rank, finite-dimensional linear operator based on multiple states of the system, effectively approximating potential nonlinear interactions between nodes in the graph. This approach allows us to capture complex dynamics within the graph accurately and efficiently. We theoretically establish a connection between the DMD-estimated operator and the original dynamic operator between system states. Building upon this foundation, we introduce a family of DMD-GNN models that effectively leverage the low-rank eigenfunctions provided by the DMD algorithm. We further discuss the potential of enhancing our approach by incorporating domain-specific constraints such as symmetry into the DMD computation, allowing the corresponding GNN models to respect known physical properties of the underlying system. Our work paves the path for applying advanced dynamical system analysis tools via GNNs. We validate our approach through extensive experiments on various learning tasks, including directed graphs, large-scale graphs, long-range interactions, and spatial-temporal graphs. We also empirically verify that our proposed models can serve as powerful encoders for link prediction tasks. The results demonstrate that our DMD-enhanced GNNs achieve state-of-the-art performance, highlighting the effectiveness of integrating DMD into GNN frameworks.
Dai Shi, Lequan Lin, Andi Han, Zhiyong Wang 0001, Yi Guo 0001, Junbin Gao
ICLR4
2025 Extended Short- and Long-Range Mesh Learning for Fast and Generalised Garment Simulation
abstract
3D garment simulation is a critical component for producing cloth-based graphics. Recent advancements in graph neural networks (GNNs) offer a promising approach for efficient garment simulation. However, GNNs require extensive message-passing to propagate information such as physical forces and maintain contact awareness across the entire garment mesh, which becomes computationally inefficient at higher resolutions. To address this, we devise a novel GNN-based mesh learning framework with two key components to extend the message-passing range with minimal overhead, namely the Laplacian-Smoothed Dual Message-Passing (LSDMP) and the Geodesic Self-Attention (GSA) modules. LSDMP enhances message-passing with a Laplacian features smoothing process, which efficiently propagates the impact of each vertex to nearby vertices. Concurrently, GSA introduces geodesic distance embeddings to represent the spatial relationship between vertices and utilises attention mechanisms to capture global mesh information. The two modules operate in parallel to ensure both short- and long-range mesh modelling. Extensive experiments demonstrate the state-of-the-art performance of our method, requiring fewer layers and lower inference latency.1
Aoran Liu, Kun Hu 0008, Clinton Mo, ChangYang Li, Zhiyong Wang 0001
ICME5
2025 Player-Team Heterogeneous Interaction Graph Transformer for Soccer Outcome Prediction
abstract
Predicting soccer match outcomes is a challenging task due to the inherently unpredictable nature of the game and the numerous dynamic factors influencing results. While it conventionally relies on meticulous feature engineering, deep learning techniques have recently shown a great promise in learning effective player and team representations directly for soccer outcome prediction. However, existing methods often overlook the heterogeneous nature of interactions among players and teams, which is crucial for accurately modeling match dynamics. To address this gap, we propose HIGFormer (Heterogeneous Interaction Graph Transformer), a novel graph-augmented transformer-based deep learning model for soccer outcome prediction. HIGFormer introduces a multi-level interaction framework that captures both fine-grained player dynamics and high-level team interactions. Specifically, it comprises (1) a Player Interaction Network, which encodes player performance through heterogeneous interaction graphs, combining local graph convolutions with a global graph-augmented transformer; (2) a Team Interaction Network, which constructs interaction graphs from a team-to-team perspective to model historical match relationships; and (3) a Match Comparison Transformer, which jointly analyzes both team and player-level information to predict match outcomes. Extensive experiments on the WyScout Open Access Dataset, a large-scale real-world soccer dataset, demonstrate that HIGFormer significantly outperforms existing methods in prediction accuracy. Furthermore, we provide valuable insights into leveraging our model for player performance evaluation, offering a new perspective on talent scouting and team strategy analysis.
Lintao Wang 0002, Shiwen Xu, Michael Horton 0001, Joachim Gudmundsson, Zhiyong Wang 0001
KDD (2)5
2025 DSACap: Enhancing Visual-Semantic Alignment with Diffusion-based Framework for Image Captioning
Liangyu Fu, Junbo Wang 0003, Qiangguo Jin, Hongsong Wang 0001, Jing Ya, Linjiang Huang, Jiangbin Zheng 0001, Zhiyong Wang 0001
ACM Multimedia11
2025 RealDTT: Towards A Comprehensive Real-World Dataset for Tampered Text Detection
Junxian Duan, Fan Ji, Zhiyong Wang 0001, Huaibo Huang
Int. J. Comput. Vis.5
2025 Graph traverse reference network for sign language corpus retrieval in the wild
abstract
Sign languages are the primary languages of the deaf community as well as hearing individuals who are unable to speak, which engage the visual-manual modality to convey meanings. In recent years, there has been an explosive growth of sign language videos available from video streaming and social media service platforms. Given the size of these corpora, sign language users often face significant challenges in effectively acquiring the information they need. Therefore, we propose a novel deep learning architecture, namely Graph Traverse Reference Network (GTRN), allowing visual signing queries to retrieve relevant sign language videos (documents) from a large corpus. GTRN introduces a traverse graph, which provides coarse-to-fine reference information in a hierarchical manner from frame-level to body-part-level observations. A reference-based attention is devised to obtain the embedding for a visual input of each level, which allows the computations to be allocated and processed at difference locations regarding local devices and central servers. A contrastive learning strategy optimizes GTRN in pursuit of a joint latent space for the queries and the documents by their meanings. Moreover, GTRN is compatible to leverage existing general visual representation foundation models, by which their resulted embeddings are used as the frame-level reference of GTRN. To the best of our knowledge, it is one of the first studies on using visual signing queries for retrieving sign language videos in a real-world setting and comprehensive experiments were conducted which demonstrated the effectiveness of our proposed method. • A deep network formulates sign language embeddings for video queries and documents. • A reference-based attention method and contrastive strategy for embedding alignment. • Comprehensive experiments and analyses on a real-world sign language corpus.
Kun Hu 0008, Fengxiang He, Adam Schembri, Zhiyong Wang 0001
Neurocomputing4
2025 Music source separation via hybrid waveform and spectrogram based generative adversarial network
Qiuxia Wu, Haipeng Deng, Kun Hu 0008, Zhiyong Wang 0001
Multim. Tools Appl.4
2025 Uncertainty-guided attention learning for malaria parasite detection in thick blood smears
abstract
Malaria may seriously threaten an individual's health and wellbeing, and early screening is pivotal for timely treatment and recovery. In malaria screening, thick blood smears are exploited to count the parasites and assess the severity of the disease. Parasites are tiny objects that can be found in high resolution blood smear images, which renders them difficult for detection. Other than using object detection based methods, prior works also applied image classification techniques to this problem. They first extracted image patches from blood smears as parasite candidates and then utilized convolutional neural networks to classify these patches as parasites or non-parasites. However, these approaches overlook the fact that the blood smear images may contain noises, errors, and background artifacts, which introduces uncertainty and makes the model predictions less stable. In this work, we propose an uncertainty-guided attention learning based network for malaria parasite detection from thick blood smears, which incorporates pixel attention mechanism to identify more fine-grained and pixel-wise informative features, to improve the classification capability of our model. We further put uncertainty estimation on channels of the feature map to guide pixel attention learning, such that the features from channels with higher uncertainty are considered unreliable and are thus restrictively exploited by pixel attention learning. To estimate channel-wise uncertainty, we introduce the Bayesian channel attention, which reformulates the traditional channel attention under the Bayesian framework. As a result, it denotes channel uncertainties with estimated variances that guide the pixel attention learning. We compared to several state-of-the-art baselines on two public datasets using parasite-level and patient-level evaluations. The proposed method demonstrates superior performance with respect to most metrics on two datasets, especially achieving highest average precision (AP) scores in both parasite and patient-level scenarios.
Hao Xiong 0001, Zhiyong Wang 0001, Roneel V. Sharan, Shlomo Berkovsky
Neural Networks2
2025 Motion-guided semantic alignment for line art animation colorization
Ning Wang 0025, Hairui Yang, Hong Zhang 0011, Zhiyong Wang 0001, Zhihui Wang 0001
Pattern Recognit.5
2025 Referring Video Object Segmentation With Cross-Modality Proxy Queries
abstract
Referring video object segmentation (RVOS) is an emerging cross-modality task that aims to generate pixel-level maps of the target objects referred by given textual expressions. The main concept involves learning an accurate alignment visual elements and language expressions within a semantic space. Recent approaches address cross-modality alignment through conditional queries, tracking the target object using a queryresponse based mechanism built upon transformer structure. However, they exhibit two limitations: (1) these conditional queries, identifying the same object across different frames through the same query, lack inter-frame dependency and variation modeling, making accurate target tracking challenging amid significant frame-to-frame variations; and (2) they handle the temporal feature of a video and build visual-language interaction sequentially, integrating textual constraints belatedly, which may cause the video features potentially focus on the non-referred objects. Therefore, we propose a novel RVOS architecture called ProxyFormer, which introduces a set of proxy queries to integrate visual and text semantics and facilitate the flow of semantics between them. By progressively updating and propagating proxy queries across multiple stages of video feature encoder, ProxyFormer ensures that the video features are as focused as much possible on the object of interest. This dynamic evolution of the queries across video also enables the proxy queries to establish inter-frame dependencies, enhancing the accuracy and coherence of object tracking throughout the video sequence. To mitigate the high computational costs associated with full spatio-temporal interactions between video and proxy queries, we propose to decouple cross-modality interactions into their temporal and spatial dimensions, respectively. Additionally, we design a Joint Semantic Consistency (JSC) training strategy to align semantic consensus between the proxy queries and the combined videotext pairs. Comprehensive experiments on four widely used RVOS benchmarks, i.e., Ref-Youtube-VOS, Ref-DAVIS17, A2D-Sentences and JHMDB-Sentences, clearly demonstrate the superiority of our ProxyFormer to the state-of-the-art methods
Baoli Sun, Xinzhu Ma, Zhihui Wang 0001, Zhiyong Wang 0001
IEEE Trans. Multim.5
2025 Dual-Attention Transformers for Class-Incremental Learning: A Tale of Two Memories
abstract
Class-incremental learning (Class-IL) aims to continuously learn a model from a sequence of tasks, which suffers from the issue of catastrophic forgetting. Recently, a few transformer based methods are proposed to address this issue by transferring self-attention into task-specific attention. However, these methods utilize shared task-specific attention modules across the whole incremental learning process, and are unable to achieve the balance between consolidation and plasticity, i.e., to remember the knowledge learned from previous tasks and absorb the knowledge from the current task simultaneously. Motivated by the mechanism of LSTM and hippocampus memory, we point out that dual attention on long and short-term memories can handle the consolidation-plasticity dilemma of Class-IL. Typically, we propose Dual-Attention Transformers (DAFormer) to learn external attention and internal attention. The former utilizes sample-dependent keys which exclusively focused on the new tasks, while the latter consolidates the knowledge from previous tasks by using sample-agnostic keys. We present two editions of DAFormer: DAFormer-S and DAFormer-M: the former utilizes shared external keys and maintains a small parameter size, while the latter utilizes multiple external keys and enhances the long-term memory. Furthermore, we propose the$K$-nearest neighbor invariant based distillation scheme, which distills knowledge from previous tasks to current task by maintaining the same neighborhood relationship of each sample over old and new models. Experimental results onCIFAR-100,ImageNet-subsetandImageNet-fulldemonstrate that DAFormer significantly outperforms all the state-of-the-art parameter-static and parameter-growing methods.
Shaofan Wang 0001, Zhiyong Wang 0001, Boyue Wang
IEEE Trans. Multim.4
2025 Frameless Graph Knowledge Distillation
abstract
Knowledge distillation (KD) has shown great potential for transferring knowledge from a complex teacher model to a simple student model in which the heavy learning task can be accomplished efficiently and without losing too much prediction accuracy. Recently, many attempts have been made by applying the KD mechanism to graph representation learning models such as graph neural networks (GNNs) to accelerate the model's inference speed via student models. However, many existing KD-based GNNs utilize multilayer perceptron (MLP) as a universal approximator in the student model to imitate the teacher model's process without considering the graph knowledge from the teacher model. In this work, we provide a KD-based framework on multiscaled GNNs, known as graph framelet, and prove that by adequately utilizing the graph knowledge in a multiscaled manner provided by graph framelet decomposition, the student model is capable of adapting both homophilic and heterophilic graphs and has the potential of alleviating the oversquashing issue with a simple yet effective graph surgery. Furthermore, we show how the graph knowledge supplied by the teacher is learned and digested by the student model via both algebra and geometry. Comprehensive experiments show that our proposed model can generate learning accuracy identical to or even surpass the teacher model while maintaining the high speed of inference.
Dai Shi, Zhiqi Shao, Junbin Gao, Zhiyong Wang 0001, Yi Guo 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Terrain Diffusion Network: Climatic-Aware Terrain Generation with Geological Sketch Guidance
abstract
Sketch-based terrain generation seeks to create realistic landscapes for virtual environments in various applications such as computer games, animation and virtual reality. Recently, deep learning based terrain generation has emerged, notably the ones based on generative adversarial networks (GAN). However, these methods often struggle to fulfill the requirements of flexible user control and maintain generative diversity for realistic terrain. Therefore, we propose a novel diffusion-based method, namely terrain diffusion network (TDN), which actively incorporates user guidance for enhanced controllability, taking into account terrain features like rivers, ridges, basins, and peaks. Instead of adhering to a conventional monolithic denoising process, which often compromises the fidelity of terrain details or the alignment with user control, a multi-level denoising scheme is proposed to generate more realistic terrains by taking into account fine-grained details, particularly those related to climatic patterns influenced by erosion and tectonic activities. Specifically, three terrain synthesisers are designed for structural, intermediate, and fine-grained level denoising purposes, which allow each synthesiser concentrate on a distinct terrain aspect. Moreover, to maximise the efficiency of our TDN, we further introduce terrain and sketch latent spaces for the synthesizers with pre-trained terrain autoencoders. Comprehensive experiments on a new dataset constructed from NASA Topology Images clearly demonstrate the effectiveness of our proposed method, achieving the state-of-the-art performance. Our code is available at https://github.com/TDNResearch/TDN.
Zexin Hu, Kun Hu 0008, Clinton Mo, Zhiyong Wang 0001
AAAI5
2024 Autoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image Generation
abstract
A 360-degree (omni-directional) image provides an all-encompassing spherical view of a scene. Recently, there has been an increasing interest in synthesising 360-degree images from conventional narrow field of view (NFoV) images captured by digital cameras and smartphones, for providing immersive experiences in various scenarios such as virtual reality. Yet, existing methods typically fall short in synthesizing intricate visual details or ensure the generated images align consistently with user-provided prompts. In this study, autoregressive omni-aware generative network (AOG-Net) is proposed for 360-degree image generation by outpainting an incomplete 360-degree image progressively with NFoV and text guidances joinly or individually. This autoregressive scheme not only allows for deriving finer-grained and text-consistent patterns by dynamically generating and adjusting the process but also offers users greater flexibility to edit their conditions throughout the generation process. A global-local conditioning mechanism is devised to comprehensively formulate the outpainting guidance in each autoregressive step. Text guidances, omni-visual cues, NFoV inputs and omni-geometry are encoded and further formulated with cross-attention based transformers into a global stream and a local stream into a conditioned generative backbone model. As AOG-Net is compatible to leverage large-scale models for the conditional encoder and the generative prior, it enables the generation to use extensive open-vocabulary text guidances. Comprehensive experiments on two commonly used 360-degree image datasets for both indoor and outdoor settings demonstrate the state-of-the-art performance of our proposed method. Our code is available at https://github.com/zhuqiangLu/AOG-NET-360.
Zhuqiang Lu, Kun Hu 0008, Lei Bai 0001, Zhiyong Wang 0001
AAAI5
2024 SurgicalSAM: Efficient Class Promptable Surgical Instrument Segmentation
abstract
The Segment Anything Model (SAM) is a powerful foundation model that has revolutionised image segmentation. To apply SAM to surgical instrument segmentation, a common approach is to locate precise points or boxes of instruments and then use them as prompts for SAM in a zero-shot manner. However, we observe two problems with this naive pipeline: (1) the domain gap between natural objects and surgical instruments leads to inferior generalisation of SAM; and (2) SAM relies on precise point or box locations for accurate segmentation, requiring either extensive manual guidance or a well-performing specialist detector for prompt preparation, which leads to a complex multi-stage pipeline. To address these problems, we introduce SurgicalSAM, a novel end-to-end efficient-tuning approach for SAM to effectively integrate surgical-specific information with SAM’s pre-trained knowledge for improved generalisation. Specifically, we propose a lightweight prototype-based class prompt encoder for tuning, which directly generates prompt embeddings from class prototypes and eliminates the use of explicit prompts for improved robustness and a simpler pipeline. In addition, to address the low inter-class variance among surgical instrument categories, we propose contrastive prototype learning, further enhancing the discrimination of the class prototypes for more accurate class prompting. The results of extensive experiments on both EndoVis2018 and EndoVis2017 datasets demonstrate that SurgicalSAM achieves state-of-the-art performance while only requiring a small number of tunable parameters. The source code is available at https://github.com/wenxi-yue/SurgicalSAM.
Wenxi Yue, Jing Zhang 0037, Kun Hu 0008, Yong Xia 0001, Jiebo Luo 0001, Zhiyong Wang 0001
AAAI6
2024 AEDNet: Adaptive Embedding and Multiview-Aware Disentanglement for Point Cloud Completion
Zhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang 0001, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun
ECCV (11)4
2024 Motion Keyframe Interpolation for Any Human Skeleton via Temporally Consistent Point Cloud Sampling and Reconstruction
Clinton Mo, Kun Hu 0008, Chengjiang Long, Dong Yuan 0001, Zhiyong Wang 0001
ECCV (82)5
2024 Identity-Consistent Diffusion Network for Grading Knee Osteoarthritis Progression in Radiographic Imaging
Wenhua Wu 0005, Kun Hu 0008, Wenxi Yue, Wei Li 0058, Milena Simic, ChangYang Li, Wei Xiang 0001, Zhiyong Wang 0001
ECCV (84)8
2024 Radio Frequency Signal based Human Silhouette Segmentation: A Sequential Diffusion Approach
abstract
Radio frequency (RF) signals have been proved to be flexible for human silhouette segmentation (HSS) under complex environments. Existing studies are mainly based on a one-shot approach, which lacks a coherent projection ability from the RF domain. Additionally, the spatio-temporal patterns have not been fully explored for human motion dynamics in HSS. Therefore, we propose a two-stage Sequential Diffusion Model (SDM) to progressively synthesize high-quality segmentation jointly with the considerations on motion dynamics. Cross-view transformation blocks are devised to guide the diffusion model in a multi-scale manner for comprehensively characterizing human related patterns in an individual frame such as directional projection from signal planes. Moreover, spatio-temporal blocks are devised to fine-tune the frame-level model to incorporate spatio-temporal contexts and motion dynamics, enhancing the consistency of the segmentation maps. Comprehensive experiments on a public benchmark - HIBER demonstrate the state-of-the-art performance of our method with an IoU 0.732. Our code is available at https://github.com/ph-w2000/SDM.
Penghui Wen, Kun Hu 0008, Dong Yuan 0001, ChangYang Li, Zhiyong Wang 0001
ICME6
2024 Bridging the Gap: Sketch-Aware Interpolation Network for High-Quality Animation Sketch Inbetweening
abstract
Hand-drawn 2D animation workflow is typically initiated with the creation of sketch keyframes. Subsequent manual inbetweens are crafted for smoothness, which is a labor-intensive process and the prospect of automatic animation sketch interpolation has become highly appealing. Yet, common frame interpolation methods are generally hindered by two key issues: 1) limited texture and colour details in sketches, and 2) exaggerated alterations between two sketch keyframes. To overcome these issues, we propose a novel deep learning method - Sketch-Aware Interpolation Network (SAIN). This approach incorporates multi-level guidance that formulates region-level correspondence, stroke-level correspondence and pixel-level dynamics. A multi-stream U-Transformer is then devised to characterize sketch inbetweening patterns using these multi-level guides through the integration of self / cross-attention mechanisms. Additionally, to facilitate future research on animation sketch inbetweening, we constructed a large-scale dataset - STD-12K, comprising 30 sketch animation series in diverse artistic styles. Comprehensive experiments on this dataset convincingly show that our proposed SAIN surpasses the state-of-the-art interpolation methods. Our code and dataset are avaliable in https://github.com/none-master/SAIN.
Kun Hu 0008, Wei Bao 0001, Chang Wen Chen, Zhiyong Wang 0001
ACM Multimedia5
2024 SITransformer: Shared Information-Guided Transformer for Extreme Multimodal Summarization
Lintao Wang 0002, Xiaogang Zhu 0001, Xuequan Lu, Zhiyong Wang 0001, Kun Hu 0008
MMAsia5
2024 T2QRM: Text-Driven Quadruped Robot Motion Generation
Kun Hu 0008, Zhiyong Wang 0001, Wenxiong Kang
MMAsia5
2024 Fast Online Adaptation of Visual SLAM via Variational Information Transfer and Preservation
Sangni Xu, Hao Xiong 0001, Qiuxia Wu, Shlomo Berkovsky, Zhiyong Wang 0001
MMAsia6
2024 Underwater Image Enhancement via Domain Adaptive Transfer Learning and Hybrid Reinforcement Model
Qing Hu 0001, Zhiyong Wang 0001
MMAsia5
2024 Point Cloud Normal Estimation via Representation Learning on Height Maps
abstract
Point Cloud Normal Estimation via Representation Learning on Height Maps
Dasith de Silva Edirimuni, Ye Zhu 0002, Shang Gao 0003, Zhiyong Wang 0001, Antonio Robles-Kelly, Xuequan Lu
MMAsia5
2024 Joint Video Denoising and Super-Resolution Network for IoT Cameras
abstract
IoT (Internet of Things) cameras have widely been deployed over the last few years. These cameras are often with limited hardware so that they can only capture noisy videos in low resolution. In this work, we propose the joint video denoising and super-resolution network for IoT cameras, which consists of the noise-robust moving-attention (NRMA) module and the noise-eliminated upsampling (NEU) module. In NRMA, we adopt a coarse-to-fine approach by first extracting the coarse flow and then refining through bi-directional feature propagation among adjacent frames. In NEU, we further utilize inner-frame features for noise-elimination and upsampling. Through this approach, we avoid the negative effects brought by applying denoising and super-resolution in tandem, and enhance the reconstruction of moving objects by the embedded attention layers in NRMA. We conduct our experiments on both synthetic datasets, which utilize existing data with additive white Gaussian noise (AWGN), and a realistic dataset captured using a pair of IoT and professional cameras. Our extensive experimental results demonstrate that our proposed method significantly reduces noise and enhances detail in both types of datasets. Notably, our approach outperforms the state-of-the-art benchmark (RealBasicVSR) by an average of 5.24 dB on the existing datasets (with noise level σ = 20) and by 0.95 dB on the realistic dataset in terms of PSNR.
Liming Ge, Wei Bao 0001, Xinyi Sheng, Dong Yuan 0001, Bing Bing Zhou, Zhiyong Wang 0001
IEEE Internet Things J.6
2024 ProtoSimi: label correction for fine-grained visual categorization
abstract
Abstract Deep models trained by using clean data have achieved tremendous success in fine-grained image classification. Yet, they generally suffer from significant performance degradation when encountering noisy labels. Existing approaches to handle label noise, though proved to be effective for generic object recognition, usually fail on fine-grained data. The reason is that, on fine-grained data, the category difference is subtle and the training sample size is small. Then deep models could easily overfit the noisy labels. To improve the robustness of deep models on noisy data for fine-grained visual categorization, in this paper, we propose a novel learning framework named ProtoSimi. Our method employs an adaptive label correction strategy, ensuring effective learning on limited data. Specifically, our approach considers the criteria of exploring the effectiveness of both global class-prototype and part class-prototype similarities in identifying and correcting labels of samples. We evaluate our method on three standard benchmarks of fine-grained recognition. Experimental results show that our method outperforms the existing label noisy methods by a large margin. In ablation studies, we also verify that our method is non-sensitive to hyper-parameters selection and can be integrated with other FGVC methods to increase the generalization performance.
Jialiang Shen, Yu Yao 0005, Shaoli Huang, Zhiyong Wang 0001, Jing Zhang 0037, Ruxing Wang, Jun Yu 0001, Tongliang Liu
Mach. Learn.4
2024 TLDW: Extreme Multimodal Summarization of News Videos
abstract
Multimodal summarisation with multimodal output is drawing increasing attention due to the rapid growth of multimedia data. While several methods have been proposed to summarise visual-text contents, their multimodal outputs are not succinct enough at an extreme level to address the information overload issue. To the end of extreme multimodal summarisation, we introduce a new task, eXtreme Multimodal Summarisation with Multimodal Output (XMSMO) for the scenario of TL;DW - Too Long; Didn’t Watch, akin to TL;DR. XMSMO aims to summarise a video-document pair into a summary with an extremely short length, which consists of one cover frame as the visual summary and one sentence as the textual summary. We propose a novel unsupervised Hierarchical Optimal Transport Network (HOT-Net) consisting of three components: hierarchical multimodal encoder, hierarchical multimodal fusion decoder, and optimal transport solver. Our method is trained, without using reference summaries, by optimising the visual and textual coverage from the perspectives of the distance between the semantic distributions under optimal transport plans. To facilitate the study on this task, we constructed a large-scale dataset, XMSMO-News, by harvesting 4,891 video-document pairs. The experimental results show that our method achieves promising performance in terms of ROUGE and IoU metrics. Our dataset and source code will be publicly available in GitHub.
Peggy Tang, Kun Hu 0008, Lei Zhang 0001, Jiebo Luo 0001, Zhiyong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Joint Spatial and Spectral Graph-Based Consistent Self-Representation for Unsupervised Hyperspectral Band Selection
abstract
Band selection (BS), which effectively reduces spectral dimensionality, stands out as a leading focus within hyperspectral image (HSI) analysis. Self-representation (SR) has surfaced as a favored technique in this domain due to its applicability to BS and unsupervised nature. However, the existing SR-based BS approaches only leverage either spatial or spectral relationships, with few integrating both while concentrating on the representation level rather than the selection level. In addition, employing all spatial pixels for spatial relationship utilization leads to considerable computational complexity. Therefore, this article proposes joint spatial and spectral graph-based consistent SR (JSSGCSR) to more effectively exploit spatial and spectral relationships for BS, which separately conducts SR to handle each view of spatial and spectral graphs to better consider two different structure characteristics, and ultimately integrates two SR results to achieve a unified and robust representative band set by imposing consistent sparsity pattern on their joint representation coefficients. In addition, the spatial and spectral relationships are integrated into different data spaces, that is, spectral graph SR and spatial graph SR are, respectively, conducted in the original HSI and the segmented and pooled HSI, which not only reduces the influence of superpixel segmentation on spectral relationships, but also improves the efficiency of spatial relationship utilization. Experimental results on three benchmark datasets have demonstrated the effectiveness of the proposed JSSGCSR in HSI classification tasks.
Mingyang Ma 0004, Fan Li 0003, Zhiyong Wang 0001, Shaohui Mei
IEEE Trans. Geosci. Remote. Sens.4
2024 TriLA: Triple-Level Alignment Based Unsupervised Domain Adaptation for Joint Segmentation of Optic Disc and Optic Cup
abstract
Cross-domain joint segmentation of optic disc and optic cup on fundus images is essential, yet challenging, for effective glaucoma screening. Although many unsupervised domain adaptation (UDA) methods have been proposed, these methods can hardly achieve complete domain alignment, leading to suboptimal performance. In this paper, we propose a triple-level alignment (TriLA) model to address this issue by aligning the source and target domains at the input level, feature level, and output level simultaneously. At the input level, a learnable Fourier domain adaptation (LFDA) module is developed to learn the cut-off frequency adaptively for frequency-domain translation. At the feature level, we disentangle the style and content features and align them in the corresponding feature spaces using consistency constraints. At the output level, we design a segmentation consistency constraint to emphasize the segmentation consistency across domains. The proposed model is trained on the RIGA+ dataset and widely evaluated on six different UDA scenarios. Our comprehensive results not only demonstrate that the proposed TriLA substantially outperforms other state-of-the-art UDA methods in joint segmentation of optic disc and optic cup, but also suggest the effectiveness of the triple-level alignment strategy.
Ziyang Chen 0003, Yongsheng Pan, Yiwen Ye, Zhiyong Wang 0001, Yong Xia 0001
IEEE J. Biomed. Health Informatics4
2024 Higher Order Polynomial Transformer for Fine-Grained Freezing of Gait Detection
abstract
Freezing of Gait (FoG) is a common symptom of Parkinson's disease (PD), manifesting as a brief, episodic absence, or marked reduction in walking, despite a patient's intention to move. Clinical assessment of FoG events from manual observations by experts is both time-consuming and highly subjective. Therefore, machine learning-based FoG identification methods would be desirable. In this article, we address this task as a fine-grained human action recognition problem based on vision inputs. A novel deep learning architecture, namely, higher order polynomial transformer (HP-Transformer), is proposed to incorporate pose and appearance feature sequences to formulate fine-grained FoG patterns. In particular, a higher order self-attention mechanism is proposed based on higher order polynomials. To this end, linear, bilinear, and trilinear transformers are formulated in pursuit of discriminative fine-grained representations. These representations are treated as multiple streams and further fused by a cross-order fusion strategy for FoG detection. Comprehensive experiments on a large in-house dataset collected during clinical assessments demonstrate the effectiveness of the proposed method, and an area under the receiver operating characteristic (ROC) curve (AUC) of 0.92 is achieved for detecting FoG.
Renfei Sun, Kun Hu 0008, Kaylena A. Ehgoetz Martens, Markus Hagenbuchner, Ah Chung Tsoi, Mohammed Bennamoun, Simon J. G. Lewis, Zhiyong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.8
2024 Digging into Depth and Color Spaces: A Mapping Constraint Network for Depth Super-Resolution
abstract
Scene depth super-resolution (DSR) poses an inherently ill-posed problem due to the extremely large space of one-to-many mapping functions from a given low-resolution (LR) depth map, which possesses limited depth information, to multiple plausible high-resolution (HR) depth maps. This characteristic renders the task highly challenging, as identifying an optimal solution becomes significantly intricate amidst this multitude of potential mappings. While simplistic constraints have been proposed to address the DSR task, the relationship between LR and HR depth maps and the color image has not been thoroughly investigated. In this paper, we introduce a novel mapping constraint network (MCNet) that incorporates additional constraints derived from both LR depth maps and color images. This integration aims to optimize the space of mapping functions and enhance the performance of DSR. Specifically, alongside the primary DSR network (DSRNet) dedicated to learning LR-to-HR mapping, we have developed an auxiliary degradation network (ADNet) that operates in reverse, generating the LR depth map from the reconstructed HR depth map to obtain depth features in LR space. To enhance the learning process of DSRNet in LR-to-HR mapping, we introduce two mapping constraints in LR space: (1) the cycle-consistent constraint, which offers additional supervision by establishing a closed loop between LR-to-HR and HR-to-LR mappings, and (2) the region-level contrastive constraint, aimed at reinforcing region-specific HR representations by explicitly modeling the consistency between LR and HR spaces. To leverage the color image effectively, we introduce a feature screening module to adaptively fuse color features at different layers, which can simultaneously maintain strong structural context and suppress texture distraction through subspace generation and image projection. Comprehensive experimental results across synthetic and real-world benchmark datasets unequivocally demonstrate the superiority of our proposed method over state-of-the-art DSR methods. Our MCNet achieves an average MAD reduction of 3.7% and 7.5% over state-of-the-art DSR method for ×8 and ×16 cases on Milddleburry dataset, respectively, without incurring additional costs during inference.
Baoli Sun, Tiantian Yan, Xinchen Ye, Zhihui Wang 0001, Zhiyong Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.7
2024 Discriminative Segment Focus Network for Fine-grained Video Action Recognition
abstract
Fine-grained video action recognition aims at identifying minor and discriminative variations among fine categories of actions. While many recent action recognition methods have been proposed to better model spatio-temporal representations, how to model the interactions among discriminative atomic actions to effectively characterize inter-class and intra-class variations has been neglected, which is vital for understanding fine-grained actions. In this work, we devise a Discriminative Segment Focus Network (DSFNet) to mine the discriminability of segment correlations and localize discriminative action-relevant segments for fine-grained video action recognition. Firstly, we propose a hierarchic correlation reasoning (HCR) module which explicitly establishes correlations between different segments at multiple temporal scales and enhances each segment by exploiting the correlations with other segments. Secondly, a discriminative segment focus (DSF) module is devised to localize the most action-relevant segments from the enhanced representations of HCR by enforcing the consistency between the discriminability and the classification confidence of a given segment with a consistency constraint. Finally, these localized segment representations are combined with the global action representation of the whole video for boosting final recognition. Extensive experimental results on two fine-grained action recognition datasets, i.e., FineGym and Diving48, and two action recognition datasets, i.e., Kinetics400 and Something-Something, demonstrate the effectiveness of our approach compared with the state-of-the-art methods.
Baoli Sun, Xinchen Ye, Tiantian Yan, Zhihui Wang 0001, Zhiyong Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2023 Multi-Scale Control Signal-Aware Transformer for Motion Synthesis without Phase
abstract
Synthesizing controllable motion for a character using deep learning has been a promising approach due to its potential to learn a compact model without laborious feature engineering. To produce dynamic motion from weak control signals such as desired paths, existing methods often require auxiliary information such as phases for alleviating motion ambiguity, which limits their generalisation capability. As past poses often contain useful auxiliary hints, in this paper, we propose a task-agnostic deep learning method, namely Multi-scale Control Signal-aware Transformer (MCS-T), with an attention based encoder-decoder architecture to discover the auxiliary information implicitly for synthesizing controllable motion without explicitly requiring auxiliary information such as phase. Specifically, an encoder is devised to adaptively formulate the motion patterns of a character's past poses with multi-scale skeletons, and a decoder driven by control signals to further synthesize and predict the character's state by paying context-specialised attention to the encoded past motion patterns. As a result, it helps alleviate the issues of low responsiveness and slow transition which often happen in conventional methods not using auxiliary information. Both qualitative and quantitative experimental results on an existing biped locomotion dataset, which involves diverse types of motion transitions, demonstrate the effectiveness of our method. In particular, MCS-T is able to successfully generate motions comparable to those generated by the methods using auxiliary information.
Lintao Wang 0002, Kun Hu 0008, Lei Bai 0001, Wanli Ouyang, Zhiyong Wang 0001
AAAI6
2023 Continuous Intermediate Token Learning with Implicit Motion Manifold for Keyframe Based Motion Interpolation
abstract
Deriving sophisticated 3D motions from sparse keyframes is a particularly challenging problem, due to continuity and exceptionally skeletal precision. The action features are often derivable accurately from the full series of keyframes, and thus, leveraging the global context with transformers has been a promising data-driven embedding approach. However, existing methods are often with inputs of interpolated intermediate frame for continuity using basic interpolation methods with keyframes, which result in a trivial local minimum during training. In this paper, we propose a novel framework to formulate latent motion manifolds with keyframe-based constraints, from which the continuous nature of intermediate token representations is considered. Particularly, our proposed framework consists of two stages for identifying a latent motion subspace, i.e., a keyframe encoding stage and an intermediate token generation stage, and a subsequent motion synthesis stage to extrapolate and compose motion data from manifolds. Through our extensive experiments conducted on both the LaFAN1 and CMU Mocap datasets, our proposed method demonstrates both superior interpolation accuracy and high visual similarity to ground truth motions.
Clinton Mo, Kun Hu 0008, Chengjiang Long, Zhiyong Wang 0001
CVPR4
2023 VAPCNet: Viewpoint-Aware 3D Point Cloud Completion
abstract
Most existing learning-based 3D point cloud completion methods ignore the fact that the completion process is highly coupled with the viewpoint of a partial scan. However, the various viewpoints of incompletely scanned objects in real-world applications are normally unknown and directly estimating the viewpoint of each incomplete object is usually time-consuming and leads to huge annotation cost. In this paper, we thus propose an unsupervised viewpoint representation learning scheme for 3D point cloud completion without explicit viewpoint estimation. To be specific, we learn abstract representations of partial scans to distinguish various viewpoints in the representation space rather than the explicit estimation in the 3D space. We also introduce a Viewpoint-Aware Point cloud Completion Network (VAPCNet) with flexible adaption to various viewpoints based on the learned representations. The proposed viewpoint representation learning scheme can extract discriminative representations to obtain accurate viewpoint information. Reported experiments on two popular public datasets show that our VAPCNet achieves state-of-the-art performance for the point cloud completion task. Source code is available at https://github.com/FZH92128/VAPCNet.
Zhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang 0001, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun
ICCV4
2023 Material-Aware Self-Supervised Network for Dynamic 3D Garment Simulation
abstract
Dynamic 3D garment simulation has various applications in many domains. Recently, self-supervised learning for this task has been studied to reduce annotation costs. However, different material characteristics of garments have been rarely explored, limiting the generalization and flexibility of existing methods. Therefore, in this paper, a novel self-supervised deep learning architecture is proposed, namely Material-aware Self-supervised Network (MSN), as a material-aware approach for dynamically simulating garments with different materials. Specifically, a material-aware parameterized regressor is introduced based on the observation that material characteristics change continuously regarding the fabric parameters. As a result, MSN realises real-time garment simulation with various material properties without model re-training. Moreover, to simulate garments of different categories (e.g., t-shirts vs. dresses), a sampling-based linear skinning strategy is studied in MSN. Comprehensive experiments on the widely used AMASS dataset demonstrated the effectiveness of MSN both quantitatively and qualitatively.
Aoran Liu, Kun Hu 0008, Wenxi Yue, Qiuxia Wu, Zhiyong Wang 0001
ICME5
2023 Online Visual SLAM Adaptation against Catastrophic Forgetting with Cycle-Consistent Contrastive Learning
abstract
Visual SLAM (Simultaneous Localisation and Mapping) aims to simultaneously estimate camera poses and depth maps from navigation videos captured. While recent deep learning based methods have achieved great success on this task, they tend to work well on source domain data and suffer from performance degradation on the unseen data of target domain. Hence, we propose an online adaptation approach to continuously adapt a pre-trained visual SLAM model to changing environments in a self-supervised manner. To preserve pre-learned knowledge against catastrophic forgetting, we perform updating on a novel adapter proposed rather than fine-tuning the whole model for adaptation. The adapter includes a cross-domain feature translation module that translates pre-learned features into translated features suitable for adaptation. Ideally, the translated new features should not only contain pre-learned knowledge but also substantially distinct from pre-learned features since these two features represent different domains. We thus introduce cycle-consistent contrastive learning to maximize the dissimilarity between these two features by enlarging the distance between them in the feature space. Besides, our contrastive learning method exploiting cycle-consistency contraint enables the translated features to be transferred back to the pre-learned ones, which helps the translated features better preserve pre-learned knowledge. Comprehensive experiments on both synthetic and real-world datasets demonstrate superior adaptation performance of our proposed method over several state-of-the-art baselines.
Sangni Xu, Hao Xiong 0001, Qiuxia Wu, Zhihui Wang 0001, Zhiyong Wang 0001
ICRA6
2023 Embedding the Self-Organisation of Deep Feature Maps in the Hamburger Framework can Yield Better and Interpretable Results
abstract
The popularity of neural networks that explicitly utilise the global correlation structure of their features have become vastly more popular ever since the Transformer architecture was developed. We propose to embed unsupervised Self-Organising Maps within neural networks as a means to model the global correlation structure. By enforcing topological preservation therein, such a neural network is able to represent more complex correlation structures and to produce interpretable visualisations as a byproduct. We validate this approach by comparing with the existing state-of-the-art attention substitute within its own ‘Hamburger’ framework and by illustrating maps learnt by the module. Overall, this paper serves as a proof of concept for integrating Self-Organising Maps within a supervised network.
Jack Humphreys, Markus Hagenbuchner, Zhiyong Wang 0001, Ah Chung Tsoi
IJCNN3
2023 Robust Audio Anti-Spoofing with Fusion-Reconstruction Learning on Multi-Order Spectrograms
abstract
Robust audio anti-spoofing has been increasingly challenging due to the recent advancements on deepfake techniques. While spectrograms have demonstrated their capability for anti-spoofing, complementary information presented in multi-order spectral patterns have not been well explored, which limits their effectiveness for varying spoofing attacks. Therefore, we propose a novel deep learning method with a spectral fusion-reconstruction strategy, namely S2pecNet, to utilise multi-order spectral patterns for robust audio anti-spoofing representations. Specifically, spectral patterns up to second-order are fused in a coarse-to-fine manner and two branches are designed for the fine-level fusion from the spectral and temporal contexts. A reconstruction from the fused representation to the input spectrograms further reduces the potential fused information loss. Our method achieved the state-of-the-art performance with an EER of 0.77% on a widely used dataset: ASVspoof2019 LA Challenge.
Penghui Wen, Kun Hu 0008, Wenxi Yue, Sen Zhang 0006, Wanlei Zhou 0001, Zhiyong Wang 0001
INTERSPEECH6
2023 Exploring Coarse-to-Fine Action Token Localization and Interaction for Fine-grained Video Action Recognition
abstract
Vision transformers have achieved impressive performance for video action recognition due to their strong capability of modeling long-range dependencies among spatio-temporal tokens. However, as for fine-grained actions, subtle and discriminative differences mainly exist in the regions of actors, directly utilizing vision transformers without removing irrelevant tokens will compromise recognition performance and lead to high computational costs. In this paper, we propose a coarse-to-fine action token localization and interaction network, namely C2F-ALIN, that dynamically localizes the most informative tokens at a coarse granularity and then partitions these located tokens to a fine granularity for sufficient fine-grained spatio-temporal interaction. Specifically, in the coarse stage, we devise a discriminative token localization module to accurately identify informative tokens and to discard irrelevant tokens, where each localized token corresponds to a large spatial region, thus effectively preserving the continuity of action regions.In the fine stage, we only further partition the localized tokens obtained in the coarse stage into a finer granularity and then characterize fine-grained token interactions in two aspects: (1) first using vanilla transformers to learn compact dependencies among all discriminative tokens; and (2) proposing a global contextual interaction module which enables each fine-grained tokens to communicate with all the spatio-temporal tokens and to embed the global context. As a result, our coarse-to-fine strategy is able to identify more relevant tokens and integrate global context for high recognition accuracy while maintaining high efficiency.Comprehensive experimental results on four widely used action recognition benchmarks, including FineGym, Diving48, Kinetics and Something-Something, clearly demonstrate the advantages of our proposed method in comparison with other state-of-the-art ones.
Baoli Sun, Xinchen Ye, Zhihui Wang 0001, Zhiyong Wang 0001
ACM Multimedia5
2023 TopicCAT: Unsupervised Topic-Guided Co-Attention Transformer for Extreme Multimodal Summarisation
abstract
The exponential growth of multimedia data has sparked a surge of interest in multimodal summarisation with multimodal output (MSMO). A relatively unexplored but essential task within this field is extreme multimodal summarisation, a process that involves creating extremely concise multimodal summaries to further address the issue of multimedia information overload. In this study, we propose a novel Unsupervised Topic-guided Co-Attention Transformer (TopicCAT) neural network to produce extreme multimodal summaries for video-document pairs. The approach consists of two learning stages for a comprehensive multimodal understanding, guided by topic-based insights: a unimodal learning stage and a cross-modal learning stage, in which a cross-modal topic model is devised to capture the overarching themes present in both documents and videos. To achieve unsupervised learning, eliminating the need for resource-expensive collection of ground-truth multimodal summaries, we propose an optimal transport-based optimisation scheme to evaluate summary coverage from a semantic distribution perspective at the topic-level. Comprehensive experiments demonstrate the effectiveness of our proposed TopicCAT method on a multimodal news dataset, achieving a BERTScore of 84.46 and an accuracy of 0.60.
Peggy Tang, Kun Hu 0008, Lei Zhang 0001, Junbin Gao, Jiebo Luo 0001, Zhiyong Wang 0001
ACM Multimedia6
2023 LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark
abstract
Large language models have emerged as a promising approach towards achieving general-purpose AI agents. The thriving open-source LLM community has greatly accelerated the development of agents that support human-machine dialogue interaction through natural language processing. However, human interaction with the world extends beyond only text as a modality, and other modalities such as vision are also crucial. Recent works on multi-modal large language models, such as GPT-4V and Bard, have demonstrated their effectiveness in handling visual modalities. However, the transparency of these works is limited and insufficient to support academic research. To the best of our knowledge, we present one of the very first open-source endeavors in the field, LAMM, encompassing a Language-Assisted Multi-Modal instruction tuning dataset, framework, and benchmark. Our aim is to establish LAMM as a growing ecosystem for training and evaluating MLLMs, with a specific focus on facilitating AI agents capable of bridging the gap between ideas and execution, thereby enabling seamless human-AI interaction. Our main contribution is three-fold: 1) We present a comprehensive dataset and benchmark, which cover a wide range of vision tasks for 2D and 3D vision. Extensive experiments validate the effectiveness of our dataset and benchmark. 2) We outline the detailed methodology of constructing multi-modal instruction tuning datasets and benchmarks for MLLMs, enabling rapid scaling and extension of MLLM research to diverse domains, tasks, and modalities. 3) We provide a primary but potential MLLM training framework optimized for modality extension. We also provide baseline models, comprehensive experimental observations, and analysis to accelerate future research. Our baseline model is trained within 24 A100 GPU hours, framework supports training with V100 and RTX3090 is available thanks to the open-source society. Codes and data are now available at https://openlamm.github.io.
Zhenfei Yin, Jianjian Cao, Zhelun Shi, Dingning Liu, Mukai Li, Xiaoshui Huang, Zhiyong Wang 0001, Lu Sheng, Lei Bai 0001, Wanli Ouyang
NeurIPS8
2023 Federated Unsupervised Cluster-Contrastive learning for person Re-identification: A coarse-to-fine approach
abstract
Person Re-identification (ReID) has attracted considerable interests in recent years, largely driven by the escalating demand for public safety measures. However, the acquisition and handling of sensitive personal data can trigger significant privacy concerns. Federated learning has been introduced as a potential solution to this problem, with the goal of limiting the exposure of sensitive data across different participating entities (clients). Existing methods often depend on labor-intensive data annotations and face difficulties in maintaining cross-domain uniformity. To tackle these challenges, we propose a Federated Unsupervised Cluster-Contrastive (FedUCC) method based on deep learning for Person ReID that follows a generic-to-specific learning strategy. First, FedUCC procures generic knowledge from a conventional federated learning scheme to aggregate and distribute parameters across local clients. Second, specialized knowledge is explored to facilitate client personalization by disentangling client-specific knowledge from generic knowledge through parameter localization. Third, to further enhance effective fine-grained patterns instead of overfitting on specialized client knowledge, we investigate two key aspects: patch-level feature alignment and camera-invariant learning. Comprehensive experiments on eight public benchmark datasets demonstrate the state-of-the-art performance of our proposed method.
Jian F. Weng, Kun Hu 0008, Jingya Wang 0001, Zhiyong Wang 0001
Comput. Vis. Image Underst.5
2023 InvolutionGAN: lightweight GAN with involution for unsupervised image-to-image translation
Haipeng Deng, Qiuxia Wu, Han Huang 0002, Xiaowei Yang 0003, Zhiyong Wang 0001
Neural Comput. Appl.5
2023 Coloring anime line art videos with transformation region enhancement network
abstract
Automatic colorization of anime line art videos aims to produce color frames given line art frames and reference color images, which is challenging due to various motions and geometric transformations across frame sequences. Existing methods usually utilize the feature maps of reference images directly and treat all the regions in an image equally. However, this may overlook the details of the regions undergoing geometric transformations . To emphasize the regions with significant transformations between the reference and target frames, we propose a Transformation Region Enhancement Network (TRE-Net) to exploit useful reference information and enhance the colorization of key transformation regions with Region Localization Module (RLM) and Feature Enhancement Module (FEM). Specifically, we propose Multi-scale Euclidean Distance Difference (Multi-scale EDD) Maps in RLM which effectively locate geometric transformation regions by contrasting the Euclidean Distance Maps of two line arts and aggregating representations at multiple scales of the network. In addition, FEM is devised to enhance feature learning in the regions with geometric transformation and to ensure proper color alignment. FEM learns locally enhanced features through an attention-gating operation at a low computational cost. With the well-represented key geometric transformation regions, our method exploits the multi-scale reference information well for color alignment, thus produces perceptually pleasing frames. Comprehensive experimental results show that our proposed method is superior to existing methods in terms of the overall quality of colorized anime line art videos.
Ning Wang 0020, Muyao Niu, Zhi Dou, Zhihui Wang 0001, Zhiyong Wang 0001, Zhaoyan Ming, Bin Liu 0040
Pattern Recognit.5
2023 A Sparse Framework for Robust Possibilistic K-Subspace Clustering
abstract
Clustering noisy, high-dimensional, and structurally complex data have always been a challenging task. As most existing clustering methods are not able to deal with both the adverse impact of noisy samples and the complex structures of data, in this article, we propose a novel robust and sparse possibilistic K-subspace (RSPKS) clustering algorithm to integrate subspace recovery and possibilistic clustering algorithms under a unified sparse framework. First, the proposed method sparsifies the membership matrix and the subspace projection vector under a dual-sparse framework to handle high-dimensional noisy data. This unifies dimensionality reduction and clustering using one objective function for which the optimization can be realized through synchronous iteration. Second, the reconstruction error of each sample in the local subspace is used as the distance metric for classification. That is, each sample itself is treated as a clustering prototype so as not to be affected by the structure of the overall data distribution. Therefore, the clustering prototype construction problem of the data with complex structures can be better addressed. Finally, to deal with nonlinear regions, our RSPKS method is further extended into a kernelized version, namely the kernelized RSPKS clustering algorithm. The experimental results on both synthetic and real-world datasets demonstrate that our proposed method outperforms state-of-the-art algorithms in terms of clustering accuracy.
Shan Zeng, Xiangjun Duan, Hao Li 0034, Yuan Yan Tang, Zhiyong Wang 0001
IEEE Trans. Fuzzy Syst.6
2023 Region Assisted Sketch Colorization
abstract
Automatic sketch colorization is a challenging task that aims to generate a color image from a sketch, primarily due to its inherently ill-posed nature. While many approaches have shown promising results, two significant challenges remain: limited color patterns and a wide range of artifacts such as color bleeding and semantic inconsistencies among relevant regions. These issues stem from the operation of traditional convolutional structures, which capture structural features in a pixel-wise manner, resulting in inadequate utilization of regional information within the sketch. Therefore, we propose the Region-Assisted Sketch Coloring (RASC) method, which introduces an intermediate representation called the 'Region Map' to explicitly characterize the regional information of the sketch. This Region Map is derived from the input sketch and is effectively formulated by our RASC architecture, enhancing the perception of region-wise features beyond the original pixel-wise features. Specifically, we start by employing the sketch encoder to extract hierarchical feature maps from the input sketches. Subsequently, we introduce a coarse-to-fine decoder comprising a series of Region-based Modulation (RM) blocks. This decoder modulates features that combine the modulation results of its previous block and the sketch features of the corresponding encoder block with our Region Formulation module. Each module explicitly formulates the sketch features in a region-wise manner. This accurately captures both the inner-region local style and inter-region global context dependency, resulting in various color patterns and fewer synthesis artifacts. Our experimental results show that our proposed method surpasses state-of-the-art methods in both synthetic and real sketch datasets.
Ning Wang 0025, Muyao Niu, Zhihui Wang 0001, Kun Hu 0008, Bin Liu 0040, Zhiyong Wang 0001
IEEE Trans. Image Process.6
2023 Multi-Level Adversarial Spatio-Temporal Learning for Footstep Pressure Based FoG Detection
abstract
Freezing of gait (FoG) is one of the most common symptoms of Parkinson's disease, which is a neurodegenerative disorder of the central nervous system impacting millions of people around the world. To address the pressing need to improve the quality of treatment for FoG, devising a computer-aided detection and quantification tool for FoG has been increasingly important. As a non-invasive technique for collecting motion patterns, the footstep pressure sequences obtained from pressure sensitive gait mats provide a great opportunity for evaluating FoG in the clinic and potentially in the home environment. In this study, FoG detection is formulated as a sequential modelling task and a novel deep learning architecture, namely Adversarial Spatio-temporal Network (ASTN), is proposed to learn FoG patterns across multiple levels. ASTN introduces a novel adversarial training scheme with a multi-level subject discriminator to obtain subject-independent FoG representations, which helps to reduce the over-fitting risk due to the high inter-subject variance. As a result, robust FoG detection can be achieved for unseen subjects. The proposed scheme also sheds light on improving subject-level clinical studies from other scenarios as it can be integrated with many existing deep architectures. To the best of our knowledge, this is one of the first studies of footstep pressure-based FoG detection and the approach of utilizing ASTN is the first deep neural network architecture in pursuit of subject-independent representations. In our experiments on 393 trials collected from 21 subjects, the proposed ASTN achieved an AUC 0.85, clearly outperforming conventional learning methods.
Kun Hu 0008, Shaohui Mei, Wei Wang 0115, Kaylena A. Ehgoetz Martens, Liang Wang 0001, Simon J. G. Lewis, David Dagan Feng, Zhiyong Wang 0001
IEEE J. Biomed. Health Informatics8
2023 Cascade Multi-Level Transformer Network for Surgical Workflow Analysis
abstract
Surgical workflow analysis aims to recognise surgical phases from untrimmed surgical videos. It is an integral component for enabling context-aware computer-aided surgical operating systems. Many deep learning-based methods have been developed for this task. However, most existing works aggregate homogeneous temporal context for all frames at a single level and neglect the fact that each frame has its specific need for information at multiple levels for accurate phase prediction. To fill this gap, in this paper we propose Cascade Multi-Level Transformer Network (CMTNet) composed of cascaded Adaptive Multi-Level Context Aggregation (AMCA) modules. Each AMCA module first extracts temporal context at the frame level and the phase level and then fuses frame-specific spatial feature, frame-level temporal context, and phase-level temporal context for each frame adaptively. By cascading multiple AMCA modules, CMTNet is able to gradually enrich the representation of each frame with the multi-level semantics that it specifically requires, achieving better phase prediction in a frame-adaptive manner. In addition, we propose a novel refinement loss for CMTNet, which explicitly guides each AMCA module to focus on extracting the key context for refining the prediction of the previous stage in terms of both prediction confidence and smoothness. This further enhances the quality of the extracted context effectively. Extensive experiments on the Cholec80 and the M2CAI datasets demonstrate that CMTNet achieves state-of-the-art performance.
Wenxi Yue, Hongen Liao, Yong Xia 0001, Vincent Lam, Jiebo Luo 0001, Zhiyong Wang 0001
IEEE Trans. Medical Imaging6
2023 Graph Fusion Network-Based Multimodal Learning for Freezing of Gait Detection
abstract
Freezing of gait (FoG) is identified as a sudden and brief episode of movement cessation despite the intention to continue walking. It is one of the most disabling symptoms of Parkinson's disease (PD) and often leads to falls and injuries. Many computer-aided FoG detection methods have been proposed to use data collected from unimodal sources, such as motion sensors, pressure sensors, and video cameras. However, there are limited efforts of multimodal-based methods to maximize the value of all the information collected from different modalities in clinical assessments and improve the FoG detection performance. Therefore, in this study, a novel end-to-end deep architecture, namely graph fusion neural network (GFN), is proposed for multimodal learning-based FoG detection by combining footstep pressure maps and video recordings. GFN constructs multimodal graphs by treating the encoded features of each modality as vertex-level inputs and measures their adjacency patterns to construct complementary FoG representations, thus reducing the representation redundancy among different modalities. In addition, since GFN is devised to process multimodal graphs of arbitrary structures, it is expected to achieve superior performance with inputs containing missing modalities, compared to the alternative unimodal methods. A multimodal FoG dataset was collected, which included clinical assessment videos and footstep pressure sequences of 340 trials from 20 PD patients. Our proposed GFN demonstrates a great promise of multimodal FoG detection with an area under the curve (AUC) of 0.882. To the best of our knowledge, this is one of the first studies to utilize multimodal learning for automated FoG detection, which offers significant opportunities for better patient assessments and clinical trials in the future.
Kun Hu 0008, Zhiyong Wang 0001, Kaylena A. Ehgoetz Martens, Markus Hagenbuchner, Mohammed Bennamoun, Ah Chung Tsoi, Simon J. G. Lewis
IEEE Trans. Neural Networks Learn. Syst.2
2022 Confidence-Calibrated Face Image Forgery Detection with Contrastive Representation Distillation
Puning Yang, Huaibo Huang, Zhiyong Wang 0001, Aijing Yu, Ran He 0001
ACCV (4)3
2022 Multi-Scale Attention based Transformer U-NET for Change Detection
abstract
In recent years, various deep learning based methods have been successfully developed for change detection, such as Convolutional Neural Network (CNN) based U-Net and its variants, and Transformer based ones. However, CNNs lack the ability to effectively learn global representations, while Transformers neglect to learn local representations. Therefore, in this paper we propose a novel deep network, namely Multi-scale Attention based Transformer U-Net (MATU), to take advantages of CNNs and Transformers for learning both local and global features effectively. The backbone of our proposed MATU is a U-Net. In the encoder, a Siamese network is used to extract features from two input images, which is followed by a transformer module to further refine the feature pairs produced by the Siamese network. The difference of the refined feature pairs is fed into an Atrous Spatial Pyramid Pooling (ASSP) module to generate a distance map. Moreover, axial-attention blocks are integrated in the decoder with the corresponding multi-level feature differences of the encoder to progressively produce and improve the change map through attention upsampling. Extensive experiments on two widely used benchmark datasets SYSU-CD and LEVIR-CD demonstrate that the proposed MATU method achieves the state-of-the-art performance. Our code is available at https://github.com/easm002/MATU.
Hengzhi Chen, Xiaofeng Wu 0001, Shan Zeng, Zhiyong Wang 0001
IGARSS4
2022 Deep Laparoscopic Stereo Matching with Transformers
Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Tom Drummond, Zhiyong Wang 0001, ZongYuan Ge
MICCAI (8)5
2022 Skin Lesion Recognition with Class-Hierarchy Regularized Hyperbolic Embeddings
Toàn D. Nguyên, Yaniv Gal, Lie Ju, Shekhar Chandra, Lei Zhang 0095, C. Paul Bonnington, Victoria Mar, Zhiyong Wang 0001, ZongYuan Ge
MICCAI (3)9
2022 Fine-grained Action Recognition with Robust Motion Representation Decoupling and Concentration
abstract
Fine-grained action recognition is a challenging task that requires identifying discriminative and subtle motion variations among fine-grained action classes. Existing methods typically focus on spatio-temporal feature extraction and long-temporal modeling to characterize complex spatio-temporal patterns of fine-grained actions. However, the learned spatio-temporal features without explicit motion modeling may emphasize more on visual appearance than on motion, which could compromise the learning of effective motion features required for fine-grained temporal reasoning. Therefore, how to decouple robust motion representations from the spatio-temporal features and further effectively leverage them to enhance the learning of discriminative features still remains less explored, which is crucial for fine-grained action recognition. In this paper, we propose a motion representation decoupling and concentration network (MDCNet) to address these two key issues. First, we devise a motion representation decoupling (MRD) module to disentangle the spatio-temporal representation into appearance and motion features through contrastive learning from video and segment views. Next, in the proposed motion representation concentration (MRC) module, the decoupled motion representations are further leveraged to learn a universal motion prototype shared across all the instances of each action class. Finally, we project the decoupled motion features onto all the motion prototypes through semantic relations to obtain the concentrated action-relevant features for each action class, which can effectively characterize the temporal distinctions of fine-grained actions for improved recognition performance. Comprehensive experimental results on four widely used action recognition benchmarks, i.e., FineGym, Diving48, Kinetics400 and Something-Something, clearly demonstrate the superiority of our proposed method in comparison with other state-of-the-art ones.
Baoli Sun, Xinchen Ye, Tiantian Yan, Zhihui Wang 0001, Zhiyong Wang 0001
ACM Multimedia6
2022 Sign Language Translation with Hierarchical Spatio-Temporal Graph Neural Network
abstract
Sign language translation (SLT), which generates text in a spoken language from visual content in a sign language, is important to assist the hard-of-hearing community for their communications. Inspired by neural machine translation (NMT), most existing SLT studies adopted a general sequence to sequence learning strategy. However, SLT is significantly different from general NMT tasks since sign languages convey messages through multiple visual-manual aspects. Therefore, in this paper, these unique characteristics of sign languages are formulated as hierarchical spatio-temporal graph representations, including high-level and fine-level graphs of which a vertex characterizes a specified body part and an edge represents their interactions. Particularly, high-level graphs represent the patterns in the regions such as hands and face, and fine-level graphs consider the joints of hands and landmarks of facial regions. To learn these graph patterns, a novel deep learning architecture, namely hierarchical spatio-temporal graph neural network (HST-GNN), is proposed. Graph convolutions and graph self-attentions with neighborhood context are proposed to characterize both the local and the global graph properties. Experimental results on benchmark datasets demonstrated the effectiveness of the proposed method.
Jichao Kan, Kun Hu 0008, Markus Hagenbuchner, Ah Chung Tsoi, Mohammed Bennamoun, Zhiyong Wang 0001
WACV6
2022 Affective Audio Annotation of Public Speeches with Convolutional Clustering Neural Network
abstract
Public speaking is a critical skill in daily communication. While more practicing such as rehearsal is helpful to improve such a skill, lack of personalized feedback limits the effectiveness of practicing. Therefore, we formulate the task of personalized feedback as an affective audio annotation problem by learning knowledge from online public speech videos. Considering the great success of deep learning techniques such as convolutional neural networks in a wide range of applications including speech recognition and object recognition, we propose a novel convolutional clustering neural network (CCNN) to solve this multi-label classification problem. Instead of aggregating the features of different channels through pooling, we introduce a novel clustering layer to derive intermediate representation for improved annotation performance. In order to evaluate the performance of our proposed method, we purposely built an affective audio annotation dataset by collecting more than 2,000 video clips from the TED website. Experimental results on this dataset demonstrate that our proposed method outperforms traditional CNN-based approaches with a lower hamming loss for affective annotation.
Jiahao Xu 0002, Zhiyong Wang 0001, Yang Wang 0002, Fang Chen 0001, Junbin Gao, David Dagan Feng
IEEE Trans. Affect. Comput.3
2022 Vision-Enhanced and Consensus-Aware Transformer for Image Captioning
abstract
Image captioning generates descriptions in a natural language for a given image. Due to its great potential for a wide range of applications, many deep learning based-methods have been proposed. The co-occurrence of words such as mouse and keyboard, constitutes commonsense knowledge, which is referred to as consensus. However, it is challenging to consider commonsense knowledge in producing captions that have rich, natural, and meaningful semantics. In this paper, a Vision-enhanced and Consensus-aware Transformer (VCT) is proposed to exploit both visual information and consensus knowledge for image captioning with three key components: a vision-enhanced encoder, consensus-aware knowledge representation generator, and consensus-aware decoder. The vision-enhanced encoder extends the vanilla self-attention module with a memory-based attention module and a visual perception module for learning better visual representation of an image. Specifically, the relationships between regions in an image and the image’s global context are leveraged with scene memory in the memory-based attention module. The visual perception module further enhances the correlation among neighboring tokens in both the spatial and channel-wise dimensions. To learn consensus-aware representations, a word correlation graph is constructed by computing the statistical co-occurrence between semantic concepts. Then consensus knowledge can be acquired using a graph convolutional network in the consensus-aware knowledge representation generator. Finally, such consensus knowledge is integrated into the consensus-aware decoder through consensus memory and a knowledge-based control module to produce a caption. Experimental results on two popular benchmark datasets (MSCOCO and Flickr30k) demonstrate that our proposed model achieves state-of-the-art performance. Extensive ablation studies also validate the effectiveness of each component.
Shan Cao 0002, Gaoyun An, Zhenxing Zheng, Zhiyong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2022 Multi-scale Features Fusion for the Detection of Tiny Bleeding in Wireless Capsule Endoscopy Images
abstract
Wireless capsule endoscopy is a modern non-invasive Internet of Medical Imaging Things that has been increasingly used in gastrointestinal tract examination. With about one gigabyte image data generated for a patient in each examination, automatic lesion detection is highly desirable to improve the efficiency of the diagnosis process and mitigate human errors. Despite many approaches for lesion detection have been proposed, they mainly focus on large lesions and are not directly applicable to tiny lesions due to the limitations of feature representation. As bleeding lesions are a common symptom in most serious gastrointestinal diseases, detecting tiny bleeding lesions is extremely important for early diagnosis of those diseases, which is highly relevant to the survival, treatment, and expenses of patients. In this article, a method is proposed to extract and fuse multi-scale deep features for detecting and locating both large and tiny lesions. A feature extracting network is first used as our backbone network to extract the basic features from wireless capsule endoscopy images, and then at each layer multiple regions could be identified as potential lesions. As a result, the features maps of those potential lesions are obtained at each level and fused in a top-down manner to the fully connected layer for producing final detection results. Our proposed method has been evaluated on a clinical dataset that contains 20,000 wireless capsule endoscopy images with clinical annotation. Experimental results demonstrate that our method can achieve 98.9% prediction accuracy and 93.5% score, which has a significant performance improvement of up to 31.69% and 22.12% in terms of recall rate and score, respectively, when compared to the state-of-the-art approaches for both large and tiny bleeding lesions. Moreover, our model also has the highest AP and the best medical diagnosis performance compared to state-of-the-art multi-scale models.
Feng Lu 0003, Wei Li 0058, Chengwangli Peng, Zhiyong Wang 0001, Bin Qian 0002, Rajiv Ranjan 0001, Hai Jin 0001, Albert Y. Zomaya
ACM Trans. Internet Things5
2022 Graph Convolutional Dictionary Selection With L₂, ₚ Norm for Video Summarization
abstract
Video Summarization (VS) has become one of the most effective solutions for quickly understanding a large volume of video data. Dictionary selection with self representation and sparse regularization has demonstrated its promise for VS by formulating the VS problem as a sparse selection task on video frames. However, existing dictionary selection models are generally designed only for data reconstruction, which results in the neglect of the inherent structured information among video frames. In addition, the sparsity commonly constrained by$L_{2,1}$norm is not strong enough, which causes the redundancy of keyframes, i.e., similar keyframes are selected. Therefore, to address these two issues, in this paper we propose a general framework called graph convolutional dictionary selection with$L_{2,p}$($0< p\leq 1$) norm (GCDS$_{2,p}$) for both keyframe selection and skimming based summarization. Firstly, we incorporate graph embedding into dictionary selection to generate the graph embedding dictionary, which can take the structured information depicted in videos into account. Secondly, we propose to use$L_{2,p}$($0< p\leq 1$) norm constrained row sparsity, in which$p$can be flexibly set for two forms of video summarization. For keyframe selection,$0< p< 1$can be utilized to select diverse and representative keyframes; and for skimming,$p=1$can be utilized to select key shots. In addition, an efficient iterative algorithm is devised to optimize the proposed model, and the convergence is theoretically proved. Experimental results including both keyframe selection and skimming based summarization on four benchmark datasets demonstrate the effectiveness and superiority of the proposed method.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, Xian-Sheng Hua 0001, David Dagan Feng
IEEE Trans. Image Process.4
2022 Action Recognition With Motion Diversification and Dynamic Selection
abstract
Motion modeling is crucial in modern action recognition methods. As motion dynamics like moving tempos and action amplitude may vary a lot in different video clips, it poses great challenge on adaptively covering proper motion information. To address this issue, we introduce a Motion Diversification and Selection (MoDS) module to generate diversified spatio-temporal motion features and then select the suitable motion representation dynamically for categorizing the input video. To be specific, we first propose a spatio-temporal motion generation (StMG) module to construct a bank of diversified motion features with varying spatial neighborhood and time range. Then, a dynamic motion selection (DMS) module is leveraged to choose the most discriminative motion feature both spatially and temporally from the feature bank. As a result, our proposed method can make full use of the diversified spatio-temporal motion information, while maintaining computational efficiency at the inference stage. Extensive experiments on five widely-used benchmarks, demonstrate the effectiveness of the method and we achieve state-of-the-art performance on Something-Something V1 & V2 that are of large motion variation.
Peiqin Zhuang, Luping Zhou, Lei Bai 0001, Ding Liang, Zhiyong Wang 0001, Yali Wang 0001, Wanli Ouyang
IEEE Trans. Image Process.7
2022 Adversarial Evolving Neural Network for Longitudinal Knee Osteoarthritis Prediction
abstract
Knee osteoarthritis (KOA) as a disabling joint disease has doubled in prevalence since the mid-20th century. Early diagnosis for the longitudinal KOA grades has been increasingly important for effective monitoring and intervention. Although recent studies have achieved promising performance for baseline KOA grading, longitudinal KOA grading has been seldom studied and the KOA domain knowledge has not been well explored yet. In this paper, a novel deep learning architecture, namely adversarial evolving neural network (A-ENN), is proposed for longitudinal grading of KOA severity. As the disease progresses from mild to severe level, ENN involves the progression patterns for accurately characterizing the disease by comparing an input image it to the template images of different KL grades using convolution and deconvolution computations. In addition, an adversarial training scheme with a discriminator is developed to obtain the evolution traces. Thus, the evolution traces as fine-grained domain knowledge are further fused with the general convolutional image representations for longitudinal grading. Note that ENN can be applied to other learning tasks together with existing deep architectures, in which the responses characterize progressive representations. Comprehensive experiments on the Osteoarthritis Initiative (OAI) dataset were conducted to evaluate the proposed method. An overall accuracy was achieved as 62.7%, with the baseline, 12-month, 24-month, 36-month, and 48-month accuracy as 64.6%, 63.9%, 63.2%, 61.8% and 60.2%, respectively.
Kun Hu 0008, Wenhua Wu 0005, Wei Li 0058, Milena Simic, Albert Y. Zomaya, Zhiyong Wang 0001
IEEE Trans. Medical Imaging6
2021 Learning Efficient Rotation Representation for Point Cloud via Local-Global Aggregation
abstract
Recently, there have been attempted to solve the problem of rotation perturbation in point cloud analysis. However, most of them fail to exploit the long-distance context and lose global location information. To address this issue, we propose a novel rotation-invariant network called LGANet, which is assembled with two key modules: local representation learning module and global alignment module. The local representation learning module is to capture local geometric features from K-nearest neighbors in both 3D Cartesian space and a latent space, while the global alignment module focuses on supplementing global location information with the adaptive selection mechanism. Extensive experiments on widely used datasets have demonstrated that our LGANet is superior to other state-of-the-art methods on the premise of ensuring rotation invariance in both classification and part segmentation.
Ruibin Gu, Qiuxia Wu, Wing W. Y. Ng, Zhiyong Wang 0001
ICME5
2021 Keyframe Extraction from Motion Capture Sequences with Graph based Deep Reinforcement Learning
abstract
Animation production workflows centred around motion capture techniques often require animators to edit the motion for various artistic and technical reasons. This process generally uses a set of keyframes. Unsupervised keyframe selection methods for motion capture sequences are highly demanded to reduce the laborious annotations. However, most existing methods are optimization-based, which cause the issues of flexibility and efficiency and eventually constrains the interactions and controls with animators. To address these limitations, we propose a novel graph based deep reinforcement learning method for efficient unsupervised keyframe selection. First, a reward function is devised in terms of reconstruction difference by comparing the original sequence and the interpolated sequence produced by the keyframes. The reward complies with the requirements of the animation pipeline satisfying: 1) incremental reward to evaluate the interpolated keyframes immediately; 2) order insensitivity for consistent evaluation; and 3) non-diminishing return for comparable rewards between optimal and sub-optimal solutions. Then by representing each skeleton frame as a graph, a graph-based deep agent is guided to heuristically select keyframes to maximize the reward. During the inference it is no longer necessary to estimate the reconstruction difference, and the evaluation time can be reduced significantly. The experimental results on the CMU Mocap dataset demonstrate that our proposed method is able to select keyframes at a high efficiency without clearly compromising the quality in comparison with the state-of-the-art methods.
Clinton Mo, Kun Hu 0008, Shaohui Mei, Zhiyong Wang 0001
ACM Multimedia5
2021 A Multi-task Kernel Learning Algorithm for Survival Analysis
Zizhuo Meng, Jie Xu 0008, Zhidong Li, Yang Wang 0002, Fang Chen 0001, Zhiyong Wang 0001
PAKDD (3)6
2021 Deep3D reconstruction: methods, data, and challenges
abstract
Three-dimensional (3D) reconstruction of shapes is an important research topic in the fields of computer vision, computer graphics, pattern recognition, and virtual reality. Existing 3D reconstruction methods usually suffer from two bottlenecks: (1) they involve multiple manually designed states which can lead to cumulative errors, but can hardly learn semantic features of 3D shapes automatically; (2) they depend heavily on the content and quality of images, as well as precisely calibrated cameras. As a result, it is difficult to improve the reconstruction accuracy of those methods. 3D reconstruction methods based on deep learning overcome both of these bottlenecks by automatically learning semantic features of 3D shapes from low-quality images using deep networks. However, while these methods have various architectures, in-depth analysis and comparisons of them are unavailable so far. We present a comprehensive survey of 3D reconstruction methods based on deep learning. First, based on different deep learning model architectures, we divide 3D reconstruction methods based on deep learning into four types, recurrent neural network, deep autoencoder, generative adversarial network, and convolutional neural network based methods, and analyze the corresponding methodologies carefully. Second, we investigate four representative databases that are commonly used by the above methods in detail. Third, we give a comprehensive comparison of 3D reconstruction methods based on deep learning, which consists of the results of different methods with respect to the same database, the results of each method with respect to different databases, and the robustness of each method with respect to the number of views. Finally, we discuss future development of 3D reconstruction methods based on deep learning.
Dehui Kong, Shaofan Wang 0001, Zhiyong Wang 0001
Frontiers Inf. Technol. Electron. Eng.4
2021 Coupling matrix manifolds assisted optimization for optimal transport problems
Dai Shi, Junbin Gao, Xia Hong 0001, S. T. Boris Choy, Zhiyong Wang 0001
Mach. Learn.5
2021 ERINet: Enhanced rotation-invariant network for point cloud classification
Ruibin Gu, Qiuxia Wu, Wing W. Y. Ng, Zhiyong Wang 0001
Pattern Recognit. Lett.5
2021 Similarity Based Block Sparse Subset Selection for Video Summarization
abstract
Video summarization (VS) is generally formulated as a subset selection problem where a set of representative keyframes or key segments is selected from an entire video frame set. Though many sparse subset selection based VS algorithms have been proposed in the past decade, most of them adopt linear sparse formulation in the explicit feature vector space of video frames, and don’t consider the local or global relationships among frames. In this paper, we first extend the conventional sparse subset selection for VS into kernel block sparse subset selection (KBS3) to utilize the advantage of kernel sparse coding and introduce a local inter-frame relationship through packing of frame blocks. Going a step further, we propose a similarity based block sparse subset selection (SB2S3) model by applying a specially designed transformation matrix on the KBS3 model in order to introduce a kind of global inter-frame relationship through the similarity. Finally, a greedy pursuit based algorithm is devised for the proposed NP-hard model optimization. The proposed SB2S3 has the following advantages: 1) through the similarity between each frame and any other frame, the global relationship among all frames can be considered; 2) through block sparse coding, the local relationship of adjacent frames is further considered; and 3) it has a wider application, since features can derive similarity, but not vice versa. It is believed that the effect of modeling such global and local relationships among frames in this paper, is similar to that of modeling the long-range and short-range dependencies among frames in deep learning based methods. Experimental results on three benchmark datasets have demonstrated that the proposed approach is superior to not only other sparse subset selection based VS methods but also most unsupervised deep-learning based VS methods.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng, Mohammed Bennamoun
IEEE Trans. Circuits Syst. Video Technol.4
2021 Keyframe Extraction From Laparoscopic Videos via Diverse and Weighted Dictionary Selection
abstract
Laparoscopic videos have been increasingly acquired for various purposes including surgical training and quality assurance, due to the wide adoption of laparoscopy in minimally invasive surgeries. However, it is very time consuming to view a large amount of laparoscopic videos, which prevents the values of laparoscopic video archives from being well exploited. In this paper, a dictionary selection based video summarization method is proposed to effectively extract keyframes for fast access of laparoscopic videos. Firstly, unlike the low-level feature used in most existing summarization methods, deep features are extracted from a convolutional neural network to effectively represent video frames. Secondly, based on such a deep representation, laparoscopic video summarization is formulated as a diverse and weighted dictionary selection model, in which image quality is taken into account to select high quality keyframes, and a diversity regularization term is added to reduce redundancy among the selected keyframes. Finally, an iterative algorithm with a rapid convergence rate is designed for model optimization, and the convergence of the proposed method is also analyzed. Experimental results on a recently released laparoscopic dataset demonstrate the clear superiority of the proposed methods. The proposed method can facilitate the access of key information in surgeries, training of junior clinicians, explanations to patients, and archive of case files.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, ZongYuan Ge, Vincent Lam, David Dagan Feng
IEEE J. Biomed. Health Informatics4
2021 Short-Term Lesion Change Detection for Melanoma Screening With Novel Siamese Neural Network
abstract
Short-term monitoring of lesion changes has been a widely accepted clinical guideline for melanoma screening. When there is a significant change of a melanocytic lesion at three months, the lesion will be excised to exclude melanoma. However, the decision on change or no-change heavily depends on the experience and bias of individual clinicians, which is subjective. For the first time, a novel deep learning based method is developed in this paper for automatically detecting short-term lesion changes in melanoma screening. The lesion change detection is formulated as a task measuring the similarity between two dermoscopy images taken for a lesion in a short time-frame, and a novel Siamese structure based deep network is proposed to produce the decision: changed (i.e. not similar) or unchanged (i.e. similar enough). Under the Siamese framework, a novel structure, namely Tensorial Regression Process, is proposed to extract the global features of lesion images, in addition to deep convolutional features. In order to mimic the decision-making process of clinicians who often focus more on regions with specific patterns when comparing a pair of lesion images, a segmentation loss (SegLoss) is further devised and incorporated into the proposed network as a regularization term. To evaluate the proposed method, an in-house dataset with 1,000 pairs of lesion images taken in a short time-frame at a clinical melanoma centre was established. Experimental results on this first-of-a-kind large dataset indicate that the proposed model is promising in detecting the short-term lesion change for objective melanoma screening.
Zhiyong Wang 0001, Junbin Gao, Chantal Rutjes, Kaitlin Nufer, Dacheng Tao, David Dagan Feng, Scott W. Menzies
IEEE Trans. Medical Imaging2
2021 Patch Based Video Summarization With Block Sparse Representation
abstract
In recent years, sparse representation has been successfully utilized for video summarization (VS). However, most of the sparse representation based VS methods characterize each video frame with global features. As a result, some important local details could be neglected by global features, which may compromise the performance of summarization. In this paper, we propose to partition each video frame into a number of patches and characterize each patch with global features. Instead of concatenating the features of each patch and utilizing conventional sparse representation, we formulate the VS problem with such video frame representation as block sparse representation by considering each video frame as a block containing a number of patches. By taking the reconstruction constraint into account, we devise a simultaneous version of block-based OMP (Orthogonal Matching Pursuit) algorithm, namely SBOMP, to solve the proposed model. The proposed model is further extended to a neighborhood based model which considers temporally adjacent frames as a super block. This is one of the first sparse representation based VS methods taking both spatial and temporal contexts into account with blocks. Experimental results on two widely used VS datasets have demonstrated that our proposed methods present clear superiority over existing sparse representation based VS methods and are highly comparable to some deep learning ones requiring supervision information for extra model training.
Shaohui Mei, Mingyang Ma 0004, Shuai Wan, Junhui Hou, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Multim.5
2021 Joint Input and Output Space Learning for Multi-Label Image Classification
abstract
Multi-label image classification aims to predict the labels associated with a given image. While most existing methods utilize unified image representations, extracting label-specific features through input space learning would improve the discriminative power of the learned features. On the other hand, most feature learning studies often ignore the learning in the output label space, although taking advantage of label correlations can boost the classification performance. In this paper, we propose a deep learning framework that incorporates flexible modules which can learn from both input and output spaces for multi-label image classification. For the input space learning, we devise a label-specific feature pooling method to refine convolutional features for obtaining features specific to each label. For the output space learning, we design a Two-Stream Graph Convolutional Network (TSGCN) to learn multi-label classifiers by mapping spatial object relationships and semantic label correlations. More specifically, we build object spatial graphs to characterize the spatial relationships among objects in an image, which supplements the label semantic graphs modelling the semantic label correlations. Experimental results on two popular benchmark datasets (i.e., Pascal VOC and MS-COCO) show that our proposed method achieves superior performance over the state-of-the-arts.
Jiahao Xu 0002, Hongda Tian, Zhiyong Wang 0001, Yang Wang 0002, Wenxiong Kang, Fang Chen 0001
IEEE Trans. Multim.3
2020 Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition
abstract
Spatial-temporal graphs have been widely used by skeleton-based action recognition algorithms to model human action dynamics. To capture robust movement patterns from these graphs, long-range and multi-scale context aggregation and spatial-temporal dependency modeling are critical aspects of a powerful feature extractor. However, existing methods have limitations in achieving (1) unbiased long-range joint relationship modeling under multi-scale operators and (2) unobstructed cross-spacetime information flow for capturing complex spatial-temporal dependencies. In this work, we present (1) a simple method to disentangle multi-scale graph convolutions and (2) a unified spatial-temporal graph convolutional operator named G3D. The proposed multi-scale aggregation scheme disentangles the importance of nodes in different neighborhoods for effective long-range modeling. The proposed G3D module leverages dense cross-spacetime edges as skip connections for direct information propagation across the spatial-temporal graph. By coupling these proposals, we develop a powerful feature extractor named MS-G3D based on which our model outperforms previous state-of-the-art methods on three large-scale datasets: NTU RGB+D 60, NTU RGB+D 120, and Kinetics Skeleton 400.
Hongwen Zhang 0001, Zhiyong Wang 0001, Wanli Ouyang
CVPR4
2020 Correlation-Aware Next Basket Recommendation Using Graph Attention Networks
Yuanzhe Zhang, Ling Luo 0002, Jianjia Zhang, Yang Wang 0002, Zhiyong Wang 0001
ICONIP (4)6
2020 Speaker-Aware Monaural Speech Separation
Jiahao Xu 0002, Kun Hu 0008, Tran Duc Chung, Zhiyong Wang 0001
INTERSPEECH5
2020 FCP Filter: A Dynamic Clustering-Prediction Framework for Customer Behavior
Yuanzhe Zhang, Ling Luo 0002, Yang Wang 0002, Zhiyong Wang 0001
PAKDD (1)4
2020 3D Hand Pose Estimation with Disentangled Cross-Modal Latent Space
abstract
Estimating 3D hand pose from a single RGB image is a challenging task because of its ill-posed nature (i.e., depth ambiguity). Recently, various generative approaches have been proposed to predict the 3D joints of an RGB hand image by learning a unified latent space between two modalities (i.e., RGB image and 3D joints). However, projecting multi-modal data (i.e., RGB images and 3D joints) into a unified latent space is difficult as the modality-specific features usually interfere the learning of the optimal latent space. Hence in this paper, we propose to disentangle the latent space into two sub-latent spaces: modality- specific latent space and pose-specific latent space for 3D hand pose estimation. Our proposed method, namely Disentangled Cross-Modal Latent Space (DCMLS), consists of two variational autoencoder networks and auxiliary components which connect the two VAEs to align underlying hand poses and transfer modality-specific context from RGB to 3D. For the hand pose latent space, we align it with the two modalities by using a cross-modal discriminator with an adversarial learning strategy. For the context latent space, we learn a context translator to gain access to the cross-modal context. Experimental results on two widely used public benchmark datasets RHD and STB demonstrate that our proposed DCMLS method is able to clearly outperform the state-of-the-art ones on single image based 3D hand pose estimation.
Jiajun Gu, Zhiyong Wang 0001, Wanli Ouyang, Jiafeng Li 0001, Li Zhuo 0001
WACV2
2020 Video summarization via block sparse dictionary selection
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Junhui Hou, Zhiyong Wang 0001, David Dagan Feng
Neurocomputing5
2020 Learning visual relationship and context-aware attention for image captioning
Junbo Wang 0003, Wei Wang 0115, Liang Wang 0001, Zhiyong Wang 0001, David Dagan Feng, Tieniu Tan
Pattern Recognit.4
2020 Real-time hand posture recognition using hand geometric features and Fisher Vector
Linpu Fang, Ningxin Liang, Wenxiong Kang, Zhiyong Wang 0001, David Dagan Feng
Signal Process. Image Commun.4
2020 Graph Sequence Recurrent Neural Network for Vision-Based Freezing of Gait Detection
abstract
Freezing of gait (FoG) is one of the most common symptoms of Parkinson's disease (PD), a neurodegenerative disorder which impacts millions of people around the world. Accurate assessment of FoG is critical for the management of PD and to evaluate the efficacy of treatments. Currently, the assessment of FoG requires well-trained experts to perform time-consuming annotations via vision-based observations. Thus, automatic FoG detection algorithms are needed. In this study, we formulate vision-based FoG detection, as a fine-grained graph sequence modelling task, by representing the anatomic joints in each temporal segment with a directed graph, since FoG events can be observed through the motion patterns of joints. A novel deep learning method is proposed, namely graph sequence recurrent neural network (GS-RNN), to characterize the FoG patterns by devising graph recurrent cells, which take graph sequences of dynamic structures as inputs. For the cases of which prior edge annotations are not available, a data-driven based adjacency estimation method is further proposed. To the best of our knowledge, this is one of the first studies on vision-based FoG detection using deep neural networks designed for graph sequences of dynamic structures. Experimental results on more than 150 videos collected from 45 patients demonstrated promising performance of the proposed GS-RNN for FoG detection with an AUC value of 0.90.
Kun Hu 0008, Zhiyong Wang 0001, Wei Wang 0115, Kaylena A. Ehgoetz Martens, Liang Wang 0001, Tieniu Tan, Simon J. G. Lewis, David Dagan Feng
IEEE Trans. Image Process.2
2020 Vision-Based Freezing of Gait Detection With Anatomic Directed Graph Representation
abstract
Parkinson's disease significantly impacts the life quality of millions of people around the world. While freezing of gait (FoG) is one of the most common symptoms of the disease, it is time consuming and subjective to assess FoG for well-trained experts. Therefore, it is highly desirable to devise computer-aided FoG detection methods for the purpose of objective and time-efficient assessment. In this paper, in line with the gold standard of FoG clinical assessment, which requires video or direct observation, we propose one of the first vision-based methods for automatic FoG detection. To better characterize FoG patterns, instead of learning an overall representation of a video, we propose a novel architecture of graph convolution neural network and represent each video as a directed graph where FoG related candidate regions are the vertices. A weakly-supervised learning strategy and a weighted adjacency matrix estimation layer are proposed to eliminate the resource expensive data annotation required for fully supervised learning. As a result, the interference of visual information irrelevant to FoG, such as gait motion of supporting staff involved in clinical assessments, has been reduced to improve FoG detection performance by identifying the vertices contributing to FoG events. To further improve the performance, the global context of a clinical video is also considered and several fusion strategies with graph predictions are investigated. Experimental results on more than 100 videos collected from 45 patients during a clinical assessment demonstrated promising performance of our proposed method with an AUC of 0.887.
Kun Hu 0008, Zhiyong Wang 0001, Shaohui Mei, Kaylena A. Ehgoetz Martens, Simon J. G. Lewis, David Dagan Feng
IEEE J. Biomed. Health Informatics2
2020 A Residual Based Attention Model for EEG Based Sleep Staging
abstract
Sleep staging is to score the sleep state of a subject into different sleep stages such as Wake and Rapid Eye Movement (REM). It plays an indispensable role in the diagnosis and treatment of sleep disorders. As manual sleep staging through well-trained sleep experts is time consuming, tedious, and subjective, many automatic methods have been developed for accurate, efficient, and objective sleep staging. Recently, deep learning based methods have been successfully proposed for electroencephalogram (EEG) based sleep staging with promising results. However, most of these methods directly take EEG raw signals as input of convolutional neural networks (CNNs) without considering the domain knowledge of EEG staging. Apart from that, to capture temporal information, most of the existing methods utilize recurrent neural networks such as LSTM (Long Short Term Memory) which are not effective for modelling global temporal context and difficult to train. Therefore, inspired by the clinical guidelines of sleep staging such as AASM (American Academy of Sleep Medicine) rules where different stages are generally characterized by EEG waveforms of various frequencies, we propose a multi-scale deep architecture by decomposing an EEG signal into different frequency bands as input to CNNs. To model global temporal context, we utilize the multi-head self-attention module of the transformer model to not only improve performance, but also shorten the training time. In addition, we choose residual based architecture which makes training end-to-end. Experimental results on two widely used sleep staging datasets, Montreal Archive of Sleep Studies (MASS) and sleep-EDF datasets, demonstrate the effectiveness and significant efficiency (up to 12 times less training time) of our proposed method over the state-of-the-art.
Zhiyong Wang 0001, Hong Hong 0001, Zheru Chi, David Dagan Feng, Ronald R. Grunstein, Christopher James Gordon
IEEE J. Biomed. Health Informatics2
2020 Non-Contact Sleep Stage Detection Using Canonical Correlation Analysis of Respiratory Sound
abstract
Respiratory sound is able to differentiate sleep stages and provide a non-contact and cost-effective solution for the diagnosis and treatment monitoring of sleep-related diseases. While most of the existing respiratory sound-based methods focus on a limited number of sleep stages such as sleep/wake and wake/rapid eye movement (REM)/non-REM, it is essential to detect sleep stages at a finer level for sleep quality evaluation. In this paper, we for the first time study a sleep stage detection method aiming at classifying sleep states into four sleep stages: wake, REM, light sleep, and deep sleep from the respiratory sound. In addition to extracting time-domain features, frequency-domain features of respiratory sound, non-linear features of snoring sound are devised to better characterize snoring-related signals of respiratory sound. To effectively fuse the three sets of features, a novel feature fusion technique combining the generalized canonical correlation analysis with the ReliefF algorithm is proposed for discriminative feature selection. Final stage detection is achieved with popular classifiers including decision tree, support vector machines, K-nearest neighbor, and the ensemble classifier. To evaluate our proposed method, we built an in-house dataset, which is comprised of 13 nights of sleep audio data from a sleep laboratory. Experimental results indicate that our proposed method outperforms the existing related ones and is promising for large-scale non-contact sleep monitoring.
Biao Xue, Boya Deng, Hong Hong 0001, Zhiyong Wang 0001, Xiaohua Zhu 0001, David Dagan Feng
IEEE J. Biomed. Health Informatics4
2019 Stacked Memory Network for Video Summarization
abstract
In recent years, supervised video summarization has achieved promising progress with various recurrent neural networks (RNNs) based methods, which treats video summarization as a sequence-to-sequence learning problem to exploit temporal dependency among video frames across variable ranges. However, RNN has limitations in modelling the long-term temporal dependency for summarizing videos with thousands of frames due to the restricted memory storage unit. Therefore, in this paper we propose a stacked memory network called SMN to explicitly model the long dependency among video frames so that redundancy could be minimized in the video summaries produced. Our proposed SMN consists of two key components: Long Short-Term Memory (LSTM) layer and memory layer, where each LSTM layer is augmented with an external memory layer. In particular, we stack multiple LSTM layers and memory layers hierarchically to integrate the learned representation from prior layers. By combining the hidden states of the LSTM layers and the read representations of the memory layers, our SMN is able to derive more accurate video summaries for individual video frames. Compared with the existing RNN based methods, our SMN is particularly good at capturing long temporal dependency among frames with few additional training parameters. Experimental results on two widely used public benchmark datasets: SumMe and TVsum, demonstrate that our proposed model is able to clearly outperform a number of state-of-the-art ones under various settings.
Junbo Wang 0003, Wei Wang 0115, Zhiyong Wang 0001, Liang Wang 0001, David Dagan Feng, Tieniu Tan
ACM Multimedia3
2019 IntersectGAN: Learning Domain Intersection for Generating Images with Multiple Attributes
abstract
Generative adversarial networks (GANs) have demonstrated great success in generating various visual content. However, images generated by existing GANs are often of attributes (e.g., smiling expression) learned from one image domain. As a result, generating images of multiple attributes requires many real samples possessing multiple attributes which are very resource expensive to be collected. In this paper, we propose a novel GAN, namely IntersectGAN, to learn multiple attributes from different image domains through an intersecting architecture. For example, given two image domains $X_1$ and $X_2$ with certain attributes, the intersection $X_1 \cap X_2$ denotes a new domain where images possess the attributes from both $X_1$ and $X_2$ domains. The proposed IntersectGAN consists of two discriminators $D_1$ and $D_2$ to distinguish between generated and real samples of different domains, and three generators where the intersection generator is trained against both discriminators. And an overall adversarial loss function is defined over three generators. As a result, our proposed IntersectGAN can be trained on multiple domains of which each presents one specific attribute, and eventually eliminates the need of real sample images simultaneously possessing multiple attributes. By using the CelebFaces Attributes dataset, our proposed IntersectGAN is able to produce high quality face images possessing multiple attributes (e.g., a face with black hair and a smiling expression). Both qualitative and quantitative evaluations are conducted to compare our proposed IntersectGAN with other baseline methods. Besides, several different applications of IntersectGAN have been explored with promising results.
Zehui Yao, Zhiyong Wang 0001, Wanli Ouyang, Dong Xu 0001, David Dagan Feng
ACM Multimedia3
2019 A study on multi-kernel intuitionistic fuzzy C-means clustering with multiple attributes
Shan Zeng, Zhiyong Wang 0001, Rui Huang 0001, David Dagan Feng
Neurocomputing2
2019 3D human pose estimation from range images with depth difference and geodesic distance
Dehui Kong, Shaofan Wang 0001, Zhiyong Wang 0001
J. Vis. Commun. Image Represent.4
2019 Robust video summarization using collaborative representation of adjacent frames
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng
Multim. Tools Appl.4
2019 Feature covariance matrix-based dynamic hand gesture recognition
Linpu Fang, Guile Wu, Wenxiong Kang, Qiuxia Wu, Zhiyong Wang 0001, David Dagan Feng
Neural Comput. Appl.5
2018 Vision-Based Freezing of Gait Detection with Anatomic Patch Based Representation
Kun Hu 0008, Zhiyong Wang 0001, Kaylena A. Ehgoetz Martens, Simon J. G. Lewis
ACCV (1)2
2018 Video Summarization via Weighted Neighborhood Based Representation
abstract
The recent explosive growth of multimedia data has posed a new set of challenges in computer vision, and video summarization (VS) techniques are increasingly important to automatically summarize a large amount of multimedia data in an effective and efficient manner. Recent years have witnessed the rise and developments of sparse representation based approaches for VS. While the existing methods select keyframes according to the information contained in the single frame, and such a selection based solely on single-frame information may not be robust. Therefore, in this paper, the information of the single frame's neighborhood is taken into consideration, and different weights are assigned to these neighbouring frames. We formulate the VS problem as a weighted neighborhood based representation model, and design a greedy pursuit algorithm to extract keyframes. Experimental results on a benchmark dataset demonstrate that the proposed method can outperform the state of the arts.
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, Ah Chung Tsoi, David Dagan Feng
ICIP4
2018 Exploiting spatial-temporal context for trajectory based action video retrieval
Lelin Zhang, Zhiyong Wang 0001, Shin'ichi Staoh, Tao Mei 0001, David Dagan Feng
Multim. Tools Appl.2
2018 Real-Time Long-Term Tracking With Prediction-Detection-Correction
abstract
Real-time long-term visual tracking is one of the most challenging problems in computer vision due to various factors such as occlusion and motion ambiguity. To achieve robust long-term tracking, most state-of-the-art methods typically construct an online detector in each frame. However, they fail to achieve real-time performance due to high computational complexity. In this paper, we propose a novel real-time long-term tracking algorithm by exploiting a joint Prediction-Detection-Correction Tracking framework (PDCT). We utilize a superpixel optical flow to construct a predictor to estimate the target motion and internal scale variation. To locate the target at a finer level, we develop an improved kernelized correlation detector with an adaptive online learning rate and translation-scale parameters from the predictor. To refine the tracking result and redetect the target in the case of a tracking failure, we devise a corrector utilizing dual online SVMs with dense sampling and reliable history samples. The SVMs are trained with passive-aggressive learning and online retraining strategies. In addition, we employ a selection mechanism for the correlation responses to maintain reliable samples effectively. As a result, our proposed tracker is able to refine tracking results via the corrector and detector and maintains reliable tracking results for subsequent tracking. Extensive experiments on the widely used object tracking benchmark show that the proposed tracker is superior to state-of-the-art trackers in terms of both effectiveness and efficiency, and the integration of each component is effective under the PDCT framework.
Ningxin Liang, Guile Wu, Wenxiong Kang, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Multim.4
2017 Exploring the influence of feature representation for dictionary selection based video summarization
abstract
Dictionary selection based video summarization (VS) algorithms, in which keyframes are considered as a dictionary to reconstruct all the video frames, have been demonstrated to be effective and efficient for video summarization. It has been noticed that the feature representation of video plays a great impact of the performance of VS. In this paper, the influence of feature representation of video frames on the performance of dictionary selection-based VS is for the first time investigated. In addition to the traditional hand-crafted features used in VS, such as color histogram, the deep features learned through deep neural networks are firstly used to represent video frames for dictionary selection-based VS. The impact of dimensionality reduction to the high-dimensional deep learning features on VS is further discussed. Experimental results on a benchmark video dataset demonstrate that deep learning features are able to achieve better performance than traditional hand-crafted features for dictionary selection-based VS. Moreover, the dimensionality of deep learning features can be reduced to decrease the computational cost without the degradation of VS performance.
Mingyang Ma 0004, Shaohui Mei, Jingyu Ji, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng
ICIP5
2017 Dictionary learning-based image compression
abstract
Dictionary learning based image compression has attracted a lot of research efforts due to the inherent sparsity of image contents. Most algorithms in the literature, however, suffer from two drawbacks. First, the atoms selected for image patch reconstruction scatter over the entire dictionary, which leads to a high coding cost. Second, the sparse representation of image patches is performed independently from the quantization of sparse coefficients, which may result in a sub-optimal solution. In this paper, we propose the entropy based orthogonal matching pursuit (EOMP) algorithm and quantization KSVD (QKSVD) algorithm for dictionary learning-based image compression. An entropy regularization term is utilized in EOMP to restrict atom selection, and hence reduces the coding cost, and an adaptive quantization method is incorporated into the dictionary learning procedure in QKSVD to minimize the reconstruction error and quantization error simultaneously. Experimental results on 10 standard benchmark images demonstrate that our proposed approach achieves better performance than several state-of-the-art ones at low bit rate, such as KSVD based compression approach, JPEG, and JPEG-2000.
Yong Xia 0001, Zhiyong Wang 0001
ICIP3
2017 Nonlinear kernel sparse dictionary selection for video summarization
abstract
Sparse dictionary selection (SDS) has demonstrated to be an effective solution for keyframe based video summarization (VS), which generally assumes a linear relation among similar video frames. However, such a linear assumption is not always true for videos. In this paper, the nonlinearity among frames is taken into consideration and a nonlinear SDS model is formulated for VS, in which the nonlinearity is transformed to linearity by projecting a video to a high dimensional feature space induced by a kernel function. Moreover, a kernel simultaneous orthogonal matching pursuit (KSOMP) is proposed to solve the problem. In order to achieve an intuitive and flexible configuration of the VS process, an adaptive criterion is devised to produce video summaries with different lengths for different video content. Experimental results on benchmark video datasets demonstrate that the proposed algorithm outperforms several state-of-the-art VS algorithms.
Mingyang Ma 0004, Shaohui Mei, Junhui Hou, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng
ICME5
2017 Matrix Neural Networks
Junbin Gao, Yi Guo 0001, Zhiyong Wang 0001
ISNN (1)3
2017 Visual tracking utilizing robust complementary learner and adaptive refiner
Guile Wu, Wenxiong Kang, Zhiyong Wang 0001, David Dagan Feng
Neurocomputing4
2017 Learning universal multiview dictionary for human action recognition
Zhiyong Wang 0001, Zhao Xie, Jun Gao 0006, David Dagan Feng
Pattern Recognit.2
2017 Eye tracking data guided feature selection for image classification
Xin Gao 0003, Zhiyong Wang 0001, Zheru Chi
Pattern Recognit.5
2016 Atmospheric turbulence mitigation based on turbulence extraction
abstract
A video taken under the influence of atmospheric turbulence suffers from serious distortion caused by the variation of optical refractive index. In order to reduce geometric distortion and time-space-varying blur, and recover both coarse structure and fine details, a novel turbulence extraction based approach for recovering a latent image from an atmospheric turbulence degraded imagery sequence is proposed. Firstly, a non-rigid image registration method is applied as a preprocessing to reduce geometric deformation. Secondly, the registered image sequence is decomposed into a low-rank background scene component and a sparse turbulent component via matrix decomposition. Different from other approaches, which intend to remove turbulence directly, we manage to extract information of distortion position from the sparse turbulent component to indicate the sharpest turbulence patches. The selected sharpest turbulence patches are then enhanced and fused to generate an enhanced detail layer. Finally, the output image is generated by fusing the deblurred background scene layer and the enhanced detail layer together. Experiments indicate that our approach is capable of significantly alleviating atmospheric turbulence blur and geometric distortion.
Zhiyong Wang 0001, Yangyu Fan, David Dagan Feng
ICASSP2
2016 Robust foreground object segmentation from handheld camera videos with occlusion map
Hao Xiong 0001, Zhiyong Wang 0001, David Dagan Feng
Multim. Tools Appl.2
2016 Investigating the impact of frame rate towards robust human action recognition
Fredro Harjanto, Zhiyong Wang 0001, Shiyang Lu, Ah Chung Tsoi, David Dagan Feng
Signal Process.2
2016 A Scalable Approach for Content-Based Image Retrieval in Peer-to-Peer Networks
abstract
Peer-to-peer networking offers a scalable solution for sharing multimedia data across the network. With a large amount of visual data distributed among different nodes, it is an important but challenging issue to perform content-based retrieval in peer-to-peer networks. While most of the existing methods focus on indexing high dimensional visual features and have limitations of scalability, in this paper we propose a scalable approach for content-based image retrieval in peer-to-peer networks by employing the bag-of-visual-words model. Compared with centralized environments, the key challenge is to efficiently obtain a global codebook, as images are distributed across the whole peer-to-peer network. In addition, a peer-to-peer network often evolves dynamically, which makes a static codebook less effective for retrieval tasks. Therefore, we propose a dynamic codebook updating method by optimizing the mutual information between the resultant codebook and relevance information, and the workload balance among nodes that manage different codewords. In order to further improve retrieval performance and reduce network cost, indexing pruning techniques are developed. Our comprehensive experimental results indicate that the proposed approach is scalable in evolving and distributed peer-to-peer networks, while achieving improved retrieval accuracy.
Lelin Zhang, Zhiyong Wang 0001, Tao Mei 0001, David Dagan Feng
IEEE Trans. Knowl. Data Eng.2
2015 Spatial-temporal correlation for trajectory based action video retrieval
abstract
The bag-of-visual-words model has been widely utilized for content based image and video retrieval due to its scalability. In this paper, we extend this model for human action video retrieval. We adopt dense trajectory features which are able to achieve the state-of-the-art performance on action recognition, while most of the existing video retrieval methods utilize descriptors of local interest points. In order to improve similarity measurement between bag-of-visual-words model based representation, we propose to discover and incorporate spatial-temporal correlation (STC) among the trajectories in a given query video. The spatial-temporal correlation consists of spatial proximity and temporal consistence among trajectories, which is capable of strengthening discriminative power among visual words. Note that such query focused spatial-temporal correlation makes our method dynamic for different queries and is able to improve retrieval performance without significantly increasing the size of a visual vocabulary. The experimental results on an action video dataset demonstrate that our proposed method outperforms other similar methods.
Lelin Zhang, Zhiyong Wang 0001, David Dagan Feng
MMSP3
2015 Resource restricted on-line Video Summarization with Minimum Sparse Reconstruction
abstract
Video Summarization (VS) techniques have been widely utilized to produce a concise video content representation, such that the video content can be quickly explored and the complexity of video based analysis and retrieval applications can be highly reduced. However, little attention has been paid for on-line applications, especially for resource restricted applications, such as onboard VS. In this paper, our previous on-line Minimum Sparse Reconstruction (OnMSR) based VS algorithm is improved for resources restricted applications by confining the size of keyframes for reconstruction. Specially, an on-line reconstruction keyframe set update strategy is designed to meet the requirement of real-time resource restricted situation. Experimental results on various types of videos demonstrate the performance of OnMSR does not vary much by imposing resource constraint in the proposed resource restricted OnMSR (RR-onMSR) algorithm. As a result, the proposed RR-onMSR is very effective for real-time onboard VS applications.
Shaohui Mei, Zhiyong Wang 0001, Mingyi He, David Dagan Feng
PCS2
2015 Video summarization via minimum sparse reconstruction
Shaohui Mei, Genliang Guan, Zhiyong Wang 0001, Shuai Wan, Mingyi He, David Dagan Feng
Pattern Recognit.3
2015 Exploratory Product Image Search With Circle-to-Search Interaction
abstract
Exploratory search is emerging as a new form of information-seeking activity in the research community, which generally combines browsing and searching content together to help users gain additional knowledge and form accurate queries, thereby assisting the users with their seeking and investigation activities. However, there have been few attempts at addressing integrated exploratory search solutions when image browsing is incorporated into the exploring loop. In this paper, we investigate the challenges of understanding users' search interests from the product images being browsed and inferring their actual search intentions. We propose a novel interactive image exploring system for allowing users to lightly switch between browse and search processes, and naturally complete visual-based exploratory search tasks in an effective and efficient way. This system enables users to specify their visual search interests in product images by circling any visual objects in web pages, and then the system automatically infers users' underlying intent by analyzing the browsing context and by analyzing the same or similar product images obtained by large-scale image search technology. Users can then utilize the recommended queries to complete intent-specific exploratory tasks. The proposed solution is one of the first attempts to understand users' interests for a visual-based exploratory product search task by integrating the browse and search activities. We have evaluated our system performance based on five million product images. The evaluation study demonstrates that the proposed system provides accurate intent-driven search results and fast response to exploratory search demands compared with the conventional image search methods, and also, provides users with robust results to satisfy their exploring experience.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2014 Iterative keyframe selection by orthogonal subspace projection
abstract
Recent developments on sparse dictionary selection have demonstrated promising results for Video Summarization (VS). However, the convex relaxation based solution cannot ensure the sparsity of the dictionary directly. In this paper, a selection matrix is proposed to model the VS problem, according to which the L0norm of this selection matrix is imposed to ensure sparsity directly. As a result, a computational efficient Orthogonal Subspace Projection (OSP) based Iterative Keyframe Selection (IKS) algorithm is proposed for VS. In addition, a Percentage Of Reconstruction (POR) criterion is proposed to provide an intuitive and flexible control of the length of final video summaries even without prior knowledge of a given video. Experimental results on a popular benchmark dataset demonstrate that our proposed algorithm outperforms the state-of-the-art methods.
Shaohui Mei, Genliang Guan, Zhiyong Wang 0001, Mingyi He, Shuai Wan, David Dagan Feng
ICIP3
2014 L2, 0 constrained sparse dictionary selection for video summarization
abstract
The ever increasing volume of video content has created profound challenges for developing efficient video summarization (VS) techniques to access the data. Recent developments on sparse dictionary selection have demonstrated promising results for VS, however, the convex relaxation based solution cannot ensure the sparsity of the dictionary directly and it selects keyframes in a local point of view. In this paper, an L2,0constrained sparse dictionary selection model is proposed to reformulate the problem of VS. In addition, a simultaneous orthogonal matching pursuit (SOMP) based method is proposed to obtain an approximate solution for the proposed model without smoothing the penalty function, and thus selects keyframes in a global point of view. In order to allow for intuitive and flexible configuration of VS process, a percentage of residuals (POR) criterion is also developed to produce video summaries in different lengths. Experimental results demonstrate that our proposed method outperforms the state-of-the-art.
Shaohui Mei, Genliang Guan, Zhiyong Wang 0001, Mingyi He, Xian-Sheng Hua 0001, David Dagan Feng
ICME3
2014 Spectral embedding based facial expression recognition with multiple features
Kaimin Yu, Zhiyong Wang 0001, Markus Hagenbuchner, David Dagan Feng
Neurocomputing2
2014 A Bag-of-Importance Model With Locality-Constrained Coding Based Feature Learning for Video Summarization
abstract
Video summarization helps users obtain quick comprehension of video content. Recently, some studies have utilized local features to represent each video frame and formulate video summarization as a coverage problem of local features. However, the importance of individual local features has not been exploited. In this paper, we propose a novel Bag-of-Importance (BoI) model for static video summarization by identifying the frames with important local features as keyframes, which is one of the first studies formulating video summarization at local feature level, instead of at global feature level. That is, by representing each frame with local features, a video is characterized with a bag of local features weighted with individual importance scores and the frames with more important local features are more representative, where the representativeness of each frame is the aggregation of the weighted importance of the local features contained in the frame. In addition, we propose to learn a transformation from a raw local feature to a more powerful sparse nonlinear representation for deriving the importance score of each local feature, rather than directly utilize the hand-crafted visual features like most of the existing approaches. Specifically, we first employ locality-constrained linear coding (LCC) to project each local feature into a sparse transformed space. LCC is able to take advantage of the manifold geometric structure of the high dimensional feature space and form the manifold of the low dimensional transformed space with the coordinates of a set of anchor points. Then we calculate the l2 norm of each anchor point as the importance score of each local feature which is projected to the anchor point. Finally, the distribution of the importance scores of all the local features in a video is obtained as the BoI representation of the video. We further differentiate the importance of local features with a spatial weighting template by taking the perceptual difference among spatial regions of a frame into account. As a result, our proposed video summarization approach is able to exploit both the inter-frame and intra-frame properties of feature representations and identify keyframes capturing both the dominant content and discriminative details within a video. Experimental results on three video datasets across various genres demonstrate that the proposed approach clearly outperforms several state-of-the-art methods.
Shiyang Lu, Zhiyong Wang 0001, Tao Mei 0001, Genliang Guan, David Dagan Feng
IEEE Trans. Multim.2
2014 Browse-to-Search: Interactive Exploratory Search with Visual Entities
abstract
With the development of image search technology, users are no longer satisfied with searching for images using just metadata and textual descriptions. Instead, more search demands are focused on retrieving images based on similarities in their contents (textures, colors, shapes etc.). Nevertheless, one image may deliver rich or complex content and multiple interests. Sometimes users do not sufficiently define or describe their seeking demands for images even when general search interests appear, owing to a lack of specific knowledge to express their intents. A new form of information seeking activity, referred to as exploratory search, is emerging in the research community, which generally combines browsing and searching content together to help users gain additional knowledge and form accurate queries, thereby assisting the users with their seeking and investigation activities. However, there have been few attempts at addressing integrated exploratory search solutions when image browsing is incorporated into the exploring loop. In this work, we investigate the challenges of understanding users' search interests from the images being browsed and infer their actual search intentions. We develop a novel system to explore an effective and efficient way for allowing users to seamlessly switch between browse and search processes, and naturally complete visual-based exploratory search tasks. The system, called Browse-to-Search enables users to specify their visual search interests by circling any visual objects in the webpages being browsed, and then the system automatically forms the visual entities to represent users' underlying intent. One visual entity is not limited by the original image content, but also encapsulated by the textual-based browsing context and the associated heterogeneous attributes. We use large-scale image search technology to find the associated textual attributes from the repository. Users can then utilize the encapsulated visual entities to complete search tasks. The Browse-to-Search system is one of the first attempts to integrate browse and search activities for a visual-based exploratory search, which is characterized by four unique properties: (1) in session—searching is performed during browsing session and search results naturally accompany with browsing content; (2) in context—the pages being browsed provide text-based contextual cues for searching; (3) in focus—users can focus on the visual content of interest without worrying about the difficulties of query formulation, and visual entities will be automatically formed; and (4) intuitiveness—a touch and visual search-based user interface provides a natural user experience. We deploy the Browse-to-Search system on tablet devices and evaluate the system performance using millions of images. We demonstrate that it is effective and efficient in facilitating the user's exploratory search compared to the conventional image search methods and, more importantly, provides users with more robust results to satisfy their exploring experience.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, Shipeng Li 0001
ACM Trans. Inf. Syst.5
2014 A Top-Down Approach for Video Summarization
abstract
While most existing video summarization approaches aim to identify important frames of a video from either a global or local perspective, we propose a top-down approach consisting of scene identification and scene summarization. For scene identification, we represent each frame with global features and utilize a scalable clustering method. We then formulate scene summarization as choosing those frames that best cover a set of local descriptors with minimal redundancy. In addition, we develop a visual word-based approach to make our approach more computationally scalable. Experimental results on two benchmark datasets demonstrate that our proposed approach clearly outperforms the state-of-the-art.
Genliang Guan, Zhiyong Wang 0001, Shaohui Mei, Maximilian Ott, Mingyi He, David Dagan Feng
ACM Trans. Multim. Comput. Commun. Appl.2
2013 A supervised multiview spectral embedding method for neuroimaging classification
abstract
The multi-view/multi-modal features are commonly used in neuroimaging classification because they could provide complementary information to each other and thus result in better classification performance than single-view features. However, it is very challenging to effectively integrate such rich features, since straightforward concatenation or singleview spectral embedding methods rarely leads to physically meaningful integration. In this paper, we present a supervised multi-view/multi-modal spectral embedding method (SMSE) for neuroimaging classification. This method embeds the high dimensional multi-view features derived from multi-modal neuroimaging data into a low dimensional feature space and preserves the optimal local embeddings among different views. The proposed SMSE algorithm, validated using three groups of neuroimaging data, is able to achieve significant classification improvement over the state-of-the-art multi-view spectral embedding methods.
Sidong Liu, Lelin Zhang, Tom Weidong Cai, Yang Song 0001, Zhiyong Wang 0001, Lingfeng Wen, David Dagan Feng
ICIP5
2013 Graph cuts based relevance feedback in image retrieval
abstract
Relevance feedback (RF) allows users to be actively involved in the information retrieval process and has been widely used in various information retrieval tasks. While most existing RF methods in content-based image retrieval (CBIR) focus on visual features of individual images only, in this paper we formulate the relevance feedback process as an energy minimization problem. The energy function takes into account both the feature aspect of each image and the manifold structure among individual images. The solution of labelling images as relevant or irrelevant is obtained with the graph cuts method. As a result, our method enables flexibly partitioning the feature space and labelling of images and is capable of handling challenging scenarios (or queries). Experimental results demonstrate that our proposed method outperforms the popular RF methods.
Lelin Zhang, Sidong Liu, Zhiyong Wang 0001, Tom Weidong Cai, Yang Song 0001, David Dagan Feng
ICIP3
2013 Fast human action classification and VOI localization with enhanced sparse coding
Shiyang Lu, Jian Zhang 0002, Zhiyong Wang 0001, David Dagan Feng
J. Vis. Commun. Image Represent.3
2013 Discriminative two-level feature selection for realistic human action recognition
Qiuxia Wu, Zhiyong Wang 0001, Feiqi Deng, Yong Xia 0001, Wenxiong Kang, David Dagan Feng
J. Vis. Commun. Image Represent.2
2013 Semantic context based refinement for news video annotation
Zhiyong Wang 0001, Genliang Guan, Li Zhuo 0001, David Dagan Feng
Multim. Tools Appl.1
2013 Learning realistic facial expressions from web images
Kaimin Yu, Zhiyong Wang 0001, Li Zhuo 0001, Zheru Chi, David Dagan Feng
Pattern Recognit.2
2013 Keypoint-Based Keyframe Selection
abstract
Keyframe selection has been crucial for effective and efficient video content analysis. While most of the existing approaches represent individual frames with global features, we, for the first time, propose a keypoint-based framework to address the keyframe selection problem so that local features can be employed in selecting keyframes. In general, the selected keyframes should both be representative of video content and containing minimum redundancy. Therefore, we introduce two criteria, coverage and redundancy, based on keypoint matching in the selection process. Comprehensive experiments demonstrate that our approach outperforms the state of the art.
Genliang Guan, Zhiyong Wang 0001, Shiyang Lu, Jeremiah D. Deng, David Dagan Feng
IEEE Trans. Circuits Syst. Video Technol.2
2013 Realistic Human Action Recognition With Multimodal Feature Selection and Fusion
abstract
Although promising results have been achieved for human action recognition under well-controlled conditions, it is very challenging to recognize human actions in realistic scenarios due to increased difficulties such as dynamic backgrounds. In this paper, we propose to take multimodal (i.e., audiovisual) characteristics of realistic human action videos into account in human action recognition for the first time, since, in realistic scenarios, audio signals accompanying an action generally provide a cue to the nature of the action, such as phone ringing to answering the phone . In order to cope with diverse audio cues of an action in realistic scenarios, we propose to identify effective features from a large number of audio features with the generalized multiple kernel learning algorithm. The widely used space-time interest point descriptors are utilized as visual features, and a support vector machine is employed for both audio- and video-based classifications. At the final stage, fuzzy integral is utilized to fuse recognition results of both audio and visual modalities. Experimental results on the challenging Hollywood-2 Human Action data set demonstrate that the proposed approach is able to achieve better recognition performance improvement than that of integrating scene context. It is also discovered how audio context influences realistic action recognition from our comprehensive experiments.
Qiuxia Wu, Zhiyong Wang 0001, Feiqi Deng, Zheru Chi, David Dagan Feng
IEEE Trans. Syst. Man Cybern. Syst.2
2012 Unsupervised Spectral Mixture Analysis with Hopfield Neural Network for hyperspectral images
abstract
Spectral Mixture Analysis (SMA) has been widely utilized to address the mixed-pixel problem in the quantitative analysis of hyperspectral remote sensing images. Recently Nonnegative Matrix Factorization (NMF) has been successfully utilized to simultaneously perform endmember extraction (EE) and abundance estimation (AE). In this paper, we formulate the solution of NMF by performing EE and AE iteratively. Based on our previous Hopfield Neural Network (HNN) based AE algorithm, an HNN is also constructed for EE to solve the multiplicative updating problem of NMF for SMA. As a result, SMA is conducted in an unsupervised manner and our algorithm is able to extract virtual endmembers without assuming the presence of spectrally pure constituents in hyperspectral scenes. We further extend such strategy to solve the constrained NMF (cNMF) models for SMA, where extra constraints are imposed to better model the mixed-pixel problem. Experimental results on both synthetic and real hyperspectral images demonstrate the effectiveness of our proposed HNN based unsupervised SMA algorithms.
Shaohui Mei, Mingyi He, Zhiyong Wang 0001, David Dagan Feng
ICIP3
2012 Browse-to-search
abstract
This demonstration presents a novel interactive online shopping application based on visual search technologies. When users want to buy something on a shopping site, they usually have the requirement of looking for related information from other web sites. Therefore users need to switch between the web page being browsed and other websites that provide search results. The proposed application enables users to naturally search products of interest when they browse a web page, and make their even causal purchase intent easily satisfied. The interactive shopping experience is characterized by: 1) in session---it allows users to specify the purchase intent in the browsing session, instead of leaving the current page and navigating to other websites; 2) in context---the browsed web page provides implicit context information which helps infer user purchase preferences; 3) in focus---users easily specify their search interest using gesture on touch devices and do not need to formulate queries in search box; 4) natural-gesture inputs and visual-based search provides users a natural shopping experience. The system is evaluated against a data set consisting of several millions commercial product images.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, David Dagan Feng, Jian-Tao Sun, Shipeng Li 0001
ACM Multimedia5
2012 What is happening: annotating images with verbs
abstract
Image annotation has been widely investigated to discover the semantics of an image. However, most of the existing algorithms focus on noun tags (e.g. concepts and objects). Since an image is a snapshot of the real world event, annotating images with verbs will enable richer understanding of an image. In this paper, we propose a data-driven approach to verb oriented image annotation. At first, we obtain verb candidates by generating search queries for a given image with initial noun tags and establishing a sentence corpus from those queries. We utilize visualness to filter tags which are not visually presentable (e.g. pain) and differentiate tags into two categories (i.e. scene based and object based) to impose linguistic rules in verb extraction. Then we further re-rank the candidate verbs with the tag context discovered from the images which are both semantically and visually similar to the given image in the MIRFlickr dataset. Our experimental results from user study demonstrate that our proposed approach is promising.
Gang Tian, Genliang Guan, Zhiyong Wang 0001, David Dagan Feng
ACM Multimedia3
2011 StoryImaging: a media-rich presentation system for textual stories
abstract
In this demo, we develop the StoryImaging system to illustrate a textual story with both images harvested from the Web and synthesized speech. At the backend, a story is firstly processed to identify key terms such as named entities and to obtain the story summary. With the aid of commercial search engines, images are then collected from the Web for those key terms and re-ranked by taking the summary as context. At last, images are clustered to provide an overview of the story. At the web-based frontend, the user interface has been tailored to both improve information comprehension and provide engaging and explorative experiences for users by closely bridging textual and visual modalities.
Genliang Guan, Zhiyong Wang 0001, Xian-Sheng Hua 0001, David Dagan Feng
ACM Multimedia2
2011 Improving Spatial-Spectral Endmember Extraction in the Presence of Anomalous Ground Objects
abstract
Endmember extraction (EE) has been widely utilized to extract spectrally unique and singular spectral signatures for spectral mixture analysis of hyperspectral images. Recently, spatial–spectral EE (SSEE) algorithms have been proposed to achieve superior performance over spectral EE (SEE) algorithms by taking both spectral similarity and spatial context into account. However, these algorithms tend to neglect anomalous endmembers that are also of interest. Therefore, in this paper, an improved SSEE (iSSEE) algorithm is proposed to address such limitation of conventional SSEE algorithms by accounting for both anomalous and normal endmembers. By developing simplex projection and simplex complementary projection, all the hyperspectral pixels are projected into a simplex determined by the normal endmembers extracted in conventional SSEE algorithms. As a result, anomalous endmembers are identified iteratively by utilizing the$l_{2}^{\infty}$norm to find the maximum simplex complementary projection. In order to determine how many anomalous endmembers are to be extracted, a novel Residual-be-Noise Probability-based algorithm is also proposed by elegantly utilizing the spatial-purity map generated in the previous SSEE step. Experimental results on both synthetic and real datasets demonstrate that simplex projection errors can be significantly reduced by identifying both anomalous and normal endmembers in the proposed iSSEE algorithm. It is also confirmed that the performance of the proposed iSSEE algorithm clearly outperforms that of SEE algorithms since both spatial context and spectral similarity are utilized.
Shaohui Mei, Mingyi He, Yifan Zhang 0006, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Geosci. Remote. Sens.4
2010 Adaptive reference frame selection for near-duplicate video shot detection
abstract
Near-duplicate video shots provide critical visual link between videos and detecting such video shots efficiently and effectively is of paramount importance in many applications such as detecting copyright infringement. In this paper, we propose an improved near-duplicate video shot detection approach by adaptively selecting reference frames for more effective shot representation. The correlation between adjacent frames is measured with Pearson's Correlation Coefficient (PCC) so that a set of compact yet representative reference frames can be selected adaptively in terms of content variation within video shots. Interest points are further extracted from the selected frames to effectively represent shot contents for similarity matching. Comprehensive experimental results on TRECVID-2008 corpus demonstrate that our proposed approach outperforms the state-of-the-art method effectively.
Shiyang Lu, Zhiyong Wang 0001, Maximilian Ott, David Dagan Feng
ICIP2
2010 Two-step similarity matching for Content-Based Video Retrieval in P2P, networks
abstract
Multimedia data, particularly, video data, has dominated peer-to-peer (P2P) networks. Therefore, it is demanding to provide content based retrieval in P2P networks. Similarity matching is one of the challenging issues. In this paper, we present a novel two-step method to reduce computational complexity of similarity matching in P2P networks. In the first step, an efficient maximum matching (MM) technique is employed to obtain an initial set of similar video candidates. In the second step, these candidates are further selected with a more accurate, but more computationally expensive optimal matching (OM) technique. In order to further improve the computational efficiency of the proposed method, four other matching techniques are proposed to replace MM technique. Various experimental results indicate that the proposed approach is more effective for CBVR while achieving significantly computational saving.
Jin Niu, Zhiyong Wang 0001, David Dagan Feng
ICME2
2010 Mixture Analysis by Multichannel Hopfield Neural Network
abstract
Due to the spatial-resolution limitation, mixed pixels containing energy reflected from more than one type of ground objects are widely present in remote sensing images, which often results in inefficient quantitative analysis. To effectively decompose such mixtures, a fully constrained linear unmixing algorithm based on a multichannel Hopfield neural network (MHNN) is proposed in this letter. The proposed MHNN algorithm is actually a Hopfield-based architecture which handles all the pixels in an image synchronously, instead of considering a per-pixel procedure. Due to the synchronous unmixing property of MHNN, a noise energy percentage (NEP) stopping criterion which utilizes the signal-to-noise ratio is proposed to obtain optimal results for different applications automatically. Experimental results demonstrate that the proposed multichannel structure makes the Hopfield-based mixture analysis feasible for real-world applications with acceptable time cost. It has also been observed that the proposed MHNN-based mixture-analysis algorithm outperforms the other two popular linear mixture-analysis algorithms and that the NEP stopping criterion can approach optimal unmixing results adaptively and accurately.
Shaohui Mei, Mingyi He, Zhiyong Wang 0001, David Dagan Feng
IEEE Geosci. Remote. Sens. Lett.3
2010 An efficient retinex-like brightness normalization method for coding camera flashes and strong brightness variation in videos
Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Zhiyong Wang 0001
Signal Process. Image Commun.4
2010 Spatial Purity Based Endmember Extraction for Spectral Mixture Analysis
abstract
Spectral mixture analysis (SMA) has been widely utilized to address the mixed-pixel problem in the quantitative analysis of hyperspectral remote sensing images, in which endmember extraction (EE) plays an extremely important role. In this paper, a novel algorithm is proposed to integrate both spectral similarity and spatial context for EE. The spatial context is exploited from two aspects. At first, initial endmember candidates are identified by determining the spatial purity (SP) of pixels in their spatial neighborhoods (SNs). Several SP measurements are investigated at both intensity level and feature level. In order to alleviate local spectra variability, the average of the pixels in pure SNs are voted as endmember candidates. Then, the spatial connectivity is utilized to merge spatially related endmember candidates by finding connection paths in a graph so that the number of endmember candidates is further reduced, which results in computational efficiency and better performance in SMA by alleviating global spectral variability. Experimental results on both synthetic and real hyperspectral images demonstrate that the proposed SP based EE (SPEE) algorithm outperforms the other popular EE algorithms. It is also observed that feature-level SP measurements are more distinguishable than intensity-level SP measurements to discriminate pure SNs from mixed SNs.
Shaohui Mei, Mingyi He, Zhiyong Wang 0001, David Dagan Feng
IEEE Trans. Geosci. Remote. Sens.3
2009 Hierarchical Gaussian Mixture Model for Image Annotation via PLSA
abstract
In order to mimic the representation of textual documents, some approaches have recently been proposed to represent visual contents in terms of visual words in many applications such as object recognition and image annotation. In this paper, we propose to build an effective visual vocabulary by using Hierarchical Gaussian Mixture model instead of traditional clustering methods. In addition, Probabilistic Latent Semantic Analysis is employed to explore semantic aspects of visual concepts and to discover topic clusters among documents and visual words so that every image is projected on to a lower dimensional topic space for more efficient and effective annotation. Experimental results obtained on TRECVID 2005 dataset demonstrate that the Hierarchical Gaussian Mixture model can achieve better annotation performance than hierarchical k-means clustering even by using simple k-NN annotation scheme.
Zhiyong Wang 0001, Huaibin Yi, David Dagan Feng
ICIG1
2009 Improved concept similarity measuring in the visual domain
abstract
Exploring semantic similarity between concepts in visual domain has a wide range of applications such as natural language processing and multimedia retrieval, which in general requires both a large pool of sample images for each concept and a model to capture its visual characteristics. Instead of relying on high quality and large quantity sample data which is very difficult to obtain, in this paper, a novel method is proposed to achieve improvement in measuring concept similarity by incorporating concept modeling technique into data pruning process. At first, a number of sampling concept models are obtained by sampling a subset from the sample dataset of each concept. Then noisy samples are discarded in terms of their probabilities to the sampling concept models. Experimental results on 31,275 Web images of 38 concepts defined in LSCOM indicate that the concept similarity obtained through our proposed approach is more coherent to human cognition. A concept hierarchy tree built from the 38 concepts with their similarity further demonstrates the effectiveness of our proposed method.
Genliang Guan, Zhiyong Wang 0001, Qi Tian 0001, David Dagan Feng
MMSP2
2009 Reliable object recognition using SIFT features
abstract
SIFT (scale invariant feature transform) features have been one of the most efficient descriptors for object recognition. However, the excessive number of key points and high dimensionality has limited its capacity in object recognition. In this paper we present a novel method based on SIFT features for reliable object recognition. At first, a matching tree is constructed to eliminate non-essential key points. In order to achieve viewpoint independence, a 3D model is constructed for each object in the filtered SIFT feature space. Experimental results on both Caltech 101 and COIL 100 datasets indicate the effectiveness of our proposed algorithm.
Florin Alexandru Pavel, Zhiyong Wang 0001, David Dagan Feng
MMSP2
2009 Two-level indexing for high-dimensional range queries in peer-to-peer networks
abstract
Supporting complex and efficient lookup queries in peer-to-peer networks is challenging, though simple keyword based lookup queries are well supported by most deployed systems. This paper presents a two-level indexing structure built on distributed hash table (DHT) aiming to support range queries on high-dimensional feature space in peer-to-peer network. Unlike most existing systems, where every node is responsible for a data partition, our design only utilizes a small part of the nodes to manage partitions. These partition nodes form the first level index. The second level index consists of one or more server nodes, which maintains links to each partition node. Additionally, a merge and split mechanism is designed to dynamically adjust the workload among nodes. Experimental results indicate that our system offers promising performance in terms of workload balance in churn networks. The flexibility to work with any DHT and the capability to support multiple feature spaces further make our proposed approach a feasible extension for file sharing networks.
Lelin Zhang, Zhiyong Wang 0001, David Dagan Feng
MMSP2
2008 Retinex based motion estimation for sequences with brightness variations and its application to H.264
abstract
Conventional motion estimation does not take inter-frame brightness variations into consideration, which causes inefficient video coding for sequences involving brightness variations. H.264 provides a specific mode called weighted prediction targeting to improve the coding efficiency for this case. In this paper, we propose a Retinex based motion estimation scheme which effectively removes the inter-frame de-correlation factor resulting from brightness variations. We also propose to use some DCT techniques to generate the Retinex images for both current and reference images and apply conventional motion estimation and compensation procedures for coding. We applied the scheme to the H.264 testing the efficiency in the multiple reference frame motion compensation environment. Experimental results show that our proposed scheme outperforms the H.264 system with weighted prediction enabled. It allows the system to use a smaller number of reference frames for coding, e.g. 2, to achieve a similar (or slightly better) compression efficiency of the H.264 system using 5 reference frames.
Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Zhiyong Wang 0001
ICASSP4
2008 Windowing technique for the DCT based retinex algorithm to handle videos with brightness variations coded using the H.264
abstract
Conventional block based motion estimators assume constant inter-frame object brightness. Pixel discrepancy is resulted primarily from motion factor without considering the influence of brightness changes. In this paper, we propose a simple and efficient windowing technique using the Hamming window and integrate it to our previously proposed algorithm. The algorithm is based on retinex approach using the DCT technique and designed to handle brightness variations. The new technique manages to greatly reduce the influence of ripple effect and further increase the compression efficiency without adding any extra overhead bits to the bit-stream. We applied the scheme to H.264 for testing. Experimental results show that the retinex based approach is an effective technique to handle inter-frame brightness variations and outperforms the H.264 system with weighted prediction enabled for sequence involving brightness variations. With our proposed windowing technique using the Hamming window, the coding efficiency can be further improved by a maximum of 0.17dB.
Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Zhiyong Wang 0001
ICIP4
2008 Measuring semantic similarity between concepts in visual domain
abstract
Concept similarity has been intensively researched in the natural language processing domain due to its important role in many applications such as language modeling and information retrieval. There are few studies on measuring concept similarity in visual domain, though concept based multimedia information retrieval has attracted a lot of attentions. In this paper, we present a scalable framework for such a purpose, which is different from traditional approaches to exploring correlation among concepts in image/video annotation domain. For each concept, a model based on feature distribution is built using sample images collected from the Internet. And similarity between concepts is measured with the similarity between their models. Hereby, a Gaussian Mixture Model (GMM) is employed to model each concept and two similarity measurements are investigated. Experimental results on 13,974 images of 16 concepts collected through image search engines have demonstrated that the similarity between concepts is very close to human perception. In addition, the entropy of GMM cluster distributions can be a good indication of selecting concepts for image/video annotation.
Zhiyong Wang 0001, Genliang Guan, David Dagan Feng
MMSP1
2008 Image annotation with parametric mixture model based multi-class multi-labeling
abstract
Image annotation, which labels an image with a set of semantic terms so as to bridge the semantic gap between low level features and high level semantics in visual information retrieval, is generally posed as a classification problem. Recently, multi-label classification has been investigated for image annotation since an image presents rich contents and can be associated with multiple concepts (i.e. labels). In this paper, a parametric mixture model based multi-class multi-labeling approach is proposed to tackle image annotation. Instead of building classifiers to learn individual labels exclusively, we model images with parametric mixture models so that the mixture characteristics of labels can be simultaneously exploited in both training and annotation processes. Our proposed method has been benchmarked with several state-of-the-art methods and achieved promising results.
Zhiyong Wang 0001, Wan-Chi Siu, David Dagan Feng
MMSP1
2007 Concept Constrained Image Region Annotation
abstract
Annotating image regions has been a challenging open issue in many areas such as image content understanding and image retrieval. In this paper, rather than solely rely on visual features of image regions, a novel approach is proposed to improve region annotation by taking concept constraints into account, since high level conceptual information such as image categories can increase the confidence of possible region labels as well as decrease the confidence of impossible region labels. We employ statistical models to learn the relationships among visual features, image concepts, and region labels. As a result, a set of possible region labels can be derived from a set of visual feature vectors of a given image so as to refine the annotation output obtained by using visual feature only. Promising experimental results have been demonstrated on 8462 regions of the University of Washington image dataset with diverse concepts for the proposed approach.
Zhiyong Wang 0001, Kelly Lam, Li Zhuo 0001, David Dagan Feng
MMSP1
2006 Annotating Image Regions Using Spatial Context
abstract
Image annotation plays an important role in bridging the semantic gap between low level features and high level semantic contents in image access. In this paper, such a task is tackled by annotating regions which are primitives of a visual scene. We propose a probabilistic model to characterize spatial context for region annotation. Such a model provides a unifying framework integrating both feature distribution models and spatial context models. A wide range of advanced modeling techniques can be utilized to further extend this framework. The approach is also potentially scalable to a large number of semantic concepts and a large number of images. Experimental results based on simple parametric models demonstrate promising results of our approach by investigating the impacts of neighbors, segmentation, and visual features
Zhiyong Wang 0001, David Dagan Feng, Zheru Chi
ISM1
2004 Comparison of image partition methods for adaptive image categorization based on structural image representation
abstract
Image categorization is very helpful for organizing large image databases efficiently, however, it is yet very challenging due to lack of effective image representations. Our previous work showed that structural representations were good at characterizing image contents, since image contents could be exploited from coarse to fine scales through the structures representation and fewer visual features are required. In this paper, several popular image partition methods are investigated for adaptive image categorization based on structural representation. Experimental results on seven categories of scenery images show that both the structure and node attributes are important to categorize image contents. In addition, the more similar the structures of each category, the better the categorization performance.
Zhiyong Wang 0001, David Dagan Feng, Zheru Chi
ICARCV1
2003 Region-of-interest based flower images retrieval
abstract
Flower image retrieval is a very important step for computer-aided plant species identification. We propose an efficient segmentation method based on color clustering and domain knowledge to extract flower regions from flower images. For flower retrieval, we use the color histogram of a flower region to characterize the color features of a flower and two shape-based sets of features, centroid-contour-distance (CCD) and angle code histogram (ACH), to characterize the shape features of a flower contour. Experimental results show that our flower region extraction approach based on color clustering and domain knowledge can achieve accurate flower regions. The retrieval results on a database of 885 flower images collected from 14 plant species show that our region-of-interest (ROI) based retrieval approach can perform better than the Swain method based on the global color histogram (Swain, M.J. and Ballard, D.H., Int. J. of Computer Vision, vol.7, no.1, p.11-32, 1991).
Anxiang Hong, Zheru Chi, Gang Chen 0006, Zhiyong Wang 0001
ICASSP (3)4
2003 Efficient Learning in Adaptive Processing of Data Structures
Siu-Yeung Cho, Zheru Chi, Zhiyong Wang 0001, Wan-Chi Siu
Neural Process. Lett.3
2002 Fuzzy integral for leaf image retrieval
abstract
Generally, the more features utilized, the better the retrieval performance. However, it is a very challenging task to combine different feature sets in a way reflecting human perception. This paper presents the combination of different shape based feature sets using fuzzy integral for leaf image retrieval. The feature sets used in our system include centroid-contour distance curve, eccentricity, and angle code histogram. The fuzzy integral approach can release the user's burden from tuning the combination parameters. In order to reduce the matching time in the retrieval process, a thinning based method is proposed to locate the start point of a leaf contour. Experimental results on 440 leaf images from 44 plant species (10 samples from each plant species) show that the fuzzy integral approach can achieve a comparable retrieval performance with the best case of the weighted summation combination. The results also indicate that our approach, which are more efficient, can achieve a better retrieval performance than both the curvature scale space (CSS) method and the modified Fourier descriptor (MFD) method.
Zhiyong Wang 0001, Zheru Chi, David Dagan Feng
FUZZ-IEEE1