VLDB 2026 Research / reviewers in the wild / expert
Hui Fang 0003
dblp:03/2511-3
· DBLP profile ↗
76ranked-venue papers
9as first author
56since 2021 · last 2026
0000-0001-9365-7420ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 5 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 7 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ACID-Style: An Adaptive Condition Injection Diffusion Model for Arbitrary Style TransferabstractArbitrary style transfer (AST), a popular AI-powered photo editing function, aims to strike an optimal balance between content and style injection from two images in order to generate a novel high-fidelity stylised image. Recently, diffusion models have been applied to AST due to their high generation quality as well as flexibility to embed conditions. However, these models are still not satisfactory and may exhibit inferior performance compared to non-diffusion based methods. This is due to the diffusion process not being purposely designed for AST, leading to suboptimal solutions to trade-off content preservation and style embedding. In this paper, we propose ACID-Style, a novel adaptive condition injection diffusion-based AST framework for improved content/style feature injection to address this research challenge. Using two lightweight adapters, a content and a style injection module, and an adaptive injection mechanism, our approach is able to fully exploit a pre-trained stable diffusion model for AST-specific adaptation and our diffusion model thus learns the most effective timing for content and style injection in the diffusion sampling process. Comprehensive evaluations demonstrate that our method achieves superior style transfer performance, both quantitatively and qualitatively, compared to other state-of-the-art style transfer methods. Ting Yang 0009, Siyu Yang 0005, Xiyao Liu 0001, Songtao Wu, Gerald Schaefer, Kuanhong Xu, Hui Fang 0003 |
AAAI | 7 |
| 2026 | A gated recurrent unit-based soft actor-critic approach with social force model crowd simulation for improved mobile robot path planning
Dezhen Zhang, Guoxu Wang, Gerald Schaefer, Hui Fang 0003 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | An adaptive multimodal semantic knowledge enhanced framework for sarcasm detection
Jing Dong 0009, Yu Sui, Qiang Zhang 0008, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015, Xiaoyong Fang |
Expert Syst. Appl. | 4 |
| 2026 | Fine-grained face personalisation using a text-guided multi-attribute embedded diffusion model
Jing Dong 0009, Qiang Zhang 0008, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015, Xiaoyong Fang |
Expert Syst. Appl. | 4 |
| 2026 | UniStyleDiff: A unified diffusion-driven framework for image and video style transfer
Siyu Yang 0005, Chunchen Ke, Jian Zhang 0048, Chunwei Miao, Xiyao Liu 0001, Songtao Wu, Kuanhong Xu, Da Huang 0002, Hui Fang 0003 |
Expert Syst. Appl. | 9 |
| 2026 | Refining pseudo-labels through iterative mix-up for weakly supervised semantic segmentationabstractWeakly supervised semantic segmentation (WSSS) aims to provide accurate pixel-level annotation based on only weak guidance, primarily derived from image-level labels. Recent WSSS methods exploit pseudo-labels generated from improved class activation maps (CAMs) to train a fine-grained classification model for semantic segmentation. However, these pseudo-labels are unreliable because they tend to either miss parts of the objects or include irrelevant regions due to weak guidance from individual images. In this paper, we propose a simple yet effective iterative mix-up strategy, Pseudo-Label-based Mix (PL-Mix), that refines pseudo-labels iteratively, thereby further enhancing WSSS performance. During each iteration, we migrate object regions from pseudo-labels produced in previous steps and render them with new contexts in a mix-up fashion. Due to model consistency enforcement across varied backgrounds and new combinations of multiple objects from enriched image samples, these pseudo-labels progressively become more accurate and reliable. Further enhanced by a masking strategy and a CAM-based earth mover’s distance loss, we achieve state-of-the-art performance on the PASCAL VOC2012 and MS COCO2014 benchmark datasets. Yifan Wang 0008, Kunhao Yuan, Gerald Schaefer, Xiyao Liu 0001, Linglin Jing, Kehua Guo, James Z. Wang 0001, Hui Fang 0003 |
Pattern Recognit. | 8 |
| 2026 | ReE3D: Boosting Novel View Synthesis for Monocular Images Using Residual EncodersabstractIn recent years, novel view synthesis from a monocular image has become a research hot-spot that attracts significant attention. Some recent work identifies latent vectors for high-quality view generation via iterative optimisation, which is a time-consuming process. In contrast, some others utilise an encoder learning a mapping function to approximately estimate optimal latent codes, which significantly reduces its processing time but sacrifices reconstruction quality. Consequently, how to balance synthesis quality and its generation efficiency still remains challenging. In this paper, we propose a residual-based encoder to incorporate with a 3D Generative Adversarial Networks (GAN), named ReE3D, for novel view synthesis. It applies an iterative prediction of latent codes to ensure much higher quality of novel view synthesis with an insignificant increase of processing time when compared to existing encoder-based 3D GAN inversion methods. Additionally, we enforce a novel geometric loss constraint on the encoder to predict view-invariant latent codes, thus effectively mitigating the trade-off between geometric and texture quality in 3D GAN inversion. Extensive experimental results demonstrate that our extended encoder-based method has achieved best trade-off performance in terms of novel view synthesis quality and its execution time. Our method has gained comparable synthesis quality with exponentially decreased processing time when compared to iterative optimisation methods, while improved synthesis performance of encoder-based methods significantly. Kehua Guo, Tianyu Chen 0004, Bin Hu 0021, Zheng Wu 0004, Shaojun Guo, Hui Fang 0003 |
IEEE Trans. Multim. | 7 |
| 2025 | Recoverable Facial Identity Protection via Adaptive Makeup Transfer Adversarial AttacksabstractUnauthorised face recognition (FR) systems have posed significant threats to digital identity and privacy protection. To alleviate the risk of compromised identities, recent makeup transfer-based attack methods embed adversarial signals in order to confuse unauthorised FR systems. However, their major weakness is that they set up a fixed image unrelated to both the protected and the makeup reference images as the confusion identity, which in turn has a negative impact on both attack success rate and visual quality of transferred photos. In addition, the generated images cannot be recognised by authorised FR systems once attacks are triggered. To address these challenges, in this paper, we propose a Recoverable Makeup Transferred Generative Adversarial Network (RMT-GAN) which has the distinctive feature of improving its image-transfer quality by selecting a suitable transfer reference photo as the target identity. Moreover, our method offers a solution to recover the protected photos to their original counterparts that can be recognised by authorised systems. Experimental results demonstrate that our method provides significantly improved attack success rates while maintaining higher visual quality compared to state-of-the-art makeup transfer-based adversarial attack methods. Our code and supplementary materials are available on Github. Xiyao Liu 0001, Junxing Ma, Xinda Wang 0006, Qianyu Lin, Jian Zhang 0048, Gerald Schaefer, Cagatay Turkay, Hui Fang 0003 |
AAAI | 8 |
| 2025 | Enhancing robustness of backdoor attacks against backdoor defenses
Bin Hu 0021, Kehua Guo, Hui Fang 0003 |
Expert Syst. Appl. | 4 |
| 2025 | An adversarial contrastive learning based cross-modality zero-watermarking scheme for DIBR 3D video copyright protectionabstractCopyright protection of depth image-based rendering (DIBR) videos has raised significant concerns due to their increasing popularity. Zero-watermarking, emerging as a powerful tool to protect the copyright of DIBR 3D videos, mainly relies on traditional feature extraction methods, thus necessitating improvements in robustness against complex geometric attacks and its ability to strike a balance between robustness and distinguishability. This paper presents a novel zero-watermarking scheme based on cross-modality feature fusion within a contrastive learning framework. Our approach integrates complementary information from 2D frames and depth maps using a cross-modality attention feature fusion mechanism to obtain discriminative features. Moreover, our features achieve a better trade-off between robustness and distinguishability by leveraging a designed contrastive learning strategy with an adversarial distortion simulator. Experimental results demonstrate our remarkable performance by reducing the false negative rates to around 0.2% when the false positive rate is equal to 0.5%, which is superior to the state-of-the-art zero-watermarking methods. • Use contrastive learning to balance watermarking robustness and distinguishability. • Employ an adversarial distortion simulator to enhance robustness against various attacks. • Design cross-modality fusion mechanism to achieve better feature representation. Xiyao Liu 0001, Qingyu Dang, Xiaoheng Deng, Xunli Fan, Cundian Yang, Hui Fang 0003 |
Neurocomputing | 8 |
| 2025 | A memory-based conditional neural process for video instance segmentationabstractVideo instance segmentation (VIS) is an evolving research topic in computer vision that aims to simultaneously detect, segment, and track semantic objects across multiple video frames. However, existing VIS methods are typically unaware of the reliability of the training samples from insufficient and imbalanced datasets, leading to suboptimal performance. To address this challenge, we propose a memory-based conditional neural process (MemCNP) module to exploit the strengths of both memory networks and the CNP model which handles heterogeneous latent space distributions for reliable modelling with insufficient data. Our MemCNP utilises predicted uncertainty to regularise VIS predictions as well as to identify reliable samples for effective training. Notably, our MemCNP is model-agnostic and can thus be seamlessly integrated into various VIS models to improve their performance. Extensive experiments on the YouTube-VIS and OVIS datasets demonstrate the effectiveness of MemCNP regardless of the underlying model architecture. • A memory-based conditional neural process. • Reliability modelling for object detection. • Uncertainty-based dynamic training sample selection. • Contrastive instance tracking. Kunhao Yuan, Gerald Schaefer, Yukun Lai, Xiyao Liu 0001, Hui Fang 0003 |
Neurocomputing | 6 |
| 2025 | A dual-aligned knowledge self-distillation framework for visible-infrared cross-modal person re-identificationabstract• Dual alignment knowledge self-distillation to better capture modality-invariant/specific features for VI-ReID • Temperature-modulated alignment and confidence-based selective masking to enhance model reliability. • CutSwap augmentation to improve model robustness against intra-class variations and modality discrepancies. • State-of-the-art performance on SYSU-MM01 and RegDB benchmarks. Visible-infrared person re-identification (VI-ReID) significantly enhances identity retrieval across different illumination conditions by matching visible and infrared modalities. However, existing contrastive-learning-based approaches predominantly focus on cross-modal feature alignment, thus undermining model reliability in complex scenarios. To address this challenge, we introduce a Dual Alignment Knowledge Distillation (DAKD) framework that leverages comprehensive self-distillation at both instance and class levels. Our framework incorporates a temperature-modulated alignment strategy, capturing rich modality-invariant generalities as well as modality-specific discriminative details. Additionally, we propose a confidence-based selective masking mechanism that guides the distillation towards confident and informative teacher predictions. To further enhance robustness against modality discrepancies and intra-class variations, we develop a dedicated augmentation technique, CutSwap, which exchanges image channels to simulate realistic cross-modality variations. Extensive experiments on the benchmark SYSU-MM01 and RegDB datasets demonstrate superior performance compared to other state-of-the-art methods, achieving rank-1 accuracies of 76.31% and 94.83%, respectively and validating the efficacy of DAKD in maintaining robust cross-modal alignment while preserving essential identity-specific discriminative information. Siyuan Deng, Kunhao Yuan, Gerald Schaefer, Shihua Zhou, George Vogiatzis, Yifan Wang 0008, Hui Fang 0003 |
Knowl. Based Syst. | 7 |
| 2025 | A sequential mixing fusion network for enhanced feature representations in multimodal sentiment analysis
Qiang Zhang 0008, Jing Dong 0009, Hui Fang 0003, Gerald Schaefer, Rui Liu 0015 |
Knowl. Based Syst. | 4 |
| 2025 | Class activation map guided level sets for weakly supervised semantic segmentation
Yifan Wang 0008, Gerald Schaefer, Xiyao Liu 0001, Jing Dong 0009, Linglin Jing, Xianghua Xie, Hui Fang 0003 |
Pattern Recognit. | 8 |
| 2025 | HSE-GNN: A hierarchical skeleton embedded graph neural network for 3D human pose estimation
Jing Dong 0009, Hui Fang 0003, Rui Liu 0015, Yu Sui |
Pattern Recognit. Lett. | 3 |
| 2025 | Attack-Defending Contrastive Learning for Volumetric Medical Image Zero-WatermarkingabstractZero-watermarking is an emerging distortion-free copyright protection method for volumetric medical images. However, achieving both robustness against various malicious attacks and distinguishability between individual images remains challenging. In this article, we propose a novel attack-defending contrastive learning zero-watermarking (ADCL-ZW) scheme to tackle the above challenge using deep learning-based representations. In our approach, we design an attack-defending data enrichment mechanism to enhance the watermarking robustness by generating a large number of image samples under various watermarking attacks. Subsequently, features for both watermarking distinguishability and robustness are enhanced through application of a contrastive loss. In particular, we implement a dual-stream Siamese network architecture to effectively handle both signal attacks and geometric attacks in order to enhance the watermarking performance. Experimental results demonstrate that ADCL-ZW achieves stronger watermarking robustness and a better tradeoff between watermarking robustness and distinguishability compared with state-of-the art zero-watermarking methods. One of the highlighted metrics is that the false-negative rate of ADCL-ZW achieves 0.01 when a fixed false-positive rate is set to 1%, which is more than 13.3 times better than the benchmark methods. Xiyao Liu 0001, Cundian Yang, Hui Fang 0003, Gerald Schaefer, Jian Zhang 0048, Yuesheng Zhu, Shichao Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | CrossBind: Collaborative Cross-Modal Identification of Protein Nucleic-Acid-Binding ResiduesabstractAccurate identification of protein nucleic acid binding residues poses a significant challenge with important implications for various biological processes and drug design. Many typical computational methods for protein analysis rely on a single model that could ignore either the semantic context of the protein or the global 3D geometric information. Consequently, these approaches may result in incomplete or inaccurate protein analysis. To address the above issue, in this paper, we present CrossBind, a novel collaborative cross modal approach for identifying binding residues by exploiting both protein geometric structure and its sequence prior knowledge extracted from a large scale protein language model. Specifically, our multi modal approach leverages a contrastive learning technique and atom wise attention to capture the positional relationships between atoms and residues, thereby incorporating fine grained local geometric knowledge, for better binding residue prediction. Extensive experimental results demonstrate that our approach outperforms the next best state of the art methods, GraphSite and GraphBind, on DNA and RNA datasets by 10.8/17.3% in terms of the harmonic mean of precision and recall (F1 Score) and 11.9/24.8% in Matthews correlation coefficient (MCC), respectively. We release the code at https://github.com/BEAM-Labs/CrossBind. Linglin Jing, Yifan Wang 0008, Zhigang Ji, Hui Fang 0003, Zhen Li 0026 |
AAAI | 7 |
| 2024 | X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge TransferabstractThe field of 4D point cloud understanding is rapidly developing with the goal of analyzing dynamic 3D point cloud sequences. However, it remains a challenging task due to the sparsity and lack of texture in point clouds. Moreover, the irregularity of point cloud poses a difficulty in aligning temporal information within video sequences. To address these issues, we propose a novel cross-modal knowledge transfer framework, called X4D-SceneFormer. This framework enhances 4D-Scene understanding by transferring texture priors from RGB sequences using a Transformer architecture with temporal relationship mining. Specifically, the framework is designed with a dual-branch architecture, consisting of an 4D point cloud transformer and a Gradient-aware Image Transformer (GIT). The GIT combines visual texture and temporal correlation features to offer rich semantics and dynamics for better point cloud representation. During training, we employ multiple knowledge transfer techniques, including temporal consistency losses and masked self-attention, to strengthen the knowledge transfer between modalities. This leads to enhanced performance during inference using single-modal 4D point cloud inputs. Extensive experiments demonstrate the superior performance of our framework on various 4D point cloud video understanding tasks, including action recognition, action segmentation and semantic segmentation. The results achieve 1st places, i.e., 85.3% (+7.9%) accuracy and 47.3% (+5.0%) mIoU for 4D action segmentation and semantic segmentation, on the HOI4D challenge, outperforming previous state-of-the-art by a large margin. We release the code at https://github.com/jinglinglingling/X4D. Linglin Jing, Ying Xue 0003, Xu Yan 0005, Chaoda Zheng, Dong Wang 0028, Ruimao Zhang, Zhigang Wang 0002, Hui Fang 0003, Bin Zhao 0001, Zhen Li 0026 |
AAAI | 8 |
| 2024 | HPL-ESS: Hybrid Pseudo-Labeling for Unsupervised Event-based Semantic SegmentationabstractEvent-based semantic segmentation has gained popularity due to its capability to deal with scenarios under high-speed motion and extreme lighting conditions, which cannot be addressed by conventional RGB cameras. Since it is hard to annotate event data, previous approaches rely on event-to-image reconstruction to obtain pseudo labels for training. However, this will inevitably introduce noise, and learning from noisy pseudo labels, especially when generated from a single source, may reinforce the errors. This drawback is also called confirmation bias in pseudo-labeling. In this paper, we propose a novel hybrid pseudo-labeling framework for unsupervised event-based semantic segmentation, HPL-ESS, to alleviate the influence of noisy pseudo labels. Specifically, we first employ a plain unsupervised domain adaptation framework as our baseline, which can generate a set of pseudo labels through self-training. Then, we incorporate offline event-to-image re-construction into the framework, and obtain another set of pseudo labels by predicting segmentation maps on the re-constructed images. A noisy label learning strategy is designed to mix the two sets of pseudo labels and enhance the quality. Moreover, we propose a soft prototypical alignment (SPA) module to further improve the consistency of target domain features. Extensive experiments show that the proposed method outperforms existing state-of-the-art methods by a large margin on benchmarks (e.g., +5.88% accuracy, +10.32% mIoU on DSEC-Semantic dataset), and even surpasses several supervised methods. Linglin Jing, Zhigang Wang 0002, Xu Yan 0005, Dong Wang 0028, Gerald Schaefer, Hui Fang 0003, Bin Zhao 0001, Xuelong Li 0001 |
CVPR | 8 |
| 2024 | Towards Compact Reversible Image Representations for Neural Style Transfer
Xiyao Liu 0001, Siyu Yang 0005, Jian Zhang 0048, Gerald Schaefer, Jiya Li, Xunli Fan, Songtao Wu, Hui Fang 0003 |
ECCV (66) | 8 |
| 2024 | Multi-Strategy Adversarial Learning for Robust Face Forgery Detection Under Heterogeneous and Composite AttacksabstractFace forgery detection has recently progressed to address the threat from image synthesis technology, although robust face forgery detection under heterogeneous attacks remains challenging. When forgers leverage image post-processing techniques to manipulate forged photos, recent detection methods exhibit significant performance degradation. In this work, we propose a novel multi-strategy adversarial learning (MAL) method to extract salient features in order to achieve more reliable forgery detection under attacks. In particular, our MAL framework creates a large number of positive and negative sample pairs by designing a composite attack generation module with supervised contrastive training to ensure the attack robustness. In addition, we exploit two intuitive strategies, hard sample selection and region consistency, to enhance the contrastive losses for further strengthened feature reliability. Extensive experimental results demonstrate our proposed method to outperform recent state-of-the-art face forgery detection methods in terms of overall accuracy under various single and composite attacks. Xiyao Liu 0001, Fengkai Dong, Jian Zhang 0048, Gerald Schaefer, Hui Fang 0003 |
ICME | 8 |
| 2024 | Guided Diffusion-based Adversarial Purification Model with Denoised Prior ConstraintabstractAdversarial attack has posed a significant threat to modern deep learning based models. Recently, various adversarial defending algorithms are proposed to tackle the problem. Among them, diffusion-based adversarial purification approaches offer the most promising solutions. However, their effectiveness are limited due to the strong adversarial perturbations presented in attacked images. These adversarial signals hinder the introduction of guidance into diffusion models in order to improve the defence efficacy. In this paper, we propose a novel approach to embed reliable guidance into diffusion-based adversarial purification model to improve both its defence effectiveness and efficiency. In specific, we present a diffusion sampling guidance enhanced by a pretrained denoising network as a prior constraint to improve the adversarial defence performance. Experimental results convincingly demonstrate the superior performance of the proposed approach in terms of enhanced robustness to standard image classifiers when compared to state-of-the-art adversarial defence approaches. Xiyao Liu 0001, Ting Yang 0009, Hui Fang 0003 |
IJCNN | 5 |
| 2024 | Memory-facilitated Joint-space Shift Adaptation in Traffic ForecastingabstractTraffic forecasting, crucial for intelligent transport systems, faces significant challenges from distribution shifts due to the dynamic nature of traffic patterns. Although normalisation approaches have been proposed to address distribution shifts in other time-series forecasting tasks such as predicting electricity consumption load prediction or influenza-like illness patient number estimation, they fall short in handling the complex spatial and temporal shifts in traffic data. In this paper, we propose a novel memory-facilitated joint-space shift adaptation framework, ST-Align, to address this problem in traffic forecasting. ST-Align comprises two key components targeting the input and latent space, respectively: a memory-based data alignment module in the input space, and an end-to-end memory network structure dedicated to alignment within the latent space. This joint-space design enables our ST-Align framework to effectively capture and adapt to dynamic distribution shifts in both spatial and temporal dimensions, thus enhancing model performance. Extensive experiments on various real-world datasets and prediction backbones convincingly demonstrate the robustness and generalisability of our method. He Haitao, Gerald Schaefer, Zhigang Ji, Yifan Wang 0008, Hui Fang 0003 |
IJCNN | 6 |
| 2024 | A novel attention model across heterogeneous features for stuttering event detectionabstractStuttering is a prevalent speech disorder affecting millions worldwide. To provide an automatic and objective stuttering assessment tool, Stuttering Event Detection (SED) is under extensive investigation for advanced speech research and applications. Despite significant progress achieved by various machine learning and deep learning models, SED directly from speech signal is still challenging due to stuttering speech’s heterogeneous and overlapped nature. This paper presents a novel SED approach using multi-feature fusion and attention mechanisms. The model utilises multiple acoustic features extracted based on different pitch, time-domain, frequency domain, and automatic speech recognition feature to detect stuttering core behaviours more accurately and reliably. In addition, we exploit both spatial and temporal attention mechanisms as well as Bidirectional Long Short-Term Memory (BI-LSTM) modules to learn better representations to improve the SED performance. The experimental evaluation and analysis convincingly demonstrate that our proposed model surpasses the state-of-the-art models on two popular stuttering datasets, with 4% and 3% overall F1 scores, respectively. The superior results indicate the consistency of our proposed method, supported by both multi-feature and attention mechanisms in different stuttering events datasets. Abed Alkarim Banna, Hui Fang 0003, Eran A. Edirisinghe |
Expert Syst. Appl. | 2 |
| 2024 | Self-supervised memory learning for scene text image super-resolution
Kehua Guo, Xiangyuan Zhu, Gerald Schaefer, Rui Ding 0017, Hui Fang 0003 |
Expert Syst. Appl. | 5 |
| 2024 | A Memory-augmented Conditional Neural Process model for traffic predictionabstractThis paper presents the first neural process-based model for traffic prediction, the Memory-augmented Conditional Neural Process (MemCNP). Spatio-temporal traffic prediction involves predicting future traffic patterns based on historical traffic data and the road network structure. This problem remains a challenge due to the dynamic and heterogeneous nature of urban traffic. Existing models often struggle to capture these complexities, particularly in data-limited scenarios. To address these limitations, our model presents a novel framework for uncertainty estimation based on the conditional neural process, and further incorporates a memory network module designed to acquire a representative contextual reference, thereby improving model performance under complex data distributions. By integrating the conditional neural process and the memory network, MemCNP enables the learning of the most representative contexts through iterative updates, enhancing the model’s generalisability. This allows our model to be applicable beyond car traffic, effectively handling diverse real-world traffic scenarios, including urban non-motorised traffic such as cycling, which is essential for advancing more sustainable transportation systems. This is demonstrated by comprehensive experimental results on six benchmark datasets (PeMS04, PeMS07, PeMS08, NYCTaxi, CHIBike, and T-Drive) against existing state-of-the-art traffic prediction models, where MemCNP demonstrates superior performance. Additionally, through ablation and reliability studies, we provide a comprehensive analysis of the model’s effectiveness. • The first neural process-based model, MemCNP, for traffic prediction with limited data. • MemCNP introduces a novel framework for uncertainty estimation. • A novel memory network module acquires a representative contextual reference. • MemCNP is effective across diverse scenarios, including non-motorised traffic. He Haitao, Kunhao Yuan, Gerald Schaefer, Zhigang Ji, Hui Fang 0003 |
Knowl. Based Syst. | 6 |
| 2024 | Modal adaptive super-resolution for medical images via continual learning
Zheng Wu 0004, Feihong Zhu, Kehua Guo, Chao Liu 0058, Hui Fang 0003 |
Signal Process. | 6 |
| 2024 | Fast Aging-Aware Timing Analysis Framework With Temporal-Spatial Graph Neural NetworkabstractWith the downscaling of CMOS technology, device aging induced by hot carrier injection and bias temperature instability effects poses severe challenges to timing analysis of digital circuits. In this work, a fast aging-aware timing analysis framework based on temporal–spatial graph neural network (GNN) is proposed for the first time. The temporal–spatial GNN takes gated tanh unit (GTU) as the temporal network to extract devices’ degradation from dynamic biases, and takes inductive GraphSAGE as the spatial network to obtain whole graph information from circuit topology and output circuit aging delay. With comprehensive comparison among the network candidates, the combination of GTU and GraphSAGE presents the highest accuracy in predicting the standard cell aging delay. Owing to the superior features capture capability, this framework significantly improves the aging prediction efficiency under various operation conditions, especially facing the iterations of usage scenario, design version and process design kit. Compared with the conventional flow, the average acceleration ratio of our temporal–spatial network in predicting aging delay is more than 200 times. Furthermore, this framework is demonstrated with ADDER and FIFO circuits in timing analysis at the end of life. Thus, this work is helpful to the aging-aware circuit design in nano-scale technology. Jinfeng Ye, Pengpeng Ren, Yongkang Xue, Hui Fang 0003, Zhigang Ji |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | Watermarking in Secure Federated Learning: A Verification Framework Based on Client-Side BackdooringabstractFederated learning (FL) allows multiple participants to collaboratively build deep learning (DL) models without directly sharing data. Consequently, the issue of copyright protection in FL becomes important since unreliable participants may gain access to the jointly trained model. Application of homomorphic encryption (HE) in a secure FL framework prevents the central server from accessing plaintext models. Thus, it is no longer feasible to embed the watermark at the central server using existing watermarking schemes. In this article, we propose a novel client-side FL watermarking scheme to tackle the copyright protection issue in secure FL with HE. To the best of our knowledge, it is the first scheme to embed the watermark to models under a secure FL environment. We design a black-box watermarking scheme based on client-side backdooring to embed a pre-designed trigger set into an FL model by a gradient-enhanced embedding method. Additionally, we propose a trigger set construction mechanism to ensure that the watermark cannot be forged. Experimental results demonstrate that our proposed scheme delivers outstanding protection performance and robustness against various watermark removal attacks and ambiguity attack. Shuo Shao 0002, Yue Yang 0007, Xiyao Liu 0001, Ximeng Liu, Zhihua Xia, Gerald Schaefer, Hui Fang 0003 |
ACM Trans. Intell. Syst. Technol. | 8 |
| 2024 | A Multiscale Framework for Capturing Oscillation Dynamics of Autonomous Vehicles in Data-Driven Car-Following ModelsabstractRecent advancements in machine learning-based car-following models have shown promise in leveraging vehicle trajectory data to accurately reproduce real-world driving behaviour in simulations. However, existing data-driven car-following models only explicitly consider individual vehicle trajectories for model training, overlooking broader traffic phenomena. This limitation hinders their ability to accurately capture the oscillation dynamics of vehicle platoons, which are critical for simulating and evaluating mesoscopic and macroscopic traffic phenomena such as congestion propagation, stop-and-go, string stability and hysteresis. To fill this gap, our study introduces a hybrid physical model-driven and data-driven framework, Multiscale Car-Following (MultiscaleCF), aimed at explicitly capturing mesoscopic oscillation dynamics within data-driven car-following models. MultiscaleCF offers two methodological advancements in the development of machine learning-based car-following models: the recursive simulation of a platoon of vehicles to reduce compound error and mesoscopic feature engineering using domain-specific attributes. Evaluated using the OpenACC database, the MultiscaleCF framework exhibited a simultaneous improvement in both microscopic and mesoscopic traffic simulation patterns. It outperforms the baseline model in microscopic trajectory prediction accuracy by up to 21%. For oscillation dynamics, it outperforms the baseline model by 42%, 32%, 29% and 42% in duration, amplitude, intensity, and hysteresis magnitude, respectively. Rowan Davies, He Haitao, Hui Fang 0003 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Progressive Diversity Generation for Single Domain GeneralizationabstractSingle domain generalization (single-DG) is a realistic yet challenging domain generalization scenario where a model trained on a single domain generalization scenario where a model trained on a single domain generalizes well to multiple unseen domains. Unlike typical single-DG methods that are essentially supervised data augmentation and focus mainly on the novelty of images, we propose a simple adversarial augmentation method, termed Progressive Diversity Generation (PDG), to synthesize novel and diverse images in a fully unsupervised manner. Specifically, PDG minimizes the uncertainty coefficient to ensure that synthesized images are novel. By modeling conditional probabilities with an auxiliary network, we transfer the adversarial process from semantics to images, thus eliminating dependency on labels. To enhance diversity, we propose the$f$-diversity, a collection of correlation or similarity measures, to allow our model to generate potential images from diverse perspectives. The proposed architecture combines a multi-attribute generator with a progressive generation framework to improve model performance. PDG is the unsupervised and easy-to-implement method that solves single-DG with only synthesized (source) images. Extensive experiments on multiple single-DG benchmarks show that PDG achieves remarkable results and outperforms existing supervised and unsupervised methods by a large margin in single domain generalization. Source code and data are available:https://github.com/Ruiding1/PDG. Rui Ding 0017, Kehua Guo, Xiangyuan Zhu, Zheng Wu 0004, Hui Fang 0003 |
IEEE Trans. Multim. | 5 |
| 2023 | Gradient-Based Graph Attention for Scene Text Image Super-resolutionabstractScene text image super-resolution (STISR) in the wild has been shown to be beneficial to support improved vision-based text recognition from low-resolution imagery. An intuitive way to enhance STISR performance is to explore the well-structured and repetitive layout characteristics of text and exploit these as prior knowledge to guide model convergence. In this paper, we propose a novel gradient-based graph attention method to embed patch-wise text layout contexts into image feature representations for high-resolution text image reconstruction in an implicit and elegant manner. We introduce a non-local group-wise attention module to extract text features which are then enhanced by a cascaded channel attention module and a novel gradient-based graph attention module in order to obtain more effective representations by exploring correlations of regional and local patch-wise text layout properties. Extensive experiments on the benchmark TextZoom dataset convincingly demonstrate that our method supports excellent text recognition and outperforms the current state-of-the-art in STISR. The source code is available at https://github.com/xyzhu1/TSAN. Xiangyuan Zhu, Kehua Guo, Hui Fang 0003, Rui Ding 0017, Zheng Wu 0004, Gerald Schaefer |
AAAI | 3 |
| 2023 | A Novel Class Activation Map for Visual Explanations in Multi-Object ScenesabstractClass activation maps (CAMs) have emerged as a popular technique to improve model interpretability of deep learning-based models. While existing CAM methods are able to extract salient semantic regions to provide high-confidence pseudo-labels for downstream tasks such as semantic segmentation, they are less effective when dealing with multi-object scenes. In this paper, we design a multi-channel weight assignment scheme that learns from both positive and negative regions to yield an improved CAM model for images comprising multiple objects. We demonstrate the effectiveness of our proposed method on two new data sets, a cat-and-dog dataset and a PASCAL VOC 2012-based multi-object dataset, and show it to compare favourably with other state-of-the-art CAM methods, outperforming them in terms of both mIoU and inter-object activation ratio (IAR), a new evaluation measure proposed to evaluate CAM performance in multi-object scenes. Yifan Wang 0008, Siyuan Deng, Kunhao Yuan, Gerald Schaefer, Xiyao Liu 0001, Hui Fang 0003 |
ICIP | 6 |
| 2023 | A Multi-Modal Transformer Approach for Football Event ClassificationabstractVideo understanding has been enhanced by the use of multi-modal networks. However, recent multi-modal video analysis models have limited applicability to sports videos due to their specialised nature. This paper proposes a novel attention-based multi-modal neural network for sports event classification featuring a multi-stage fusion training strategy. The proposed multi-modal neural network integrates three modalities, including an image sequence modality, an audio modality and a newly proposed sports formation modality, to improve the sports video classification performance. Empirical results show that the proposed model outperforms the state-of-the-art transformer-based video method by 4.43% on top-1 accuracy on Soccernet-V2 dataset. Baihua Li, Hui Fang 0003, Qinggang Meng |
ICIP | 3 |
| 2023 | Robust Steganography without Embedding Based on Secure Container Synthesis and Iterative Message RecoveryabstractSynthesis-based steganography without embedding (SWE) methods transform secret messages to container images synthesised by generative networks, which eliminates distortions of container images and thus can fundamentally resist typical steganalysis tools. However, existing methods suffer from weak message recovery robustness, synthesis fidelity, and the risk of message leakage. To address these problems, we propose a novel robust steganography without embedding method in this paper. In particular, we design a secure weight modulation-based generator by introducing secure factors to hide secret messages in synthesised container images. In this manner, the synthesised results are modulated by secure factors and thus the secret messages are inaccessible when using fake factors, thus reducing the risk of message leakage. Furthermore, we design a difference predictor via the reconstruction of tampered container images together with an adversarial training strategy to iteratively update the estimation of hidden messages. This ensures robustness of recovering hidden messages, while degradation of synthesis fidelity is reduced since the generator is not included in the adversarial training. Extensive experimental results convincingly demonstrate that our proposed method is effective in avoiding message leakage and superior to other existing methods in terms of recovery robustness and synthesis fidelity. Ziping Ma 0002, Yuesheng Zhu, Guibo Luo, Xiyao Liu 0001, Gerald Schaefer, Hui Fang 0003 |
IJCAI | 6 |
| 2023 | Single Domain Generalization via Unsupervised Diversity ProbeabstractSingle domain generalization (SDG) is a realistic yet challenging domain generalization scenario that aims to generalize a model trained on a single domain to multiple unseen domains. Typical SDG methods are essentially supervised data augmentation strategies, which tend to enhance the novelty rather than the diversity of augmented samples. Insufficient diversity may jeopardize the model generalization ability. In this paper, we propose a novel adversarial method, termed Unsupervised Diversity Probe (UDP), to synthesize novel and diverse samples in fully unsupervised settings. More specifically, to ensure that samples are novel, we study SDG from an information-theoretic perspective that minimizes the uncertainty coefficients between synthesized and source samples. Considering that the variation in a single source domain is limited, we introduce a regularization imposed on the auxiliary module that synthesizes variable samples, incorporated with uncertainty coefficients in an adversarial manner to complement the diversity. Subsequently, an available region is utilized to guarantee the samples' safety. For the network architecture, we design a simple probe module that can synthesize samples in several different aspects. UDP is an unsupervised and easy-to-implement method that solves SDG using only synthetic (source) samples, thus reducing the dependence on task models. Extensive experiments on three benchmark datasets show that UDP achieves remarkable results and outperforms existing supervised and unsupervised methods by a large margin in single domain generalization. Kehua Guo, Rui Ding 0017, Tian Qiu 0002, Xiangyuan Zhu, Zheng Wu 0004, Hui Fang 0003 |
ACM Multimedia | 7 |
| 2023 | A novel temporal generative adversarial network for electrocardiography anomaly detection
Jing Qin 0007, Fujie Gao, David Wong 0001, Zhibin Zhao 0002, Samuel D. Relton, Hui Fang 0003 |
Artif. Intell. Medicine | 7 |
| 2023 | Semantics-guided generative diffusion model with a 3DMM model condition for face swappingabstractAbstract Face swapping is a technique that replaces a face in a target media with another face of a different identity from a source face image. Currently, research on the effective utilisation of prior knowledge and semantic guidance for photo‐realistic face swapping remains limited, despite the impressive synthesis quality achieved by recent generative models. In this paper, we propose a novel conditional Denoising Diffusion Probabilistic Model (DDPM) enforced by a two‐level face prior guidance. Specifically, it includes (i) an image‐level condition generated by a 3D Morphable Model (3DMM), and (ii) a high‐semantic level guidance driven by information extracted from several pre‐trained attribute classifiers, for high‐quality face image synthesis. Although swapped face image from 3DMM does not achieve photo‐realistic quality on its own, it provides a strong image‐level prior, in parallel with high‐level face semantics, to guide the DDPM for high fidelity image generation. The experimental results demonstrate that our method outperforms state‐of‐the‐art face swapping methods on benchmark datasets in terms of its synthesis quality, and capability to preserve the target face attributes and swap the source face identity. Xiyao Liu 0001, Ting Yang 0009, Jian Zhang 0048, Victoria Wang, Hui Fang 0003 |
Comput. Graph. Forum | 7 |
| 2023 | Stereoscopic image super-resolution with interactive memory learning
Xiangyuan Zhu, Kehua Guo, Tian Qiu 0002, Hui Fang 0003, Zheng Wu 0004, Xuyang Tan, Chao Liu 0058 |
Expert Syst. Appl. | 4 |
| 2023 | A multi-strategy contrastive learning framework for weakly supervised semantic segmentationabstractWeakly supervised semantic segmentation (WSSS) has gained significant popularity as it relies only on weak labels such as image level annotations rather than the pixel level annotations required by supervised semantic segmentation (SSS) methods. Despite drastically reduced annotation costs, typical feature representations learned from WSSS are only representative of some salient parts of objects and less reliable compared to SSS due to the weak guidance during training. In this paper, we propose a novel Multi-Strategy Contrastive Learning (MuSCLe) framework to obtain enhanced feature representations and improve WSSS performance by exploiting similarity and dissimilarity of contrastive sample pairs at image, region, pixel and object boundary levels. Extensive experiments demonstrate the effectiveness of our method and show that MuSCLe outperforms current state-of-the-art methods on the widely used PASCAL VOC 2012 dataset. Kunhao Yuan, Gerald Schaefer, Yukun Lai, Yifan Wang 0008, Xiyao Liu 0001, Hui Fang 0003 |
Pattern Recognit. | 7 |
| 2023 | Development and Evaluation of Two Approaches of Visual Sensitivity Analysis to Support Epidemiological ModelingabstractComputational modeling is a commonly used technology in many scientific disciplines and has played a noticeable role in combating the COVID-19 pandemic. Modeling scientists conduct sensitivity analysis frequently to observe and monitor the behavior of a model during its development and deployment. The traditional algorithmic ranking of sensitivity of different parameters usually does not provide modeling scientists with sufficient information to understand the interactions between different parameters and model outputs, while modeling scientists need to observe a large number of model runs in order to gain actionable information for parameter optimization. To address the above challenge, we developed and compared two visual analytics approaches, namely: algorithm-centric and visualization-assisted, and visualization-centric and algorithm-assisted. We evaluated the two approaches based on a structured analysis of different tasks in visual sensitivity analysis as well as the feedback of domain experts. While the work was carried out in the context of epidemiological modeling, the two approaches developed in this work are directly applicable to a variety of modeling processes featuring time series outputs, and can be extended to work with models with other types of outputs. Erik Rydow, Rita Borgo, Hui Fang 0003, Thomas Torsney-Weir, Ben Swallow, Thibaud Porphyre, Cagatay Turkay, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2022 | Towards a more efficient few-shot learning-based human gesture recognition via dynamic vision sensors
Linglin Jing, Yifan Wang 0008, Tailin Chen, Shirin Dora, Zhigang Ji, Hui Fang 0003 |
BMVC | 6 |
| 2022 | Image Disentanglement Autoencoder for Steganography without EmbeddingabstractConventional steganography approaches embed a secret message into a carrier for concealed communication but are prone to attack by recent advanced steganalysis tools. In this paper, we propose Image DisEntanglement Autoencoder for Steganography (IDEAS) as a novel steganography without embedding (SWE) technique. Instead of directly embedding the secret message into a carrier image, our approach hides it by transforming it into a synthesised image, and is thus fundamentally immune to typical steganalysis attacks. By disentangling an image into two representations for structure and texture, we exploit the stability of structure representation to improve secret message extraction while increasing synthesis diversity via randomising texture representations to enhance steganography security. In addition, we design an adaptive mapping mechanism to further enhance the diversity of synthesised images when ensuring different required extraction levels. Experimental results convincingly demonstrate IDEAS to achieve superior performance in terms of enhanced security, reliable secret message extraction and flexible adaptation for different extraction levels, compared to state-of-the-art SWE methods. Xiyao Liu 0001, Ziping Ma 0002, Junxing Ma, Jian Zhang 0048, Gerald Schaefer, Hui Fang 0003 |
CVPR | 6 |
| 2022 | LightLog: A lightweight temporal convolutional network for log anomaly detection on the edge
Jiyu Tian, Hui Fang 0003, Liming Chen 0001, Jing Qin 0007 |
Comput. Networks | 3 |
| 2022 | Multiple-feature-based zero-watermarking for robust and discriminative copyright protection of DIBR 3D videosabstractZero-watermarking is a key technique for achieving lossless and flexible copyright protection of depth image-based rendering (DIBR) videos. Existing approaches extract features of both 2D frames and depth maps via a single mechanism to protect them simultaneously. However, it is difficult for these schemes to fully satisfy the copyright protection requirements of the two components, including the remarkable discriminative capability of 3D videos and robustness against various attacks. Hence, in this paper, we propose a novel multiple-feature-based zero-watermarking scheme to protect the copyright of DIBR 3D videos. To the best of our knowledge, this is the first scheme that integrates multiple features to improve both the discriminative capability and robustness against various attacks. Specifically, dual-tree complex wavelet transform and discrete cosine transform features enhance the robustness against DIBR conversion and noise addition, respectively, while ring-partition statistical residual features ensure robustness against geometric attacks and provide sufficient discriminative capacity. In addition, we use a logistic-logistic chaotic system to encrypt these multiple features for enhanced security and design an attention-based fusion approach to offer an optimal copyright protection solution. Extensive experimental results demonstrate that our proposed scheme has stronger robustness and discriminative capacity compared to state-of-the-art zero-watermarking methods. Xiyao Liu 0001, Yayun Zhang, Yuying Sun, Gerald Schaefer, Hui Fang 0003 |
Inf. Sci. | 8 |
| 2022 | Imitation learning based decision-making for autonomous vehicle control at traffic roundaboutsabstractAbstract The essential of developing an advanced driving assistance system is to learn human-like decisions to enhance driving safety. When controlling a vehicle, joining roundabouts smoothly and timely is a challenging task even for human drivers. In this paper, we propose a novel imitation learning based decision making framework to provide recommendations to join roundabouts. Our proposed approach takes observations from a monocular camera mounted on vehicle as input and use deep policy networks to provide decisions when is the best timing to enter a roundabout. The domain expert guided learning framework can not only improve the decision-making but also speed up the convergence of the deep policy networks. We evaluate the proposed framework by comparing with state-of-the-art supervised learning methods, including conventional supervised learning methods, such as SVM and kNN, and deep learning based methods. The experimental results demonstrate that the imitation learning-based decision making framework, which ourperforms supervised learning methods, can be applied in driving assistance system to facilitate better decision-making when approaching roundabouts. Weichao Wang, Shiran Lin, Hui Fang 0003, Qinggang Meng |
Multim. Tools Appl. | 4 |
| 2022 | Hiding multiple images into a single image via joint compressive autoencodersabstractInterest in image hiding has been continually growing. Recently, deep learning-based image hiding approaches improve the hidden capacity significantly. However, the major challenges of the existing methods are that they are difficult to balance between the errors of the modified cover image and those of the recovered secret image. To solve this problem, in this paper, we develop an image hiding algorithm based on a joint compressive autoencoder framework. Further, we propose a novel strategy to enlarge the hidden capacity, i.e., hiding multi-images in one container image. Specifically, our approach provides an extremely high image hidden capacity coupled with small reconstruction errors of the secret image. More importantly, we tackle the trade-off problem of earlier approaches by mapping the image representations in the latent spaces of the joint compressive autoencoder models, leading to both high visual quality of the container image and low reconstruction error the secret image. In an extensive set of experiments, we confirm our proposed approach to outperform several state-of-the-art image hiding methods, yielding high imperceptibility and steganalysis resistance of the container images with high recovery quality of the secret images, while improving the image hidden capacity significantly (four times higher than full-image hiding capacity). Xiyao Liu 0001, Ziping Ma 0002, Fangfang Li 0004, Gerald Schaefer, Hui Fang 0003 |
Pattern Recognit. | 7 |
| 2022 | Lightweight Image Super-Resolution With Expectation-Maximization Attention MechanismabstractIn recent years, with the rapid development of deep learning, super-resolution methods based on convolutional neural networks (CNNs) have made great progress. However, the parameters and the required consumption of computing resources of these methods are also increasing to the point that such methods are difficult to implement on devices with low computing power. To address this issue, we propose a lightweight single image super-resolution network with an expectation-maximization attention mechanism (EMASRN) for better balancing performance and applicability. Specifically, a progressive multi-scale feature extraction block (PMSFE) is proposed to extract feature maps of different sizes. Furthermore, we propose an HR-size expectation-maximization attention block (HREMAB) that directly captures the long-range dependencies of HR-size feature maps. We also utilize a feedback network to feed the high-level features of each generation into the next generation’s shallow network. Compared with the existing lightweight single image super-resolution (SISR) methods, our EMASRN reduces the number of parameters by almost one-third. The experimental results demonstrate the superiority of our EMASRN over state-of-the-art lightweight SISR methods in terms of both quantitative metrics and visual quality. The source code can be downloaded athttps://github.com/xyzhu1/EMASRN. Xiangyuan Zhu, Kehua Guo, Bin Hu 0021, Min Hu 0007, Hui Fang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Cross View Capture for Stereo Image Super-ResolutionabstractStereo image super-resolution exploits additional features from cross view image pairs for high resolution (HR) image reconstruction. Recently, several new methods have been proposed to investigate cross view features along epipolar lines to enhance the visual perception of recovered HR images. Despite the impressive performance of these methods, global contextual features from cross view images are left unexplored. In this paper, we propose a cross view capture network (CVCnet) for stereo image super-resolution by using both global contextual and local features extracted from both views. Specifically, we design a cross view block to capture diverse feature embeddings from the views in stereo vision. In addition, a cascaded spatial perception module is proposed to redistribute each location in feature maps according to the weight it occupies to make the extraction of features more effective. Extensive experiments demonstrate that our proposed CVCnet outperforms the state-of-the-art image super-resolution methods to achieve the best performance for stereo image super-resolution tasks. The source code is available at https://github.com/xyzhu1/CVCnet. Xiangyuan Zhu, Kehua Guo, Hui Fang 0003, Bin Hu 0021 |
IEEE Trans. Multim. | 3 |
| 2021 | Discriminative and Geometrically Robust Zero-Watermarking Scheme for Protecting DIBR 3D VideosabstractCopyright protection of depth image-based rendering (DIBR) 3D videos is crucial due to the popularity of these videos. Despite the success of recent watermarking schemes, it is still challenging to ensure the robustness against strong geometric attacks when both lossless quality and distinguishability of protected videos are required. In this paper, we pro-pose a novel zero-watermarking scheme to improve the performance under strong geometric attacks when satisfying the other two requirements. In our scheme, CT-SVD-based features are extracted to ensure both distinguishability and robustness against signal processing and DIBR conversion at-tacks, while a SIFT-based rectication mechanism is designed to resist geometric attacks. Further, an attention-based fusion strategy is proposed to complement the robustness of rectied and unrectied CT-SVD features. Experimental results demonstrate that our scheme outperforms the existing zero-watermarking schemes in terms of distinguishability and robustness against strong geometric attacks such as rotation, cyclic translation and shearing. Xiyao Liu 0001, Yayun Zhang, Sibo Du, Jian Zhang 0048, Hui Fang 0003 |
ICME | 6 |
| 2021 | Current Advances on Deep Learning-based Human Action Recognition from Videos: a SurveyabstractHuman action recognition (HAR) from RGB videos is essential and challenging in the computer vision field due to its wide range of real-world applications in fields of human behaviour analysis, human-computer interactions, robotics and surveillance etc. Since the breakthrough and fast development of deep learning technology, the performance of HAR based on deep neural networks has been significantly improved in this decade. In this survey, we discuss the growing use of deep learning for HAR, such as representative two-stream and 3D CNNs, and particularly highlight most recent success achieved by using attention and transformers. We will provide our perspective on the new trend of designing innovative deep learning methods. In addition, we also present popular HAR datasets developed in recent years and benchmark accuracy achieved by current advancement in deep learning. This draws research attention to the challenges of HAR by identifying performance gaps when applying the deep learning methods on large HAR datasets. Further, this survey sheds light on the development of new methods and facilitates qualitative comparison with state of the art. Baihua Li, Hui Fang 0003, Qinggang Meng |
ICMLA | 3 |
| 2021 | A Novel Method for Network Traffic Prediction Using Residual Mogrifier GRUabstractNetwork traffic prediction is essential for network management and resource scheduling within Web information systems. However, existing prediction methods have difficulty fitting mutation values in traffic time-series data and are still inadequate in terms of precision. Here we describe a method for prediction using multimodal web traffic data. The method creates multi-dimensional time series on request traffic, response traffic, and abnormal code traffic, and uses the rich information contained in the different sequences in the preceding time window to make inferences about the traffic scale in subsequent time windows. In addition, we propose an improved algorithm based on the Gated Recurrent Unit (GRU) to reduce the prediction error. The algorithm introduces the residual structure into a stacked multi-layer recurrent network structure and uses the Mogrifier structure to interact the information before it is fed to the gating unit. The experimental results show that the improved method leads to a further reduction in the error between the predicted and true values, providing high usability in the field of network traffic prediction. Jinyu Tian 0004, Jing Qin 0007, Liming Chen 0001, Hui Fang 0003 |
SERVICES | 4 |
| 2021 | A Novel Information Hiding Method for H.266/VVC Based on Selections of Luminance Transform and Chrominance Prediction ModesabstractThis paper proposes a novel information hiding method designed for H.266/Versatile Video Coding (VVC) compressed video streams. In this work, we explore two exclusive tools in H.266/VVC standard, named Multiple Transform Selection (MTS) and Cross-component linear model (CCLM), to hide information. These two tools are utilized to preserve high video reconstruction quality and compression efficiency as well as enhance hidden capacity. In specific, MTS is for hiding information into luminance blocks by modifying the selections of transforms. Comparing with other tools, MTS has less significant impact on compression quality and efficiency. In addition, CCLM is further used to hide information into chrominance blocks to further enlarge the hidden capacity with little impact on the other two metrics. To our best knowledge, it is the first information hiding method exclusively designed for H.266/VVC. Experimental results show that our proposed information hiding method ensures high hidden capacity, remarkable video reconstruction quality and insignificant impact on compression efficiency, which achieves better overall performances comparing to existing methods for compressed video. Xiyao Liu 0001, Kaiyue Shi, Aihua Li, Hao Zhang 0032, Hui Fang 0003 |
SMC | 6 |
| 2021 | Secure Federated Learning Model Verification: A Client-side Backdoor Triggered Watermarking SchemeabstractFederated learning (FL) has become an emerging distributed framework to build deep learning models with collaborative efforts from multiple participants. Consequently, copyright protection of FL deep model is urgently required because too many participants have access to the joint-trained model. Recently, Secure FL framework is developed to address data leakage issue when central node is not fully trustable. This encryption process has made existing DL model watermarking schemes impossible to embed watermark at the central node. In this paper, we propose a novel client-side Federated Learning watermarking method to tackle the model verification issue under the Secure FL framework. In specific, we design a backdoor-based watermarking scheme to allow model owners to embed their pre-designed noise patterns into the FL deep model. Thus, our method provides reliable copyright protection while ensuring the data privacy because the central node has no access to the encrypted gradient information. The experimental results have demonstrated the efficiency of our method in terms of both FL model performance and watermarking robustness. Xiyao Liu 0001, Shuo Shao 0002, Yue Yang 0007, Kangming Wu, Hui Fang 0003 |
SMC | 6 |
| 2021 | Robust and discriminative zero-watermark scheme based on invariant features and similarity-based retrieval to protect large-scale DIBR 3D videos
Xiyao Liu 0001, Yifan Wang 0008, Ziqiang Sun, Lei Wang 0017, Rongchang Zhao, Yuesheng Zhu, Beiji Zou 0001, Hui Fang 0003 |
Inf. Sci. | 9 |
| 2021 | A novel zero-watermarking scheme with enhanced distinguishability and robustness for volumetric medical imaging
Xiyao Liu 0001, Yuying Sun, Cundian Yang, Yayun Zhang, Lei Wang 0017, Yan Chen 0012, Hui Fang 0003 |
Signal Process. Image Commun. | 8 |
| 2020 | Joint compressive autoencoders for full-image-to-image hidingabstractImage hiding has received significant attention due to the need of enhanced multimedia services such as multimedia security and meta-information embedding for multimedia augmentation. Recently, deep learning-based methods have been introduced that are capable of significantly increasing the hidden capacity and supporting full-size image hiding. However, these methods suffer from the necessity to balance the errors of the modified cover image and the recovered hidden image. In this paper, we propose a novel joint compressive autoencoder (J-CAE) framework to design an image hiding algorithm that achieves full-size image hidden capacity with small reconstruction errors of the hidden image. More importantly, our approach addresses the trade-off problem of previous deep learning-based methods by mapping the image representations in the latent spaces of the joint CAE models. Thus, both visual quality of the container image and recovery quality of the hidden image can be simultaneously improved. Extensive experimental results demonstrate that our proposed method outperforms several state-of-the-art deep learning-based image hiding techniques in terms of imperceptibility and recovery quality of the hidden images while maintaining full-size image hidden capacity. Xiyao Liu 0001, Ziping Ma 0002, Xingbei Guo, Jialu Hou, Lei Wang 0017, Jian Zhang 0048, Gerald Schaefer, Hui Fang 0003 |
ICPR | 8 |
| 2020 | Camouflage Generative Adversarial Network: Coverless Full-image-to-image HidingabstractImage hiding, one of the most important data hiding techniques, is widely used to enhance cybersecurity when transmitting multimedia data. In recent years, deep learning-based image hiding algorithms have been designed to improve the embedding capacity whilst maintaining sufficient imperceptibility to malicious eavesdroppers. These methods can hide a full-size secret image into a cover image, thus allowing full-image-to-image hiding. However, these methods suffer from a trade-off challenge to balance the possibility of detection from the container image against the recovery quality of secret image. In this paper, we propose Camouflage Generative Adversarial Network (Cam-GAN), a novel two-stage coverless full-image-to-image hiding method named, to tackle this problem. Our method offers a hiding solution through image synthesis to avoid using a modified cover image as the image hiding container and thus enhancing both image hiding imperceptibility and recovery quality of secret images. Our experimental results demonstrate that Cam-GAN outperforms state-of-the-art full-image-to-image hiding algorithms on both aspects. Xiyao Liu 0001, Ziping Ma 0002, Xingbei Guo, Jialu Hou, Gerald Schaefer, Lei Wang 0017, Victoria Wang, Hui Fang 0003 |
SMC | 8 |
| 2020 | Colour Quantisation using Human Mental Search and Local RefinementabstractColour quantisation is a common image processing technique to reduce the number of distinct colours in an image which are then represented by a colour palette. Selection of appropriate entries in this palette is challenging since the quality of the quantised image is directly dictated by the palette colours. In this paper, we propose a novel colour quantisation algorithm based on the human mental search (HMS) algorithm and subsequent refinement of the colour palette using k-means. HMS is a recent population-based metaheuristic algorithm that has been shown to yield good performance on a variety of optimisation problems. In the first stage, we use HMS to find a high-quality initial colour palette. In the second stage, this palette is refined using k-means to converge towards a local optimum and thus to further improve the quality of the quantised image. We evaluate our algorithm on a set of benchmark images and compare it to several conventional and soft computing-based colour quantisation algorithms to demonstrate excellent image quality, outperforming the other methods. Seyed Jalaleddin Mousavirad, Gerald Schaefer, M. Emre Celebi 0001, Hui Fang 0003, Xiyao Liu 0001 |
SMC | 4 |
| 2020 | Micro-expression Video Clip Synthesis Method based on Spatial-temporal Statistical Model and Motion Intensity Evaluation FunctionabstractMicro-expression (ME) recognition is an effective method to detect lies and other subtle human emotions. Machine learning-based and deep learning-based models have achieved remarkable results recently. However, these models are vulnerable to overfitting issue due to the scarcity of ME video clips. These videos are much harder to collect and annotate than normal expression video clips, thus limiting the recognition performance improvement. To address this issue, we propose a micro-expression video clip synthesis method based on spatial-temporal statistical and motion intensity evaluation in this paper. In our proposed scheme, we establish a micro-expression spatial and temporal statistical model (MSTSM) by analyzing the dynamic characteristics of micro-expressions and deploy this model to provide the rules for micro-expressions video synthesis. In addition, we design a motion intensity evaluation function (MIEF) to ensure that the intensity of facial expression in the synthesized video clips is consistent with those in real -ME. Finally, facial video clips with MEs of new subjects can be generated by deploying the MIEF together with the widely-used 3D facial morphable model and the rules provided by the MSTSM. The experimental results have demonstrated that the accuracy of micro-expression recognition can be effectively improved by adding the synthesized video clips generated by our proposed method. Lei Wang 0017, Jialu Hou, Xingbei Guo, Ziping Ma 0002, Xiyao Liu 0001, Hui Fang 0003 |
SMC | 6 |
| 2020 | Supervised classification of bradykinesia in Parkinson's disease from smartphone videos
Samuel D. Relton, Hui Fang 0003, Jane E. Alty, Rami Qahwaji, Christopher D. Graham, David Wong 0001 |
Artif. Intell. Medicine | 3 |
| 2019 | Supervised Classification of Bradykinesia for Parkinson's Disease Diagnosis from Smartphone VideosabstractSlowness of movement, known as bradykinesia, is an important early symptom of Parkinson's disease. This symptom is currently assessed subjectively by clinical experts. However, expert assessment has been shown to be subject to inter-rater variability. We propose a low-cost, contactless system using smarthphone videos to automatically determine the presence of bradykinesia. Using 70 videos recorded in a pilot study, we predicted the presence of bradykinesia with an estimated test accuracy of 0.79 and the presence of Parkinson's disease with estimated test accuracy 0.63. Even on a small set of pilot data this accuracy is comparable to that recorded by blinded human experts. David Wong 0001, Samuel D. Relton, Hui Fang 0003, Rami Qhawaji, Christopher D. Graham, Jane E. Alty |
CBMS | 3 |
| 2017 | Fast and reliable human action recognition in video sequences by sequential analysisabstractHuman action recognition from video sequences is a challenging topic in computer vision research. In recent years, many studies have explored the use of deep learning representations to consistently improve the analysis accuracy. Meanwhile, designing a fast and reliable framework is becoming increasingly important given the exponential growth of video data collected for many purposes (e.g. public security, entertainment, and early medical diagnosis etc.). In order to design a more efficient automatic human action annotation method, the sequential probability ratio test, one of the classical statistical sampling scheme, is adapted to solve a multi-classes hypothesis test problem in our work. With the proposed algorithm, the computational cost is reduced significantly without sacrificing the performance of the underlying system. The experimental results based on the UCF101 data set demonstrated the efficiency of the framework compared to the fixed sampling scheme. Hui Fang 0003, Jeyan Thiyagalingam, Nik Bessis, Eran A. Edirisinghe |
ICIP | 1 |
| 2017 | Categorical Colormap Optimization with Visualization Case StudiesabstractMapping a set of categorical values to different colors is an elementary technique in data visualization. Users of visualization software routinely rely on the default colormaps provided by a system, or colormaps suggested by software such as ColorBrewer. In practice, users often have to select a set of colors in a semantically meaningful way (e.g., based on conventions, color metaphors, and logological associations), and consequently would like to ensure their perceptual differentiation is optimized. In this paper, we present an algorithmic approach for maximizing the perceptual distances among a set of given colors. We address two technical problems in optimization, i.e., (i) the phenomena of local maxima that halt the optimization too soon, and (ii) the arbitrary reassignment of colors that leads to the loss of the original semantic association. We paid particular attention to different types of constraints that users may wish to impose during the optimization process. To demonstrate the effectiveness of this work, we tested this technique in two case studies. To reach out to a wider range of users, we also developed a web application called Colourmap Hospital. Hui Fang 0003, Simon J. Walton, Emily Delahaye, Dmitry A. Storchak, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2014 | Visual MultiplexingabstractAbstract The majority of display devices used in visualization are 2D displays. Inevitably, it is often necessary to overlay one piece of visual information on top of another, especially in applications such as multi‐field visualization and geo‐spatial information visualization. In this paper, we present a conceptual framework for studying the mechanisms for overlaying multiple pieces of visual information while allowing users to recover occluded information. We adopt the term ‘multiplexing’ from tele‐ and data communication to encompass all such overlapping mechanisms. We establish 10 categories of visual multiplexing mechanisms. We draw support evidence from both perception literature and existing works in visualization to support this conceptual framework. We examine the relationships between multiplexing and information theoretic measures. This new conceptual categorization provides the much‐needed theory of visualization with an integral component. Min Chen 0001, Simon J. Walton, Kai Berger, Jeyan Thiyagalingam, Brian Duffy, Hui Fang 0003, Cameron Holloway, Anne E. Trefethen |
Comput. Graph. Forum | 6 |
| 2014 | Facial expression recognition in dynamic sequences: An integrated approach
Hui Fang 0003, Neil Mac Parthaláin, Andrew J. Aubrey, Gary K. L. Tam, Rita Borgo, Paul L. Rosin, Phil W. Grant, David Marshall 0001, Min Chen 0001 |
Pattern Recognit. | 1 |
| 2013 | Recognizing Conversational Interaction Based on 3D Human Pose
Jingjing Deng 0001, Xianghua Xie, Ben Daubney, Hui Fang 0003, Phil W. Grant |
ACIVS | 4 |
| 2013 | From clamped local shape models to global shape modelabstractFacial fiducial point localization is a crucial step for most facial analysis applications, e.g., face recognition, expression recognition and facial aging simulation. Although state-of-art methods have the ability to provide good salient point location on frontal faces, finding a global solution under large variations caused by off-plane rotations and exaggerated expression changes is still a challenge. In this paper, we present a system with a two-level shape model to facilitate accurate facial fiducial point localization. In the first level, two local component models interact with each other in order to offer novel shape constraints. At the same time, the clamped local shape model provides constrained non-linear shape initialization for better convergence performance of the shape model as a whole. The experimental results confirm that the proposed method is capable of dealing with the face alignment under large shape variations. Hui Fang 0003, Jingjing Deng 0001, Xianghua Xie, Phil W. Grant |
ICIP | 1 |
| 2013 | Visualizing Natural Image StatisticsabstractNatural image statistics is an important area of research in cognitive sciences and computer vision. Visualization of statistical results can help identify clusters and anomalies as well as analyze deviation, distribution, and correlation. Furthermore, they can provide visual abstractions and symbolism for categorized data. In this paper, we begin our study of visualization of image statistics by considering visual representations of power spectra, which are commonly used to visualize different categories of images. We show that they convey a limited amount of statistical information about image categories and their support for analytical tasks is ineffective. We then introduce several new visual representations, which convey different or more information about image statistics. We apply ANOVA to the image statistics to help select statistically more meaningful measurements in our design process. A task-based user evaluation was carried out to compare the new visual representations with the conventional power spectra plots. Based on the results of the evaluation, we made further improvement of visualizations by introducing composite visual representations of image statistics. Hui Fang 0003, Gary K. L. Tam, Rita Borgo, Andrew J. Aubrey, Phil W. Grant, Paul L. Rosin, Christian Wallraven, Douglas W. Cunningham, David Marshall 0001, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2011 | Visualization of Time-Series Data in Parameter Space for Understanding Facial DynamicsabstractAbstract Over the past decade, computer scientists and psychologists have made great efforts to collect and analyze facial dynamics data that exhibit different expressions and emotions. Such data is commonly captured as videos and are transformed into feature‐based time‐series prior to any analysis. However, the analytical tasks, such as expression classification, have been hindered by the lack of understanding of the complex data space and the associated algorithm space. Conventional graph‐based time‐series visualization is also found inadequate to support such tasks. In this work, we adopt a visual analytics approach by visualizing the correlation between the algorithm space and our goal – classifying facial dynamics. We transform multiple feature‐based time‐series for each expression in measurement space to a multi‐dimensional representation in parameter space. This enables us to utilize parallel coordinates visualization to gain an understanding of the algorithm space, providing a fast and cost‐effective means to support the design of analytical algorithms. Gary K. L. Tam, Hui Fang 0003, Andrew J. Aubrey, Phil W. Grant, Paul L. Rosin, David Marshall 0001, Min Chen 0001 |
Comput. Graph. Forum | 2 |
| 2010 | Discriminant Feature Manifold for Facial Aging EstimationabstractComputerised facial aging estimation, which has the potential for many applications in human-computer interactions, has been investigated by many computer vision researchers in recent years. In this paper, a feature-based discriminant subspace is proposed to extract more discriminating and robust representations for aging estimation. After aligning all the faces by a piece-wise affine transform, orthogonal locality preserving projection (OLPP) is employed to project local binary patterns (LBP) from the faces into an age-discriminant subspace. The feature extracted from this manifold is more distinctive for age estimation compared with the features using in the state-of-the-art methods. Based on the public database FG-NET, the performance of the proposed feature is evaluated by using two different regression techniques, quadratic function and neural-network regression. The proposed feature subspace achieves the best performance based on both types of regression. Hui Fang 0003, Phil W. Grant, Min Chen 0001 |
ICPR | 1 |
| 2010 | An Evaluation of Video-to-Video Face VerificationabstractPerson recognition using facial features, e.g., mug-shot images, has long been used in identity documents. However, due to the widespread use of web-cams and mobile devices embedded with a camera, it is now possible to realize facial video recognition, rather than resorting to just still images. In fact, facial video recognition offers many advantages over still image recognition; these include the potential of boosting the system accuracy and deterring spoof attacks. This paper presents an evaluation of person identity verification using facial video data, organized in conjunction with the International Conference on Biometrics (ICB 2009). It involves 18 systems submitted by seven academic institutes. These systems provide for a diverse set of assumptions, including feature representation and preprocessing variations, allowing us to assess the effect of adverse conditions, usage of quality information, query selection, and template construction for video-to-video face authentication. Norman Poh, Chi-Ho Chan, Josef Kittler, Sébastien Marcel, Chris McCool, Enrique Argones-Rúa, José Luis Alba-Castro, Mauricio Villegas, Roberto Paredes, Vitomir Struc, Nikola Pavesic, Albert Ali Salah, Hui Fang 0003, Nicholas Costen |
IEEE Trans. Inf. Forensics Secur. | 13 |
| 2009 | From Rank-N to Rank-1 Face Recognition Based on Motion SimilarityabstractIn this paper, we present a sequential framework using facial motion information as a subsidiary to improve face recognition performance. As is generally known, reasonable static face recognition has been achieved based on subspace reduction techniques. In order to further improve performance, some extra cues, such as temporal variation, are investigated by building dynamic models. We propose a permuted similarity motion fea-ture and integrate it into a sequential recognition system. This system can select the best candidate from the Rank-N candidates picked up in the recognition step based on static appearance parameters by using motion information. The recognition rate of the motion similarity is compared with the motion feature obtained from auto-regressive models to prove its efficiency. In addition, the sequential system achieves better performance when the motion information is integrated with the static appearance information in a flexible manner. 1 Hui Fang 0003, Nicholas Costen |
BMVC | 1 |
| 2008 | 3D facial geometry recovery via group-wise optical flowabstractWe describe an algorithm for automatically finding correspondences from face video sequences. This method is useful to many applications such as face tracking, face modeling and 3D face recovery. Given a sequence of images, the face feature points are tracked by a model-constraint optical flow algorithm. By employing a minimum description length (MDL) point-refinement framework, the drift-off error caused by the optical flow algorithm can be reduced and the correspondences can be matched robustly by optimizing the statistical model. As a result, the face is able to be tracked precisely. Furthermore, it offers a new method of building an appearance model automatically. The objective root mean square error (RMSE) is used to prove the efficiency of the algorithm. At the same time, the performance is evaluated subjectively by generating 3D face models based upon it. Hui Fang 0003, Nicholas Costen, David Cristinacce, John Darby |
FG | 1 |
| 2006 | A fuzzy logic approach for detection of video shot boundaries
Hui Fang 0003, Jianmin Jiang |
Pattern Recognit. | 1 |
| 2003 | Video Extraction in Compressed DomainabstractIn this paper, we propose a video extraction algorithm directly in the compressed domain for low cost and fast content access to those compressed video data via MPEG. Extensive experiments show that such extracted images and videos not only maintain well-preserved content features, but also illustrate reasonable quality in terms of both PSNR values and visual inspection. In cases where video processing tasks do not necessarily require full resolution pixel data such as browsing, pattern recognition, and object tracking in surveillance applications, the proposed algorithm will provide superior performance in terms of computing efficiency, under the context that millions of video frames need to be accessed, yet they are stored in compressed format. Puteri Norhashimah, Hui Fang 0003, Jianmin Jiang |
AVSS | 2 |