EDBT 2026 Demo / reviewers in the wild / expert
Yuesheng Zhu
dblp:35/8757
· DBLP profile ↗
158ranked-venue papers
1as first author
94since 2021 · last 2026
0000-0003-2524-6800ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 79 · 38 since 2021Artificial intelligence and machine learning · 71 · 48 since 2021Human-computer interaction and ubiquitous computing · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Computer networks · 7 · 2 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Security and privacy · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Feature-Aware One-Shot Federated Learning via Hierarchical Token SequencesabstractOne-shot federated learning (OSFL) reduces the communication cost and privacy risks of iterative federated learning by constructing a global model with a single round of communication. However, most existing methods struggle to achieve robust performance on real-world domains such as medical imaging, or are inefficient when handling non-IID (Independent and Identically Distributed) data. To address these limitations, we introduce FALCON, a novel framework that enhances the effectiveness of OSFL over non-IID image data. The core idea of FALCON is to leverage the feature-aware hierarchical token sequences generation and knowledge distillation into OSFL. First, each client leverages a pretrained visual encoder with hierarchical scale encoding to compress images into hierarchical token sequences, which capture multi-scale semantics. Second, a multi-scale autoregressive transformer generator is used to model the distribution of these token sequences and generate the synthetic sequences. Third, clients upload the synthetic sequences along with the local classifier trained on the real token sequences to the server. Finally, the server incorporates knowledge distillation into global training to reduce reliance on precise distribution modeling. Experiments on medical and natural image datasets validate the effectiveness of FALCON in diverse non-IID scenarios, outperforming the best OSFL baselines by 9.58\% in average accuracy. Shudong Liu 0008, Hanwen Zhang 0007, Yuesheng Zhu, Guibo Luo |
AAAI | 4 |
| 2026 | DSFedMed: Dual-Scale Federated Medical Image Segmentation via Mutual Distillation Between Foundation and Lightweight ModelsabstractFoundation Models (FMs) have demonstrated strong generalization across diverse vision tasks. However, their deployment in federated settings is hindered by high computational demands, substantial communication overhead, and significant inference costs. We propose DSFedMed, a dual-scale federated framework that enables mutual knowledge distillation between a centralized foundation model and lightweight client models for medical image segmentation. To support knowledge distillation, a set of high-quality medical images is generated to replace real public datasets, and a learnability-guided sample selection strategy is proposed to enhance efficiency and effectiveness in dual-scale distillation. This mutual distillation enables the foundation model to transfer general knowledge to lightweight clients, while also incorporating client-specific insights to refine the foundation model. Evaluations on five medical imaging segmentation datasets show that DSFedMed achieves an average 2 percent improvement in Dice score while reducing communication costs and inference time by nearly 90 percent compared to existing federated foundation model baselines. These results demonstrate significant efficiency gains and scalability for resource-limited federated deployments. Hanwen Zhang 0007, Qiaojin Shen, Yuxi Liu 0005, Yuesheng Zhu, Guibo Luo |
AAAI | 4 |
| 2026 | HAR-SCP: Fine-Grained Advertising Law Auditing from Platform URLs with Multimodal Evidence Centralization and Statutory Constraint Propagation
Huayang Hsu, Yuesheng Zhu, Guibo Luo, Hongye Liu |
ICIC (16) | 2 |
| 2026 | Robust and Diversified Image Steganography Without Embedding Through a Disentanglement AutoencoderabstractImage Steganography without Embedding (SWE) is an emerging data hiding paradigm. Instead of embedding a secret message into a container image, SWE synthesises a novel image by using the secret message as a latent code. Current SWE methods have achieved high synthesis quality and strong resistance to steganalysis tools. However, it remains challenging to apply the SWE due to two reasons: (i) lack of synthesis diversity and (ii) recovery of secret messages under malicious image attacks. In this paper, we present a novel SWE framework with a disentanglement autoencoder to tackle the above challenges. Specifically, the autoencoder disentangles an image into a structure and texture representation. Then, we exploit the stability of the structure representation to improve secret message recovery reliability, while increasing synthesis diversity by randomising texture representations and employing a chaotic system for structure randomisation to enhance its security. To further achieve a robust message recovery under malicious attacks, an adversarial learning strategy is introduced into our framework, which guarantees high recovery accuracy. Our method outperforms other state-of-the-art SWE methods in terms of synthesis quality, synthesis diversity and secret message recovery accuracy under various image attacks. The source code is publicly available athttps://github.com/Lemok00/RDI-SWE. Xiyao Liu 0001, Ziping Ma 0002, Jian Zhang 0048, Gerald Schaefer, Kehua Guo, Yuesheng Zhu, Shichao Zhang 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | Federated Learning for Medical Image Classification: A Comprehensive BenchmarkabstractThe federated learning (FL) paradigm is well-suited for the field of medical image analysis, as it can effectively cope with machine learning on isolated multi-center data while protecting the privacy of participating parties. However, current research on optimization algorithms in FL often focuses on limited datasets and scenarios, primarily centered around natural images, with insufficient comparative experiments in medical contexts. In this work, we conduct a comprehensive evaluation of several state-of-the-art FL algorithms in the context of medical imaging. We conduct a fair comparison of classification models trained using various FL algorithms across multiple medical imaging datasets. Additionally, we evaluate system performance metrics, such as communication cost and computational efficiency, while considering different FL architectures. Our findings show that medical imaging datasets pose substantial challenges for current FL optimization algorithms. No single algorithm consistently delivers optimal performance across all medical FL scenarios, and many optimization algorithms may under-perform when applied to these datasets. Our experiments provide a benchmark and guidance for future research and application of FL in medical imaging contexts. Furthermore, we propose an efficient and robust method that combines generative techniques using denoising diffusion probabilistic models with label smoothing to augment datasets, widely enhancing the performance of FL on classification tasks across various medical imaging datasets. Our codes are released on GitHub, offering a reliable and comprehensive benchmark for future FL studies in medical imaging. Zhekai Zhou, Guibo Luo, Zhenyu Weng, Yuesheng Zhu |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Robust Image Hashing Based on Contrastive Masked Autoencoder with Weak-Strong Augmentation AlignmentabstractRecently, numerous robust image hashing schemes have been developed for content identification. However, many of these schemes face the challenges of maintaining discrimination while simultaneously resisting large-scale attacks. In this paper, we propose a robust image hashing scheme based on Contrastive Masked Autoencoder with weak-strong augmentation Alignment (CMAA). Leveraging contrastive learning, CMAA is designed to learn features that are robust to large-scale and hybrid attacks while maintaining the discrimination of those features. Specifically, it utilizes distribution divergence to align weak attack augmented features with strong attack augmented features, namely weak-strong augmentation alignment, to enhance the robustness to strong attacks. In addition, a masked vision transformer is incorporated to further enhance content identification performance. CMAA also includes a parameter-free quantization layer to mitigate the loss induced by binarization. Experimental results demonstrate that our method exhibits remarkable robustness against various attacks, including challenging ones such as rotation and hybrid attacks, and delivers excellent identification performance with a F1 score close to 1.0. Our code and supplementary materials are available on Github. Cundian Yang, Guibo Luo, Yuesheng Zhu, Xiyao Liu 0001 |
AAAI | 3 |
| 2025 | Identity-Preserving Face Privacy Enhancement via Diffusion Models with Cognitive-Aware Obfuscation
Haoxuan Tai, Zhaokun Zhou, Yuesheng Zhu |
CogSci | 4 |
| 2025 | A Physical Backdoor Attack Against Practical Federated LearningabstractFederated learning (FL) is a distributed machine learning paradigm designed to build a global model while safeguarding data privacy. However, the decentralized nature and data heterogeneity of FL increase its susceptibility to backdoor attacks. Existing works on backdoor attacks and defenses typically focus on digital attacks that use digitally generated patterns as backdoor triggers. While using digitally generated patterns as triggers is convenient, its practicality remains questionable. In this work, a physical backdoor attack is proposed to enhance the practicality of backdoor attacks in FL. We employ real-world objects as backdoor triggers, integrating Backdoor-Robust Training Set Selection and Prox Attack, to execute backdoor attacks across various scenarios. Extensive experiments show that the physical backdoor attack in FL pose a serious real-world threat. This finding underscores the urgent need for more robust backdoor defenses in the physical world. Yang Li 0034, Xiangtao Lu, Guibo Luo, Yuesheng Zhu |
CSCWD | 4 |
| 2025 | MFDF-IML: Multi-Feature Dynamic Fusion for Image Manipulation LocalizationabstractWith the rapid advancement of image tampering techniques, the authenticity of multimedia content is increasingly challenged, necessitating the development of robust Image Manipulation Localization (IML) technologies. This paper introduces a novel approach, Multi-Feature Dynamic Fusion for Image Manipulation Localization (MFDF-IML), which ad-dresses the limitations of existing methods by integrating Error Level Analysis (ELA) as a new forgery feature. ELA enhances feature diversity and robustness by capturing subtle variations through analyzing discrepancies at different compression levels. Additionally, MFDF-IML employs a dynamic gating mechanism to adaptively fuse multiple features, including SRM, Bayar, Noiseprint++, and ELA, adjusting their weights according to various forgery scenarios. This method also integrates features from Convolutional Neural Networks (CNN) and Vision Transformers (ViT), leveraging CNN's local feature extraction and ViT's global dependency modeling to significantly improve forgery localization precision. Extensive experiments demonstrate that MFDF-IML outperforms existing methods across diverse forgery scenarios, highlighting its potential in image forensics. Xiangtao Lu, Fuyuan Cheng, Guibo Luo, Yuesheng Zhu |
CSCWD | 4 |
| 2025 | RGB Single-Channel Frequency Domain Backdoor Attack on Image Manipulation LocalizationabstractIn the contemporary digital era, the advancement and pervasive adoption of image manipulation technologies present unprecedented challenges to the authenticity of information. The extensive application of deep learning techniques in image processing has made detecting forged images a critical task for ensuring information security. Nonetheless, research in this field often neglects a potential risk: the security of training datasets. Previous backdoor attacks typically introduce triggers through semantic or noise manipulation, causing models to focus on these elements. Models designed to locate manipulated regions, however, often need to focus on high-frequency edges or artifacts, leading to suboptimal performance when distracted by these triggers. This paper introduces a novel Single-Channel RGB Frequency-Domain Backdoor Attack (SC-FDI), designed to manipulate the RGB channels by introducing triggers into the training data. This manipulation hinders the trained model from accurately identifying forged regions in images. Specifically, the original image is first separated into its RGB channels, then the Discrete Cosine Transform (DCT) is applied to one of the channels, the DCT coefficients in the high-frequency region are amplified, and the image is reassembled after inverse transformation. Experimental results on multiple datasets indicate that our method surpasses existing approaches in both stealthiness and attack efficacy. Xiangtao Lu, Yang Li 0034, Guibo Luo, Yuesheng Zhu |
CSCWD | 4 |
| 2025 | DeformSleepNet: Adaptive Multi-Level Feature Extraction for Enhanced Sleep Stage ClassificationabstractSleep quality plays a crucial role in human health, with sleep stages serving as key indicators of overall well-being. Despite the advancements in sleep stage classification, existing methods encounter several limitations: 1) the use of fixed-size convolution kernels restricts the adaptive extraction of multi-scale features, and 2) redundancy in multi-modal signals results in inadequate attention to highly correlated features. In this paper, we introduce DeformSleepNet, a novel architecture that combines standard and deformable convolutions to enable adaptive multi-scale feature extraction across various frequency bands. The Multi-level Feature Encoder (MFE) update different features in different manner, achieving a trade-off between accuracy and efficiency, effectively balancing accuracy and computational efficiency. Additionally, the Content-Aware Fusion Module (CAFM) dynamically adjusts the importance of different features during fusion, thereby enhancing the overall classification performance. Experimental evaluations conducted on three publicly available datasets demonstrate that DeformSleepNet outperforms main-stream methods. Ablation studies further validate the contribution of each individual module, highlighting the robustness and effectiveness of the proposed approach. Haoyang Xu, Yuesheng Zhu, Guibo Luo |
CSCWD | 2 |
| 2025 | Robust Image Hashing Based on Mixture-of-Experts with Hard-Sample Mining Contrastive LearningabstractIn collaborative system, a large amount of digital image transmission requires a reliable content identification system to ensure the authenticity and copyright of these images. Robust image hashing schemes could efficiently extract robust features of images without manipulating the original image, making them an effective solution for content identification. However, existing schemes face the challenge of resisting complex attacks, such as diverse attack types and their sophisticated combinations, in the real world while maintaining discrimination. The complexity of attacks is reflected in the diversity of attack types, the variation in attack intensity and the combination of multiple attacks. In this study, we propose a robust image hashing approach based on Mixture-of-experts with Hard-sample mining contrastive learning (MiHa). In particular, MiHa utilizes a mixture-of-experts architecture to adaptively extract features from the image under various attacks, including hybrid attacks, thereby enhancing the robustness. Additionally, large-scale attacks can generate hard samples for contrastive learning. To address this, we propose a hard-sample mining contrastive loss that assigns greater weight to these hard samples, thereby further improving the performance of MiHa. Extensive experiments demonstrate that our method achieves superior robustness against various attacks while maintaining competitive discrimination. Cundian Yang, Guibo Luo, Yuesheng Zhu, Xiyao Liu 0001 |
CSCWD | 3 |
| 2025 | A New One-Shot Federated Learning Framework for Medical Imaging Classification with Feature-Guided Rectified Flow and Knowledge DistillationabstractIn multi-center scenarios, One-Shot Federated Learning (OSFL) has attracted increasing attention due to its low communication overhead, requiring only a single round of transmission. However, existing generative model-based OSFL methods suffer from low training efficiency and potential privacy leakage in the healthcare domain. Additionally, achieving convergence within a single round of model aggregation is challenging under non-Independent and Identically Distributed (non-IID) data. To address these challenges, in this paper a modified OSFL framework is proposed, in which a new Feature-Guided Rectified Flow Model (FG-RF) and Dual-Layer Knowledge Distillation (DLKD) aggregation method are developed. FG-RF on the client side accelerates generative modeling in medical imaging scenarios while preserving privacy by synthesizing feature-level images rather than pixel-level images. To handle non-IID distributions, DLKD enables the global student model to simultaneously mimic the output logits and align the intermediate-layer features of client-side teacher models during aggregation. Experimental results on three non-IID medical imaging datasets show that our new framework and method outperform multi-round federated learning approaches, achieving up to 21.73% improvement, and exceed the baseline FedISCA by an average of 21.75%. Furthermore, our experiments demonstrate that feature-level synthetic images significantly reduce privacy leakage risks compared to pixel-level synthetic images. The code is available at https://github.com/LMIAPC/one-shot-fl-medical. Hanwen Zhang 0007, Qiya Yang, Guibo Luo, Yuesheng Zhu |
ECAI | 5 |
| 2025 | FLIDS-Mamba: Multi-Center Network Intrusion Detection Based on Federated Learning and Bidirectional State Space Model
Denghua Li, Yuesheng Zhu, Guibo Luo |
ICA3PP (7) | 2 |
| 2025 | SpikingPoint: Rethinking Point as Spike for Efficient 3D Point Cloud AnalysisabstractSpiking Neural Networks (SNNs), due to their unique spike-based inference mechanism, offer low power consumption and biological plausibility. As a fundamental technology for various real-world applications, 3D point cloud analysis faces significant challenges related to high computational overhead and energy-intensive. In fact, each point can be viewed as a specialized spike data containing positional information in 3D space. Therefore, utilizing the spiking features of SNNs to represent point clouds holds significant potential. However, this exploration faces two main challenges: SNN-compatible architecture design and spike-based 3D spatial modeling. In this work, we introduce the SpikingPoint, a pure Spiking Multi-layer Perceptron (MLP) Architecture that leverages the low power consumption of SNNs and the computational efficiency of linear layers. Furthermore, we propose the spiking 3D Position Embedding (SPE) that effectively models the 3D spatial feature into spike-form features. SpikingPoint achieves competitive performance with a small number of parameters while reducing energy consumption. For example, when achieving similar accuracy, the theoretical energy consumption of our method is reduced by 97.1% compared to the mainstream ANN-based KPConv. With fewer parameters, SpikingPoint surpasses the state-of-the-art SNN-based P2SResLNet on ModelNet40 by 1.42% in accuracy. Zhaokun Zhou, Yijie Lu, Jiaqiyu Zhan, Guibo Luo, Yuesheng Zhu |
ICASSP | 5 |
| 2025 | Spiking Transformer with Spatial-Temporal Spiking Self-AttentionabstractSpiking Neural Networks are celebrated for energy efficiency and biological plausibility. Building on Spiking Self-Attention (SSA), Spiking Transformers are extensively studied due to their exceptional performance. However, SSA focuses solely on spatial dimension at each time step, overlooking the crucial features across temporal dimension. To address this, we propose the Spatial-Temporal Spiking Self-Attention (STSSA), a spike-driven mechanism that leverages both spatial and temporal information with negligible additional computational overhead. Specifically, we extract the Representative Spiking Temporal Tokens (RSTT) and apply temporal window masking to the RSTT. These tokens are inserted between the Query and Key to integrate temporal features. Furthermore, we design a Multi-dimensional Learnable Scaling Factor (MLSF) to adapt to STSSA. Our results consistently demonstrate that STSSA outperforms SSA across extensive experiments on sequential, neuromorphic, and static datasets. Notably, STSSA achieves performance improvements of 5.7% and 2.9% over SSA on Sequential CIFAR-100 and CIFAR-10DVS, respectively. STSSA provides a powerful alternative within the family of Spiking Self-Attention mechanisms. Zhaokun Zhou, Li Yuan 0007, Yuesheng Zhu |
ICASSP | 5 |
| 2025 | Optical Model-Driven Sharpness Mapping for Autofocus in Small Depth-of-Field and Severe Defocus Scenarios
Chen-Liang Fan, Mingpei Cao, Chih Chien Hung, Yuesheng Zhu |
ICCV | 4 |
| 2025 | OTTER: Optimized Training with Trustworthy Enhanced Replication via Diffusion and Federated VMUNet for Privacy-Aware Medical Segmentation
Haocheng Kan, Yuesheng Zhu, Guibo Luo, Hanwen Zhang 0007 |
ICICS (2) | 2 |
| 2025 | MedDiff-FT: Data-Efficient Diffusion Model Fine-Tuning with Structural Guidance for Controllable Medical Image Synthesis
Jianhao Xie, Zhenyu Weng, Yuesheng Zhu, Guibo Luo |
MICCAI (4) | 4 |
| 2025 | AgentFactory: Towards Automated Agentic System Design and Optimization
Enci Zhang, Yuesheng Zhu, Xiaole Cui, Guibo Luo |
PRICAI | 3 |
| 2025 | Non-IID Medical Image Segmentation Based on Cascaded Diffusion Model for Diverse Multi-Center ScenariosabstractLearning from multi-center medical datasets to obtain a high-performance global model is challenging due to the privacy protection and data heterogeneity in healthcare systems. Current federated learning approaches are not efficient enough to learn Non-Independent and Identically Distributed (Non-IID) data and require high communication costs. In this work, a practical privacy computing framework is proposed to train a Non-IID medical image segmentation model under various multi-center setting in low communication cost. Specifically, an efficient cascaded diffusion model is trained to generate image-mask pairs that have similar distribution to the training data of clients, providing rich labeled data on client side to mitigate heterogeneity. Also, a label construction module is developed to improve the quality of generated image-mask pairs. Moreover, a set of aggregation methods is proposed to achieve global model from data generated from Cascaded Diffusion model for diverse scenarios: CD-Syn, CD-Ens and its extension CD-KD. CD-Syn is a one-shot method that trains segmentation model solely on public generated datasets while CD-Ens and CD-KD maximize the utilization of local original data by an extra communication round of ensemble or knowledge distillation. In this way, the setting of our proposed framework is highly practical, providing multiple aggregation methods which can flexibly adapt to varying demands for efficiency, privacy, and accuracy. We systematically evaluated the effectiveness of our proposed framework on five Non-IID medical datasets and observe 5.38% improvement in Dice score compared with baseline method (FednnU-Net) on average. Hanwen Zhang 0007, Yuxi Liu 0005, Guibo Luo, Yuesheng Zhu |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Attack-Defending Contrastive Learning for Volumetric Medical Image Zero-WatermarkingabstractZero-watermarking is an emerging distortion-free copyright protection method for volumetric medical images. However, achieving both robustness against various malicious attacks and distinguishability between individual images remains challenging. In this article, we propose a novel attack-defending contrastive learning zero-watermarking (ADCL-ZW) scheme to tackle the above challenge using deep learning-based representations. In our approach, we design an attack-defending data enrichment mechanism to enhance the watermarking robustness by generating a large number of image samples under various watermarking attacks. Subsequently, features for both watermarking distinguishability and robustness are enhanced through application of a contrastive loss. In particular, we implement a dual-stream Siamese network architecture to effectively handle both signal attacks and geometric attacks in order to enhance the watermarking performance. Experimental results demonstrate that ADCL-ZW achieves stronger watermarking robustness and a better tradeoff between watermarking robustness and distinguishability compared with state-of-the art zero-watermarking methods. One of the highlighted metrics is that the false-negative rate of ADCL-ZW achieves 0.01 when a fixed false-positive rate is set to 1%, which is more than 13.3 times better than the benchmark methods. Xiyao Liu 0001, Cundian Yang, Hui Fang 0003, Gerald Schaefer, Jian Zhang 0048, Yuesheng Zhu, Shichao Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2024 | Better than Random: Reliable NLG Human Evaluation with Constrained Active SamplingabstractHuman evaluation is viewed as a reliable evaluation method for NLG which is expensive and time-consuming. To save labor and costs, researchers usually perform human evaluation on a small subset of data sampled from the whole dataset in practice. However, different selection subsets will lead to different rankings of the systems. To give a more correct inter-system ranking and make the gold standard human evaluation more reliable, we propose a Constrained Active Sampling Framework (CASF) for reliable human judgment. CASF operates through a Learner, a Systematic Sampler and a Constrained Controller to select representative samples for getting a more correct inter-system ranking. Experiment results on 137 real NLG evaluation setups with 44 human evaluation metrics across 16 datasets and 5 NLG tasks demonstrate CASF receives 93.18\% top-ranked system recognition accuracy and ranks first or ranks second on 90.91\% of the human metrics with 0.83 overall inter-system ranking Kendall correlation. Code and data are publicly available online. Jie Ruan, Xiao Pu 0003, Mingqi Gao 0002, Xiaojun Wan 0001, Yuesheng Zhu |
AAAI | 5 |
| 2024 | Cross-Stage Transfer in Multi-Stage Cascade Ranking and Filtering SystemsabstractUnsupervised Domain Adaptation (UDA) aims to transfer a model from a labeled source domain to an unlabeled target domain, addressing challenges of distinct data distributions, termed domain shift. Existing UDA research primarily focuses on classification-like tasks, but neglects ranking and filtering tasks essential for applications like medical diagnosis and search engines. This paper is the first to notice and identify a new real-world transfer problem: cross-stage transfer in multi-stage cascade ranking and filtering systems, a common issue in diverse applications, including information retrieval systems, medical diagnosis, and other real-world ranking/filtering systems. In this problem, we emphasize the crucial assumption of order-invariance and address the key issue named Cross-stage Class Concept Conflict (C4), highlighting potential inconsistencies in class concepts for the same sample at different stages. To tackle these challenges, we propose a novel method, Unsupervised Rank Adaptation (URA), comprising two key components: order-conditional distribution alignment, characterizing the order-conditional distribution intra-stage and aligning them across stages; and principal projection alignment, aligning the principal component’s projection matrix with classifier parameters to ensure order-invariance without guessing pseudo-labels, mitigating the influence of C4. Experimental results show that our approach reaches state-of-the-art performance in various cross-stage transfer tasks. Yifan Pan, Guibo Luo, Yuesheng Zhu |
ECAI | 3 |
| 2024 | AdaIFL: Adaptive Image Forgery Localization via a Dynamic and Importance-Aware Transformer Network
Yuxi Li 0008, Fuyuan Cheng, Wangbo Yu, Guangshuo Wang, Guibo Luo, Yuesheng Zhu |
ECCV (43) | 6 |
| 2024 | Generalizable Deepfake Detection with Unbiased Feature Extraction and Low-Level Forgery Enhancement
Zhihan Yu, Guangshuo Wang, Yuesheng Zhu, Guibo Luo |
ICANN (2) | 4 |
| 2024 | Enhanced Unsupervised Domain Adaptation with Dual-Attention Between Classification and Domain AlignmentabstractUnsupervised Domain Adaptation (UDA) deals with transferring knowledge from labeled source domains to unlabeled target domains. This addresses the challenge of different distributions across domains, commonly known as domain shift. Numerous methods attempt to align distributions across domains while learning the core tasks (e.g., classification) on source domain separately. However, limited research has explored the mutual influence between classification and domain alignment. In this paper, we discuss the conflicting optimization between domain alignment and classification tasks, emphasizing the risk of negative transfer due to conflicting optimization directions. For better optimization consistency, these tasks should concentrate on the common information of features. To address this issue, we propose an innovative framework Dual-attention between classification and Domain Alignment (DuDA). DuDA employs gradient-based saliency maps to generate interpretable attentions, concurrently enhancing both classification and domain alignment through a dual-attention mechanism. Experimental results verify the effectiveness of DuDA in mitigating negative transfer and its strong adaptability and promising performance. Yifan Pan, Guibo Luo, Bairong Li, Yuesheng Zhu |
ICASSP | 4 |
| 2024 | LoTraNet: Locality-Guided Transformer Network for Image Manipulation Localization
Fuyuan Cheng, Yuxi Li 0008, Xiangtao Lu, Guibo Luo, Yuesheng Zhu |
ICONIP (7) | 5 |
| 2024 | TFCM: Tuning-Free Facial Concept-Erasure in Text-to-Image Models Through Attention and Sample Modulation
Guibo Luo, Zhiqiang Bai, Yuesheng Zhu |
ICONIP (7) | 4 |
| 2024 | Fine-Grained Transformer Encoder of Image-Text Retrieval for Streaming Data in Cross-Modal Continual Learning
Jiaqiyu Zhan, Yuesheng Zhu, Guibo Luo |
ICONIP (9) | 3 |
| 2024 | MambaFuse: Fusing Multi-scale Mamba and CNN Features for Seizure Prediction
Ding Pan, Guibo Luo, Yuesheng Zhu |
ICONIP (4) | 3 |
| 2024 | Seizure Prediction Based on Multi-scale Fusion-Attention Transformer
Ding Pan, Guibo Luo, Yuesheng Zhu |
ICONIP (5) | 3 |
| 2024 | OKey: Towards More Controllable, Secure and Robust Diffusion Model Image Steganography Using Optimized Key
Dongxu Yue, Guibo Luo, Yuesheng Zhu |
ICONIP (6) | 3 |
| 2024 | DACOA: Diffusion-Aligned Coherent Augmentation and Consistency Constraint Strategies for Federated Domain Generalization
Guangshuo Wang, Yuesheng Zhu, Guibo Luo |
ICPR (27) | 2 |
| 2024 | Multi-Scale Image Stitching Based on Unsupervised Controllable Fusion StructureabstractTraditional feature-based image stitching techniques rely heavily on the quality of feature detection and often fail to get a high quality solution for images with fewer features. Generating high-quality stitching images with a natural structure is still a challenging task in computer vision. The current distortion-based image alignment relies on the division of the grid, which is currently mostly equally divided without reference to the information in the image itself, and thus there are some problems such as straight line retention in the alignment stage. On this basis, if the deep learning approach is added in the later reconstruction stage to limit the generation of artifacts at the pixel level, it will be greatly improved, and we can ensure the high quality of the final splicing result of the image through two stages. In this paper, we succeed in proposing a novelty model called multi-scale image stitching based on unsupervised controllable fusion module which can be divided in two stages: Multi-scale grids alignment structure based on feature points density and image reconstruction structure based on unsupervised controllable fusion structure. We develop new energy terms to adapt to limit the transformations of the new grids. Extensive experiments demonstrate that the proposed method outperforms most state-of-the-arts by effectively preserving the linear structure in the image and improving the robustness. Wangbo Yu, Yuesheng Zhu, Guibo Luo |
IJCNN | 3 |
| 2024 | High Resolution Face Privacy-Enhancing method based on Latent Optimization with Identity-Preserving Facial MaskingabstractThe widespread application of face recognition technology easily leads to the leakage of facial soft biometric privacy. Therefore, image-level face privacy-enhancing methods for face recognition scenarios aim to remove soft biometric privacy information while preserving individual identity. Existing methods suffer from performance degradation in face recognition and poor quality in privacy-enhanced images when processing high-resolution images, often resulting in artifacts and complexity in model optimization. To address these issues, we propose an identity-preserving face privacy-enhancing method based on latent optimization and face masking for high-resolution images. The proposed method generates privacy-enhanced faces by optimizing latent codes and mask blending, without relying on a well-defined encoder-decoder structure, making it more suitable for processing high-resolution images. We introduce a face masking strategy and design a facial elements identity influence prioritization module to adaptively select privacy-enhancing facial parts and attributes. As a result, the proposed method can protect multiple facial privacy attributes while improving the quality of generated images and maintaining face recognizability. Extensive experiments on two benchmark datasets demonstrate the effectiveness of the proposed method in anonymizing visual appearances for high-resolution face images while maintaining superior face recognition performance and image quality. Haoxuan Tai, Bofei Guo, Guibo Luo, Yuesheng Zhu |
IJCNN | 5 |
| 2024 | Data Augmentation in Class-Conditional Diffusion Model for Semi-Supervised Medical Image SegmentationabstractAccurate segmentation of specific organs or diseased tissues in medical images is crucial for precise diagnosis and effective treatment planning. While fully-supervised deep learning methods have demonstrated remarkable performance, their effectiveness heavily relies on the availability of a substantial number of labeled images. Unfortunately, acquiring and manually labeling a large medical dataset is often expensive and impractical, especially for rare diseases, due to challenges related to data sharing and privacy. To address this limitation, a class-conditional diffusion model is proposed to synthesize realistic medical images, thereby augmenting small-scale datasets with high-quality samples. The diffusion-based method generates an unlimited number of authentic medical images, each conditioned on specific class labels, offering a valuable contribution to dataset augmentation strategies. To evaluate the utility of the generated data, synthetic images are incorporated into the genuine labeled dataset, thereby creating augmented datasets. By utilizing the generated data as unlabeled augmented data for the original dataset, our synthetic data is effectively integrated with semi-supervised medical image segmentation algorithms. This integration successfully combines the diversity of synthetic images with the advantages of semi-supervised learning, resulting in a more comprehensive utilization of limited small-scale data. Our experimental results on two public datasets have demonstrated that our class-conditional diffusion model can generate high-fidelity synthetic medical images. Furthermore, it enables semi-supervised methods to achieve superior segmentation performance in comparison to the methods using only original real datasets. Guibo Luo, Yuesheng Zhu |
IJCNN | 4 |
| 2024 | FedFMS: Exploring Federated Foundation Models for Medical Image Segmentation
Yuxi Liu 0005, Guibo Luo, Yuesheng Zhu |
MICCAI (8) | 3 |
| 2024 | CodeDetector: Revealing Forgery Traces with Codebook for Generalized Deepfake DetectionabstractThe malicious use of deepfake technologies poses a significant threat to social security, emphasizing the urgent necessity to advance deepfake detection. Existing detection models tend to overfit specific forgery traces within the training set, resulting in a weak generalization performance on unseen data. Considering that the feature distribution of real images is limited compared to the diverse forgery patterns, capturing the real feature space and treating features outside the distribution as potential forgery traces can mitigate overfitting and enhance the detector's generalization ability. Furthermore, since facial manipulation generates artifacts in pixels and disrupts the global consistency of the data, potential forgery traces at both pixel and feature levels can contribute to detection. In this paper, we propose a novel two-stage deepfake detection model named CodeDetector, which utilizes a codebook to capture the feature space of real faces and obtain potential forgery traces by facial reconstruction for detection. In the codebook learning stage, a codebook is used to capture the real feature distribution. In the detector training stage, we obtain pixel-level and feature-level residuals as potential forgery traces through facial reconstruction on real and fake faces to guide the model's attention to forgery clues. Specifically, we propose a Quantized Residual-Guided Attention Module and a Dual Residual Attention Module, calculating residuals and utilizing the attention mechanism to enhance global feature representation. Additionally, an Indices Prediction Module is introduced to ensure the accuracy of residual guidance that enhances the robustness of reconstruction during detection. Extensive experiments have demonstrated that CodeDetector outperforms state-of-the-art in deepfake detection cross-dataset benchmark. Zhihan Yu, Guibo Luo, Yuesheng Zhu |
ICMR | 4 |
| 2024 | Spiking Transformer with Experts MixtureabstractSpiking Neural Networks (SNNs) provide a sparse spike-driven mechanism which is believed to be critical for energy-efficient deep learning.
Mixture-of-Experts (MoE), on the other side, aligns with the brain mechanism of distributed and sparse processing, resulting in an efficient way of enhancing model capacity and conditional computation.
In this work, we consider how to incorporate SNNs’ spike-driven and MoE’s conditional computation into a unified framework.
However, MoE uses softmax to get the dense conditional weights for each expert and TopK to hard-sparsify the network, which does not fit the properties of SNNs.
To address this issue, we reformulate MoE in SNNs and introduce the Spiking Experts Mixture Mechanism (SEMM) from the perspective of sparse spiking activation.
Both the experts and the router output spiking sequences, and their element-wise operation makes SEMM computation spike-driven and dynamic sparse-conditional.
By developing SEMM into Spiking Transformer, the Experts Mixture Spiking Attention (EMSA) and the Experts Mixture Spiking Perceptron (EMSP) are proposed, which performs routing allocation for head-wise and channel-wise spiking experts, respectively. Experiments show that SEMM realizes sparse conditional computation and obtains a stable improvement on neuromorphic and static datasets with approximate computational overhead based on the Spiking Transformer baselines. Zhaokun Zhou, Yijie Lu, Yanhao Jia, Kaiwei Che, Liwei Huang, Yuesheng Zhu, Guoqi Li 0002, Zhaofei Yu, Li Yuan 0007 |
NeurIPS | 8 |
| 2024 | Describe Images in a Boring Way: Towards Cross-Modal Sarcasm GenerationabstractSarcasm generation has been investigated in previous studies by considering it as a text-to-text generation problem, i.e., generating a sarcastic sentence for an input sentence. In this paper, we study a new problem of cross-modal sarcasm generation (CMSG), i.e., generating a sarcastic description for a given image. CMSG is challenging as models need to satisfy the characteristics of sarcasm, as well as the correlation between different modalities. In addition, there should be some inconsistency between the two modalities, which requires imagination. Moreover, high-quality training data is insufficient. To address these problems, we take a step toward generating sarcastic descriptions from images without paired training data and propose an Extraction-Generation-Ranking based Modular method (EGRM) for CMSG. Specifically, EGRM first extracts diverse information from an image at different levels and uses the obtained image tags, sentimental descriptive caption, and commonsense-based consequence to generate candidate sarcastic texts. Then, a comprehensive ranking algorithm, which considers image-text relation, sarcasticness, and grammaticality, is proposed to select a final text from the candidate texts. Human evaluation at five criteria on a total of 2100 generated image-text pairs and auxiliary automatic evaluation show the superiority of our method. Code and data are publicly available1. Jie Ruan, Xiaojun Wan 0001, Yuesheng Zhu |
WACV | 4 |
| 2024 | Class Incremental Learning With Deep Contrastive Learning and Attention DistillationabstractClass incremental learning can solve the issue of catastrophic forgetting when the trained model is used to learn a new task in which the model may forget part of previous knowledge learned. The key issue is to alleviate the stability-plasticity dilemma and maintain the balance between preventing old knowledge from being forgotten and learning new knowledge. In this paper, a new class incremental learning method with deep contrastive learning and attention distillation is proposed. The deep contrastive learning can optimize last layer of the model to learn strong-semantic information and the intermediate layers to learn weak-semantic information when new data is coming so that the plasticity can be improved. To ensure a balance between plasticity and stability, a new attention distillation approach is developed to learn and distill the latent attention information of the feature maps and the similarity between different samples. Our analysis and experimental results indicate that the proposed method can improve the learning and memorization ability and the average accuracy of the model compared with the previous methods, and obtain competitive results in class increment scenarios for image classification. Jitao Zhu, Guibo Luo, Baishan Duan, Yuesheng Zhu |
IEEE Signal Process. Lett. | 4 |
| 2024 | Adaptive Face Recognition for Multi-Type OcclusionsabstractDue to the prevalence of influenza outbreaks and outdoor scenarios with various obstructing decorations, recognizing faces with occlusions has become a pressing challenge to address. However, current research mainly focuses on facial recognition with one kind of occlusion and does not provide compatible solutions for different kinds of common occlusions like glasses, sunglasses, and masks. Therefore, an Adaptive Multi-Type Occluded Face Recognition Model (AMOFR) is proposed to effectively handle multiple occlusion types simultaneously in this paper. In AMOFR, a generator is developed to produce diverse occluded face images for training, achieved by simulating various occlusion types on unoccluded face images. Subsequently, an occlusion type-based adapter is formulated to address a range of occlusion scenarios, guided by prompts from a Visual-Language model. To enhance overall performance by leveraging complete facial information, a feature-level knowledge distillation loss function is implemented, facilitating joint learning of unoccluded-face and occluded-face features. Furthermore, a new sunglasses-wearing dataset (CALFW-SUNGLASSES) is generated for more comprehensive test for AMOFR and further occlusion recognition research. Experimental results on datasets containing different types of occlusions have demonstrated that AMOFR achieves significantly higher accuracy compared to other advanced face recognition models. The implementation codes of AMOFR is available athttps://github.com/LIU-YUXI/Adaptive-Multi-occlusion-Face-Recognition. Yuxi Liu 0005, Guibo Luo, Zhenyu Weng, Yuesheng Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Temporal Weighted Heterogeneous Multigraph Embedding for Ethereum Phishing Scams DetectionabstractEthereum phishing scams have proven to be highly profitable in recent years, and pose a serious risk to the security of the blockchain ecosystem. Existing techniques for detecting phishing scams mostly model the transaction network at a very coarse-grained level. These methods rarely take into account the heterogeneity of the network, and do not consider multiple transactions over time between pairs of accounts. To this end, we model the Ethereum transaction network as a heterogeneous muiltidigraph and propose a novel graph embedding technique. Specifically, we use a temporal-weighted biased walking method based on the Jump-Stay strategy, which not only captures the properties of the dynamic transaction network more comprehensively, but also elegantly balances the distribution of different types of nodes. The superior performance of our model is shown through classification experiments on a real-world dataset of Ethereum transactions. Mingpei Cao, Yuesheng Zhu |
CSCWD | 5 |
| 2023 | Zero-Shot Offline Handwritten Chinese Character Recognition with Graph EmbeddingabstractHandwritten Chinese character recognition (HCCR) is a challenging topic in the field of computer vision due to its numerous categories, complex structure, shape similarity, casual writing style, and lack of training data. Some studies have attempted to develop radicals-based approaches with zero-shot recognition capacity to alleviate the data dependency problem of deep learning. However, previous studies tend to treat Chinese characters as isolated individuals and focus solely on the structural characteristics, ignoring the correlation information among Chinese characters. In this paper, we construct a graph to represent the correlation information and propose a novelty zero-shot HCCR method based on graph embedding. In addition, we present a pre-training approach for the encoder based on contrastive learning. The experimental results show that our method outperforms the other radical-based zero-shot recognition methods and also achieves a competitive performance on traditional experiment setting. Zhenyu Weng, Yuesheng Zhu |
CSCWD | 5 |
| 2023 | A New Targeted Online Password Guessing Algorithm Based on Old PasswordabstractPassword authentication is a widely used identity authentication method for computer supported cooperative systems. However, the frequent occurrence of password leakage incidents has become a universal problem, and the leaked passwords seriously threaten the security of users’ unleaked passwords. In order to gain a deeper understanding of the relationship between users’ old passwords and new passwords and help users choose a securer new password when their old passwords are leaked, we propose a new targeted online guessing algorithm, Targuess-II+, based on old password in this article. As a new probabilistic algorithm, Targuess-II+not only supports the application of strong transformation rules at any positions in a password, but also shows the transformation process from one password to another. Our analysis and experimental results have demonstrated that Targuess-II+obtains better performance in terms of crack rate and efficiency compared with other existing algorithms. Yuesheng Zhu |
CSCWD | 4 |
| 2023 | DeepVecFont-v2: Exploiting Transformers to Synthesize Vector Fonts with Higher QualityabstractVector font synthesis is a challenging and ongoing problem in the fields of Computer Vision and Computer Graphics. The recently-proposed DeepVecFont [27] achieved state-of-the-art performance by exploiting information of both the image and sequence modalities of vector fonts. However, it has limited capability for handling long sequence data and heavily relies on an image-guided outline refinement post-processing. Thus, vector glyphs synthesized by DeepVecFont still often contain some distortions and artifacts and cannot rival human-designed results. To address the above problems, this paper proposes an enhanced version of DeepVecFont mainly by making the following three novel technical contributions. First, we adopt Transformers instead of RNNs to process sequential data and design a relaxation representation for vector outlines, markedly improving the model's capability and stability of synthesizing long and complex outlines. Second, we propose to sample auxiliary points in addition to control points to precisely align the generated and target Bézier curves or lines. Finally, to alleviate error accumulation in the sequential generation process, we develop a context-based self-refinement module based on another Transformer-based decoder to remove artifacts in the initially synthesized glyphs. Both qualitative and quantitative results demonstrate that the proposed method effectively resolves those intrinsic problems of the original DeepVecFont and outperforms existing approaches in generating English and Chinese vector fonts with complicated structures and diverse styles. Yuqing Wang 0006, Longhui Yu, Yuesheng Zhu, Zhouhui Lian |
CVPR | 4 |
| 2023 | Imbalanced Conditional Conv-Transformer for Mathematical Expression Recognition
Shuaijian Ji, Zhaokun Zhou, Yuqing Wang 0006, Baishan Duan, Zhenyu Weng, Yuesheng Zhu |
ICANN (6) | 7 |
| 2023 | Style Expansion Without Forgetting for Handwritten Character Recognition
Jie Ruan, Zhenyu Weng, Jian Zhang 0018, Yuqing Wang 0006, Longhui Yu, Qiankun Gao, Yuesheng Zhu |
ICANN (3) | 7 |
| 2023 | Spikformer: When Spiking Neural Network Meets Transformer
Zhaokun Zhou, Yuesheng Zhu, Yaowei Wang 0001, Shuicheng Yan, Yonghong Tian 0001, Li Yuan 0007 |
ICLR | 2 |
| 2023 | Trapdoor Normalization with Irreversible Ownership VerificationabstractThis paper introduces a deep model watermark with an irreversible ownership verification scheme: Trapdoor Normalization (TdN), inspired by the trapdoor function in traditional cryptography. To protect intellectual property within deep models, the proposed method is able to embed ownership information into normalization layers during training. We argue and empirically validate that relevant methods are vulnerable to ambiguity attacks, where the forged watermarks can cast ambiguity over the ownership verification. The primary trait that distinguishes this work from previous ones, is its design of a bidirectional connection between watermarks and deep models. Thereby, TdN enables an irreversible ownership verification scheme that is difficult for the adversary to compromise. In this way, the proposed TdN can effectively defeat ambiguity attacks. Extensive experiments demonstrate that the proposed method is not only superior to previous state-of-the-art methods in robustness, but also has better efficiency. Zhenyu Weng, Yuesheng Zhu, Yadong Mu |
ICML | 3 |
| 2023 | Machine Unlearning with Affine Hyperplane Shifting and Maintaining for Image Classification
Mengda Liu, Guibo Luo, Yuesheng Zhu |
ICONIP (13) | 3 |
| 2023 | Incomplete Multi-view Subspace Clustering Using Non-uniform Hyper-graph for High-Order Information
Jiaqiyu Zhan, Yuesheng Zhu |
ICONIP (13) | 2 |
| 2023 | bt-vMF Contrastive and Collaborative Learning for Long-Tailed Visual RecognitionabstractReal-world data often exhibit long tail distributions with heavy class imbalance, where the majority (head) classes can dominate the training process and alter the decision boundaries of the minority (tail) classes, leading to biased feature spaces. Recently, researchers have investigated the potential of contrastive learning for long-tailed visual recognition and introduced a class-balanced factor in loss function engineering. Although this method can help improve performance, it harms head performance due to undesirable bias, resulting in poor separability of minority samples in feature spaces. In this paper, we target the logit adjustment and propose balanced student-t von Mises-Fisher (bt-vMF) contrastive learning, encouraging a large margin between the head and tail classes and providing better generalization. In addition, the network trained on long-tailed datasets suffers from great uncertainty in predictions. To alleviate this issue, we build mutual supervision among multiple experts via proposed bilateral collaborative learning (BCL), in which the collaboration is conducted from both bt-vMF similarity and relationship distillation. Simply put, our designs focus on the generalization power of a single expert and the knowledge transfer among multiple experts to alleviate the biased feature space and uncertainty in long-tailed learning, respectively. Experiments on multiple datasets show that our method achieves competitive performance on long-tailed visual recognition task. Jinhao Du, Guibo Luo, Yuesheng Zhu, Zhiqiang Bai |
ICTAI | 3 |
| 2023 | Robust Steganography without Embedding Based on Secure Container Synthesis and Iterative Message RecoveryabstractSynthesis-based steganography without embedding (SWE) methods transform secret messages to container images synthesised by generative networks, which eliminates distortions of container images and thus can fundamentally resist typical steganalysis tools. However, existing methods suffer from weak message recovery robustness, synthesis fidelity, and the risk of message leakage. To address these problems, we propose a novel robust steganography without embedding method in this paper. In particular, we design a secure weight modulation-based generator by introducing secure factors to hide secret messages in synthesised container images. In this manner, the synthesised results are modulated by secure factors and thus the secret messages are inaccessible when using fake factors, thus reducing the risk of message leakage. Furthermore, we design a difference predictor via the reconstruction of tampered container images together with an adversarial training strategy to iteratively update the estimation of hidden messages. This ensures robustness of recovering hidden messages, while degradation of synthesis fidelity is reduced since the generator is not included in the adversarial training. Extensive experimental results convincingly demonstrate that our proposed method is effective in avoiding message leakage and superior to other existing methods in terms of recovery robustness and synthesis fidelity. Ziping Ma 0002, Yuesheng Zhu, Guibo Luo, Xiyao Liu 0001, Gerald Schaefer, Hui Fang 0003 |
IJCAI | 2 |
| 2023 | TRMER: Transformer-Based End to End Printed Mathematical Expression RecognitionabstractAs a fundamental task of transcribing formula images into structural mathematical expressions, Printed Mathematical Expression Recognition (PMER) is wildly used in many fields. However, there is still a lack of an end-to-end approach toward fully exploring the spatial structure and semantic information in the formula to achieve high recognition accuracy. In this work, a Transformer-based Mathematical Expression Recognition (TRMER) model, is proposed to enhance the recognition accuracy. A Dual-Branch Encoder (DBE) is developed to extract multi-scaled feature maps from a formula image so that the spatial and semantic information can be obtained synchronously, and the different feature maps are fused with a Fusion Enhancement Module (FEM) by merging and reinforcing the spatial-semantic information. A standard transformer-based decoder is developed to decode the rich spatial-semantic information of the image and output a recognized mathematical expression in LaTex sequence. The experimental results have illustrated that the TRMER has achieved state-of-the-art recognition performance. Zhaokun Zhou, Shuaijian Ji, Yuqing Wang 0006, Zhenyu Weng, Yuesheng Zhu |
IJCNN | 5 |
| 2023 | Curriculum Graph PoisoningabstractDespite the success of graph neural networks (GNNs) over the Web in recent years, the typical transductive learning setting for node classification requires GNNs to be retrained frequently, making them vulnerable to poisoning attacks by corrupting the training graph. Poisoning attacks on graphs are, however, non-trivial as the attack space is potentially large, and the discrete graph structure makes the poisoning function non-differentiable. In this paper, we revisit the bi-level optimization problem in graph poisoning and propose a novel graph poisoning method, termed Curriculum Graph Poisoning (CuGPo), inspired by curriculum learning. In contrast to other poisoning attacks that use heuristics or directly optimize the graph, our method learns to generate poisoned graphs from basic adversarial knowledge first and advanced knowledge later. Specifically, for the outer optimization, we utilize the slightly perturbed graphs which represent the easy poisoning task at the beginning, and then enlarge the attack space until the final; for the inner optimization, we firstly exploit the knowledge from the clean graph and then adapt quickly to perturbed graphs to obtain the adversarial knowledge. Extensive experiments demonstrate that CuGPo achieves state-of-the-art performance in graph poisoning attacks. Peilin Zhao, Tingyang Xu, Yatao Bian, Junzhou Huang, Yuesheng Zhu, Yadong Mu |
WWW | 6 |
| 2023 | Separately Guided Context-Aware Network for Weakly Supervised Temporal Action Detection
Bairong Li, Yifan Pan, Ruixin Liu, Yuesheng Zhu |
Neural Process. Lett. | 4 |
| 2023 | Superframe-Based Temporal Proposals for Weakly Supervised Temporal Action DetectionabstractThe weakly supervised Temporal Action Detection (TAD) by using the video-level annotations can lighten the burden of labor consumption. However, the current methods for weakly supervised TAD do not take full advantage of the short-term consistency between consecutive frames and the long-term continuity inside an action, resulting in less accurate detecting boundaries of actions in untrimmed videos. In this paper, the SuperFrame-based Temporal Proposal (SFTP) is proposed, in which superframes are formed for representing a series of consecutive frames with high temporal consistency and their features are pooled from the features of frames through the integration function. Then, the temporal proposal is built based on the multiple consecutive superframes and the features of all proposals are generated from a pyramidal feature hierarchy. This hierarchy consists of the designed Structured Outer-Inner Context (SOIC) features formed from superframe features and is able to explicitly characterize the temporal continuity inside a proposal. Furthermore, a novel Scale-Wise Normalization Strategy (SWNS) is proposed to identify proposals, which can effectively detect multiple actions with different duration in one untrimmed video. Extensive experiments are conducted on two public datasets: THUMOS14 and ActivityNet1.2 for performance evaluation. Our experimental results have demonstrated that the proposed approach is able to detect the boundaries of actions more effectively and obtain competitive mAP (mean average precision) compared with other approaches. Bairong Li, Yuesheng Zhu, Jianfeng Yin, Xiangli Ji |
IEEE Trans. Multim. | 3 |
| 2023 | DREAMT: Diversity Enlarged Mutual Teaching for Unsupervised Domain Adaptive Person Re-IdentificationabstractPseudo-label-based methods of unsupervised domain adaption (UDA) can transfer the knowledge learned from a labeled source domain to an unlabeled target domain and have recently achieved significant progress in the application of person reidentification (re-ID). However, these methods suffer from serious label noise problems that downgrade the retrieval performance in UDA person re-ID. The mutual teaching framework (MTF) with dual networks attempts to tackle this problem by generating reliable soft pseudo labels but results in a mutual convergence problem. In this paper, a novel DiveRsity EnlArged Mutual Teaching framework (DREAMT) is proposed to solve the problem mentioned above. Based on the primary mutual-mean-teaching mechanism two strategies are developed in DREAMT, that is, GAN-based source domain augmentation (GSDA) and cross-branch mutual supervision (CBMS) for dual networks. Specifically, GSDA exploits two GANs to augment source domain datasets in different ways for pre-training to improve the pre-trained models’ performance and enlarge the diversity at the beginning of target domain adaption. During target adaption, each network in MTF adopts two branches to extract different features. CBMS based on hard and soft pseudo labels is across branches and networks and can help to maintain the diversity between training peers in the whole training process. Extensive experiments have demonstrated that our proposed DREAMT framework achieves better mAP and CMC performance than the existing mutual teaching methods and outperforms various state-of-the-art methods in UDA person re-ID tasks. Yusheng Tao, Jian Zhang 0018, Jiajing Hong, Yuesheng Zhu |
IEEE Trans. Multim. | 4 |
| 2022 | Unsupervised Online Hashing with Multi-Bit Quantization
Zhenyu Weng, Yuesheng Zhu |
ACCV (7) | 2 |
| 2022 | Dual Branch Network Towards Accurate Printed Mathematical Expression Recognition
Yuqing Wang 0006, Zhenyu Weng, Zhaokun Zhou, Shuaijian Ji, Zhongjie Ye, Yuesheng Zhu |
ICANN (4) | 6 |
| 2022 | Multi-Teacher Knowledge Distillation for Incremental Implicitly-Refined ClassificationabstractIncremental learning methods can learn new classes continually by distilling knowledge from the last model (as a teacher model) to the current model (as a student model) in the sequentially learning process. However, these methods cannot work for Incremental Implicitly-Refined Classification (IIRC), an incremental learning extension where the incoming classes could have two granularity levels, a superclass label and a subclass label. This is because the previously learned superclass knowledge may be occupied by the sub-class knowledge learned sequentially. To solve this problem, we propose a novel Multi-Teacher Knowledge Distillation (MTKD) strategy. To preserve the subclass knowledge, we use the last model as a general teacher to distill the previous knowledge for the student model. To preserve the superclass knowledge, we use the initial model as a superclass teacher to distill the superclass knowledge as the initial model contains abundant superclass knowledge. However, distilling knowledge from two teacher models could result in the student model making some redundant predictions. We further propose a post-processing mechanism, called as Top-k prediction restriction to reduce the redundant predictions. Our experimental results on IIRC-ImageNet120 and IIRC-CIFAR100 show that the proposed method can achieve better classification accuracy compared with existing state-of-the-art methods. Longhui Yu, Zhenyu Weng, Yuqing Wang 0006, Yuesheng Zhu |
ICME | 4 |
| 2022 | Class-Incremental Learning with Multiscale Distillation for Weakly Supervised Temporal Action Localization
Tianquan Chen, Bairong Li, Yusheng Tao, Yuqing Wang 0006, Yuesheng Zhu |
ICONIP (1) | 5 |
| 2022 | Interactive Image Inpainting Using Semantic GuidanceabstractImage inpainting approaches have achieved significant progress with the help of deep neural networks. How-ever, existing approaches mainly focus on leveraging the priori distribution learned by neural networks to produce a single inpainting result or further yielding multiple solutions, where the controllability is not well studied. This paper develops a novel image inpainting approach that enables users to customize the inpainting result by their own preference or memory. Specifically, our approach is composed of two stages that utilize the prior of neural network and user’s guidance to jointly inpaint corrupted images. In the first stage, an autoencoder based on a novel external spatial attention mechanism is deployed to produce reconstructed features of the corrupted image and a coarse inpainting result that provides semantic mask as the medium for user interaction. In the second stage, a semantic decoder that takes the reconstructed features as prior is adopted to synthesize a fine inpainting result guided by user’s customized semantic mask, so that the final inpainting result will share the same content with user’s guidance while the textures and colors reconstructed in the first stage are preserved. Extensive experiments demonstrate the superiority of our approach in terms of inpainting quality and controllability. Wangbo Yu, Jinhao Du, Ruixin Liu, Yuesheng Zhu |
ICPR | 5 |
| 2022 | Transformer-based Contrastive Learning for Unsupervised Person Re-IdentificationabstractUnsupervised Re-identification (Re-ID) methods have been dominated by convolutional neural networks (CNN) for many years. Most of these current methods apply pseudo-label-based contrastive learning (CL) and achieve great progress. However, they have limited capacity to represent global fea-tures, suffer from severe performance drops when training with limited computing resources, and are unable to effectively use pseudo-label information when training with CL. To tackle these problems, we propose a Transformer-based Contrastive Learning (TransCL) method to enhance the performance of CL and improve the feature representation ability of Re-ID, in which a batch and memory contrast (BMC) strategy is developed to optimize multi-level CL tasks concurrently to fully use the pseudo-label information. Additionally, a GCN aggregated clustering (GAC) scheme is designed to assist in generating more effective pseudo labels for CL. Extensive experimental results indicate that GAC and BMC work with vision transformer (ViT) achieves better training performance and enhances the representation ability of the Re-ID model. TransCL surpasses the state-of-the-art CNN method by 8.0% in mAP on the challenging MSMT17 dataset. Yusheng Tao, Jian Zhang 0018, Tianquan Chen, Yuqing Wang 0006, Yuesheng Zhu |
IJCNN | 5 |
| 2022 | A Secure and Efficient Data Deduplication Scheme with Dynamic Ownership Management in Cloud ComputingabstractEncrypted data deduplication is an important technique for saving storage space and network bandwidth, which has been widely used in cloud storage. Recently, a number of schemes that solve the problem of data deduplication with dynamic ownership management have been proposed. However, these schemes suffer from low efficiency when the dynamic ownership changes a lot. To this end, in this paper, we propose a novel server-side deduplication scheme for encrypted data in a hybrid cloud architecture, where a public cloud (Pub-CSP) manages the storage and a private cloud (Pri-CSP) plays a role as the data owner to perform deduplication and dynamic ownership management. Further, to reduce the communication overhead we use an initial uploader check mechanism to ensure only the first uploader needs to perform encryption, and adopt an access control technique that verifies the validity of the data users before they download data. Our security analysis and performance evaluation demonstrate that our proposed server-side deduplication scheme has better performance in terms of security, effectiveness, and practicability compared with previous schemes. Meanwhile, our method can efficiently resist collusion attacks and duplicate faking attacks. Xuewei Ma, Yuesheng Zhu, Zhiqiang Bai |
IPCCC | 3 |
| 2022 | A Multiple Positives Enhanced NCE Loss for Image-Text Retrieval
Yuesheng Zhu |
MMM (1) | 3 |
| 2022 | TokenAuditor: Detecting Manipulation Risk in Token Smart Contract by FuzzingabstractDecentralized cryptocurrencies are influential smart contract applications in the blockchain, drawing interest from industry and academia. The capacity to govern and manage token behavior provided by the token smart contract adds to thriving decentralized applications. However, token smart contracts face security challenges in technology weakness and manipulation risks. In this work, we briefly describe the manipulation risk and propose TokenAuditor, a fuzzing framework detecting those risks in token smart contracts. TokenAuditor constructs basic blocks based on the contract bytecodes and adopts the rarity selection and mutation strategy to generate test cases. The main idea is to select the test cases that have hit rare basic blocks since the fuzzing started as candidates and perform mutation operations on them. In our evaluation, TokenAudiotr discovered 664 manipulation risks of four types in 4021 real-world token contracts. Mingpei Cao, Yueze Zhang, Zhenxuan Feng, Yuesheng Zhu |
QRS | 5 |
| 2022 | DEMA: a distance-bounded energy-field minimization algorithm to model and layout biomolecular networks with quantitative featuresabstractSUMMARY: In biology, graph layout algorithms can reveal comprehensive biological contexts by visually positioning graph nodes in their relevant neighborhoods. A layout software algorithm/engine commonly takes a set of nodes and edges and produces layout coordinates of nodes according to edge constraints. However, current layout engines normally do not consider node, edge or node-set properties during layout and only curate these properties after the layout is created. Here, we propose a new layout algorithm, distance-bounded energy-field minimization algorithm (DEMA), to natively consider various biological factors, i.e., the strength of gene-to-gene association, the gene's relative contribution weight and the functional groups of genes, to enhance the interpretation of complex network graphs. In DEMA, we introduce a parameterized energy model where nodes are repelled by the network topology and attracted by a few biological factors, i.e., interaction coefficient, effect coefficient and fold change of gene expression. We generalize these factors as gene weights, protein-protein interaction weights, gene-to-gene correlations and the gene set annotations-four parameterized functional properties used in DEMA. Moreover, DEMA considers further attraction/repulsion/grouping coefficient to enable different preferences in generating network views. Applying DEMA, we performed two case studies using genetic data in autism spectrum disorder and Alzheimer's disease, respectively, for gene candidate discovery. Furthermore, we implement our algorithm as a plugin to Cytoscape, an open-source software platform for visualizing networks; hence, it is convenient. Our software and demo can be freely accessed at http://discovery.informatics.uab.edu/dema. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhenyu Weng, Zongliang Yue, Yuesheng Zhu, Jake Yue Chen |
Bioinform. | 3 |
| 2022 | Mutually activated residual linear modeling GAN for pose-guided person image generation
Yuesheng Zhu |
Neurocomputing | 2 |
| 2022 | Precise Correspondence Enhanced GAN for Person Image Generation
Yuesheng Zhu |
Neural Process. Lett. | 2 |
| 2022 | Weakly Supervised Temporal Action Detection With Temporal Dependency LearningabstractWeakly supervised temporal action detection aims at localizing temporal positions of action instances in untrimmed videos with only action class labels. In general, previous methods individually classify each frame based on the appearance information and the short-term motion information, and then integrate consecutive high-response action frames into entities which serve as detected action instances. However, the long-range temporal dependencies between action frames are not fully utilized, and the detection results are more likely to be trapped in the most discriminative action segments. To alleviate this issue, we propose a novel two-branch (i.e., the coarse detection branch and the refining detection branch) detection framework with learning the long-range temporal dependencies for obtaining more accurate detection results, where only action class labels are required. The coarse detection branch is used to localize the most discriminative segments of action instances based on a typical multi-instance learning paradigm under the supervision of action class labels, whereas the refining detection branch is expected to localize the less discriminative segments of action instances via learning the long-range temporal dependencies between frames based on the proposed Transformer-style architecture and learning strategies. This collaboration mechanism takes full advantage of complementary information from the provided action class labels and the natural temporal dependencies between action frames, forming a more comprehensive solution. Consequently, our method obtains more precise detection results. Expectedly, the proposed method outperforms recent weakly supervised temporal action detection methods on dataset THUMOS14 and ActivityNet measured by mAP@tIoU and AR@AN. Bairong Li, Ruixin Liu, Tianquan Chen, Yuesheng Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Temporal Group Fusion Network for Deep Video InpaintingabstractVideo inpainting is a task of synthesizing spatio-temporal coherent content in missing regions of the given video sequence, which has recently drawn increasing attention. To utilize the temporal information across frames, most recent deep learning-based methods align reference frames to target frame firstly with explicit or implicit motion estimation and then integrate the information from the aligned frames. However, their performance relies heavily on the accuracy of frame-to-frame alignment. To alleviate the above problem, in this paper, a novel Temporal Group Fusion Network (TGF-Net) is proposed to effectively integrate temporal information through a two-stage fusion strategy. Specifically, the input frames are reorganized into different groups, where each group is followed by an intra-group fusion module to integrate information within the group. Different groups provide complementary information for the missing region. A temporal attention model is further designed to adaptively integrate the information across groups. Such a temporal information fusion way gets rid of the dependence on alignment operations, greatly improving the visual quality and temporal consistency of the inpainted results. In addition, a coarse alignment model is introduced at the beginning of the network to handle videos with large motion. Extensive experiments on DAVIS and Youtube-VOS datasets demonstrate the superiority of our proposed method in terms of PSNR/SSIM values, visual quality and temporal consistency, respectively. Ruixin Liu, Bairong Li, Yuesheng Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Semantic-Aware Context Aggregation for Image InpaintingabstractRecent attention-based image inpainting methods have made inspiring progress by propagating distant contextual information into holes. However, they tend to generate blurry contents since the propagation process is always misled by preliminarily-recovered holes features which are not well-inferred. To handle this problem, we propose a novel semantic-aware context aggregation module (SACA) that aggregates distant contextual information from a semantic perspective by exploiting the internal semantic similarity of the input feature map. Compared with existing attention mechanisms that model the relation of all pixel-pairs, SACA can suppress the impact of misleading holes features in context aggregation and significantly reduce computation burden by learning the relation between pixels and semantics. Also, we apply SACA to both high-level and low-level feature maps in our model for generating both semantically and visually plausible results. Extensive experiments on Outdoor Scenes, CelebA and Paris StreetView datasets validate the superiority of our method compared with existing methods. Zhilin Huang, Chujun Qin, Ruixin Liu, Zhenyu Weng, Yuesheng Zhu |
ICASSP | 5 |
| 2021 | SERN: Stance Extraction and Reasoning Network for Fake News DetectionabstractFake news brings us panic and misunderstanding against the truth, especially under some unusual circumstances, such as the outbreak of COVID-19. It’s crucial to detect fake news on social media early to avoid further propagation. Previous methods manually label the stances implied in post-reply pairs to aid fake news detection, which costs much time and effort. To solve this problem, a novel Stance Extraction and Reasoning Network (SERN) is proposed to extract the stances implied in post-reply pairs implicitly and integrate the stance representations for fake news detection without manually labeling stances, which saves much time and effort. Besides, the adequate utilization of multimodal content in the news is beneficial for complementing information for unimodal representation and jointly improving decision confidence. Thus, a sentence-guided visual attention mechanism is proposed in the text-image fusion module that leverages text-image content for better fake news detection. Encouraging empirical results on Fakeddit and PHEME demonstrate that our method outperforms the state-of-the-art methods. Jianhui Xie, Ruixin Liu, Yinghong Zhang, Yuesheng Zhu |
ICASSP | 5 |
| 2021 | GI-AEE: GAN Inversion Based Attentive Expression Embedding Network For Facial Expression EditingabstractFacial expression editing aims to modify facial expression by specific conditions. Existing methods adopt an encoder-decoder architecture under the guidance of expression condition to process the desired expression. However, these methods always tend to produce artifacts and blurs in expression-intensive regions due to simultaneously modifying images in expression changed regions and ensuring the consistency of other attributes with the source image. To address these issues, we propose a GAN inversion based Attentive Expression Embedding Network (GI-AEE) for facial expression editing, which decouples this task utilizing GAN inversion to alleviate the strong effect of the source image on the target image and produces high-quality expression editing results. Furthermore, different from existing methods that directly embed the expression condition into the network, we propose an Attentive Expression Embedding module to embed corresponding expression vectors into different facial regions, producing more plausible results. Qualitative and quantitative experiments demonstrate our method outperforms the state-of-the-art expression editing methods. Ruixin Liu, Yifan Pan, Yuesheng Zhu, Zhiqiang Bai |
ICIP | 5 |
| 2021 | Multi-Dimensional Attentive Hierarchical Graph Pooling Network for Video-Text RetrievalabstractVideo-text retrieval task has raised increasing attention due to the rapid growth of videos on the Internet. Existing works adopt various networks to encode videos and texts into a common latent space and calculate their similarities. However, most works ignore mining significant frames of videos and the difference among different dimensions in word representations, leading to unsatisfactory retrieval results. In this paper, we propose a Multi-Dimensional Attentive Hierarchical Graph Pooling Network (MAGP) to learn improved representations for video-text retrieval. Specifically, we design a novel hierarchical graph pooling method to extract significant frames in videos and discard unrelated frames, hence the model can learn hierarchical and discriminative video representations. Moreover, a multi-dimensional attention mechanism is utilized in text encoder to strengthen representation ability by dimension-level attention. Experimental results on three video-text datasets demonstrate our MAGP model out-performs the state-of-the-art models. Yinghong Zhang, Yuesheng Zhu |
ICME | 4 |
| 2021 | Watermarking Deep Neural Networks with Greedy ResidualsabstractDeep neural networks (DNNs) are considered as intellectual property of their corresponding owners and thus are in urgent need of ownership protection, due to the massive amount of time and resources invested in designing, tuning and training them. In this paper, we propose a novel watermark-based ownership protection method by using the residuals of important parameters. Different from other watermark-based ownership protection methods that rely on some specific neural network architectures and during verification require external data source, namely ownership indicators, our method does not explicitly use ownership indicators for verification to defeat various attacks against DNN watermarks. Specifically, we greedily select a few and important model parameters for embedding so that the impairment caused by the changed parameters can be reduced and the robustness against different attacks can be improved as the selected parameters can well preserve the model information. Also, without the external data sources for verification, the adversary can hardly cast doubts on ownership verification by forging counterfeit watermarks. The extensive experiments show that our method outperforms previous state-of-the-art methods in five tasks. Zhenyu Weng, Yuesheng Zhu |
ICML | 3 |
| 2021 | A Multi-interaction Model with Cross-Branch Feature Fusion for Video-Text Retrieval
Junting Li, Yuesheng Zhu, Zhiqiang Bai |
ICONIP (6) | 3 |
| 2021 | Manipulation-Invariant Fingerprints for Cross-Dataset Deepfake Detection
Zuoyan Li, Ruixin Liu, Yuesheng Zhu |
ICONIP (4) | 4 |
| 2021 | Sequence-Aware Graph Neural Network for Session-based RecommendationabstractSession-based recommendation (SBR) nowadays plays a vital role in many online services, aiming to predict users' next action based on anonymous sessions. Recent research of GNNs-based methods models a session as a graph via investigating complex transitions of items in a session. However, these methods do not consider sequential information of the session when aggregating item embeddings to form a session-level embedding. Most methods consider not all previous but the last one item as the interest of a user, which restricts the performance of the model. To address this problem, we propose a model named Sequence-Aware Graph Neural Network (SA-GNN) for session-based recommendation. In SA-GNN, we design a sequence-aware attention to adaptively weigh the previous items to generate a session-level embedding, which greatly improves the representation ability of the model. Also, to improve the representation ability of the item embeddings, SA-GNN harnesses the power of self-attention within the GNN layer to capture both transitions between adjacent items and long-range dependencies among all items in a session. In empirical evaluations on three public recommendation datasets, our method consistently outperforms an extensive of state-of-the-art session-based recommendation methods. Zhencheng Huang, Zhenyu Weng, Yuesheng Zhu, Zhiqiang Bai |
IJCNN | 4 |
| 2021 | Bi-encoder Network with Structure-texture Consistency for Image InpaintingabstractExisting image inpainting methods have shown their potential in filling corrupted regions with plausible contents. However, these methods tend to produce results with distorted structures or unnatural textures since they neglect the difference between structures and textures in images and jointly process these two different types of information. To solve this problem, we propose a bi-encoder network (BE-Net) that seeks to handle structure and texture information separately, and fuse them to reconstruct completed images. Specifically, BE-Net first uses two parallel encoders to infer structure and texture features of the input images respectively. Then a structure-texture consistency module (STCM) is designed to weaken artifacts and enhance visual coherency of the output images by keeping the texture features consistent with the structure features. Finally, the structure features and the texture features are fused at each level of the decoder to recover images with reasonable structures and realistic textures. Extensive experiments on Paris StreetView and CelebA datasets show the proposed approach is effective in generating realistic and visually plausible results and outperforms several state-of-the-art methods. Chujun Qin, Zhilin Huang, Ruixin Liu, Zhenyu Weng, Yuesheng Zhu |
IJCNN | 5 |
| 2021 | Homogeneous Symptom Graph Attentive Reasoning Network for Herb RecommendationabstractThe herb recommendation system aiming for recommending a set of herb for patients is a significant task for Traditional Chinese Medicine (TCM). Recent works apply a graph convolutional network to model the relations among symptoms and herbs, showing promising performance. However, they typically suffer from two limitations: (1) The learning of the relations of symptoms and herbs from symptom-herb heterogeneous graphs would be disturbed by the semantic gap and the weak correlations between symptoms and herbs. (2) They ignore the complex diagnosis and systemic relations of a patient's multi-symptom, resulting in the lack of effectiveness and personalization in syndrome diagnosis. To overcome these limitations, we propose a novel Homogeneous Symptom Graph Attentive Reasoning Network (HSGARN). Firstly, to alleviate the noisy semantic gap and weak correlations of heterogeneous graphs, we propose a homogeneous graph embedding module to comprehensively model the semantic relations of symptoms and herbs. Secondly, we propose a symptom attentive reasoning module to generate syndrome representation for patients, which can sufficiently exploit the interrelation of a patient's symptoms and model the individual difference. Experimental results on two TCM datasets demonstrate the advantages of HSGARN over the state-of-the-arts. Yinghong Zhang, Jianhui Xie, Ruixing Liu, Yuesheng Zhu, Zhiqiang Bai |
IJCNN | 5 |
| 2021 | MM-Net: Learning Adaptive Meta-metric for Few-Shot Biometric Recognition
Qinghua Gu, Zhengding Luo, Wanyu Zhao, Yuesheng Zhu |
MMM (1) | 4 |
| 2021 | Confidence-Based Global Attention Guided Network for Image Inpainting
Zhilin Huang, Chujun Qin, Ruixin Liu, Yuesheng Zhu |
MMM (1) | 5 |
| 2021 | An Adaptive Face-Iris Multimodal Identification System Based on Quality Assessment Network
Zhengding Luo, Qinghua Gu, Guoxiong Su, Yuesheng Zhu, Zhiqiang Bai |
MMM (1) | 4 |
| 2021 | Knowledge Compensation Network with Divisible Feature Learning for Unsupervised Domain Adaptive Person Re-identification
Jiajing Hong, Yang Zhang 0070, Yuesheng Zhu |
PRICAI (2) | 3 |
| 2021 | Learning frame-level affinity with video-level labels for weakly supervised temporal action detection
Bairong Li, Yuesheng Zhu, Ruixin Liu, Zhenyu Weng |
Neurocomputing | 2 |
| 2021 | A secure heuristic semantic searching scheme with blockchain-based verification
Boyu Sun, Yuesheng Zhu, Dehao Wu 0004 |
Inf. Process. Manag. | 3 |
| 2021 | Robust and discriminative zero-watermark scheme based on invariant features and similarity-based retrieval to protect large-scale DIBR 3D videos
Xiyao Liu 0001, Yifan Wang 0008, Ziqiang Sun, Lei Wang 0017, Rongchang Zhao, Yuesheng Zhu, Beiji Zou 0001, Hui Fang 0003 |
Inf. Sci. | 6 |
| 2021 | A Deep Feature Fusion Network Based on Multiple Attention Mechanisms for Joint Iris-Periocular Biometric RecognitionabstractJoint iris-periocular recognition based on feature fusion can overcome some inherent drawbacks of unimodal biometrics, but most of the prior works are limited by conventional feature extraction approaches and fixed fusion schemes. To achieve more accurate and adaptive recognition, an end-to-end deep feature fusion network for joint iris-periocular recognition is proposed in this paper. Multiple attention mechanisms including self-attention and co-attention mechanisms are integrated into the network. Specifically, two forms of self-attention mechanisms, spatial attention and channel attention, are inserted into the feature extraction module, aiming to effectively learn the most important features and suppress unnecessary ones. Also, co-attention mechanism is introduced in the feature fusion module, which can adaptively fuse features to obtain more representative iris-periocular features. Additionally, in order to further enhance the discriminative power of the learned features, the proposed network is trained with a joint supervision of softmax loss and center loss. On two publicly available datasets, the proposed network with a small number of parameters outperforms unimodal biometrics and several iris-periocular recognition approaches. Zhengding Luo, Junting Li, Yuesheng Zhu |
IEEE Signal Process. Lett. | 3 |
| 2021 | A Verifiable Semantic Searching Scheme by Optimal Matching Over Encrypted Data in Public CloudabstractSemantic searching over encrypted data is a crucial task for secure information retrieval in public cloud. It aims to provide retrieval service to arbitrary words so that queries and search results are flexible. In existing semantic searching schemes, the verifiable searching does not be supported since it is dependent on the forecasted results from predefined keywords to verify the search results from cloud, and the queries are expanded on plaintext and the exact matching is performed by the extended semantically words with predefined keywords, which limits their accuracy. In this paper, we propose a secure verifiable semantic searching scheme. For semantic optimal matching on ciphertext, we formulate word transportation (WT) problem to calculate the minimum word transportation cost (MWTC) as the similarity between queries and documents, and propose a secure transformation to transform WT problems into random linear programming (LP) problems to obtain the encrypted MWTC. For verifiability, we explore the duality theorem of LP to design a verification mechanism using the intermediate data produced in matching process to verify the correctness of search results. Security analysis demonstrates that our scheme can guarantee verifiability and confidentiality. Experimental results on two datasets show our scheme has higher accuracy than other schemes. Yuesheng Zhu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Online Hashing With Bit Selection for Image RetrievalabstractOnline hashing methods have been intensively investigated in semantic image retrieval due to their efficiency in learning the hash functions with one pass through the streaming data. Among the online hashing methods, those based on the target codes are usually superior to others. However, the target codes in these methods are generated heuristically in advance and cannot be learned online to capture the characteristics of the data. In this paper, we propose a new online hashing method in which the target codes are constructed according to the data characteristics and are used to learn the hash functions online. By designing a metric to select the effective bits online for constructing the target codes, the learned hash functions are resistant to the bit-flipping error. At the same time, the correlation between the hash functions is also considered in the designed metric. Hence, the hash functions have low redundancy. Extensive experiments show that our method can achieve comparable or better performance than other online hashing methods on both the static database and the dynamic database. Zhenyu Weng, Yuesheng Zhu |
IEEE Trans. Multim. | 2 |
| 2020 | Efficient Querying from Weighted Binary CodesabstractBinary codes are widely used to represent the data due to their small storage and efficient computation. However, there exists an ambiguity problem that lots of binary codes share the same Hamming distance to a query. To alleviate the ambiguity problem, weighted binary codes assign different weights to each bit of binary codes and compare the binary codes by the weighted Hamming distance. Till now, performing the querying from the weighted binary codes efficiently is still an open issue. In this paper, we propose a new method to rank the weighted binary codes and return the nearest weighted binary codes of the query efficiently. In our method, based on the multi-index hash tables, two algorithms, the table bucket finding algorithm and the table merging algorithm, are proposed to select the nearest weighted binary codes of the query in a non-exhaustive and accurate way. The proposed algorithms are justified by proving their theoretic properties. The experiments on three large-scale datasets validate both the search efficiency and the search accuracy of our method. Especially for the number of weighted binary codes up to one billion, our method shows a great improvement of more than 1000 times faster than the linear scan. Zhenyu Weng, Yuesheng Zhu |
AAAI | 2 |
| 2020 | Online Hashing with Efficient Updating of Binary CodesabstractOnline hashing methods are efficient in learning the hash functions from the streaming data. However, when the hash functions change, the binary codes for the database have to be recomputed to guarantee the retrieval accuracy. Recomputing the binary codes by accumulating the whole database brings a timeliness challenge to the online retrieval process. In this paper, we propose a novel online hashing framework to update the binary codes efficiently without accumulating the whole database. In our framework, the hash functions are fixed and the projection functions are introduced to learn online from the streaming data. Therefore, inefficient updating of the binary codes by accumulating the whole database can be transformed to efficient updating of the binary codes by projecting the binary codes into another binary space. The queries and the binary code database are projected asymmetrically to further improve the retrieval accuracy. The experiments on two multi-label image databases demonstrate the effectiveness and the efficiency of our method for multi-label image retrieval. Zhenyu Weng, Yuesheng Zhu |
AAAI | 2 |
| 2020 | A Two-Stream Network with Image-to-Class Deep Metric for Few-Shot Classification
Qinghua Gu, Zhengding Luo, Yuesheng Zhu |
ECAI | 3 |
| 2020 | Real-Time Multiple Object Tracking with Discriminative FeaturesabstractTracking-by-detection methods track multiple objects by detecting the objects of interest in each frame and associating the detected objects with the tracks. By allowing object detection and appearance embedding to be learned in a shared network, recent tracking-by-detection methods can implement the tracking task in real time with the power of deep neural networks. However, they just focus on the detection stage and do not take advantage of the embedding features well in the association stage. In this paper, we exploit the discriminative embedding features in the association stage to improve the tracking performance. By combing the embedding features with the bounding boxes to associate the detected objects with the tracks, the number of identity switches during tracking can be reduced. Further, after associating the detected objects with the tracks, the embedding feature of each track is not only updated according to the associated object, but also learned to distinguish the similar detected objects. The experiments show that our method can achieve competitive tracking performance in real time compared to the state-of-the-art tracking methods. Zhenyu Weng, Yuesheng Zhu, Zhiping Lin 0001, Haizhou Li 0001 |
ICARCV | 2 |
| 2020 | Arnet: Attention-Based Refinement Network for Few-Shot Semantic SegmentationabstractSemantic segmentation is a challenging task for computer vision which aims to classify the objects from the pixel level. Previous methods based on deep learning have made some progress but the labeling work is very time-consuming. Few-shot semantic segmentation can alleviate this problem. In this paper, we propose an Attention-based Refinement Network (ARNet) for few-shot semantic segmentation, which consists of three branches: the guidance branch, the segmentation branch and the refinement branch. The Residual Attention Module (RAM) can highlight the features from segmentation branch, giving a better guidance to refinement branch. And the Parallel Dilated Convolution Module (PDCM) in the end of refinement branch can refine the segmentation results. Experiments on PASCAL VOC 2012 dataset show that our model achieves a mean Intersection-over-Union (mIoU) score of 48.1% for one-shot segmentation and 49.1% for five-shot segmentation, outperforming state-of-the-art methods by 1.8% and 2.0%, respectively. Rusheng Li, Hanhui Liu, Yuesheng Zhu, Zhiqiang Bai |
ICASSP | 3 |
| 2020 | Unsupervised Person Re-Identification Using Multi-Branch Feature Compensation Network and Link-Based Cluster Dissimilarity MetricabstractFeature extraction and label estimation are critical in unsupervised person re-identification (re-ID). Most previous works focus on acquiring high-layer semantic features and reckon without the lower-layer details lost in the learning process, which causes the extracted features to be less comprehensive and may degrade re-ID performance. Therefore in this paper, a Multi-branch Feature Compensation Network (MFC-Net) is developed in which the significant parts of lower-layer features are learned and fused with high-layer feature as compensation. Moreover, to accurately conduct label estimation through applying hierarchical clustering on individual samples, a Link-based Cluster Dissimilarity Metric (LCDM) is proposed to discover the inner correlation between clusters. These two methods are later integrated as a MFC-LCDM scheme to improve unsupervised re-ID performance. Extensive experiments on large-scale image-based and video-based datasets ulteriorly demonstrate the superiority and effectiveness of MFC-LCDM. Gege Qi, Yuesheng Zhu |
ICASSP | 4 |
| 2020 | Residual Encoder-Decoder Network For Deep Subspace ClusteringabstractSubspace clustering aims to cluster unlabeled data that lies in a union of low-dimensional linear subspaces. Deep subspace clustering approaches based on auto-encoders have become very popular to learn the linear representation coefficients from data. However, the training of current deep methods converges slowly, which is extremely expensive. We propose a novel Residual Encoder-Decoder network for deep Subspace Clustering (RED-SC) with skip-layer connections to accelerate the convergence, using a new strategy to generate the linear coefficients by learning the linearity of data in multiple latent spaces. Experiments show the superiority of RED-SC in training efficiency and clustering accuracy. Yuesheng Zhu |
ICIP | 3 |
| 2020 | Deep Self-Learning Hashing for Image RetrievalabstractWith the advances in deep learning, deep-hashing methods have achieved promising results for image retrieval. However, the problem of the distribution gap between training data and test data remains unsolved. Existing methods rely too much on manually labeled information to construct similarity matrices as supervision signals and focus less on pre-trained networks that can extract semantic information. This limits the generalization performance of the network and produces less discriminative hash codes. In this paper, we propose a novel hashing method, deep self-learning hashing (DSLH), that uses a self-learning strategy with labels constructed using the pre-trained features to enhance the embedded representation of the hash codes. Furthermore, we develop an improved loss function that preserves the similarity of the hash codes while reducing the quantization loss and ensuring the balance of the hash codes. Our analysis and experimental results demonstrate that, compared with recent image-retrieval methods, our method can achieve greater retrieval performance on two benchmark datasets: CIFAR-10 and NUS-WIDE. Jiawei Zhan, Zhaoguo Mo, Yuesheng Zhu |
ICIP | 3 |
| 2020 | An Expression-Reinforced Sparse Subspace Clustering By Orthogonal Matching PursuitabstractSparse Subspace Clustering (SSC) is an efficient method for clustering high-dimensional data. Traditional SSC calculates the sparse representation of each point separately to obtain coefficient matrix, might undermine the clustering results due to the potential of representation for interaction being neglected. In this paper, based on orthogonal matching pursuit (OMP), a new module for SSC is developed to release potential of interaction among representations of similar data points for expression reinforcement, and an Expression-Reinforced Max Chain for Interaction (ER-MCI) algorithm is proposed to strengthen the effectiveness of affinity matrix. By finding a specific atom of each point and tracking these atoms to build the chain iteratively on which all atoms would lie in the same subspace with high probability. Experimental results show that our approach achieves better clustering performances compared with other SSC algorithms in terms of clustering accuracy, anti-noise ability and subspace-preserving property, and keeps time efficiency. Jiaqiyu Zhan, Yuesheng Zhu, Zhiqiang Bai |
ICIP | 2 |
| 2020 | Targeted Incorporating Spatial Information in Sparse Subspace Clustering of Hyperspectral Remote Sensing ImagesabstractMethods based on sparse subspace clustering (SSC) have shown great potential for hyperspectral image (HSI) clustering. However their performance is limited due to the complex spatial-spectral structure in HSIs. In this paper, a spatial best-fit direction (SBFD) algorithm is proposed to update the coefficients obtained from sparse representation to more discriminant features by integrating the spatial-contextual information given by the best-fit pixel of each target pixel. Also, SBFD is more targeted by searching for the best-fit direction than directly using the local window to do max pooling. The proposed SBFD was tested on two widely used hyperspectral dataset, the experimental results indicate its improvement in the clustering accuracy and spatial homogeneity. Jiaqiyu Zhan, Yuesheng Zhu, Zhiqiang Bai |
ICIP | 2 |
| 2020 | Seg-Hashnet: Semantic Segmentation Based Unsupervised HashingabstractHashing based multimedia retrieval methods have been widely studied because of the advantages of computation efficiency and low storage cost. Generally, there are two ways of doing hashing: Supervised and unsupervised hashing. Supervised hashing was known to be more efficient and popular. However, in many real applications, only unsupervised method work because there are no labels. But unsupervised hashing method still face some challenges such as how to get semantic information from unlabeled data. To address this problem, in this paper, we propose a deep unsupervised hashing method based on semantic segmentation (Seg-HashNet). Specifically, a pre-trained semantic segmentation model is used to mine informative semantic information to guide the learning process and generate highquality hash codes. Extensive experiments on two public datasets have been performed, which showed that Unsupervised Semantic Segmentation hashing is superior to the latest unsupervised hashing method in image retrieval tasks. Yuesheng Zhu, Zhiqiang Bai, Jinlong Lin |
ICIP | 3 |
| 2020 | Multi-Similarity Semantic Correctional Hashing For Cross Modal RetrievalabstractGiven the benefits of their low storage requirements and high retrieval efficiency, hashing methods have attracted considerable attention for large scale cross-modal retrieval and significant progress has been made recently. However, the existing methods generally use the label-guided similarity matrix to measure the similarities of sample pairs, which limits their semantic representation capability. Moreover, the sample imbalance of different classes would bias the learning process toward majority classes and affect the retrieval performance. To boost the semantic representation, to alleviate the impact of data imbalance, and to obtain a high-ranking correlation of hash code pairs, we propose a novel hashing method that uses a semantic correctional similarity matrix to enhance the embedded representation of sample pairs. Furthermore, we propose a novel cross-modal multi-similarity loss based on the general pair weighting framework to collect and weight informative pairs efficiently and accurately, thus improving the retrieval performance. Our analysis and experimental results demonstrate that, compared with recent cross-modal retrieval methods, our methods achieve greater retrieval performance on two datasets MIRFlickr-25K and NUS-WIDE. Jiawei Zhan, Zhaoguo Mo, Yuesheng Zhu |
ICME | 4 |
| 2020 | Sparse-Dense Subspace ClusteringabstractSubspace clustering refers to the problem of clustering high-dimensional data into a union of low-dimensional subspaces. Current subspace clustering approaches are usually based on a two-stage framework. In the first stage, an affinity matrix is generated from data. In the second one, spectral clustering is applied on the affinity matrix. However, the affinity matrix produced by two-stage methods cannot fully reveal the similarity between data points from the same subspace, resulting in inaccurate clustering. Besides, most approaches fail to solve large-scale clustering problems due to poor efficiency. In this paper, we first propose a new scalable sparse method called Iterative Maximum Correlation (IMC) to learn the affinity matrix from data. Then we develop Piecewise Correlation Estimation (PCE) to densify the intra-subspace similarity produced by IMC. Finally we extend our work into a Sparse-Dense Subspace Clustering (SDSC) framework with a dense stage to optimize the affinity matrix for two-stage methods. We show that IMC is efficient for large-scale tasks, and PCE ensures better performance for IMC. We show the universality of our SDSC framework for current two-stage methods as well. Experiments on benchmark data sets demonstrate the effectiveness of our approaches. Yuesheng Zhu |
ICPR | 3 |
| 2020 | Temporal Adaptive Alignment Network for Deep Video InpaintingabstractVideo inpainting aims to synthesize visually pleasant and temporally consistent content in missing regions of video. Due to a variety of motions across different frames, it is highly challenging to utilize effective temporal information to recover videos. Existing deep learning based methods usually estimate optical flow to align frames and thereby exploit useful information between frames. However, these methods tend to generate artifacts once the estimated optical flow is inaccurate. To alleviate above problem, we propose a novel end-to-end Temporal Adaptive Alignment Network(TAAN) for video inpainting. The TAAN aligns reference frames with target frame via implicit motion estimation at a feature level and then reconstruct target frame by taking the aggregated aligned reference frame features as input. In the proposed network, a Temporal Adaptive Alignment (TAA) module based on deformable convolutions is designed to perform temporal alignment in a local, dense and adaptive manner. Both quantitative and qualitative evaluation results show that our method significantly outperforms existing deep learning based methods. Ruixin Liu, Zhenyu Weng, Yuesheng Zhu, Bairong Li |
IJCAI | 3 |
| 2020 | TAM-Net: Temporal Enhanced Appearance-to-Motion Generative Network for Video Anomaly DetectionabstractVideo anomaly detection is a challenging task due to the diversity of anomaly. Existing GAN-based approaches model normal motion pattern through transforming a single image to optical flow map, which tends to learn the mapping between two adjacent frames instead of motion evolution in normal scenes. Therefore, this paper proposes a Temporal enhanced Appearance-to-Motion generative Network (TAM-Net) to model evolution of appearance and motion for normal events. In the motion generative branch, the corresponding optical flow map is generated by a ConvLSTM-based generative adversarial network from consecutive frames to learn normal motion pattern. In order to learn appearance pattern, consecutive frames are reconstructed by a auto-encoder in the reconstruction branch. Temporal encoded features of consecutive frames are shared by these two branches to represent changes of normal appearance along with time. By modeling spatio-temporal evolution of normal events, our network can effectively highlight abnormal regions with high generation errors of the predicted optical flow map and reconstructed frame. Experimental results on three independent datasets, UCSD Ped1, Ped2 and Avenue, demonstrate the competitive performance of the proposed method with the other approaches. Xiangli Ji, Bairong Li, Yuesheng Zhu |
IJCNN | 3 |
| 2020 | Improved Hierarchical Clustering with Non-locally Enhanced Features for Unsupervised Person Re-identificationabstractDue to the high cost of data annotation in supervised person re-identification (re-ID) methods, unsupervised person re-ID methods have attracted more and more attention. The unsupervised person re-ID methods based on deep clustering have achieved good performance. However, the distance metrics used in existing unsupervised clustering methods ignore intra-cluster distance and are likely to cause some wrong merging situations and uneven distribution within clusters. Besides, these models based on deep clustering usually ignore the importance of global features for person re-ID. In this paper, we address the above problems by proposing an improved hierarchical clustering approach with non-locally enhanced features. To improve the clustering performance, we design a new metric which consists of intermediate distance as inter-cluster distance and compactness degree as intra-cluster distance. The former one can prevent some wrong merging situations and the latter one can promote the uniform distribution within clusters. In addition, we develop a non-locally enhanced feature network to take advantage of global features of images. Extensive experiments on Market-1501, DukeMTMC-reID, MARS and DukeMTMC-VideoReID demonstrate that our method obtains significant improvement over the state-of-the-art unsupervised methods. Wanyu Zhao, Bairong Li, Qinghua Gu, Yuesheng Zhu |
IJCNN | 4 |
| 2020 | An Encrypted Traffic Classification Method Combining Graph Convolutional Network and AutoencoderabstractThe increase in the source and size of encrypted network traffic brings significant challenges for network traffic analysis. The challenging problem in the encrypted traffic classification field is obtaining high classification accuracy with small number of labeled samples. To solve this problem, we propose a novel encryption traffic classification method that learns the feature representation from the traffic structure and the traffic flow data in this paper. We construct a K-Nearest Neighbor (KNN) traffic graph to represent the structure of traffic data, which contains more similarity information about the traffic. We utilize a two-layer Graph Convolutional Network (GCN) architecture for flows feature extraction and encrypted traffic classification. We further use the autoencoder to learn the representation of the flow data itself and integrate it into the GCN-learned representation to form a more complete feature representation. The proposed method leverages the benefits of the GCN and the autoencoder, which can obtain higher classification performance with only very few labeled data. The experimental results on two public datasets demonstrate that our method achieves impressive results compared to the state-of-the-art competitors. Boyu Sun, Mengqi Yan, Yuesheng Zhu, Zhiqiang Bai |
IPCCC | 5 |
| 2020 | BDTF: A Blockchain-Based Data Trading Framework with Trusted Execution EnvironmentabstractThe need for data trading promotes the emergence of data market. However, in conventional data markets, both data buyers and data sellers have to use a centralized trading platform which might be dishonest. A dishonest centralized trading platform may steal and resell the data seller's data, or may refuse to send data after receiving payment from the data buyer. It seriously affects the fair data transaction and harm the interests of both parties to the transaction. To address this issue, we propose a novel blockchain-based data trading framework with Trusted Execution Environment (TEE) to provide a trusted decentralized platform for fair data trading. In our design, a blockchain network is proposed to realize the payments from data buyers to data sellers, and a trusted exchange is built by using a TEE for the first time to achieve fair data transmission. With these help, data buyers and data sellers can conduct transactions directly. We implement our proposed framework on Ethereum and Intel SGX, security analysis and experimental results have demonstrated that the framework proposed can effectively guarantee the fair completion of data tradings. Guoxiong Su, Zhengding Luo, Yinghong Zhang, Zhiqiang Bai, Yuesheng Zhu |
MSN | 6 |
| 2020 | A Disocclusion Inpainting Framework for Depth-Based View SynthesisabstractThis paper proposes a disocclusion inpainting framework for depth-based view synthesis. It consists of four modules: foreground extraction, motion compensation, improved background reconstruction, and inpainting. The foreground extraction module detects the foreground objects and removes them from both depth map and rendered video; the motion compensation module guarantees the background reconstruction model to suit for moving camera scenarios; the improved background reconstruction module constructs a stable background video by exploiting the temporal correlation information in both 2D video and its corresponding depth map; and the constructed background video and inpainting module are used to eliminate the holes in the synthesized view. The analysis and experiment indicate that the proposed framework has good generality, scalability and effectiveness, which means most of the existing background reconstruction methods and image inpainting methods can be employed or extended as the modules in our framework. Our comparison results have demonstrated that the proposed framework achieves better synthesized quality, temporal consistency, and has lower running time compared to the other methods. Guibo Luo, Yuesheng Zhu, Zhenyu Weng, Zhaotian Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Concatenation hashing: A relative position preserving method for learning binary codes
Zhenyu Weng, Yuesheng Zhu |
Pattern Recognit. | 2 |
| 2019 | A New Data Selection Strategy for One-Shot Video-Based Person Re-IdentificationabstractPerson re-identification under one-shot setting is an under-researched semi-supervised task in essence, in which label estimation and data selection are two key modules. In this paper, we propose an effective strategy named Collaboration between the Labeled and Unlabeled(CLU) to estimate pseudo labels and select reliable data progressively. In CLU, label estimation is an iterative process where class center representations are estimated by high-scored unlabeled data and corrected by corresponding labeled data. Then reliable data are selected based on confidence scores decided by both labeled and unlabeled data actually. Moreover, a stepwise learning framework for one-shot video-based person re-identification named SCLU is implemented to validate the effectiveness of CLU. Extensive experiments conducted on MARS and DukeMTMC-VideoReID illustrate that CLU generates pseudo labels of high quality and facilitates the subsequent model training. Meanwhile, the developed scheme SCLU yields the state-of-the-art results on the two datasets. Jianfeng Yin, Bairong Li, Yuesheng Zhu |
ICIP | 4 |
| 2019 | Deep Captioning Hashing Network for Complex Scene Image RetrievalabstractHashing methods have been widely applied to approximate nearest neighbor search for large-scale image retrieval, due to its computation efficiency and retrieval quality. Deep hashing can improve the retrieval quality by representation learning and hash coding. Existing deep hashing methods only take image spatial features into account and result in the lack of accurate semantic similarities of images pairs. In this paper, a novel deep hashing network, Deep Captioning Hashing Network (DCHN), is proposed to enhance semantic similarities of hash codes. In DCHN, the binary hash codes are generated in a Bayesian learning framework by fusing deep spatial representation and deep content captioning representation obtained by image captioning. Our analysis and simulation results have demonstrated that DCHN can achieve better retrieval performance in complex scene images compared with other supervised hashing methods and unsupervised methods on two complex scene image datasets MS COCO and NUS-WIDE. Jiawei Zhan, Zhengding Luo, Gege Qi, Zhiqiang Bai, Yuesheng Zhu |
ICTAI | 6 |
| 2019 | A Robust Single-Sensor Face and Iris Biometric Identification System Based on Multimodal Feature Extraction NetworkabstractJoint face-iris identification can integrate complementary information from face and iris to fulfill the requirement of performance improvement and security. However, most of the current face-iris multimodal biometric systems acquire face and iris with different sensors which brings about the increase of capturing complexity and device cost. Besides, they are limited by the identification performance degradation under non-ideal scenarios. In order to address these problems, a robust single-sensor face and iris biometric identification system based on multimodal feature extraction (MFE) network is proposed. Only a single sensor is needed to obtain face and iris images in the proposed system, with the goal of improving recognition performance while minimizing sensor cost and acquisition time. The MFE network is designed as a general network module to extract both face and iris features and it is trained with a triplet framework to reduce intra-class variations and enlarge inter-class variations. Our experimental results on CASIA.v4-distance and FRGC v2.0 non-ideal datasets show that the proposed system achieves better identification performance in terms of Equal Error Rate (EER) and False Reject Rate (FFR), etc. compared with other unimodal and multimodal biometric systems. Zhengding Luo, Qinghua Gu, Gege Qi, Yuesheng Zhu, Zhiqiang Bai |
ICTAI | 5 |
| 2019 | Self-Learned Feature Reconstruction and Offset-Dilated Feature Fusion for Real-Time Semantic SegmentationabstractRecent approaches for real-time semantic segmentation usually employ the encoder-decoder architecture as the backbone to generate a high-quality segmentation prediction. There has been a lot of research on designing efficient encoding methods. However, enhancing the performance of components in decoder is also crucial for pixel-level recognition. In this paper, we propose a self-learned feature reconstruction (SFR) method and an offset-dilated feature fusion (ODFF) module to improve the prediction reconstruction capability of the decoder. Concretely, SFR can effectively reconstruct the high-resolution feature maps by recombining feature space, in which the space transformation matrix implicitly contained in a convolution layer can selectively highlight features at each position by leveraging the knowledge of label space in a self-learned way. Moreover, ODFF module can effectively fuse multilevel features with multiscale contextual information by feeding the feature maps into designed parallel offset-dilated convolutions, which enhances the feature representation capability of the decoder. Experiments on Cityscapes and CamVid datasets demonstrate the superior performance of our proposed methods embedded in ESPNet. Gege Qi, Zhengding Luo, Yuesheng Zhu |
ICTAI | 5 |
| 2019 | Data-to-Text Generation with Attention Recurrent UnitabstractRecurrent Neural Networks (RNNs) have shown promising results in many text generation tasks with their ability in modeling complex data distribution. However, the text generation model in their encoder or decoder RNNs still can not use the context efficiently. In this paper, we propose a novel Attention Recurrent Unit (ARU) to generate short descriptive texts conditioned on database records. Different from conventional approaches Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU), ARU allows the context information from the encoder to be aligned first inside the unit, which can improve the ability of content selection and surface realization for the model. And we also design a method called DoubleAtten to enhance the attention distribution when computing the generation probabilities. On the recently released ROTOWIRE dataset, extensive experimental results demonstrate that the ARU and DoubleAtten can efficiently improve the model performance for data-to-text generation task. Hechong Wang, Yuesheng Zhu, Zhiqiang Bai |
IJCNN | 3 |
| 2019 | A Novel Two-stage Separable Deep Learning Framework for Practical Blind WatermarkingabstractAs a vital copyright protection technology, blind watermarking based on deep learning with an end-to-end encoder-decoder architecture has been recently proposed. Although the one-stage end-to-end training (OET) facilitates the joint learning of encoder and decoder, the noise attack must be simulated in a differentiable way, which is not always applicable in practice. In addition, OET often encounters the problems of converging slowly and tends to degrade the quality of watermarked images under noise attack. In order to address the above problems and improve the practicability and robustness of algorithms, this paper proposes a novel two-stage separable deep learning (TSDL) framework for practical blind watermarking. Precisely, the TSDL framework is composed of noise-free end-to-end adversary training (FEAT) and noise-aware decoder-only training (ADOT). A redundant multi-layer feature encoding network is developed in FEAT to obtain the encoder, while ADOT is used to get the decoder which is robust and practical enough to accept any type of noise. Extensive experiments demonstrate that the proposed framework not only exhibits better stability, greater performance and faster convergence speed compared with current state-of-the-art OET methods, but is also able to resist high-intensity noises that have not been tested in previous works. Yang Liu 0308, Mengxi Guo, Jian Zhang 0018, Yuesheng Zhu |
ACM Multimedia | 4 |
| 2019 | An Energy Efficient Cooperation Design for Multi-UAVs Enabled Wireless Powered Communication NetworksabstractThis paper studies a multi-Unmanned Aerial Vehicles (UAVs) -enabled wireless powered communication network (WPCN). We present a novel use of UAVs with energy harvesting module. Thus, UAVs are used as flying base stations to collecting data and charging ground IoT devices via radio frequency (RF). Due to energy constraint and limited coverage of single UAV, its applications are limited. So we investigate a Multi-UAVs solution to enhance system transmitting performance via UAVs' cooperation. In particular, we focus on finding a tradeoff solution to prolong devices lifetime and improve the average uplink throughput efficiency among all ground users in a given time. To tackle this problem, we first consider a relaxed problem in which IoT devices are distributed in different residential blocks. A Work-Gain algorithm is designed for the relaxed problem. The performances comparison are presented in numerical results. Our performance analysis and simulation results have demonstrated that the proposed approach can improve the uplink performance approximately by 20% compared to other methods in certain circumstances. Yuxuan Wei, Zhiqiang Bai, Yuesheng Zhu |
VTC Fall | 3 |
| 2019 | Annular Sector Model for tracking multiple indistinguishable and deformable objects in occlusions
Guibo Luo, Zhenyu Weng, Yuesheng Zhu |
Neurocomputing | 4 |
| 2019 | A fast online spherical hashing method based on data sampling for large scale image retrieval
Zhenyu Weng, Yuesheng Zhu, Yinhe Lan, Long-Kai Huang |
Neurocomputing | 2 |
| 2019 | Restricted Connection Orthogonal Matching Pursuit for Sparse Subspace ClusteringabstractSparse Subspace Clustering (SSC) is one of the most popular methods for clustering a group of data points lying in a union of low-dimensional subspaces. However, SSC may suffer from heavy computational burden. Applying Orthogonal Matching Pursuit (OMP) on SSC may accelerate the computation but the trade-off is the loss of clustering accuracy. In this letter, we propose a noise-robust algorithm, Restricted Connection Orthogonal Matching Pursuit for Sparse Subspace Clustering (RCOMP-SSC), to improve the clustering accuracy and remain computationally efficient by restricting the number of connections of each data point during the iteration of OMP. Also, we develop a framework of control matrix to realize RCOMP-SSC. The proposed control matrix can be applied to other selection strategies for data point. Our analysis and experiments on synthetic data and two real-world databases (EYaleB & Usps) have demonstrated that our algorithm outperforms other methods in clustering accuracy. Yuesheng Zhu |
IEEE Signal Process. Lett. | 3 |
| 2018 | A New Temporal Deconvolutional Pyramid Network for Action Detection
Xiangli Ji, Guibo Luo, Yuesheng Zhu |
ACCV (4) | 3 |
| 2018 | Classification of Bone Tumor on CT Images Using Deep Convolutional Neural Network
Yang Li 0034, Wenyu Zhou, Guiwen Lv, Guibo Luo, Yuesheng Zhu |
ICANN (2) | 5 |
| 2018 | A New Sparse Subspace Clustering by Rotated Orthogonal Matching PursuitabstractSparse Subspace Clustering (SSC) is one of the most popular clustering methods in computer vision. However, SSC solved by convex programming tools may suffer from noise and outliers. Orthogonal Matching Pursuit (OMP) can solve these problems effectively but may lose correct information representation and accuracy. To overcome these problems, a new Sparse Subspace Clustering method by Rotated Orthogonal Matching Pursuit (SSC-ROMP) is proposed in this paper, in which once some vector chooses another vector as its neighbor, it has to be rotated to avoid it being chosen when the selected vector chooses neighbors. Also, Nonnegative Matrix Factorization is used for dimension reduction in SSC-ROMP. The analysis and experiment results on synthetic data and several open datasets have demonstrated that our new approach can achieve better performances in terms of accuracy and information representation than other subspace clustering algorithms and keep a good sparsity of clustering representation. Yuesheng Zhu, Guibo Luo |
ICIP | 2 |
| 2018 | Fast MRF-Based Hole Filling for View SynthesisabstractHole filling is one of the key issues in generating virtual view from video-plus-depth sequence by depth-image-based rendering. Hole filling method based on Markov random fields (MRF) is a practical way for view synthesis, but the traditional ones might introduce some foreground textures to the hole regions, and suffer from high computational complexity. In this letter, a fast MRF-based hole filling method is proposed for view synthesis, which is formulated as an energy minimization problem and is solved with loopy belief propagation (LBP). The energy function is optimized by employing the depth information to prevent the foreground textures filling holes. Furthermore, the LBP process maintains the visual consistency in the synthesized view by reserving all useful candidate labels. In addition, efficient belief propagation strategy is developed to optimize the LBP process, whose computational complexity is reduced to be linear with the number of candidate labels. Experimental results demonstrate the effectiveness of the proposed method with low running time and good visual consistency. Guibo Luo, Yuesheng Zhu |
IEEE Signal Process. Lett. | 2 |
| 2017 | Using personal information to aid in guessing passwords of Chinese websabstractIn order to remember easily, human beings may use personal information as part of their passwords. In recent years, incidents of Internet information leakage emerge incessantly, which provides more materials for password guessing, such as password database and personal information database. If we have known part of the users' personal information, how to use these data to guess passwords, the research related is still rare. In this paper, on the basis of more than 200 million leaked password accounts and 20 million personal information records in China, we analyze these data statistically and find out that at least 37.20% of these passwords contain personal information. Based on the latest Probabilistic Context-Free Grammars (PCFG) model, we propose a novel method to import personal information and generate passwords containing personal information. In offline attacks, experiments show that the efficiency of password guessing increases by 12.41% compared to PCFG 3.1 and 8.56% to order-4 Markov model after importing personal information in batches. In online attacks, the cracking probability is also improved significantly after importing personal information one by one. Yuesheng Zhu |
ICC | 2 |
| 2017 | Integrated metric learning with adaptive constraints for person re-identificationabstractPerson re-identification is an important technique to search a probe person against a set of gallery persons and metric learning methods have shown their effectiveness in matching person images. In this paper, an Integrated Metric Learning with Adaptive Constraints (IMLAC) method is proposed to promote the performance for person re-identification. In the method, the difference and commonness of an image pair are combined to define a novel integrated metric. Considering the complex variations of pedestrian images, a rule of adaptive pairwise constraints is extended for the integrated metric to further enhance separation and reunion between image pairs. Extensive experiments conducted on three person re-identification datasets including VIPeR, PRID450S and GRID indicate that the proposed method outperforms the state-of-the-art methods. Wenbin Yao, Chao Pei, Yuesheng Zhu |
ICIP | 4 |
| 2017 | Pedestrian detection with dynamic iterative bootstrappingabstractRecent years have seen the increasing importance of pedestrian detection, which is a key problem in computer vision. In this paper, we propose a novel pedestrian detection approach based on Faster R-CNN. In order to obtain high-quality candidate regions, relevant adjustments with more precise anchors are made for region proposal network. To resolve the data imbalance issue in the classifier training, we propose a dynamic iterative bootstrapping method where the hard negative examples are automatically selected and the weights of the network are updated iteratively by them to make the training more effective. The square method is used to optimize the multi-task loss in our approach, which can accelerate convergence and reduce sensitivity. Experimental results on different widely used benchmark datasets show that the proposed approach achieves better performance in comparison with other common methods. Chao Pei, Yuesheng Zhu |
ICIP | 3 |
| 2017 | Learning biased distance metrics with diversity regularizer for person re-identificationabstractMatching certain person across views, known as person re-identification, has attracted much attention in computer vision community. For this challenging problem, metric learning methods have shown their effectiveness in matching person images. In this paper, a Biased Metric Learning with Diversity Regularizer (BMLDR) method is proposed to promote the performance for person re-identification. The adaptive rule in the method assigns biases to different image pairs when the harder the pairs are, the larger weight they will be. By treating images pairs differently, the BMLDR method can exploit more discriminative information provided by the hard pairs thus can effectively distinguish between pairs. We validated the proposed method on three person re-identification datasets including VIPeR, PRID450S and GRID obtaining comparative performance compared to the state-of-the-art methods. Daiying Wang, Yuesheng Zhu |
VCIP | 3 |
| 2017 | A new spherical hashing method in a low-dimensional isotropic spaceabstractBy using a hypersphere to group the spatially coherent data points into the same bit, spherical hashing (SPH) can achieve a good performance in approximate nearest neighbor (ANN) search. However, when the data dimensionality rises, the data becomes sparse and the hypersphere needs to increase its radius to maintain the same coverage, which makes the data points in the hypersphere less coherent. To alleviate the effect brought from the high dimensionality of the data, a new hypersphere-based hashing method is proposed. By constructing a low-dimensional isotropic space where the variance of projection along each component is equal, both the similarity and the distribution of the original data can be preserved in this space. And then, the hashing functions are learnt by SPH in this space. The experiments on SIFT1M and GIST1M datasets show that the performance of SPH can be improved by our method and is superior to other state-of-the-art hashing methods in terms of recall and mAP performance. Yinhe Lan, Zhenyu Weng, Yuesheng Zhu |
VCIP | 3 |
| 2017 | An aligned bidirectional feature representation for person re-identificationabstractPerson re-identification plays an important role in intelligent surveillance analysis, in which it is critical to design a robust feature representation. In this paper, a novel feature representation named Aligned Bidirectional Maximum Occurrence (ABMO) is proposed. First, in order to handle background interference and spatial misalignment, Multi-layer Cellular Automata (MCA) is introduced for foreground segmentation and the alignment operation is put forward. Then, the bidirectional maximum occurrence feature representation is presented, which enhances representation completeness and robustness to illumination variations and viewpoint changes. Our approach is evaluated on an existing re-id dataset (PRID450s) as well as a new, more realistic re-id dataset (PRID365s), where the original surveillance materials are available. The results show that the proposed approach can obtain better re-identification performance compared with other state-of-the-art methods. Daiyin Wang, Yuesheng Zhu |
VCIP | 3 |
| 2017 | Depth estimation for outdoor image using couple dictionary learning and region detectionabstractDepth estimation from a single image is a significant and challenging task in computer vision. It is difficult to represent the correspondence between depth and RGB image without any prior information. Unlike previous approaches that only map the RGB images to the corresponding depth locally, we propose to combine the local depth with region-level and global scene structures. Firstly, the global layout is retrieved from similar images. Secondly, local depth estimated by the coupled dictionary learning (DL) formulation is combined with the global layout to maintain the global result. Finally, in order to further refine the depth estimated, we propose to detect the sky region in outdoor scene. In addition, several edge-preserving strategies are taken to clearly distinguish different objects. The experimental results demonstrate that our method represents a more realistic depth map than other methods on the popular public dataset Make3D. Qiqi Yao, Guibo Luo, Yuesheng Zhu |
VCIP | 3 |
| 2017 | A robust and synthesized-unseen watermarking for the DRM of DIBR-based 3D video
Xiyao Liu 0001, Fangfang Li 0004, Jingyu Du, Yang Guan, Yuesheng Zhu, Beiji Zou 0001 |
Neurocomputing | 5 |
| 2017 | A near-duplicate 3D video detection algorithm by using hypercomplex representations
Ziqiang Sun, Yuesheng Zhu, Xiaomei Xing, Guibo Luo, Xiyao Liu 0001 |
Multim. Tools Appl. | 2 |
| 2017 | Foreground Removal Approach for Hole Filling in 3D Video and FVV SynthesisabstractThe depth-image-based rendering is a key technique for 3D video and free viewpoint video synthesis. One of the critical problems in current synthesis methods is that the background (BG) occluded by the foreground objects might be exposed in the new view, and some holes are produced in the synthesized video. However, most of the traditional hole-filling approaches may bring some blurry effect or artifacts in the virtual view. In this paper, a foreground removal approach for hole filling is proposed, in which the foreground objects are removed from both the 2D video and its corresponding depth map, and then a BG video and its depth map are generated before the 3D warping and used to eliminate the holes in the synthesized video. Moreover, a BG extension method is applied in the reference view to prevent the large holes occurring along the border areas in the virtual view. Our analysis and experimental results have indicated that the proposed approach has better performance compared with the other methods in terms of the quality of synthesized video, computational complexity, and running time in multiview synthesis or multiframe synthesis. Guibo Luo, Yuesheng Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | WS-Rank: Bringing Sentences into Graph for Keyword Extraction
Yuesheng Zhu, Yu-Jia Ma |
APWeb (2) | 2 |
| 2016 | A Hole Filling Approach Based on Background Reconstruction for View Synthesis in 3D VideoabstractThe depth image based rendering (DIBR) plays a key role in 3D video synthesis, by which other virtual views can be generated from a 2D video and its depth map. However, in the synthesis process, the background occluded by the foreground objects might be exposed in the new view, resulting in some holes in the synthetized video. In this paper, a hole filling approach based on background reconstruction is proposed, in which the temporal correlation information in both the 2D video and its corresponding depth map are exploited to construct a background video. To construct a clean background video, the foreground objects are detected and removed. Also motion compensation is applied to make the background reconstruction model suitable for moving camera scenario. Each frame is projected to the current plane where a modified Gaussian mixture model is performed. The constructed background video is used to eliminate the holes in the synthetized video. Our experimental results have indicated that the proposed approach has better quality of the synthetized 3D video compared with the other methods. Guibo Luo, Yuesheng Zhu, Zhaotian Li, Liming Zhang 0002 |
CVPR | 2 |
| 2016 | Asymmetric distance for spherical hashingabstractUsually, most of hashing methods for information retrieval have a two-step procedure, embedding the data into a low-dimensional intermediate space and then quantizing them into binary codes. In the hyperplane-based hashing methods, the distance between the data in the intermediate space can replace the Hamming distance to improve the retrieval accuracy. In this paper, a novel asymmetric distance for the hypersphere-based method is proposed to improve the accuracy of similarity search. By showing that the distance in the intermediate space can approximate the Euclidean distance between the data points, more useful information can be taken to improve the retrieval accuracy. According to the characteristics of the hypersphere-based hashing method, various asymmetric distance models are developed and described. Our experiments with two datasets have demonstrated that the proposed method can improve the retrieval accuracy of the hypersphere-based hashing methods significantly and achieve the state-of-the-art recall performance. Zhenyu Weng, Wenbin Yao, Ziqiang Sun, Yuesheng Zhu |
ICIP | 4 |
| 2016 | Diversity regularized metric learning for person re-identificationabstractMetric learning is an effective method for person re-identification. It utilizes latent factors to find a suitable space for measuring distances. In general, a small number of factors are not powerful enough to match the pedestrians while a large number of factors cause high computational cost. In this paper, to balance this trade-off, a novel diversity regularized distance metric learning method is proposed. For feature representation, the local discriminative features are extracted from the source image and an adjacency maximal constraint is developed to handle the misaligned issue. Then a diversity regularizer is used to learn a metric by making the latent factors uncorrelated so that a small amount of latent factors can preserve effectiveness in measuring distances while reducing the computational burden. Our experimental results show that the proposed method with a small amount of factors can obtain comparative or even better performance compared to the state-of-art methods. Wenbin Yao, Zhenyu Weng, Yuesheng Zhu |
ICIP | 3 |
| 2016 | Asymmetric hashing with multi-bit quantization for image retrieval
Zhenyu Weng, Ziqiang Sun, Yuesheng Zhu |
Neurocomputing | 3 |
| 2016 | Analysis of minimum transmit sum power and scheduling power gain for multi-user MIMO-OFDM networks with rate constraintsabstractABSTRACT Minimum transmit sum power (MTSP) is of high theoretical and practical value in multi‐user rate‐constrained systems; it is, however, quite difficult to be numerically characterized in complex channels for the prohibitively high computational power required. In this paper, we present a computationally efficient method to approximate the MTSP in multi‐user multiple‐input multiple‐output orthogonal frequency division multiplexing (MU‐MIMO‐OFDM) wireless networks. Specifically, we propose both lower and upper bounds of the MTSP, which are asymptotically accurate in the limit of large K, the number of users. Then, we develop two iterative water‐filling algorithms to numerically solve the proposed bounds. These algorithms are with low complexity, that is, linear in K, and therefore enable the analysis of MTSP in complex channels even if K is large. Numerical results demonstrate the effectiveness of the bounds in approximating the MTSP and the high computational efficiency of the proposed iterative water‐filling algorithms. With the proposed bounds, we further numerically study scheduling power gain (SPG), which is defined as MTSP reduction achieved by scheduling resources over multiple channel blocks in time domain. We simulate the SPG in different wireless environments defined in Third Generation Partnership Project spatial channel extended model and find insignificant SPG in some cases, indicating that the benefit from scheduling over multiple channel blocks is limited and simply allocating resources within the present channel is sufficient. Our analysis on the MTSP and SPG provides guidelines on the design of resource schedulers in MU‐MIMO‐OFDM networks. Copyright © 2014 John Wiley & Sons, Ltd. Yuesheng Zhu, Li Ping 0001 |
Wirel. Commun. Mob. Comput. | 2 |
| 2015 | A novel robust video fingerprinting-watermarking hybrid scheme based on visual secret sharing
Xiyao Liu 0001, Yuesheng Zhu, Ziqiang Sun, Mengge Diao, Liming Zhang 0002 |
Multim. Tools Appl. | 2 |
| 2015 | An enhanced Kerberos protocol with non-interactive zero-knowledge proofabstractAbstract As one of the most important trusted third‐party‐based authentication protocols, Kerberos is widely used to provide authentication service in distributed networks. However, it is vulnerable to common brute force password‐guessing attacks because of its password‐based mechanism. Some enhanced Kerberos protocols based on public key cryptography were proposed as solutions, but they require excessive computation and communication resources. In this paper, a new enhanced Kerberos protocol with non‐interactive zero‐knowledge proof is proposed, in which the clients and the authentication server can mutually authenticate each other without revealing any information during the authentication process. Our security analysis and experimental results have shown that the proposed scheme can resist password‐guessing attacks and is more convenient and efficient than previous schemes. Copyright © 2014 John Wiley & Sons, Ltd. Yuesheng Zhu, Jinjiang Zhang |
Secur. Commun. Networks | 1 |
| 2014 | A low-complexity visual tracking approach with single hidden layer neural networksabstractVisual tracking algorithms based on deep learning have robust performance against variations in a complex environment because deep learning can learn generic features from numerous unlabeled images. However, due to the multilayer architecture, the deep learning trackers suffer from expensive computational costs and are not suitable for real-time applications. In this paper, a low-complexity visual tracking scheme with single hidden layer neural network is proposed based on denoising autoencoder. To further reduce the computational costs, feature selection is applied to simplify the networks and two optimization methods are used during the online tracking process. The experimental results have demonstrated that the proposed algorithm is about six times faster than the trackers based on deep nets and rapid enough for real-time applications with encouraging accuracy. Yuesheng Zhu, Guibo Luo |
ICARCV | 2 |
| 2014 | A digital blind watermarking scheme based on quantization index modulation in depth map for 3D videoabstract3D video provides an immersive experience to viewers and is getting more and more popular. The solution to create 3D video from 2D video is low-cost compared with that captures 3D video directly, and the generation of depth map from 2D video is a key in the 2D-3D video conversion systems. Therefore, protection of depth map is vital for 3D video. In this paper, a digital blind watermarking scheme based on Quantization Index Modulation (QIM) algorithm is proposed in which the copyright information is embedded in the DCT coefficients of depth map imperceptibly. The experimental results show that the proposed scheme has good robustness against video attacks such as salt noise, median filtering, wiener filtering, and scaling. In the meanwhile, the stereo video embedded watermarking can accomplish zero distortion in comparison with the original one. Yang Guan, Yuesheng Zhu, Xiyao Liu 0001, Guibo Luo, Ziqiang Sun, Liming Zhang 0002 |
ICARCV | 2 |
| 2014 | Linear transceiver design with intercarrier interference reduction for multiple-input-multiple-output with orthogonal frequency division multiplexing systemsabstractIn this study, a joint design of precoder and equaliser of a linear transceiver for multiple‐input–multiple‐output system with orthogonal frequency division multiplexing in the presence of intercarrier interference (ICI) is presented. The matrix structures of the precoder and equaliser are banded for the sake of reducing the computational complexity and feedback overhead from the receiver to the transmitter. The design criterion is to minimise the mean‐squared error subject to a total transmitted power constraint of which the power is allocated over space and frequency domains in the precoder. The authors use the Karush–Kuhn–Tucker conditions to derive an iterative procedure to obtain a convergent solution and a closed‐form procedure for the optimal full transceiver. Numerical results show that the banded precoder is an efficient scheme to improve the bit error rate (BER) of the transceiver by simply increasing its band size and can provide better BER performance than that of the existing jointly designed full transceiver in the presence of ICI. With small band sizes, the proposed transceiver can give performance close to that of the jointly designed full transceiver but with lower implementation complexity and feedback overhead to the transmitter. Fengyong Qian, Shu Hung Leung, Ruikai Mai, Yuesheng Zhu |
IET Commun. | 4 |
| 2014 | A robust video watermarking algorithm based on singular value decomposition and slope-based embedding technique
Hongyuan Chen, Yuesheng Zhu |
Multim. Tools Appl. | 2 |
| 2013 | An improved background subtraction approach in target detection and trackingabstractIn this paper, a novel background subtraction approach is proposed to avoid stationary foreground objects being merged into the background in target detection and tracking, in which an improved background model is designed by using virtual frames and the blur can be attenuated with this model when an object moves again after it stays for a long time. Moreover, the proposed model is fused with the eigenbackgrounds to improve the environmental adaptability. Our experimental results indicate that the proposed approach enhances the performance of target detection and tracking in intelligent surveillance and is superior to some state-of-the-art methods according to the precision-recall measurement. Hao Lai, Yuesheng Zhu, Zhenming Nong |
ICMV | 2 |
| 2013 | An Analytical Formula of Spatial Correlation Based on the Hierarchical Angle Structure for 3GPP Spatial Channel ModelabstractIn this paper, an analytical expression of normalized spatial correlation of a hierarchical angle structure for 3GPP Spatial Channel Model (SCM) is derived. Most of the existing formulas of the normalized spatial correlation were derived by simplifying the angle structure. In the paper, the hierarchical angle is considered as a function of multiple random variables, which can include those simplified angle structures as special cases. A closed-form expression of the probability density function of the hierarchical angle for uniform linear array is developed that can yield the analytical formula of normalized spatial correlation. Computer evaluation shows that the derived formula matches well with the simulated correlations generated from the 3GPP SCM. Lin Zhang 0025, Yuesheng Zhu, Shu Hung Leung |
VTC Fall | 2 |
| 2012 | Damped sinusoidal signals parameter estimation in frequency domain
Fengyong Qian, Shu Hung Leung, Yuesheng Zhu, Waiki Wong, Derek Chi-Wai Pao, Wing Hong Lau |
Signal Process. | 3 |
| 2012 | Simplified Precoder Design for MIMO Systems With Receive Correlation in Ricean ChannelsabstractIn this letter, a simplified algorithm for designing the precoder at the transmitter for multiple antenna systems with receive correlation over Ricean fading channels is proposed. From the common property of the asymptotic solutions of high and low SNR's, a simple power allocation structure is proposed, which greatly reduce the computational complexity in comparison with the existing precoder schemes. For the case of two transmit antennas, it becomes an optimal precoder. A simple enhancement procedure, which refines the power allocation to further improve the symbol error rate (SER) performance, is also derived. Unlike most of the existing precoder schemes which need expensive computational complexity but cannot guarantee the iterative algorithms to converge, the proposed algorithm is computationally efficient and can converge to the solution fast. Simulation results show that the proposed scheme achieves SER performance close to that of the optimal one, but requires less computational complexity with fast convergence. Lin Zhang 0025, Shu Hung Leung, Yuesheng Zhu |
IEEE Signal Process. Lett. | 4 |
| 2011 | A new multidirectional extrapolation hole-filling method for Depth-Image-Based RenderingabstractDepth-Image-Based Rendering (DIBR) is widely used to generate virtual view of a scene from a known view with associated depth map in 3D video applications. However, disocclusion arises in image warping of DIBR. Many hole-filling methods have been proposed such as constant color, horizontal interpolation, horizontal extrapolation, and variational inpainting, but they cause different types of annoying artifact for large holes with complex texture background. In this paper, a novel multidirectional extrapolation hole-filling method is proposed to enhance visual quality for large hole-filling with complex texture background. The proposed method uses neighbor pixels' texture features to estimate hole-filling direction in a pixel-by-pixel manner. Experimental results demonstrated that the proposed method could provide better visual quality compared with conventional methods for virtual views synthesis with high-quality depth map. Lai-Man Po, Shihang Zhang, Xuyuan Xu, Yuesheng Zhu |
ICIP | 4 |
| 2011 | A Controllable Error-Drift Elimination Scheme for Watermarking Algorithm in H.264/AVC StreamabstractEmbedding watermark into H.264/AVC streams directly can reduce computational complexity compared to encoder-based watermarking algorithms. However, it would cause intra error propagation and decrease the video quality. To improve the video quality and reduce the computational complexity, an improved compensation scheme is developed in this paper. In this scheme, only some of the integer transform (IT) coefficients are processed as opposed to process all the IT-coefficients as seen in other compensation methods. This approach helps to reduce the computation complexity of this compensation scheme. The simulation results have shown that our proposed method has lower computational complexity compared to other algorithms. Weijing Huo, Yuesheng Zhu, Hongyuan Chen |
IEEE Signal Process. Lett. | 2 |
| 2011 | String Searching Engine for Virus ScanningabstractA memory-efficient hardware string searching engine for antivirus applications is presented. The proposed QSV method is based on quick sampling of the input stream against fixed-length pattern prefixes, and on-demand verification of variable-length pattern suffixes. Patterns handled by the QSV method are required to have at least 16 bytes, and possess distinct 16-byte prefixes. The latter requirement can be fulfilled by a preprocessing procedure. The search engine uses the pipelined Aho-Corasick (P-AC) architecture developed by the first author to process 4 to 15-byte short patterns and a small number of exception cases. Our design was evaluated using the ClamAV virus database having 82,888 strings with a total size that exceeds 8 Mbyte. In terms of byte count, 99.3 percent of the pattern set is handled by the QSV method and 0.7 percent of the pattern set is handled by P-AC. A pattern with distinct 16-byte prefix only occupies up to three lookup table entries in QSV. The overall memory cost of our system is about 1.4 Mbyte, i.e., 1.4 bit per character of the ClamAV pattern set. The proposed method is memory-based, hence, updates to the pattern set can be accommodated by modifying the contents of the lookup tables without reconfiguring the hardware circuits. Derek Chi-Wai Pao, Yuesheng Zhu |
IEEE Trans. Computers | 5 |
| 2010 | A novel watermarking scheme with compensation in bit-stream domain for H.264/AVCabstractCurrently, most of the watermarking algorithms for H.264/AVC video coding standard are encoder-based due to their high perceptual quality. However, for the compressed video, they increase the computational burden to decode the video, embed the watermark, and then re-encode it. Obviously it is a bottleneck for real-time applications. Conventional watermarking algorithms in the bit-stream domain can reduce the computation complexity but result in the error propagation and PSNR loss. In this paper, a new fast watermarking scheme with compensation in bit-stream compressed video domain is proposed, in which the watermark is directly embedded into the quantized residual coefficients. A texture-based perceptual model is employed to decide whether a 4×4-block is appropriate for watermarking or not. A secret key is used to decide the actual location to be embedded in a 4×4-block. With the proposed compensation method, the video watermarking scheme can achieve high robustness and good visual quality without much bit-rate increase. The simulation results have shown that a high PSNR is achieved with little increase in bit rate, and the watermarks can resist common attacks. Yuesheng Zhu, Lai-Man Po |
ICASSP | 2 |