VLDB 2026 Research / reviewers in the wild / expert
Yuanman Li
dblp:158/9418 · also Yuan-Man Li
· DBLP profile ↗
57ranked-venue papers
13as first author
47since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 7 first-author · 24 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 14 since 2021Security and privacy · 8 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deferred Poisoning: Making the Model More Vulnerable via Hessian SingularizationabstractRecent studies have shown that deep learning models are very vulnerable to poisoning attacks. Many defense methods have been proposed to address this issue. However, traditional poisoning attacks are not as threatening as commonly believed. This is because they often cause differences in how the model performs on the training set compared to the validation set. Such inconsistency can alert defenders that their data has been poisoned, allowing them to take the necessary defensive actions. In this paper, we introduce a more threatening type of poisoning attack called the Deferred Poisoning Attack. This new attack allows the model to function normally during the training and validation phases but makes it very sensitive to evasion attacks or even natural noise. We achieve this by ensuring the poisoned model's loss function has a similar value as a normally trained model at each input sample but with a large local curvature. A similar model loss ensures that there is no obvious inconsistency between the training and validation accuracy, demonstrating high stealthiness. On the other hand, the large curvature implies that a small perturbation may cause a significant increase in model loss, leading to substantial performance degradation, which reflects a worse robustness. We fulfill this purpose by making the model have singular Hessian information at the optimal point via our proposed Singularization Regularization term. We have conducted both theoretical and empirical analyses of the proposed method and validated its effectiveness through experiments on image classification tasks. Furthermore, we have confirmed the hazards of this form of poisoning attack under more general scenarios using natural noise, offering a new perspective for research in the field of security. Yuhao He 0001, Jinyu Tian 0001, Xianwei Zheng, Li Dong 0006, Yuanman Li, Jiantao Zhou 0001 |
AAAI | 5 |
| 2026 | Green Industrial Engineering on the Web: Agent-Driven Ant Colony Optimization Tuning for Energy-Efficient 3D Pipe Routing
Xuanhan Fan, Jibin Zhou, Han Liu 0008, Yuanman Li, Wei Wang 0077 |
WWW | 6 |
| 2026 | Detecting diffusion-based text tampering in scene images based on multi-scale feature fusion
Qingwen Zhu, Li Dong 0006, Yuanman Li, Haiwei Wu, Yushu Zhang 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Deep Robust Reversible WatermarkingabstractRobust Reversible Watermarking (RRW) enables perfect recovery of cover images and watermarks in lossless channels while ensuring robust watermark extraction under lossy channels. However, existing RRW methods, mostly non-deep learning-based, suffer from complex designs, high computational costs, and poor robustness limiting their practical applications. To address these issues, this paper proposes Deep Robust Reversible Watermarking (DRRW), a deep learning-based RRW scheme. DRRW introduces an Integer Invertible Watermark Network (iIWN) to achieve an invertible mapping between integer data distributions, fundamentally addressing the limitations of conventional RRW approaches. Unlike traditional RRW methods requiring task-specific designs for different distortions, DRRW adopts an encoder-noise layer-decoder framework, enabling adaptive robustness against various distortions through end-to-end training. During inference, the cover image and watermark are mapped into an overflowed stego image and latent variables. Arithmetic coding efficiently compresses these into a compact bitstream, which is embedded via reversible data hiding to ensure lossless recovery of both the image and watermark. To reduce pixel overflow, we introduce an overflow penalty loss, significantly shortening the auxiliary bitstream while improving both robustness and stego image quality. Additionally, we propose an adaptive weight adjustment strategy that eliminates the need to manually preset the watermark loss weight, ensuring improved training stability and performance. Experiments on multiple datasets demonstrate that the proposed DRRW addresses key challenges in current RRW methods and significantly advances the practical deployment of RRW. Wei Wang 0077, Chongyang Shi 0001, Li Dong 0006, Yuanman Li, Xiping Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Emotions Like Human: Self-Supervised Emotion Label Augmentation for Emotion Recognition in ConversationabstractThe rapid development of natural language processing (NLP) has enabled automatic emotion recognition in conversations (ERC). However, most existing models train on isolated one-hot emotion labels that overlook emotional connections between utterances and cannot effectively convey the extent and complexity of nuanced emotional perception similar to human understanding, resulting in reduced performance. To address this, we propose a self-supervisedSituation-awareEmotionLabel (SEL) generation method, which can automatically and effectively compute weights of all emotions for every utterance in a self-supervised way, utilizing no extra emotional annotated data. To generate SELs, we propose a self-supervised situation-aware method to first fusesituationinformation into utterances to generate emotion-perceptive representations.Situationis defined as the combination of the topic and surrounding people, with which emotions are widely thought to be more predictable. Then, we combine similarities of the emotion-perceptive representations with prior human knowledge to compute SELs. SEL can be deployed in most existing ERC models that train on one-hot emotion labels to enhance their performance. Experiments on six ERC models and four benchmark datasets indicate that SEL can effectively boost the performance of base models and attain state-of-the-art outcomes. Ablation study, case study, and quantitative analysis further show how SEL works. Qifeng Lai, Han Liu 0008, Yuanman Li, Zhiguo Gong, Wei Wang 0077 |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | Fast and Effective Video Inpainting via Implicit Motion-Guided Propagation and Sparse AttentionabstractVideo inpainting aims to reconstruct missing or corrupted regions in video frames, with applications in video editing, restoration, and special effects. Current deep video inpainting methods rely on optical flow to guide the propagation of effective features and spatiotemporal attention mechanisms to model relationships between frames. However, as an explicit motion representation, the optical flow extracted offline in preceding steps often suffers from instability and errors during estimation. These errors accumulate during subsequent content hallucination, resulting in artifacts and blurring. Meanwhile, although traditional spatiotemporal attention effectively captures frame relationships, its dense computational nature introduces redundant information, disrupting inpainting tasks and reducing efficiency. To address these issues, we propose an implicit motion-guided approach for efficient video inpainting. Instead of relying on optical flow, our method uses implicit motion in the latent feature space to guide the dual-domain propagation of images and features end-to-end, avoiding error accumulation from the independent optical flow estimation process. Additionally, we introduce a self-correcting module that enables feedback between image and feature propagation, reducing errors during propagation. Furthermore, we design an adaptive sparse video attention mechanism to focus on highly relevant regions, minimizing the impact of irrelevant information. Experimental results demonstrate that the proposed method outperforms state-of-the-art approaches both qualitatively and quantitatively, while also delivering superior efficiency. Yuanman Li, Bin Li 0011, Yanshan Li, Jiantao Zhou 0001, Xia Li 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Universal and Quality-Preserving Watermark Removal Based on Unpaired LearningabstractInvisible image watermarking plays a critical role in safeguarding AI-generated images, yet current removal methods face practical limitations in real-world settings. They compromise image quality, are tailored to specific watermarking schemes, or depend on original-watermarked image pairs. These limitations hinder reliable evaluations of watermark robustness. In this work, we propose a universal and quality-preserving watermark removal method based on unpaired learning. Specifically, we implement a three-stage training framework in which we first pre-train the remover to denoise corrupted images. Then the discriminator is trained to distinguish watermarked images from original ones. Finally, the remover and the discriminator are jointly trained in an adversarial manner to further strengthen the watermark elimination capability of the remover while enhancing image quality. Evaluated across various watermarking schemes, our method achieves watermark extraction error rates close to random guessing while maintaining high visual quality. This work reveals that most existing watermarking methods lack sufficient robustness. Fangjun Yan, Xiaojian Ji, Li Dong 0006, Weiwei Sun 0009, Yuanman Li |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2026 | FaceReclaim: Deep Traceability of Face-Swapped Images Through Feature Decoupling
Yuanman Li, Yuanchen Niu, Haiwei Wu, Yushu Zhang 0001, Jiantao Zhou 0001, Bin Li 0011 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | DAFN: A Robust Dual-Modal Framework for Online Video Source Platform IdentificationabstractIdentifying the source platform of an online video is a critical yet challenging task in digital forensics, complicated by proprietary transcoding and adversarial post-processing. The preceding forensic methods are limited to a single modality, analyzing either fragile container metadata or easily distorted spatiotemporal artifacts severely compromise their robustness. This paper pioneers a multi-modal framework that deeply integrates both information domains. We contend that naive fusion is insufficient as it fails to address cross-modal ambiguity-instances where modalities provide conflicting evidence. To resolve this, we propose DAFN, a novel Dual-modality Ambiguity-Aware Fingerprinting Network which extract features from both container structure and video content. At its core, DAFN introduces an adaptive fusion mechanism guided by a measure of cross-modal ambiguity. This mechanism, which incorporates a variational module to quantify the discrepancy between modalities, enables the model to intelligently arbitrate between evidence sources. We further contribute CNSNVD, a large-scale dataset with nine major platforms and six post-processing types. Extensive experiments show that DAFN significantly outperforms existing baselines, establishing a new state-of-the-art by effectively resolving modal ambiguity to achieve superior accuracy and resilience. Yulong Zheng, Xia Li 0006, Yuanman Li |
CloudCom | 4 |
| 2025 | Flow-Aware Dynamic Fusion for Video Inpainting DetectionabstractThe malicious misuse of deep learning-based video inpainting techniques poses significant security risks, highlighting the critical importance of accurately detecting inpainted regions in digital video content. However, existing methods suffer from limited accuracy in complex motion scenarios, and their cross-modal feature fusion efficiency is low due to simplistic integration strategies. To address these challenges, we propose FAD-Net, an end-to-end dual-branch framework that performs video inpainting localization by jointly modeling optical flow and RGB modalities to complement each other's capabilities. The optical flow branch employs a flow consistency error constraint and flow anomaly awareness module to mitigate the impact of inaccurate optical flow estimation, while the parallel RGB branch utilizes an encoder for spatial texture feature extraction and cross-frame attention for long-term temporal modeling. A bidirectional dynamic fusion module then adaptively integrates complementary features using motion-aware weights. Experimental results demonstrate that FAD-Net outperforms existing methods in accurately localizing the inpainted regions, particularly exhibiting excellent performance in complex dynamic scenarios and with unknown inpainting types, enabling reliable forensic analysis of tampered videos. Wenze Zheng, Xia Li 0006, Yuanman Li |
CloudCom | 4 |
| 2025 | Frequency-Enhanced Multi-Scale Progressive Detection for Document Image ManipulationabstractImages playa vital role in modern communication, cultural exchange, and information dissemination. Document images, as a digital form of textual content, are widely used in government, finance, and e-commerce scenarios. However, tampered with document images has become increasingly common and is often exploited for identity fraud and financial scams, posing serious threats to platform credibility and public interests. Compared to natural images, document forgeries typically involve small, visually inconspicuous regions against uniform backgrounds, making RGB-based detection methods less effective. To address these challenges, we propose MFPD-Net (Multi-scale Frequency Progressive Detector), which integrates a Transformer-based frequency feature enhancement module combined with an adaptive feature fusion strategy. Additionally, a Progressive multi-scale feedback decoder is designed to refine mask generation and improve localization accuracy. Extensive experiments on multiple document tampered detection tasks demonstrate that our method outperforms existing approaches in both accuracy and robustness, showing strong practicality and generalization capabilities. Qiyue Zhong, Xia Li 0006, Yuanman Li |
CloudCom | 3 |
| 2025 | HQA-VLAttack: Towards High Quality Adversarial Attack on Vision-Language Pre-Trained ModelsabstractBlack-box adversarial attack on vision-language pre-trained models is a practical and challenging task, as text and image perturbations need to be considered simultaneously, and only the predicted results are accessible. Research on this problem is in its infancy, and only a handful of methods are available. Nevertheless, existing methods either rely on a complex iterative cross-search strategy, which inevitably consumes numerous queries, or only consider reducing the similarity of positive image-text pairs but ignore that of negative ones, which will also be implicitly diminished, thus inevitably affecting the attack performance. To alleviate the above issues, we propose a simple yet effective framework to generate high-quality adversarial examples on vision-language pre-trained models, named HQA-VLAttack, which consists of text and image attack stages. For text perturbation generation, it leverages the counter-fitting word vector to generate the substitute word set, thus guaranteeing the semantic consistency between the substitute word and the original word. For image perturbation generation, it first initializes the image adversarial example via the layer-importance guided strategy, and then utilizes contrastive learning to optimize the image adversarial perturbation, which ensures that the similarity of positive image-text pairs is decreased while that of negative image-text pairs is increased. In this way, the optimized adversarial images and texts are more likely to retrieve negative examples, thereby enhancing the attack success rate. Experimental results on three benchmark datasets demonstrate that HQA-VLAttack significantly outperforms strong baselines in terms of attack success rate. Han Liu 0008, Zhi Xu 0008, Xiaotong Zhang 0003, Xiaoming Xu 0003, Fenglong Ma, Yuanman Li, Hong Yu 0005 |
NeurIPS | 7 |
| 2025 | TransCMFD: An adaptive transformer for copy-move forgery detection
Enji Liang, Zhongyun Hua, Yuanman Li, Xiaohua Jia |
Neurocomputing | 4 |
| 2025 | A memory-augmented multi-task collaborative framework for unsupervised traffic anomaly detection in driving videos
Rongqin Liang, Yuanman Li, Yingxin Yi, Jiantao Zhou 0001, Xia Li 0006 |
Pattern Recognit. | 2 |
| 2025 | Dual-Stream Image Sharing Chain Detection via Dynamic Information CompensationabstractImage Sharing Chain Detection (ISCD) aims to reconstruct the complete trajectory of an image's dissemination across social platforms and is an important task in multimedia forensics. Current methods using DCT histograms are insufficient in uncovering platform compression traces and exhibit limitations in detecting weak trace platforms. In this letter, we propose an innovative dual-stream ISCD framework via dynamic information compensation. This framework integrates features from both the frequency domain and the residual domain to extract compression characteristics. Unlike existing methods, we employ binary stereo DCT in the frequency domain to focus on the spatiality of compression operations. Additionally, we design a dynamic information compensation mechanism to enhance platform traces by storing compensation fingerprints of the sharing chains. Furthermore, we develop a new dataset, F-4OSN-SC, encompassing 4 platforms to simulate more realistic social networking scenarios. Experimental results demonstrate that our model outperforms existing methods across multiple datasets. Xinyi Su, Yuanman Li, Yulong Zheng, Xia Li 0006 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Rethinking Image Forgery Detection via Soft Contrastive Learning and Unsupervised ClusteringabstractImage forgery detection aims to detect and locate forged regions in an image. Most existing forgery detection algorithms formulate classification problems to classify pixels into forged or pristine. However, the definition of forged and pristine pixels is only relative within one single image, e.g., a forged region in image A is actually a pristine one in its source image B (splicing forgery). Such a relative definition has been severely overlooked by existing methods, which unnecessarily mix forged (pristine) regions across different images into the same category. To resolve this dilemma, we propose the FOrensic ContrAstive cLustering (FOCAL) method, a novel, simple yet very effective paradigm based on soft contrastive learning and unsupervised clustering for the image forgery detection. Specifically, FOCAL 1) designs a soft contrastive learning (SCL) to supervise the high-level forensic feature extraction in an image-by-image manner, explicitly reflecting the above relative definition; 2) employs an on-the-fly unsupervised clustering algorithm (instead of a trained one) to cluster the learned features into forged/pristine categories, further suppressing the cross-image influence from training data; and 3) allows to further boost the detection performance via simple feature-level concatenation without the need of retraining. Extensive experimental results over six public testing datasets demonstrate that our proposed FOCALsignificantlyoutperforms the state-of-the-art competitors by big margins: +24.8% onCoverage, +18.9% onColumbia, +17.3% onFF++, +15.3% onMISD, +15.0% onCASIAand +10.5% onNISTin terms of IoU (see also Fig. 1). The paradigm of FOCAL could bring fresh insights and serve as a novel benchmark for the image forgery detection task. The code is available athttps://github.com/HighwayWu/FOCAL. Haiwei Wu, Jiantao Zhou 0001, Yuanman Li |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Image Copy-Move Forgery Detection via Deep PatchMatch and Pairwise Ranking LearningabstractRecent advances in deep learning algorithms have shown impressive progress in image copy-move forgery detection (CMFD). However, these algorithms lack generalizability in practical scenarios where the copied regions are not present in the training images, or the cloned regions are part of the background. Additionally, these algorithms utilize convolution operations to distinguish source and target regions, leading to unsatisfactory results when the target regions blend well with the background. To address these limitations, this study proposes a novel end-to-end CMFD framework that integrates the strengths of conventional and deep learning methods. Specifically, the study develops a deep cross-scale PatchMatch (PM) method that is customized for CMFD to locate copy-move regions. Unlike existing deep models, our approach utilizes features extracted from high-resolution scales to seek explicit and reliable point-to-point matching between source and target regions. Furthermore, we propose a novel pairwise rank learning framework to separate source and target regions. By leveraging the strong prior of point-to-point matches, the framework can identify subtle differences and effectively discriminate between source and target regions, even when the target regions blend well with the background. Our framework is fully differentiable and can be trained end-to-end. Comprehensive experimental results highlight the remarkable generalizability of our scheme across various copy-move scenarios, significantly outperforming existing methods. Yuanman Li, Yingjie He 0003, Changsheng Chen 0001, Li Dong 0006, Bin Li 0011, Jiantao Zhou 0001, Xia Li 0006 |
IEEE Trans. Image Process. | 1 |
| 2025 | DuPMAM: An Efficient Dual Perception Framework Equipped With a Sharp Testing Strategy for Point Cloud AnalysisabstractThe challenges in point cloud analysis are primarily attributed to the irregular and unordered nature of the data. Numerous existing approaches, inspired by the Transformer, introduce attention mechanisms to extract the 3D geometric features. However, these intricate geometric extractors incur high computational overhead and unfavorable inference latency. To tackle this predicament, in this paper, we propose a lightweight and faster attention-based network, named Dual Perception MAM (DuPMAM), for point cloud analysis. Specifically, we present a novel simple Point Multiplicative Attention Mechanism (PMAM). It is implemented solely through single feed-forward fully connected layers, hence leading to lower model complexity and superior inference speed. Based on that, we further devise a dual perception strategy by constructing both a local attention block and a global attention block to learn fine-grained geometric and overall representational features, respectively. Consequently, compared to the existing approaches, our method has excellent perception of local details and global contours of the point cloud objects. In addition, we ingeniously design a Graph-Multiscale Perceptual Field (GMPF) testing strategy for model performance enhancement. It has significant advantage over the traditional voting strategy and is generally applicable to point cloud tasks, encompassing classification, part segmentation and indoor scene segmentation. Empowered by the GMPF testing strategy, DuPMAM delivers the new State-of-the-Art on the real-world dataset ScanObjectNN, the synthetic dataset ModelNet40 and the part segmentation dataset ShapeNet, and compared to the recent GB-Net, our DuPMAM trains 6 times faster and tests 2 times faster. Xianwei Zheng, Zhulun Yang, Xutao Li 0004, Jiantao Zhou 0001, Yuanman Li |
IEEE Trans. Multim. | 6 |
| 2025 | Cascaded Adaptive Graph Representation Learning for Image Copy-Move Forgery DetectionabstractIn the realm of image security, there has been a burgeoning interest in harnessing deep learning techniques for the detection of digital image copy-move forgeries, resulting in promising outcomes. The generation process of such forgeries results in a distinctive topological structure among patches, and collaborative modeling based on these underlying topologies proves instrumental in enhancing the discrimination of ambiguous pixels. Despite the attention received, existing deep learning models predominantly rely on convolutional neural networks, falling short in adequately capturing correlations among distant patches. This limitation impedes the seamless propagation of information and collaborative learning across related patches. To address this gap, our work introduces an innovative framework for image copy-move forensics rooted in graph representation learning. Initially, we introduce an adaptive graph learning approach to foster collaboration among related patches, dynamically learning the inherent topology of patches. The devised approach excels in promoting efficient information flow among related patches, encompassing both short-range and long-range correlations. Additionally, we formulate a cascaded graph learning framework, progressively refining patch representations and disseminating information to broader correlated patches based on their updated topologies. Finally, we propose a hierarchical cross-attention mechanism facilitating the exchange of information between the cascaded graph learning branch and a dedicated forgery detection branch. This equips our method with the capability to jointly grasp the homology of copy-move correspondences and identify inconsistencies between the target region and the background. Comprehensive experimental results validate the superiority of our proposed scheme, providing a robust solution to security challenges posed by digital image manipulations. Yuanman Li, Lanhao Ye, Haokun Cao, Wei Wang 0077, Zhongyun Hua |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | A Unified Environmental Network for Pedestrian Trajectory PredictionabstractAccurately predicting pedestrian movements in complex environments is challenging due to social interactions, scene constraints, and pedestrians' multimodal behaviors. Sequential models like long short-term memory fail to effectively integrate scene features to make predicted trajectories comply with scene constraints due to disparate feature modalities of scene and trajectory. Though existing convolution neural network (CNN) models can extract scene features, they are ineffective in mapping these features into scene constraints for pedestrians and struggle to model pedestrian interactions due to the loss of target pedestrian information. To address these issues, we propose a unified environmental network based on CNN for pedestrian trajectory prediction. We introduce a polar-based method to reflect the distance and direction relationship between any position in the environment and the target pedestrian. This enables us to simultaneously model scene constraints and pedestrian social interactions in the form of feature maps. Additionally, we capture essential local features in the feature map, characterizing potential multimodal movements of pedestrians at each time step to prevent redundant predicted trajectories. We verify the performance of our proposed model on four trajectory prediction datasets, encompassing both short-term and long-term predictions. The experimental results demonstrate the superiority of our approach over existing methods. Yuanman Li, Wei Wang 0077, Jiantao Zhou 0001, Xia Li 0006 |
AAAI | 2 |
| 2024 | Transformer-Based Image Inpainting Detection via Label Decoupling and Constrained Adversarial TrainingabstractImage inpainting based on generative adversarial networks (GANs) has achieved great success in producing visually plausible images and plays an important role in many real tasks. However, the techniques of image inpainting might also be maliciously used, e.g., altering or removing interesting objects to report fake news. Despite the promising performance of recently developed inpainting detection algorithms, they are built on convolutional neural networks (CNNs) with limited receptive fields. Consequently, they fail to fully capture the disparity between the inpainted regions and untouched regions and thus are ineffective in obtaining fine-grained detection results. In this work, we develop a new image inpainting detection approach. First, we propose a locally enhanced transformer architecture tailored for image inpainting detection. Unlike previous CNN-based methods, our approach leverages both the short-range and long-range dependencies of pixels, enabling the learning of diverse statistical behaviors of inpainted and untouched regions. Second, to mitigate the distraction caused by near-edge pixels with a mixed nature during training, we propose decoupling the label into a body map and a soft-edge map, and then a cross-modality attention module is designed to propagate their information interactively. It demonstrates that our decoupling strategy outperforms the conventional edge supervision in enhancing detection accuracy. Finally, we devise a constrained adversarial training methodology in consideration of the confrontational generation procedure of deep image inpainting methods. It shows that our constrained adversarial training further enhances the detection performance by adaptively introducing interference noise in the inpainted regions. Extensive experiments validate the superiority of our scheme compared to existing CNN-based methods, showcasing its desirable detection generalizability for both deep inpainting and traditional inpainting algorithms. Yuanman Li, Liangpei Hu, Li Dong 0006, Haiwei Wu, Jinyu Tian 0001, Jiantao Zhou 0001, Xia Li 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Text-Driven Traffic Anomaly Detection With Temporal High-Frequency Modeling in Driving VideosabstractTraffic anomaly detection (TAD) in driving videos is critical for ensuring the safety of autonomous driving and advanced driver assistance systems. Previous single-stage TAD methods primarily rely on frame prediction, making them vulnerable to interference from dynamic backgrounds induced by the rapid movement of the dashboard camera. While two-stage TAD methods appear to be a natural solution to mitigate such interference by pre-extracting background-independent features (such as bounding boxes and optical flow) using perceptual algorithms, they are susceptible to the performance of first-stage perceptual algorithms and may result in error propagation. In this paper, we introduce TTHF, a novel single-stage method aligning video clips with text prompts, offering a new perspective on traffic anomaly detection. Unlike previous approaches, the supervised signal of our method is derived from languages rather than orthogonal one-hot vectors, providing a more comprehensive representation. Further, concerning visual representation, we propose to model the high frequency of driving videos in the temporal domain. This modeling captures the dynamic changes of driving scenes, enhances the perception of driving behavior, and significantly improves the detection of traffic anomalies. In addition, to better perceive various types of traffic anomalies, we carefully design an attentive anomaly focusing mechanism that visually and linguistically guides the model to adaptively focus on the visual context of interest, thereby facilitating the detection of traffic anomalies. It is shown that our proposed TTHF achieves promising performance, outperforming state-of-the-art competitors by +5.4% AUC on the DoTA dataset and achieving high generalization on the DADA dataset. Rongqin Liang, Yuanman Li, Jiantao Zhou 0001, Xia Li 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | STGlow: A Flow-Based Generative Framework With Dual-Graphormer for Pedestrian Trajectory PredictionabstractThe pedestrian trajectory prediction task is an essential component of intelligent systems. Its applications include but are not limited to autonomous driving, robot navigation, and anomaly detection of monitoring systems. Due to the diversity of motion behaviors and the complex social interactions among pedestrians, accurately forecasting their future trajectory is challenging. Existing approaches commonly adopt generative adversarial networks (GANs) or conditional variational autoencoders (CVAEs) to generate diverse trajectories. However, GAN-based methods do not directly model data in a latent space, which may make them fail to have full support over the underlying data distribution. CVAE-based methods optimize a lower bound on the log-likelihood of observations, which may cause the learned distribution to deviate from the underlying distribution. The above limitations make existing approaches often generate highly biased or inaccurate trajectories. In this article, we propose a novel generative flow-based framework with a dual-graphormer for pedestrian trajectory prediction (STGlow). Different from previous approaches, our method can more precisely model the underlying data distribution by optimizing the exact log-likelihood of motion behaviors. Besides, our method has clear physical meanings for simulating the evolution of human motion behaviors. The forward process of the flow gradually degrades complex motion behavior into simple behavior, while its reverse process represents the evolution of simple behavior into complex motion behavior. Furthermore, we introduce a dual-graphormer combined with the graph structure to more adequately model the temporal dependencies and the mutual spatial interactions. Experimental results on several benchmarks demonstrate that our method achieves much better performance compared to previous state-of-the-art approaches. Rongqin Liang, Yuanman Li, Jiantao Zhou 0001, Xia Li 0006 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | Uformer-ICS: A U-Shaped Transformer for Image Compressive Sensing ServiceabstractMany service computing applications require real-time dataset collection from multiple devices, necessitating efficient sampling techniques to reduce bandwidth and storage pressure. Compressive sensing (CS) has found wide-ranging applications in image acquisition and reconstruction. Recently, numerous deep-learning methods have been introduced for CS tasks. However, the accurate reconstruction of images from measurements remains a significant challenge, especially at low sampling rates. In this paper, we propose Uformer-ICS as a novel U-shaped transformer for image CS tasks by introducing inner characteristics of CS into transformer architecture. To utilize the uneven sparsity distribution of image blocks, we design an adaptive sampling architecture that allocates measurement resources based on the estimated block sparsity, allowing the compressed results to retain maximum information from the original image. Additionally, we introduce a multi-channel projection (MCP) module inspired by traditional CS optimization methods. By integrating the MCP module into the transformer blocks, we construct projection-based transformer blocks, and then form a symmetrical reconstruction model using these blocks and residual convolutional blocks. Therefore, our reconstruction model can simultaneously utilize the local features and long-range dependencies of image, and the prior projection knowledge of CS theory. Experimental results demonstrate its significantly better reconstruction performance than state-of-the-art deep learning-based CS methods. Zhongyun Hua, Yuanman Li, Yushu Zhang 0001, Yicong Zhou |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Image Sharing Chain Detection VIA Sequence-To-Sequence ModelabstractImage sharing chain detection aims to recover the sharing history of an image downloaded from online social networks (OSNs), including the ever-shared OSNs and their orders, which is an important task in the multimedia forensics community. Most of the existing algorithms directly treat the sharing chain detection as a classification problem by simply assigning a unique label to each sharing chain. Such a strategy though seems straightforward, it ignores the inherent properties of the sharing chain which can be regarded as a time sequence that carries the sharing history of an online image. In this paper, we suggest a new sharing chain detection framework via Sequence-to-Sequence (Seq2Seq) model. Different from previous classification based approaches, our model detects the sharing chain of online image progressively via a decoder. This progressive manner can fully utilize the decoded chain, which is embedded into a series of learned representations. Experimental results show that our method can detect sharing chains involving up to three OSNs, and exhibits much better performance than conventional ones. Jiaxiang You, Yuanman Li, Rongqin Liang, Yuxuan Tan, Jiantao Zhou 0001, Xia Li 0006 |
ICASSP | 2 |
| 2023 | Image Copy-Move Forgery Detection via Deep Cross-Scale PatchMatchabstractThe recently developed deep algorithms achieve promising progress in the field of image copy-move forgery detection (CMFD). However, they have limited generalizability in some practical scenarios, where the copy-move objects may not appear in the training images or cloned regions are from the background. To address the above issues, in this work, we propose a novel end-to-end CMFD framework by integrating merits from both conventional and deep methods. Specifically, we design a deep cross-scale patchmatch method tailored for CMFD to localize copy-move regions. In contrast to existing deep models, our scheme aims to seek explicit and reliable point-to-point matching between source and target regions using features extracted from high-resolution scales. Further, we develop a manipulation region location branch for source/target separation. The proposed CMFD framework is completely differentiable and can be trained in an end-to-end manner. Extensive experimental results demonstrate the high generalizability of our method to different copy-move contents, and the proposed scheme achieves significantly better performance than existing approaches. Yingjie He 0003, Yuanman Li, Changsheng Chen 0001, Xia Li 0006 |
ICME | 2 |
| 2023 | Multiple degraded image restoration via degradation history estimationabstractImage restoration is a fundamental task in low-level computer vision. Most existing algorithms assume that the input image has a single known degradation type. In reality, images usually contain multiple degradations, making the restoration challenging. Though recent works restore the multiple degraded images, they assume that the degradation history is known. Obviously, such an ideal assumption often does not hold in real applications. This work proposes a novel restoration framework for multiple degraded images via degradation history estimation. Specifically, we first develop a sequential model to estimate the degradation history, including both the degradation operation chain and the corresponding parameters. By resorting to designed self-attention and cross-attention mechanisms, our method can effectively model the correlation of the input image, degradation operation chain, and parameters. Then, we apply our estimation framework for the multiple degraded image restoration, without requiring the degradation history. Experiment results demonstrate much better performance than existing approaches. Minhua Liu, Yuanman Li, Rongqin Liang, Jiaxiang You, Xia Li 0006 |
ICME | 2 |
| 2023 | Multi-scale Target-Aware Framework for Constrained Splicing Detection and LocalizationabstractConstrained image splicing detection and localization (CISDL) is a fundamental task of multimedia forensics, which detects splicing operation between two suspected images and localizes the spliced region on both images. Recent works regard it as a deep matching problem and have made significant progress. However, existing frameworks typically perform feature extraction and correlation matching as separate processes, which may hinder the model's ability to learn discriminative features for matching and can be susceptible to interference from ambiguous background pixels. In this work, we propose a multi-scale target-aware framework to couple feature extraction and correlation matching in a unified pipeline. In contrast to previous methods, we design a target-aware attention mechanism that jointly learns features and performs correlation matching between the probe and donor images. Our approach can effectively promote the collaborative learning of related patches, and perform mutual promotion of feature learning and correlation matching. Additionally, in order to handle scale transformations, we introduce a multi-scale projection method, which can be readily integrated into our target-aware framework that enables the attention process to be conducted between tokens containing information of varying scales. Our experiments demonstrate that our model, which uses a unified pipeline, outperforms state-of-the-art methods on several benchmark datasets and is robust against scale transformations. Yuxuan Tan, Yuanman Li, Limin Zeng, Jiaxiong Ye, Wei Wang 0077, Xia Li 0006 |
ACM Multimedia | 2 |
| 2023 | Enabling Large-Capacity Reversible Data Hiding Over Encrypted JPEG BitstreamsabstractCloud computing offers advantages in handling the exponential growth of images but also entails privacy concerns on outsourced private images. Reversible data hiding (RDH) over encrypted images has emerged as an effective technique for securely storing and managing confidential images in the cloud. Most existing schemes only work on uncompressed images. However, almost all images are transmitted and stored in compressed formats such as JPEG. Recently, some RDH schemes over encrypted JPEG bitstreams have been developed, but these works have some disadvantages such as a small embedding capacity (particularly for low quality factors), damage to the JPEG format, and file size expansion. In this study, we propose a permutation-based embedding technique that allows the embedding of significantly more data than existing techniques. Using the proposed embedding technique, we further design a large-capacity RDH scheme over encrypted JPEG bitstreams, in which a grouping method is designed to boost the number of embeddable blocks. The designed RDH scheme allows a content owner to encrypt a JPEG bitstream before uploading it to a cloud server. The cloud server can embed additional data (e.g., copyright and identification information) into the encrypted JPEG bitstream for storage, management, or other processing purpose. A receiver can losslessly recover the original JPEG bitstream using a decryption key. Comprehensive evaluation results demonstrate that our proposed design can achieve approximately twice the average embedding capacity compared to the best prior scheme while preserving the file format without file size expansion. Zhongyun Hua, Yifeng Zheng 0001, Yongyong Chen, Yuanman Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Boundary-Sensitive Loss Function With Location Constraint for Hard Region SegmentationabstractIn computer-aided diagnosis and treatment planning, accurate segmentation of medical images plays an essential role, especially for some hard regions including boundaries, small objects and background interference. However, existing segmentation loss functions including distribution-, region- and boundary-based losses cannot achieve satisfactory performances on these hard regions. In this paper, a boundary-sensitive loss function with location constraint is proposed for hard region segmentation in medical images, which provides three advantages: i) our Boundary-Sensitive loss (BS-loss) can automatically pay more attention to the hard-to-segment boundaries (e.g., thin structures and blurred boundaries), thus obtaining finer object boundaries; ii) BS-loss also can adjust its attention to small objects during training to segment them more accurately; and iii) our location constraint can alleviate the negative impact of the background interference, through the distribution matching of pixels between prediction and Ground Truth (GT) along each axis. By resorting to the proposed BS-loss and location constraint, the hard regions in both foreground and background are considered. Experimental results on three public datasets demonstrate the superiority of our method. Specifically, compared to the second-best method tested in this study, our method improves performance on hard regions in terms of Dice similarity coefficient (DSC) and 95% Hausdorff distance (95%HD) of up to 4.17% and 73% respectively. In addition, it also achieves the best overall segmentation performance. Hence, we can conclude that our method can accurately segment these hard regions and improve the overall segmentation performance in medical images. Jie Du 0001, Peng Liu 0070, Yuanman Li, Tianfu Wang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Image Operation Chain Detection with Machine Translation FrameworkabstractThe aim of operation chain detection for a given manipulated image is to reveal the operations involved and the order in which they were applied, which is significant for image processing and multimedia forensics. Currently,allexisting approaches simply treat image operation chain detection as a classification problem and consider only chains of at most two operations. Considering the complex interplay between operations and the exponentially increasing solution space, detecting longer operation chains is extremely challenging. To address this issue, in this work, we devise a new methodology for image operation chain detection. Different from existing approaches based on classification modeling, we strategically conduct operation chain detection within a machine translation framework. Specifically, the chain in our work is modeled as a sentence in a target language, with each possible operation represented by a word in that language. When executing chain detection, we propose first transforming the input image into a sentence in a latent source language from the learned deep features. Then, we propose translating the latent language into the target language within a machine translation framework and finally decoding all operations, arranged in order. Besides, a chain inversion strategy and a bi-directional modeling mechanism are developed to improve the detection performance. We further design a weighted cross-entropy loss to alleviate the problems presented by imbalance among chain lengths and chain categories. Our method can detect operation chains containing up to seven operations and obtains very promising results in various scenarios for the detection of both short and long chains. Yuanman Li, Jiaxiang You, Jiantao Zhou 0001, Wei Wang 0077, Xin Liao 0001, Xia Li 0006 |
IEEE Trans. Multim. | 1 |
| 2023 | AMS-Net: Adaptive Multi-Scale Network for Image Compressive SensingabstractRecently, deep convolutional neural networks have been applied to image compressive sensing (CS) to improve reconstruction quality while reducing computation cost. Existing deep learning-based CS methods can be divided into two classes: sampling image at single scale and sampling image across multiple scales. However, these existing methods treat the image low-frequency and high-frequency components equally, which is an obstruction to get a high reconstruction quality. This paper proposes an adaptive multi-scale image CS network in wavelet domain called AMS-Net, which fully exploits the different importance of image low-frequency and high-frequency components. First, the discrete wavelet transform is used to decompose an image into four sub-bands, namely the low-low (LL), low-high (LH), high-low (HL), and high-high (HH) sub-bands. Considering that the LL sub-band is more important to the final reconstruction quality, the AMS-Net allocates it a larger sampling ratio, while allocating the other three sub-bands a smaller one. Since different blocks in each sub-band have different sparsity, the sampling ratio is further allocated block-by-block within the four sub-bands. Then a dual-channel scalable sampling model is developed to adaptively sample the LL and the other three sub-bands at arbitrary sampling ratios. Finally, by unfolding the iterative reconstruction process of the traditional multi-scale block CS algorithm, we construct a multi-stage reconstruction model to utilize multi-scale features for further improving the reconstruction quality. Experimental results demonstrate that the proposed model outperforms both the traditional and state-of-the-art deep learning-based methods. Zhongyun Hua, Yuanman Li, Yongyong Chen, Yicong Zhou |
IEEE Trans. Multim. | 3 |
| 2022 | Synchronous Bi-directional Pedestrian Trajectory Prediction with Error Compensation
Ce Xie, Yuanman Li, Rongqin Liang, Li Dong 0006, Xia Li 0006 |
ACCV (6) | 2 |
| 2022 | Watermark-Preserving Keypoint Enhancement for Screen-Shooting Resilient WatermarkingabstractScreen-shooting resilient (SSR) watermark is a special kind of robust watermarking. One can extract the watermark message even the embedded image communicates via a physical screen to the camera channel. The keypoint-based SSR watermarking is one promising solution to realize such screen-to-camera communication. The enhanced keypoints were used to locate the embedding region and then perform watermark embedding. However, the keypoint-based SSR watermarking treats the critical two steps, keypoint enhancement and watermark embedding, independently, neglecting their inter-play. This work proposes a watermark-preserving keypoint enhancement algorithm for SSR watermarking. Specifically, we resort to a convex constrained optimization framework to unify keypoint enhancement and watermark embedding. Multiple constraints are imposed to simultaneously ensure the watermark validity and blind synchronization of embedding regions. Our method enables jointly optimizing the watermarking distortion and keypoint enhancement. The proposed method achieves superior watermark extraction accuracy while retaining better watermarked image quality when compared with previous works. Li Dong 0006, Chengbin Peng 0001, Yuanman Li, Weiwei Sun 0009 |
ICME | 4 |
| 2022 | Robust Document Image Forgery Localization Against Image BlendingabstractDigital documents, as a twin of hard copy, are increasingly being used as credible evidence. Unfortunately, digital document images easily suffer forgery or malicious manipulation, with the availability of sophisticated image editing tools. To verify and detect the possible forgeries for a given document, a number of forensic schemes have been developed. However, in the real-world scenario, the doctored image could be further processed or transmitted over a channel with unknown distortion, which dramatically degrade the forgery detection performance. In this work, we make the first step towards designing a robust document image forgery localization against image blending. Specifically, we propose an encoder-decoder neural network architecture consisting of three modules. The first module is responsible for capturing the multi-scale features from the high-level feature maps, and the remaining two attention-based modules aim to extract low-level local features and high-level global features. For training the model, we construct a dedicated forgery document database processed by several recent image blending procedures. Extensive experiments demonstrate the effectiveness and superiority of the proposed method in detecting the forgery that undergoes image blending. The source code, models and the constructed image dataset are publicly available at https://github.com/lwp0201/Image-Forgery-Localization-Against-Image-Blending. Weipeng Liang, Li Dong 0006, Rangding Wang, Diqun Yan, Yuanman Li |
TrustCom | 5 |
| 2022 | Privacy-preserving and verifiable deep learning inference based on secret sharing
Jia Duan, Jiantao Zhou 0001, Yuanman Li, Caishi Huang |
Neurocomputing | 3 |
| 2022 | Robust Matrix Factorization via Minimum Weighted Error Entropy CriterionabstractLearning the intrinsic low-dimensional subspace from high-dimensional data is a key step for many social systems of artificial intelligence. In practical scenarios, the observed data are usually corrupted by many types of noise, which brings a great challenge for social systems to analyze data. As a commonly utilized subspace learning technique, robust low-rank matrix factorization (LRMF) focuses on recovering the underlying subspaces in a noisy environment. However, most of the existing approaches simply assume that the noise contaminating the data is independent identically distributed (i.i.d.), such as Gaussian and Laplacian noises. This assumption, though greatly simplifies the underlying learning problem, may not hold for more complex non-i.i.d. noise widely existed in social systems. In this work, we suggest a robust LRMF approach to deal with various types of noise in a unified manner. Different from traditional algorithms, noise in our framework is modeled using an independent and piecewise identically distributed (i.p.i.d.) source, which employs a collection of distributions, instead of a single one to characterize the statistical behavior of the underlying noise. Assisted by the generic noise model, we then design a robust LRMF algorithm under the information-theoretic learning (ITL) framework through a new minimization criterion. By adopting the half-quadratic optimization paradigm, we further deliver an optimization strategy for our proposed method. Experimental results on both synthetic and real data are provided to demonstrate the superiority of our proposed scheme. Yuanman Li, Jiantao Zhou 0001, Junyang Chen 0001, Jinyu Tian 0001, Li Dong 0006, Xia Li 0006 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2022 | Trajectory Forecasting Based on Prior-Aware Directed Graph Convolutional Neural NetworkabstractPredicting the motion trajectories of moving agents in complex traffic scenes, such as crossroads and roundabouts, plays an important role in cooperative intelligent transportation systems. Nevertheless, accurately forecasting the motion behavior in a dynamic scenario is challenging due to the complex cooperative interactions between moving agents. Graph Convolutional Neural Network has recently been employed to deal with the cooperative interactions between agents. Despite the promising performance of resulting trajectory prediction algorithms, many existing graph-based approaches model interactions with an undirected graph, where the strength of influence between agents is assumed to be symmetric. However, such an assumption often does not hold in reality. For example, in pedestrian or vehicle interaction modeling, the moving behavior of a pedestrian or vehicle is highly affected by the ones ahead, while the ones ahead usually pay less attention to the ones behind. To fully exploit the asymmetric attributes of the cooperative interactions in intelligent transportation systems, in this work, we present a directed graph convolutional neural network for multiple agents trajectory prediction. First, we propose three directed graph topologies, i.e., view graph, direction graph, and rate graph, by encoding different prior knowledge of a cooperative scenario, which endows the capability of our framework to effectively characterize the asymmetric influence between agents. Then, a fusion mechanism is devised to jointly exploit the asymmetric mutual relationships embedded in constructed graphs. Furthermore, a loss function based on Cauchy distribution is designed to generate multimodal trajectories. Experimental results on complex traffic scenes demonstrate the superior performance of our proposed model when compared with existing approaches. Jie Du 0001, Yuanman Li, Xia Li 0006, Rongqin Liang, Zhongyun Hua, Jiantao Zhou 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Deep Generative Model for Image Inpainting With Local Binary Pattern Learning and Spatial AttentionabstractDeep learning (DL) has demonstrated its powerful capabilities in the field of image inpainting. The DL-based image inpainting approaches can produce visually plausible results, but often generate various unpleasant artifacts, especially in the boundary and highly textured regions. To tackle this challenge, in this work, we propose a new end-to-end, two-stage (coarse-to-fine) generative model through combining a local binary pattern (LBP) learning network with an actual inpainting network. Specifically, the first LBP learning network using U-Net architecture is designed to accurately predict the structural information of the missing region, which subsequently guides the second image inpainting network for better filling the missing pixels. Furthermore, an improved spatial attention mechanism is integrated into the image inpainting network, by considering the consistency not only between the known region with the generated one, but also within the generated region itself. Extensive experiments on public datasets includingCelebA-HQ,PlacesandParis StreetViewdemonstrate that our model generates better inpainting results than the state-of-the-art competing algorithms, both quantitatively and qualitatively. The source code and trained models are available athttps://github.com/HighwayWu/ImageInpainting. Haiwei Wu, Jiantao Zhou 0001, Yuanman Li |
IEEE Trans. Multim. | 3 |
| 2022 | Weighted Error Entropy-Based Information Theoretic Learning for Robust Subspace RepresentationabstractIn most of the existing representation learning frameworks, the noise contaminating the data points is often assumed to be independent and identically distributed (i.i.d.), where the Gaussian distribution is often imposed. This assumption, though greatly simplifies the resulting representation problems, may not hold in many practical scenarios. For example, the noise in face representation is usually attributable to local variation, random occlusion, and unconstrained illumination, which is essentially structural, and hence, does not satisfy the i.i.d. property or the Gaussianity. In this article, we devise a generic noise model, referred to as independent and piecewise identically distributed (i.p.i.d.) model for robust presentation learning, where the statistical behavior of the underlying noise is characterized using a union of distributions. We demonstrate that our proposed i.p.i.d. model can better describe the complex noise encountered in practical scenarios and accommodate the traditional i.i.d. one as a special case. Assisted by the proposed noise model, we then develop a new information-theoretic learning framework for robust subspace representation through a novel minimum weighted error entropy criterion. Thanks to the superior modeling capability of the i.p.i.d. model, our proposed learning method achieves superior robustness against various types of noise. When applying our scheme to the subspace clustering and image recognition problems, we observe significant performance gains over the existing approaches. Yuanman Li, Jiantao Zhou 0001, Jinyu Tian 0001, Xianwei Zheng, Yuan Yan Tang |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Temporal Pyramid Network for Pedestrian Trajectory Prediction with Multi-SupervisionabstractPredicting human motion behavior in a crowd is important for many applications, ranging from the natural navigation of autonomous vehicles to intelligent security systems of video surveillance. All the previous works model and predict the trajectory with a single resolution, which is relatively ineffective and difficult to simultaneously exploit the long-range information (e.g., the destination of the trajectory), and the short-range information (e.g., the walking direction and speed at a certain time) of the motion behavior. In this paper, we propose a temporal pyramid network for pedestrian trajectory prediction through a squeeze modulation and a dilation modulation. Our hierarchical framework builds a feature pyramid with increasingly richer temporal information from top to bottom, which can better capture the motion behavior at various tempos. Furthermore, we propose a coarse-to-fine fusion strategy with multi-supervision. By progressively merging the top coarse features of global context to the bottom fine features of rich local context, our method can fully exploit both the long-range and short-range information of the trajectory. Experimental results on two benchmarks demonstrate the superiority of our method. Our code and models will be available upon acceptance. Rongqin Liang, Yuanman Li, Xia Li 0006, Yi Tang 0008, Jiantao Zhou 0001, Wenbin Zou |
AAAI | 2 |
| 2021 | Detecting Adversarial Examples from Sensitivity Inconsistency of Spatial-Transform DomainabstractDeep neural networks (DNNs) have been shown to be vulnerable against adversarial examples (AEs), which are maliciously designed to cause dramatic model output errors. In this work, we reveal that normal examples (NEs) are insensitive to the fluctuations occurring at the highly-curved region of the decision boundary, while AEs typically designed over one single domain (mostly spatial domain) exhibit exorbitant sensitivity on such fluctuations. This phenomenon motivates us to design another classifier (called dual classifier) with transformed decision boundary, which can be collaboratively used with the original classifier (called primal classifier) to detect AEs, by virtue of the sensitivity inconsistency. When comparing with the state-of-the-art algorithms based on Local Intrinsic Dimensionality (LID), Mahalanobis Distance (MD), and Feature Squeezing (FS), our proposed Sensitivity Inconsistency Detector (SID) achieves improved AE detection performance and superior generalization capabilities, especially in the challenging cases where the adversarial perturbation levels are small. Intensive experimental results on ResNet and VGG validate the superiority of the proposed SID. Jinyu Tian 0001, Jiantao Zhou 0001, Yuanman Li, Jia Duan |
AAAI | 3 |
| 2021 | Task Scheduling Game Optimization for Mobile Edge ComputingabstractTask scheduling on edge computing servers is an important issue that affects user experience. Existing scheduling methods require centralized control to achieve the best overall performance. However, it is impractical to force all users to act according to centralized control. We propose a distributed edge computing server task scheduling model based on game theory. Our method comprehensively considers the link quality from the mobile device to the server and the server's computing resource allocation when selecting edge computing servers, and achieves a balance between link quality and computing resources. Once the Nash equilibrium is reached, our model can provide different QoS for users of different priorities. Acceleration methods are proposed to achieve the Nash equilibrium faster. The simulation results show that the proposed model can provide differentiated services while optimizing the scheduling of computing resources, and ensure that the algorithm achieves an approximate Nash equilibrium in polynomial time. Wei Wang 0077, Bingxian Lu, Yuanman Li, Wei Wei 0006, Jianqing Li 0001, Shahid Mumtaz, Mohsen Guizani |
ICC | 3 |
| 2021 | A Transformer based Approach for Image Manipulation Chain DetectionabstractImage manipulation chain detection aims to identify the existence of involved operations and also their orders, playing an important role in multimedia forensics and image analysis. However,all the existing algorithms model the manipulation chain detection as a classification problem, and can only detect chains containing up to two operations. Due to the exponentially increased solution space and the complex interactions among operations, how to reveal a long chain from a processed image remains a long-standing problem in the multimedia forensic community. To address this challenge, in this paper, we propose a new direction for manipulation chain detection. Different from previous works, we treat the manipulation chain detection as a machine translation problem rather than a classification one, where we model the chains as the sentences of a target language, and each word serves as one possible image operation. Specifically, we first transform the manipulated image into a deep feature space, and further model the traces left by the manipulation chain as a sentence of a latent source language. Then, we propose to detect the manipulation chain through learning the mapping from the source language to the target one under a machine translation framework. Our method can detect manipulation chains consisting of up to five operations, and we obtain promising results on both the short-chain detection and the long-chain detection. Jiaxiang You, Yuanman Li, Jiantao Zhou 0001, Zhongyun Hua, Weiwei Sun 0009, Xia Li 0006 |
ACM Multimedia | 2 |
| 2021 | Visually secure image encryption using adaptive-thresholding sparsification and parallel compressive sensing
Zhongyun Hua, Yuanman Li, Yicong Zhou |
Signal Process. | 3 |
| 2021 | Robust High-Capacity Watermarking Over Online Social Network Shared ImagesabstractIn recent years, online social networks (OSNs) have become extremely popular and been one of the most common ways for storing and distributing images. Naturally, such widespread availability of OSN makes it a viable channel for transmitting additional data along with the image sharing. However, various lossy operations, e.g., resizing and compression, conducted by OSN platforms impose great challenges for designing a robust watermarking scheme over OSN shared images. In this paper, we tackle this challenge and propose a robust high-capacity watermarking technique, by using Facebook as a representative OSN. To achieve the satisfactory robustness, we first probe into Facebook and recover the image manipulation mechanism via a deep convolutional neural network (DCNN) approach. Assisted with the precise knowledge on the lossy channel offered by Facebook, we then suggest a DCT-domain image watermarking method that is highly robust against the lossy operations on Facebook, even without any error correcting codes (ECC). The proposed technique is also extended to other popular OSNs, e.g., Wechat and Twitter. Extensive experimental results are provided to show the superior performance of our method in terms of the embedding capacity, data extraction accuracy, and quality of the reconstructed images. Weiwei Sun 0009, Jiantao Zhou 0001, Yuanman Li, Ming Cheung 0001, James She |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Secure and Verifiable Outsourcing of Large-Scale Nonnegative Matrix Factorization (NMF)abstractNowadays, cloud computing platforms are becoming increasingly prevalent and readily available, providing alternative and economic services for resource-constrained clients to perform large-scale computations. This work addresses the problem of secure outsourcing of large-scale nonnegative matrix factorization (NMF) to a cloud in a way that the client can verify the correctness of the results with small overhead. The protection of the input matrix is achieved by a random permutation and scaling encryption mechanism. By exploiting the iterative nature of NMF computation, we propose a single-round verification strategy, which can be proved to be quite effective. Theoretical and experimental results are provided to show the superior performance of the proposed scheme. Jia Duan, Jiantao Zhou 0001, Yuanman Li |
IEEE Trans. Serv. Comput. | 3 |
| 2020 | Privacy-Preserving distributed deep learning based on secret sharing
Jia Duan, Jiantao Zhou 0001, Yuanman Li |
Inf. Sci. | 3 |
| 2019 | Robust Subspace Clustering With Independent and Piecewise Identically Distributed Noise ModelingabstractMost of the existing subspace clustering (SC) frameworks assume that the noise contaminating the data is generated by an independent and identically distributed (i.i.d.) source, where the Gaussianity is often imposed. Though these assumptions greatly simplify the underlying problems, they do not hold in many real-world applications. For instance, in face clustering, the noise is usually caused by random occlusions, local variations and unconstrained illuminations, which is essentially structural and hence satisfies neither the i.i.d. property nor the Gaussianity. In this work, we propose an independent and piecewise identically distributed (i.p.i.d.) noise model, where the i.i.d. property only holds locally. We demonstrate that the i.p.i.d. model better characterizes the noise encountered in practical scenarios, and accommodates the traditional i.i.d. model as a special case. Assisted by this generalized noise model, we design an information theoretic learning (ITL) framework for robust SC through a novel minimum weighted error entropy (MWEE) criterion. Extensive experimental results show that our proposed SC scheme significantly outperforms the state-of-the-art competing algorithms. Yuanman Li, Jiantao Zhou 0001, Xianwei Zheng, Jinyu Tian 0001, Yuan Yan Tang |
CVPR | 1 |
| 2019 | Fast and Effective Image Copy-Move Forgery Detection via Hierarchical Feature Point MatchingabstractCopy-move forgery is one of the most commonly used manipulations for tampering digital images. Keypoint-based detection methods have been reported to be very effective in revealing copy-move evidence due to their robustness against various attacks, such as large-scale geometric transformations. However, these methods fail to handle the cases when copy-move forgeries only involve small or smooth regions, where the number of keypoints is very limited. To tackle this challenge, we propose a fast and effective copy-move forgery detection algorithm through hierarchical feature point matching. We first show that it is possible to generate a sufficient number of keypoints that exist even in small or smooth regions by lowering the contrast threshold and rescaling the input image. We then develop a novel hierarchical matching strategy to solve the keypoint matching problems over a massive number of keypoints. To reduce the false alarm rate and accurately localize the tampered regions, we further propose a novel iterative localization technique by exploiting the robustness properties (including the dominant orientation and the scale information) and the color information of each keypoint. Extensive experimental results are provided to demonstrate the superior performance of our proposed scheme in terms of both efficiency and accuracy. Yuanman Li, Jiantao Zhou 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2017 | SIFT Keypoint Removal via Directed Graph Construction for Color ImagesabstractAs one of the most successful feature extraction algorithms, scale invariant feature transform (SIFT) has been widely employed in many applications. Recently, the security of SIFT against malicious attack has been attracting increasing attention, and several techniques have been devised to remove SIFT keypoints intentionally. However, most of the existing methods still suffer from the following three problems: 1) the keypoint removal rate achieved by many techniques is unsatisfactory when removing keypoints within multiple octaves; 2) noticeable artifacts are introduced in the processed image, especially in those highly textured regions; and 3) the color information is totally neglected, precluding the widespread adoption of those methods. To tackle these challenges, in this paper, we propose a novel SIFT keypoint removal framework. By modeling the difference of Gaussian space as a directed weighted graph, we derive a set of strict inequality constraints to remove a SIFT keypoint along a pre-constructed acyclic path. To minimize the incurred distortion, the path is strategically designed over the directed graph. Furthermore, we propose a simple yet effective optimization framework for recovering the color information of the keypoint-removed image. Extensive experiments are provided to show the superior performance of our proposed scheme over the state-of-the-art techniques, in both the scenarios of removing keypoints in a single octave and in multiple octaves. Yuanman Li, Jiantao Zhou 0001, An Cheng |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2016 | Secure and Verifiable Outsourcing of Nonnegative Matrix Factorization (NMF)abstractCloud computing platforms are becoming increasingly prevalent and readily available nowadays, providing us alternative and economic services for resource-constrained clients to perform large-scale computation. In this work, we address the problem of secure outsourcing of large-scale nonnegative matrix factorization (NMF) to a cloud in a way that the client can verify the correctness of results with small overhead. The input matrix protection is achieved by a lightweight, permutation-based encryption mechanism. By exploiting the iterative nature of NMF computation, we propose a single-round verification strategy, which can be proved to be effective. Both theoretical and experimental results are given to demonstrate the superior performance of our scheme. Jia Duan, Jiantao Zhou 0001, Yuanman Li |
IH&MMSec | 3 |
| 2016 | SIFT Keypoint Removal and Injection via Convex RelaxationabstractScale invariant feature transform (SIFT), as one of the most popular local feature extraction algorithms, has been widely employed in many computer vision and multimedia security applications. Although SIFT has been extensively investigated from various perspectives, its security against malicious attacks has rarely been discussed. In this paper, we show that the SIFT keypoints can be effectively removed with minimized distortion on the processed image. The SIFT keypoint removal is formulated as a constrained optimization problem, where the constraints are carefully designed to suppress the existence of local extrema and prevent generating new keypoints within a local cuboid in the scale space. To hide the traces of performing SIFT keypoint removal, we then propose to inject a large number of fake SIFT keypoints into the previously cleaned image with minimized distortion. As demonstrated experimentally, our proposed SIFT removal and injection algorithms significantly outperform the state-of-the-art techniques. Furthermore, it is shown that the combined SIFT keypoint removal and injection attack strategy is capable of defeating the most powerful forensic detector designed for SIFT keypoint removal. Our results suggest that an authorization mechanism is required for SIFT-based systems to verify the validity of the input data, so as to achieve high reliability. Yuanman Li, Jiantao Zhou 0001, An Cheng, Xianming Liu 0005, Yuan Yan Tang |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2015 | Sift keypoint removal via convex relaxationabstractDue to the high robustness against various image transformations, Scale Invariant Feature Transform (SIFT) has been widely employed in many computer vision and multimedia security areas to extract image local features. Though SIFT has been extensively studied from various perspectives, its security against malicious attack has rarely been addressed. In this work, we demonstrate that the SIFT keypoints can be effectively removed, without introducing serious distortion on the image. This is achieved by formulating the SIFT keypoint removal as a constrained optimization problem, where the constraints are well-designed to suppress the existence of local extremum and prevent generating new keypoints within a local cuboid in the scale space. We show that such optimization problem in the ideal case is non-convex. To make the computation feasible, we propose a relaxation technique to convexify the original problem, while maximally preserving the solution space. As demonstrated experimentally, our proposed SIFT removal algorithm significantly outperforms the state-of-the-arts in terms of keypoint removal rate-distortion (KRR-D) performance. Our results imply that an authorization mechanism is required for SIFT-based systems to verify the validity of the input data, so as to achieve high reliability. An Cheng, Yuanman Li, Jiantao Zhou 0001 |
ICME | 2 |
| 2015 | Ciphertext-Only Attack on an Image Homomorphic Encryption Scheme with Small Ciphertext ExpansionabstractThe paper "An Efficient Image Homomorphic Encryption Scheme with Small Ciphertext Expansion" In Proc. ACM MM'13, pp.803--812) presented a novel image homomorphic encryption approach achieving significant reduction of the ciphertext expansion. In the current work, we study the security of this cryptosystem under a ciphertext-only attack (COA). We show that our proposed COA is effective in generating a sketch of great fidelity of the original image. Experimental results are provided to verify the validity of the proposed attack strategy. Yunyu Li, Jiantao Zhou 0001, Yuanman Li |
ACM Multimedia | 3 |
| 2015 | Anti-Forensics of Lossy Predictive Image CompressionabstractImage compression evidence has been utilized as an important forensic feature to justify image authenticity. However, some recent studies showed that the compression evidence of block transform-based image coding, e.g., JPEG and JPEG2000, can be effectively erased by adding designed dither noise in the transform domain. In this paper, we demonstrate that it is also feasible to hide the compression evidence of lossy predictive image coding, a class of compression paradigm widely employed in critical scenarios. To tackle the challenging issue of error propagation inherent to predictive coding, we design a prediction-direction preserving strategy, allowing us to add dither noise in the prediction error (PE) domain, while minimizing the incurred distortion. Extensive experimental results are provided to verify the effectiveness of the proposed anti-forensic algorithm for lossy predictive image coding. Yuanman Li, Jiantao Zhou 0001 |
IEEE Signal Process. Lett. | 1 |
| 2014 | Sparsity-driven reconstruction of ℓ∞-decoded imagesabstractIn this paper, we propose a sparsity-driven restoration technique to improve the coding performance of the ℓ∞-decoded images. This is achieved by incorporating a ℓ1minimization term into a ℓ2optimization framework, where the weighting vectors balancing the relative contribution of each term are appropriately determined. The ℓ∞constraints inherent to ℓ∞-constrained predictive coding are also included to narrow the solution space, leading to more accurate estimation. Experimental results show that our proposed scheme significantly improves the ℓ2performance of the ℓ∞-decoded images, while still preserving a tight error bound on every single pixel. In addition, when comparing with the existing scheme of restoring the ℓ∞-decoded images, the PSNR gain can be up to 1 dB. Yuanman Li, Jiantao Zhou 0001 |
ICIP | 1 |