VLDB 2026 Research / reviewers in the wild / expert
Yuanjing Luo
dblp:255/1256
· DBLP profile ↗
13ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0001-9924-1100ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Video plot segmentation
Xichen Tan, Yuanjing Luo, Yunfan Ye, Chengyu Wang 0008, Fang Liu 0002, Zhiping Cai |
Expert Syst. Appl. | 2 |
| 2026 | SSRW-INN: An invertible neural network for screen shooting robust watermarking
Jiaxing Liao, Jiaohua Qin, Yuanjing Luo, Xuyu Xiang |
Pattern Recognit. | 3 |
| 2026 | DCL-Net: Decoupled Contrastive Learning Network for Image Manipulation LocalizationabstractExtracting discriminative forensic artifacts from high-dimensional latent spaces is pivotal for Image Manipulation Localization (IML). Nevertheless, the intrinsic feature entanglement arising from subtle structural discrepancies continues to impede precise localization. Specifically, manipulation artifacts are frequently overwhelmed by coherent background textures or ’soft boundaries’, rendering the trace-rich features inextricably mixed with intrinsic image content. To disentangle these intertwined representations, we propose the Decoupled Contrastive Learning Network (DCL-Net) for robust image tampering localization. Leveraging a Vision Transformer (ViT) backbone integrated with a Simple Feature Pyramid Network (SFPN), DCL-Net incorporates a novel Feature Decoupling Module (FDM). The FDM explicitly disentangles the feature space into foreground, background, and uncertainty regions-thereby effectively isolating manipulation cues from coherent background textures and capturing the transitional nature of ambiguous boundaries. Furthermore, to align the optimization objective with the intrinsic structure of manipulation traces, we introduce a prior-guided contrastive learning strategy that explicitly ’pushes away’ manipulated features from authentic and uncertain components. Finally, the Contrast Feature Aggregation Module (CFAM) employs a two-stage attention mechanism to refine these disentangled features, further suppressing redundant background details. Extensive experiments across five public benchmarks demonstrate that DCL-Net delivers state-of-the-art localization performance and exhibits strong robustness against common distortions. Xuyu Xiang, Jiaohua Qin, Wenyan Pan, Yuanjing Luo, Yun Tan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | ALLVB: All-in-One Long Video Understanding BenchmarkabstractFrom image to video understanding, the capabilities of Multi-modal LLMs (MLLMs) are increasingly powerful. However, most existing video understanding benchmarks are relatively short, which makes them inadequate for effectively evaluating the long-sequence modeling capabilities of MLLMs. This highlights the urgent need for a comprehensive and integrated long video understanding benchmark to assess the ability of MLLMs thoroughly. To this end, we propose ALLVB (ALL-in-One Long Video Understanding Benchmark). ALLVB's main contributions include: 1) It integrates 9 major video understanding tasks. These tasks are converted into video QA formats, allowing a single benchmark to evaluate 9 different video understanding capabilities of MLLMs, highlighting the versatility, comprehensiveness, and challenging nature of ALLVB. 2) A fully automated annotation pipeline using GPT-4o is designed, requiring only human quality control, which facilitates the maintenance and expansion of the benchmark. 3) It contains 1,376 videos across 16 categories, averaging nearly 2 hours each, with a total of 252k QAs. To the best of our knowledge, it is the largest long video understanding benchmark in terms of the number of videos, average duration, and number of QAs. We have tested various mainstream MLLMs on ALLVB, and the results indicate that even the most advanced commercial models have significant room for improvement. This reflects the benchmark's challenging nature and demonstrates the substantial potential for development in long video understanding. Xichen Tan, Yuanjing Luo, Yunfan Ye, Fang Liu 0002, Zhiping Cai |
AAAI | 2 |
| 2025 | BlkInfoM: versatile blockchain-based mapping mechanism for secure information transmissionabstractAbstract Information mapping is a widely adopted strategy in information hiding, leveraging the inherent features of carriers to convey hidden information without altering the carrier itself, thus maintaining integrity and avoiding detection. However, current mapping-based techniques face significant challenges, including potential data loss during transmission, limited capacity of carriers like images, and reliance on costly third-party storage solutions. To address these limitations, we introduce BlkInfoM, an innovative algorithm that utilizes blockchain’s decentralized, immutable, and traceable properties as a novel data source for secure information hiding. BlkInfoM leverages blockchain transaction data, such as Merkle hash values, timestamps, and locations, in combination with a reversible ASCII-based binary encoding to enable precise information-to-block matching. Experimental results demonstrate that BlkInfoM not only improves the success rate and efficiency of information mapping compared to traditional methods but also reduces operational costs by eliminating the need for third-party storage. This work highlights the potential of blockchain technology to revolutionize information hiding, offering enhanced security, scalability, and cost-effectiveness. Yuanjing Luo, Xichen Tan, Jiaohua Qin, Zhiping Cai |
Comput. J. | 1 |
| 2025 | Spatiotemporal attention-based real-time video watermarking
Quan Yan, Yuanjing Luo, Zhangdong Wang, Junhua Xi, Geming Xia, Zhiping Cai |
Data Min. Knowl. Discov. | 2 |
| 2025 | Post-encoding enhancement: A screen-shooting resistant watermarking scheme with feature enhancement and hybrid distortion simulation
Zhuangjifei Liu, Jiaohua Qin, Xuyu Xiang, Yuanjing Luo, Yun Tan |
Knowl. Based Syst. | 4 |
| 2025 | CLME: Robust Screen-Shooting Watermarking With Contrastive Learning and Mask-Guided EmbeddingabstractScreen-shooting watermarking technology plays a critical role in copyright protection and traceability. However, existing methods often lack sufficient robustness under strong noise interference and tend to introduce noticeable visual artifacts when embedding watermarks in smooth image regions, thereby degrading visual quality and increasing the risk of watermark exposure. To address these limitations, this paper proposes a Contrastive Learning and Mask-guided Embedding (CLME) framework for robust screen-shooting watermarking. The framework comprises two key components: (1) a mask-guided watermark embedding module that utilizes a Residual Dense Feature Extraction Block (RDFEB) and an Attention Mask Generation Block (AMGB) to adaptively embed watermarks into texture-rich regions, improving watermark invisibility; and (2) a contrastive learning-based watermark decoding network that employs contrastive loss to enhance the consistency of decoded features by treating features from the same watermarked image under different noise conditions as positive samples and features from different watermarked images as negative samples, thereby improving the robustness of watermark extraction. Experimental results demonstrate that the proposed CLME framework outperforms existing methods in terms of both robustness and visual quality. Specifically, at a shooting distance of 100 cm and a shooting angle of 40°, the watermark extraction accuracy reaches 99.58%, and the peak signal-to-noise ratio (PSNR) of the watermarked images reaches 42.624 dB, highlighting the framework’s strong potential for real-world applications. Jiaxing Liao, Jiaohua Qin, Yuanjing Luo, Wenyan Pan, Xuyu Xiang, Yun Tan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | PPIDM: Privacy-Preserving Inference for Diffusion Model in the CloudabstractCloud environments enhance diffusion model efficiency but introduce privacy risks, including intellectual property theft and data breaches. As AI-generated images gain recognition as copyright-protected works, ensuring their security and intellectual property protection in cloud environments has become a pressing challenge. This paper addresses privacy protection in diffusion model inference under cloud environments, identifying two key characteristics—denoising-encryption antagonism and stepwise generative nature—that create challenges such as incompatibility with traditional encryption, incomplete input parameter representation, and inseparability of the generative process. We propose PPIDM (Privacy-PreservingInference forDiffusionModels), a framework that balances efficiency and privacy by retaining lightweight text encoding and image decoding on the client while offloading computationally intensive U-Net layers to multiple non-colluding cloud servers. Client-side aggregation reduces computational overhead and enhances security. Experiments show PPIDM offloads 67% of Stable Diffusion computations to the cloud, reduces image leakage by 75%, and maintains high output quality (PSNR = 36.9, FID = 4.56), comparable to standard outputs. PPIDM offers a secure and efficient solution for cloud-based diffusion model inference. Zhangdong Wang, Zhihuang Liu, Yuanjing Luo, Tongqing Zhou, Jiaohua Qin, Zhiping Cai |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Fixing the Double Agent Vulnerability of Deep Watermarking: A Patch-Level Solution Against Artwork PlagiarismabstractIncreasing artwork plagiarism incidents stresses the urgent need for proper copyright protection on behalf of the creators. The latest development in this context focuses on embedding watermarks via deep encoder-decoder networks. However, we find that deep watermarking has a serious vulnerability on its robustness when facing deliberate plagiarism. To manifest it, we construct an attack that misuses watermarking encoder as a plagiarism lookout for bypassing copyright detection. As a remedy, we propose a patch-level deep watermarking framework (DIPW) to retain copyright evidence in essential patches with plagiarism resistance, inspired by a user study observation that subject elements in artworks are the principal plagiarism entities. Technically, DIPW adaptively finds the embedding patches by identifying a subset of non-overlapping and feature-rich objects; and tailors the model with dual-distortion losses and adversarial plagiarism noise injection for robustness. Experimental results demonstrate the superiority of DIPW in facilitating better robustness, secrecy, and imperceptibility with acceptable time burden. Yuanjing Luo, Tongqing Zhou, Shenglan Cui, Yunfan Ye, Fang Liu 0002, Zhiping Cai |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | IRWArt: Levering Watermarking Performance for Protecting High-quality Artwork ImagesabstractIncreasing artwork plagiarism incidents underscores the urgent need for reliable copyright protection for high-quality artwork images. Although watermarking is helpful to this issue, existing methods are limited in imperceptibility and robustness. To provide high-level protection for valuable artwork images, we propose a novel invisible robust watermarking framework, dubbed as IRWArt. In our architecture, the embedding and recovery of the watermark are treated as a pair of image transformations’ inverse problems, and can be implemented through the forward and backward processes of an invertible neural networks (INN), respectively. For high visual quality, we embed the watermark in high-frequency domains with minimal impact on artwork and supervise image reconstruction using a human visual system(HVS)-consistent deep perceptual loss. For strong plagiarism-resistant, we construct a quality enhancement module for the embedded image against possible distortions caused by plagiarism actions. Moreover, the two-stagecontrastive training strategy enables the simultaneous realization of the above two goals. Experimental results on 4 datasets demonstrate the superiority of our IRWArt over other state-of-the-art watermarking methods. Code: https://github.com/1024yy/IRWArt. Yuanjing Luo, Tongqing Zhou, Fang Liu 0002, Zhiping Cai |
WWW | 1 |
| 2021 | Coverless Image Steganography Based on Multi-Object RecognitionabstractMost of the existing coverless steganography approaches have poor robustness to geometric attacks, because these approaches use features of the entire image to map information, and these features are easy to be lost when being attacked. In order to improve the robustness against geometric attacks, we propose a coverless image steganography method based on multi-object recognition. In this scheme, we firstly use Faster RCNN to detect objects in the image data set, establish a mapping dictionary between object labels and binary sequence. Then we propose a novel mapping rule based on the filtered robust object labels for sequence generation. Therefore, an image can generate robust binary sequence through multi-objects recognition. In the transmission process, the transmitted image has not been modified, so our method can fundamentally resist steganalysis tools and avoid the attacker’s suspicions. In addition, the capacity and hiding rate of the proposed method are both considerable. Evaluations with under geometric attacks shows, on average,$3.1\times $robustness increase over other five coverless steganography methods. Moreover, evaluations under ten noise attacks shows, on average, the robustness of our method is also excellent, which reaches 83%. Yuanjing Luo, Jiaohua Qin, Xuyu Xiang, Yun Tan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Coverless steganography based on image retrieval of DenseNet features and DWT sequence mapping
Qiang Liu 0004, Xuyu Xiang, Jiaohua Qin, Yun Tan, Junshan Tan, Yuanjing Luo |
Knowl. Based Syst. | 6 |